Skip to content

session-eval

Provenance

  • Source: .spec/spexcode/spec-eval/session-eval/spec.md
  • Source SHA-256: 0c8c5e624063112d92dc4823d5ba7a20a0649bdf6e5376e9dfe4cb7b3deb5074

raw source

A human deciding whether to merge — or just wanting to see what a session has done so far — shouldn't have to hand-read the diff and hunt the evidence. Give them one proof of work — the session's measured eval readings evidence, its diff, and the merge gates, in a single beautiful page, available for any session (it comes into its own at review, but the human can open it any time). It is fully DERIVED: it costs the agent nothing and can never go stale, because it is built from what the system already knows and generated on the fly each time it is opened. This is the optimizer's measured loss, marshaled at the moment a human decides. And it is not a separate world: the human's directive re-homed un-merged worktree evals into the ONE Evals route family — a session's evaluation is the same pages as the project's, scoped.

expanded spec

One engine, thin faces (the [[eval-history]] / buildBoard pattern). The engine is sessioneval.ts in [[spec-eval]] — the marshaled evaluation lives with the evaluation package and is the one place the eval engine reaches into the full review state ([[manager-cockpit]]'s reviewPayload), and only the self-contained export needs that full bundle. It runs ONLY on the backend: buildExportModel(id) joins the payload's diff (grouped per spec node) with each affected scenario's [[eval-tab]] timeline (latest reading per scenario — verdict, expected, the content-addressed evidence) and the gates; renderExportHtml(model) emits ONE self-contained HTML document, evidence inlined as data-URIs ([[eval-core]]'s cache) so it stands alone as a plain file. The whole model is rooted at the session's worktree — the eval timelines (freshness and readings reflect that branch, not the backend's checkout) AND the spec tree itself (loadSpecs at the worktree root): a node the branch added is a first-class changed node here — present with its declared scenarios and filed readings while the branch is still un-merged — because the worktree's .spec is the branch's pending proposal ([[source-of-truth]]), not invisible. A session with no worktree reads the backend checkout unchanged. The headline is DERIVED (the node, else the branch) — there is no agent-authored claim, manifest, or narrative. A frontend node with no eval.md shows as an honest blind spot, never hidden.

The session scope is scenario-shaped, not node-shaped. Three independent axes feed the model and never stand in for one another. Declared comes from the current worktree's eval.md. Affected is derived against the session merge-base: a scenario enters when its own code axis (else its node's inherited code: axis) intersects the worktree diff, when that scenario's semantic contract (description + expected, the SAME scenarioHash projection [[eval-core]] uses) changed, or when this session actually filed a reading for it. Fresh remains the live [[eval-core]] judgment on the latest reading after the scenario has entered scope. The impact predicate lives once in sessioneval.ts; every session face consumes the already-scoped model. Touching a node's spec.md, eval.md, or one sibling scenario therefore never sweeps the eval.md's other scenarios into the session. Contract comparison follows an eval.md rename/reparent back to its merge-base path, so a pure move changes no scenario; the measurement axis reads the live worktree sidecar, so a reading this session just filed is visible before its eventual evidence commit even when the session changed no code. An affected scenario stays visible whether its latest reading is fresh, stale, legacy, or missing: stale and missing are review work, not reasons to hide it. A declared affected scenario with no effective reading is the precise blind/unmeasured case (the contract is known, the measurement is absent). A changed frontend node with no eval.md is unknown coverage (there is no declared contract to count as a scenario) and remains called out separately; it never inflates the scenario accounting or the list's scenario filter counts. The session toolbar renders that accounting once, as its mutually exclusive categories: fresh pass, fresh fail, measured-but-stale/unscored work needing review, and blind declarations add back to the affected total; unknown remains outside it. It does not repeat a measured/declared fraction beside that complete decomposition.

That predicate is also a public exact-revision impact projection, not a session-UI heuristic. Given a repository plus base/head selectors, it resolves both selectors once, pins every declaration and diff read to those object ids, then verifies that the selectors still resolve to the same ids before publishing. Its output carries the resolved base, head, a content revision, each node's review causes, and every scenario delta: code|contract|measurement impact reasons, selector-hit commits/symbols, base/head scenarioHash, and separate semantic versus metadata movement. A scenario rename is one removed declaration plus one added declaration; description/expected movement is semantic and moves the hash, while test/code/related/tags are metadata and do not. A scenario's effective code is its own code relation when present, otherwise the node's complete code relation including selectors. Node/scenario related movement remains node review context only and never spreads impact to every scenario. Its full graph is the explicit GET /api/evals/impact?scope=<id> read; a paged list item keeps its own canonical reason, but the list never carries the scope-sized graph as metadata.

The range contract is ancestry, not an arbitrary two-tree comparison: base must be an ancestor of head, normally the session merge-base and its exact HEAD. A divergent pair is explicitly unavailable; the engine never substitutes another merge-base behind the caller's back. A live session adds one immutable worktree overlay captured under the existing content-revision double read. That overlay includes staged, unstaged, renamed, deleted, and untracked paths, exact old/new hunk ranges and bytes, plus the complete worktree spec/eval declarations, so dirty semantic/metadata edits and branch-new nodes are first-class impact input rather than a reason to disable the face. Snapshot or Git transport failure is unavailable; ordinary dirty work is not.

Selector-aware code impact reuses [[eval-core]]'s scenarioCodeAxis and [[code-anchor]]'s ONE parse -> extractor -> resolve -> revision-hunk/range intersection engine. A changed shared file therefore selects only the scenario whose named unit was hit; no impact-local parser, extractor registry, or path matcher exists. Removed declarations resolve and intersect on base; added declarations do so on head. A structurally invalid, dead, ambiguous, unavailable, or unextractable selector makes the projection explicitly unavailable with the selector and repair named, even when the changed-path set is empty. A caller's symbolic base/head moving during direct projection is an explicit retry error; session routes pass immutable OIDs and use their existing outer content-revision fence to retry live HEAD movement. Neither condition may publish zero impact or certify a falsely current result.

Committed code impact reads the same complete Git-derived event fold as anchor drift, but projects a plain ancestry window without acknowledgement filtering. Base and head declarations each bind their path identity at that revision, so deletion, rename chains, path reuse, and incomparable forks retain the same meaning as the gate. The live overlay derives both path images from one rename-aware change set; a pure rename owns no line, while a rename with edits intersects the old and new units on their respective sides.

This is the one product predicate. Scoped /api/evals, its list/detail model, summaries, and export consume the same projection; none retain a path-only scenario candidate fallback. Each projection reads the base/head changed-path set once, batch-reads each distinct exact .spec tree once for both spec and eval declarations, and shares selector source/window/hunk results inside that build. The scope's reading timelines obey the same rule: the whole affected node set is read through ONE batched timeline pass, so the off-history content probes and the anchor probes union across every node in scope and issue one child per probe kind. Reading a timeline per node inside the scope loop is the same defect one level down — it multiplies a build's Git children by the node count, and a single open's cost must grow with the selected session's node closure, never with a per-node or per-reading spawn. It adds no second resident cache, generation, or gate: the result enters the existing content-revision/projection cache like every other session model.

When that one model needs history or anchor-hunk facts, its build joins [[source-of-truth]]'s one ledger build context before either demand runs. If its writer lock is free, the model retains the ordinary transaction: one decoded integrity-checked snapshot, lock, and possible replacement instead of an event-stream transaction followed by a hunk-fact transaction. If an unrelated live writer already owns that lock, the foreground model reads the ledger's current atomic integrity-checked snapshot and derives missing facts through the same Git adapters without waiting; only those contended additions are discarded. It never adds a second cache or substitutes an empty verdict, and corrupt input or Git-interpretation movement remains loud ([[event-ledger-demand]]).

The toolbar summary is a coherent projection, not a small fetch. sessionEvalSummary lives beside the affected selector in this engine and reduces the already-scoped model to seven useful counts: measured, affected total, fresh pass, fresh fail, measured-but-needing-review, blind, and unknown. The toolbar renders the four mutually exclusive scenario categories plus unknown, while measured and total remain projection facts for coherence, paging, exports, and other consumers rather than a second visible aggregate. Each paged list and bounded detail response carries that exact projection too, so a consumer never re-implements the fold. Each stable build also carries one content revision over every input that can move the result: the session HEAD, the base branch HEAD and merge-base, the staged and unstaged diff (renames and untracked content included), scenario declarations and their semantic hashes/code axes, reading/retraction sidecars, and the trunk remark track that participates in freshness. A build reads the revision before and after the fold; a mismatch is discarded and recomputed. Thus a summary and a demand projection bearing the same revision are the same evaluation cut, not two coincidentally similar reads.

That same content-addressed cut carries the derived full eval model, not only the summary. Manager review chrome is deliberately outside the cut: it is current session state, not eval input, and including it would make every summary and exact-impact read synchronously buy questions they do not publish. Every build deposits its model beside the summary under the one session + content revision key, so a repeat open at an unmoved revision replays it instead of re-deriving it from Git — the observable difference between opening a session's evaluation once and opening it again is a read, not a rebuild. It is the same cache owner, key and invalidation: a moved input yields a different key, only the newest key per session is retained, and summary and model are dropped together so they can never describe different revisions.

There is ONE builder, and the graph's fold IS the demand's fold. These were two: the graph's kept only the latest reading per scenario and deposited counts alone, so opening the page re-derived the whole model to recover history the graph fold had held one line earlier and discarded. The trim never made that fold cheaper — the cost is the freshness pass over every node in scope, and it ran identically either way — and it never mattered to the counts, because the summary reduction folds latest-per-scenario itself. So the second build bought nothing. Its price is paid in MEMORY instead: a cut retains complete A/B history rather than latest-per-scenario, bounded by the same one-cut-per-session rule. A replay still passes the stability, observer and generation fences, and it never weakens the validity contract: an unavailable projection deposits nothing, so a dead or unextractable selector re-derives and raises again rather than being masked by an earlier successful model. No TTL, patrol, second generation, extra gate, or client-side copy participates.

The backend retains a content-addressed, per-session projection cache. Each entry has a process epoch and a monotonic input generation g, a single in-flight build, and the last stable projection. A cache miss is loading; a relevant canonical input event increments g synchronously and becomes updating while preserving the last-known value; a stable build publishes ready only when both its generation and content revision still match. A changed generation or revision discards the result and the runner follows the newest generation. An error is explicit and also preserves last-known. Burst events may coalesce into one build/publication for the latest generation, but no older compute can overwrite it. The graph snapshot only batch-reads these cached lean projections; it never runs buildSessionEvals once per session row. A single-session READ of that same cache exists beside the batch one — it mints no entry, schedules no build, and moves no generation — so a consumer that must NOT build can report the existing projection with its phase. [[manager-cockpit]]'s review payload is exactly that consumer: this engine calls that payload, so the dependency runs one way only, and the cockpit reports loading/updating/error/absent as itself rather than as a computed result. Initial cache misses for live sessions may start one independent build per eligible entry. There is no cross-session concurrency cap or publication cohort: one session's Git/history work and one session's summary have no shared evaluation cut. Each entry publishes its own stable/error outcome as soon as its own generation and content revision still match; an unrelated slow, failed, or newer session never keeps that entry at loading/updating. Entry identity and generation are compared again at publication, so this independence cannot let an old result overwrite newer input. The graph stream emits the resulting session-unit deltas through its existing debounce; no summary-specific transport, cache, timer, or aggregate is introduced. Dormant offline history is intentionally demand-only: the graph emits its loading/last-known projection without scheduling a summary build for every retained session, while opening that session's scoped Evals route builds only the requested worktree model. This keeps the toolbar projection useful for active work without turning historical session count into a cold-start fan-out.

Eager builds are enabled only while a delta graph subscriber owns the current stream era; plain HTTP/CLI reads therefore expose loading or last-known summaries without starting work for retained records. Within an enabled era, every eligible current-generation projection starts independently. The resulting Git/node child count may scale with the active session set; that is the honest cost of removing an unrelated session's place in line from the toolbar's freshness contract. Each session still has exactly one in-flight summary for its own generation.

The demand path shares this entry-local flight. A selected session cancels an as-yet-unstarted summary for its current generation; when that summary is already running, demand waits only for that SAME entry to settle, then replays the stable content-addressed model rather than duplicating its work. A second demand for the same id joins the first promise. Unrelated eager summaries neither delay nor acquire a special lane around demand; every entry is already independently launchable. A rejected demand rejects only its own waiter, while other entries continue to settle and a later generation restores ordinary eager eligibility.

Freshness is event-driven, and an event NAMES its scope. The one graph stream owns invalidation: refs cover session/main HEAD and merge-base moves (including CLI remark commits). A session's fingerprint reads exactly three refs — its own tip, the base tip, and their merge-base — so the ref that moved decides which projections it can possibly have moved: the base branch invalidates every session, a session's own branch invalidates that session, and a tag, a remote-tracking ref, or a branch no session owns invalidates none. A packed-refs rewrite or a HEAD flip names nothing, since one event there can carry many refs, so it stays broad. Discarding the ref name and invalidating everything is what makes a busy repository permanently cold — observed on an adopter as an input generation of 1208 against a last-known 254, i.e. no session ever reaching ready. Narrowing this scope is a correctness change wearing performance clothes ([[taste]] 19): its failure mode is a stale answer that says nothing about being stale, so the derivation is pinned ref shape by ref shape rather than trusted to inspection; server remark/eval writes nudge it atomically; each linked worktree is watched recursively for dirty source, rename, scenario and sidecar edits, and its gitdir index is watched for stage/reset-only changes. Watch failure or a pathless/overflow-like event increments the generation and places a keyed observer hold on the affected projection: it stays updating(lastKnown) and no compute or demand read may certify it current while that input axis is unobservable. Holds compose, so restoring one source cannot mask a second failed source. A successful resubscription removes only its own hold, advances the generation again, and then performs one authoritative double-read rebuild with the replacement watcher already attached; an edit in the unwatched interval is therefore inside that rescan. A cold demand route initializes these same canonical observers even when no graph request preceded it. If it arrives during an observer hold, it waits boundedly for that recovery transition and only then performs its double-read build; a transient attach failure therefore stays loading and reaches a current response after authoritative resubscription instead of escaping as a generic 500. A persistent attach failure remains explicitly unavailable/non-current instead of falling through to a cache build. No TTL, periodic fingerprint scan, or patrol makes a summary current. The patrol may still repair unrelated graph state, but never advances an eval generation or certifies this projection. The guarantee is necessarily over events the OS watcher delivers: an operating system that silently drops an event without an error exposes no fact a purely event-driven process can detect, and the UI must not claim otherwise.

Every changed file — spec.md included — is a drill-down: its row expands to the unified diff (base..HEAD), and further to the full original ↔ new content side by side, all derived from git and inlined behind native toggles (capped so a huge changeset can't bloat the page). File grouping is complete independently of scenario impact: a node whose spec.md changed but whose scenarios did not still carries that file row and says that no declared scenario is affected. Nothing is hidden — the whole diff and both file versions are there to jump into, no extra fetch.

The interactive face is the Evals route family, session-scoped ([[evals-view]]): the canonical address of a session's evaluation is #/evals?q=is:eval scope:<id> (the list — the same [[evals-feed]] row grammar — its toolbar leading with the icon-only terminal door as its first focusable control, labelled by the short localized Back to session terminal / 返回会话终端 command ([[evals-view]]) — filed readings ordered newest-first across source ownership, with this session's measurements ✦-marked and inherited measurements legible by the absent ✦ rather than a privileged position, then blind spots as inert unmeasured rows — all bounded by the backend's affected-scenario set; unknown coverage is reported separately from those scenario rows and counts. A reading is the session's own iff THIS session filed it or its codeSha is one of the branch's commits, derived, never hand-tagged) and #/evals/<node>/<scenario>?q=scope:<id> (the [[event-detail]] page whose A/B history walks the WORKTREE-rooted readings — the live, remarkable reading of a still-open branch, what a CI/MR note links; merging first is not required, and the inert ?format=html export is not the link). The face fetches /api/evals pages for lists and /api/evals/detail for the selected scenario's complete history plus at most five lightweight neighbors. A live scope:<id> is worktree-rooted; when a historical detail's worktree no longer exists, that bounded detail explicitly resolves to trunk instead of pretending the scenario disappeared. It carries no diff enrichment or inlined evidence bytes. It rides the tiered loading every eval face shares: rows first, evidence streamed from /api/evidence only on the detail page. The console and phone session surfaces expose DOORS that are REAL ANCHORS — the console tab bar's eval ↗ entry and the phone session header's eval button carry the canonical scoped-list address as their literal href ([[address-routing]]'s one projection; copy-link and middle-click work for free) and clicking one is a single ordinary hash push landing directly on the final address — never a console-local eval pane, never a JS-only button, never the legacy ?session param; the typed /eval board command opens the same door. The LEGACY address #/sessions/<id>/eval[/<node>/<scenario>] normalizes to the canonical form at the route layer ([[side-nav]] — replace, old links keep working). The scoped detail exposes no second terminal door: its one small back arrow, plus load-failed and not-found list links, return to this scoped list; only from the list does the terminal door leave the Evals hierarchy. The page-bounded list carries no manager gate strip. Conflict, lint, ahead, and committed remain together on explicit manager review and the self-contained export: none is a property of the selected scenario page, and none may delay it. Soundness is proven by measuring the real product, not by a language-specific checker; a session with no worktree/diff shows a clean empty state.

Interactive full rows are not a transport. A scoped list receives one 25-row page; a scoped detail receives only its selected row, that scenario's complete history, and at most five lightweight neighbors.

A detail open measures what it renders, not what is in scope. The scope's expense is the freshness pass, and it grows with the node closure — but a detail publishes one verdict and at most five neighboring states while still owing the population's index and total. So the engine reads that population's sequence with no probes at all (identity and filed time answer it), lets the caller name the few nodes whose verdicts will be published, and runs the freshness pass over only those. The model this returns is deliberately PARTIAL — its nodes are the rendered window, not the scope — and a partial model may never enter the content-addressed cut, because the list page and the graph read that cut as the session's whole evaluation. It still prefers a cached FULL model when one exists: a complete answer already paid for beats a cheap incomplete one. Rows from the sequence pass carry no freshness claim and are refused outright if asked for a verdict, so the saving can never be spent on a guess.

It does not buy manager review chrome either. Whether this branch conflicts with the base, how far ahead it is, what is uncommitted, and whether the backend checkout lints clean are questions about review, not about the scenario page being read. List and detail models therefore read the session's IDENTITY from its record (free) rather than its review payload and carry no gates. This is the same rule as the freshness scope, applied one layer out: a response buys the questions it answers. Each response carries the same generation and content revision as the graph field. It carries the sessionEvalSummary projection too WHEN it has one honestly: a focused detail measured the rows it renders, not the scope, so it omits the counts rather than publishing a fold over its window — a fold that would report a 1336-scenario session as a 17-scenario one. Nothing is lost by the omission, because an equal content revision already IS the same evaluation cut; the counts were a redundant confirmation of an identity the revision had proven. If the client has already observed a newer graph generation, it rejects the old response and reloads; equal generations must have equal content revisions. This fence keeps a slow demand read from repainting newer loss while preserving the tier split: summary on graph, scenarios/readings on Evals open, evidence bytes on detail expansion.

The self-contained HTML (renderExportHtml: evidence inlined as data-URIs, every changed file's diff + before/after drill-down) remains as the export artifact — CI attachments, sharing, a bare browser — behind the session-scoped list's export ↗ link (labelled as the export it is — GET /api/sessions/:id/evals?format=html; the bare route rejects because interactive JSON uses the paged review routes), and spex eval ls --session <SEL> --export (--out/--open, a backend client that works against a remote backend unchanged). Its cards and denominator project the scoped declarations, not only readings that happen to exist: each affected missing scenario renders as unmeasured and still counts, while stale readings remain visible. Inlining everything is the right shape for a file that must stand alone, and the wrong shape for an interactive page — that is the whole split.

The CLI mirrors the vocabulary, not just the artifact. spex eval ls --session <SEL> is the session-scoped list's CLI twin: it walks the same /api/evals pages and renders the same attention order as text by projecting the returned item sequence directly — filed readings newest-first across nodes and source ownership, the session's own readings ✦-marked and inherited readings distinguished by the absent ✦, then blind spots. A row retains its node label as context; the node is never a grouping key that can move that row out of the global sequence. The paged model's empty gates array is machine-shape compatibility, not a visible empty section: human text omits a blank gates : line, while --json preserves the exact response. Full conflict/lint/ahead/committed gates remain visible in explicit review and the export artifact. An uncovered frontend node remains flagged — all over the same affected-scenario set, so a terminal-bound manager reads the measured loss without the dashboard. proof is no longer a user-facing word at all: the export rides the eval read as its --export flag, and the old spex review proof spelling is gone — a signpost names the canonical form and exits non-zero, never running ([[cli-surface]]). The read/write split stays intact: spex eval ls --session READS a session's evaluation; filing a reading remains spex eval add.

The impact snapshot carries PARSED relation entries, not a pair of projections to be reassembled. A node's relation is one list of {path, selectors}; code/related (bare paths) and codeScoped/relatedScoped (the selector-bearing subset) are views derived from it for consumers that want exactly those, and the snapshot ships the entries themselves alongside. It used to ship only the two views, so the exact-revision projection minted path#selector STRINGS back out of them and handed those to the relation parser to recover the entries the loader had all along — a serialize/reparse round-trip through a form nobody ever stored, at five call sites. Nothing validated by that reparse was load-bearing: a snapshot's relation problems are already carried on the snapshot and already throw before any of it runs, and a reparse of rows minted from parsed entries cannot surface a problem the original parse did not. One parse, one shape, and the ordinary loader's and the fixed-revision snapshot's history/window semantics stay distinct — sharing the relation projection is not licence to collapse those.