eval-tab¶
Provenance¶
- Source:
.spec/spexcode/spec-eval/eval-tab/spec.md - Source SHA-256:
ffb145aacb59e1a9ab1f525a94ed336d5136a9d112ae1c833953adec279dbb18
raw source¶
The eval/loss engine ([[spec-eval]], built by [[eval-core]]) records readings; this is the surface that
reads them back. Realize the founding "Evidence — one timeline, two sources" contract's first source: a
node's eval tab lists its measurements chronologically, each carrying its verdict, the scenario's
expected, and the freshness signal spex eval lint reports, with the captured evidence — a gallery of the
reading's whole evidence list (N images plus a video clip that plays inline, and/or a transcript) —
expanding inline. LOCAL readings only for now — the forge issue-events source is a later
sibling; leave a clean seam for it.
expanded spec¶
Two halves behind one tab. The read engine ([[spec-cli]], in evaltab.ts) computes what only a live
read knows. A node's measurement timeline is every reading from its evals.ndjson sidecar (scenario,
the read's codeSha, an evidence LIST — each {hash, kind} resolved to its live blob state — verdict, ts) joined with the scenario's expected
(from the live eval.md — what zero loss looks like) and a freshness flag, derived live from git by the
same freshness machinery the lint uses ([[eval-core]]'s scenario-freshness derivation): a reading is current until its governed code or its scenario
moved past the sha it was taken at, otherwise stale (and which axis moved); a code-stale
reading also carries the code axis's per-file drift detail (codeDrift — which governed files moved, +how many
commits) so the [[event-detail]] stale readout can name it, never re-deriving git in the browser;
newest-first.
The board carries only state counts ([[graph-lean]]), so tile/stat/caption glances need no scenario or
reading rows. Opening the Eval tab requests the first 25 node-timeline rows from /api/evals through the
ONE [[paged-review]] contract with view=timeline and the fixed node: qualifier. The server performs the
same filter/facet/count semantics across the whole node timeline before slicing. A scenario detail loads
that one scenario's full A/B history and at most five lightweight ordered neighbors through
[[paged-review]]'s bounded detail contract; the popup never downloads every timeline row as fallback.
Bytes are never folded anywhere: /api/evidence serves each evidence entry by its content hash from the
shared cache, fetched lazily on expand, with a per-entry miss original file signal when the bytes are gone, MIME
sniffed from the content.
The eval tab ([[spec-dashboard]]) is a fourth face on the node popup beside spec/history/issues, on the
same panesFor registry. It is a thin consumer of the chronological-timeline
scaffold the history tab uses (see [[node-popup]]): newest expanded, older reveal on the down gesture, an
individual row-header toggle, no bulk-expand control, and — on a long timeline — an extremely compact
embedded face of the canonical Evals filter. The popup and list page share one query parser, conjunctive
filter engine, and field semantics through domain configuration/data adapters; the popup does not grow a
modal-only state machine or a second eval query dialect ([[review-filters]]). Blind-spot rows and dangling tracks participate in
the same honest field rules as the canonical list, so a filtered view stays one coherent set. Popup filter
state survives tab switches while the node popup remains open, but owns no canonical address; opening a
reading still follows the Evals route family. Each
row's header names its scenario, the verdict badge (✓ pass / ✗ fail, optional note
beside; legacy for a pre-verdict or note-only reading), and the per-reading score circle
([[eval-score-badge]]), then its codeSha and time. Beside every reading row — a SIBLING of the expand
toggle, never nested inside it — sits a real anchor onto the scenario's canonical routed detail
(#/evals/<node>/<scenario>, [[address-routing]]), so the popup timeline links out to the full-page
review surface; blind-spot and dangling rows, having no reading to open, stay inert.
Its evidence is the scenario's expected over the captured proof — a gallery mapping the reading's
evidence list (N screenshots, a video, a transcript), each entry showing miss original file when its blob
was pruned.
The tab's server population is the whole declared set, not only the readings, while the popup receives
only its first 25 filtered timeline rows. Its compact summary reads the server's full total; when more
rows exist, a true View all anchor opens the canonical node-filtered Evals list instead of silently presenting
page 1 as complete. A declared scenario
with no reading leads that list as a blind-spot row — the empty score ring over its name, its
expected, and the files it tracks. The ring is the only distinction (no fenced-off band, no second
scrollbar): an unmeasured scenario is outstanding loss and belongs where the attention is, so a node's
intent is legible inside the popup before a reading lands. No reading at all → those rows under a hint; some measured, some not → those
rows lead the timeline. The one presence-distinct empty state survives: a node with no scenarios
(no eval.md → no evals field) shows nothing.
The seam / out of scope: the forge issue-events half of the timeline — each tracked issue appearing twice (open, close) and linking out to its forge-hosted image rather than a local blob — arrives with the [[needs-eval]] forge node; the tab joins it at read time then. Backend and computer-use measuring hands, and the cache cleanup surface, stay with their own nodes.