evals-feed¶
The Evals ListView rows and filters through [[review-chrome]] — latest result per scenario, structured state/title/filer/time/kind rows, Fail/Pass/Unmeasured quick filters, secondary human-review/freshness/evidence/presence builders, and token-only node/filer/scope over ONE visible query; media stays strictly lazy.
raw source¶
The Evals list page (evals-view) is where a human reads the project's current measured loss — the leading review surface. A feed of every result ever filed grows without bound; a feed of the project's current loss does not. The unit is the scenario, not the result: the eval engine already defines the latest result per scenario as the current score. That population is structural and slow-growing but still paged at the request layer, so a large project never pays or renders its whole declared set. Review attends to what still counts; and in the GitHub navigation model each row is a LINK to the eval's own page, not a selection in a pane.
The node inspector's Eval pane is a bounded preview of that same current-loss population, never a second history feed. It always carries one compact icon-only door to the node-filtered Evals ListView, so a reviewer can reach scenarios beyond its first page without interpreting a pagination count as a command. The scenario detail page, not either list, is where that scenario's complete measurement history lives.
expanded spec¶
Default view: latest result per scenario, newest first — fresh and stale, reviewed and unreviewed,
scored and blind mixed honestly. Filed readings sort by filed time descending across source ownership;
blind scenarios, which have no filed time, follow measured rows with a deterministic tie-break and are
never promoted as a leading block. A stale result is real measured loss and remains visible; freshness is
an honest facet, never a default hide. The leading quick-filter axis is Fail / Pass / Unmeasured: each
action wears the ONE review-chrome ReviewState icon + tone + count and surgically toggles
verdict:fail|pass|unmeasured in the visible query. Unmeasured names exactly a declared scenario with no
reading, not an unscored or unknown reading. It is deliberately a named pressed-button group, not a
tablist: the axis is still not exhaustive, so the default is:eval list may show unscored or unknown rows
while no button is pressed.
Counts are computed under the rest of the query, excluding the active verdict token, and a second click
clears it back to the honest whole list.
A measured chip LEADS with fresh and names its stale debt. Fail and Pass show the fresh half of
paged-review's already-split count, then the stale half as one quieter suffix beside it — the count the
review-filters freshness token means, spelled with that same localized stale word, so a remeasurement
campaign reads its backlog off the header instead of a second query. The two numbers are the verdict's whole
population and clicking the chip still selects all of it; Unmeasured has no reading and shows one number.
The feed only READS that fold — a 25-row page could never re-derive it — so under freshness:fresh the
stale half is already zero and the suffix simply does not render, no token special-case. The debt is
visible at every width. The phone condenses the WORDING, never the number — the suffix shows its bare
+N while the full +N stale stays the accessible name and tooltip — and where even that does not fit,
the header takes a second contained line rather than clipping or dropping a control (review-chrome owns
that geometry). Reaching the count through the Filters menu is not a substitute: a default view that cannot
distinguish fresh current work from stale work awaiting re-measurement has lost the thing this split
exists to show. A fresh human-ok'd result is state:reviewed; everything else is
state:current. That lifecycle remains transparent in the query and editable through the secondary
Human review builder (Needs review / Reviewed), but no longer occupies the top visual hierarchy.
When the worktree scope contributes its terminal/gates strip, the feed hands it to ListPage as leading
content inside the shared page-scroll; it never wraps the list with a second shell or moves the track.
Every filter is a token in review-chrome's ONE visible query (review-query),
and matching travels through review-filters's Eval adapter — page code only bridges the parsed text
into the shared engine, so the embedded node list cannot acquire different parsing or matching semantics:
verdict:, freshness:, and evidence: (values exactly video | image | all; the default is all,
with NO data-dependent fallback) keep low-cardinality menus that are pure query builders; the
source-session presence facet is session:present|missing (live-session-filter); node: and
filer: are HIGH-cardinality token-only dimensions — hand-typed or completed from the input's bounded
autocomplete, never an enumerating dropdown — and scope: sources the worktree model (evals-view).
Bare words search scenario/node/filer/evaluator; an unknown qualifier matches nothing, honestly. Common
menus stay visible on desktop, low-frequency/width-displaced ones move
to the semantic secondary Filters menu, and only the primary facet survives beside the tabs at 390px. Every pick is
URL-query state: a human's action pushes ?q=<raw text> (the default view stays bare), reload/Back
fully replay, and no local filter state survives. If live data contracts, an active menu value keeps its
All off-switch — and the visible text is always the canonical release, whatever state the scope is in.
An inert blind-spot row participates in that SAME conjunctive contract: it can match its real node,
verdict:unmeasured, and query text, but a selected Fail/Pass quick filter, evidence kind, freshness, filer,
or source-session presence value excludes it
because an unmeasured scenario owns none of those result facts. Blind rows never leak into a filtered
result population and never gain an href just to satisfy list structure. Filed results and non-result
rows form one tagged set through the shared result-kind field; the canonical list and embedded node pane
consume that same discriminator, with no legacy-name compatibility branch.
Kinds are honest — and a result carries a SET of them. Evidence is a LIST: a result's kinds are
every entry it holds (video/image/transcript; a legacy scalar blob with no kind is an image), plus
note when it holds no blob. A MIXED result belongs to EVERY media filter it contains and its tag
lists its kinds video-first (vid·img); it never advertises media it lacks. note and transcript are
data-level kinds only, never filter options — they surface under all.
Rows use the shared two-level primitive, and each row is a REAL <a> to
#/evals/<node>/<scenario> (the worktree scope's rows carry ?q=scope:<id> and nothing else) — shared verdict visual +
wrapping scenario title; node, filer, and filed time below; evidence kind/scope at the right (joining the
secondary line at 390px). No media request of any kind occurs in the list, no per-row write affordance
(reviewing + signing lives on the detail page). A human-ok'd row adds the one settled certification mark
(the shared stroke check in a quiet green ring, signer/time as its accessible name). Clicking a row is a
history PUSH onto the detail page; j/k move the cursor and Enter opens it (review-chrome).
One data path, one computation. Trunk and worktree scope both request paged-review with the visible
query and route page. The server owns latest-per-scenario selection, shared matching, full-population
counts/facets, stable ordering/revision, and the one 25-row slice; the feed receives only that page and
owns no local matcher or full-list prop. scope:<id> selects the worktree snapshot through the same
contract, row grammar, and newest-first order — ✦ marking the in-session rows, blind spots as inert
unmeasured lines after filed readings — while
the session toolbar consumes only its canonical lean graph summary (session-eval), never the list or
a row model. Unknown coverage stays in the scoped leading strip and cannot enter the result adapter. Loading,
empty, and failed models do not replace the list shell: the scope
and kind controls stay mounted, with the appropriate empty note or explicit error beneath them.
The embedded node preview names its escape explicitly. The node Eval pane uses the same server page and
therefore shows at most 25 latest-per-scenario rows; it never unfolds historical readings to fill its
space. Its compact filter row retains showing X of Y as status only, while a permanent icon-only anchor
(with tooltip and accessible name) opens the canonical is:eval node:<id> ListView at every cardinality.
There is no conditional trailing showing 25 of N link: an item count is not an action, and hiding the
door for short lists gives the same review surface two different navigation grammars. The full ListView
keeps its own bounded pagination; opening a scenario there reaches its detail, whose history is scoped to
that one scenario.