You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Split out of #201, which shipped the query-level half of D2 (PR #205) and left this half
untouched by design. Step 1 answers does the vault hold evidence for this query at all. This is
the other question: a real query whose list runs out before limit does — the "results 7–10 are
filler" complaint, the per-hit sibling of the same dishonesty.
The evidence exists and is measured; the labels do not. evals/queries.json names the relevant note for a query — it says nothing about the irrelevance of ranks 5–10, which is exactly
what a tail rule must be judged on. So the provenance is reported and no rule is drawn from it:
HitProvenance per served row (bm25_rank, vector_rank, distance), reaching a caller through Vault::search_evidence → EvidencedResult.
just eval's search evidence bake-off prints the dense_only reading — served rows the lexical
half never ranked. At the shipped bar: 0 of 410 served positive rows are dense-only, against 20
of 50 negative ones. That separation is the reason to think a tail rule is findable; it is not
evidence that any particular one is right.
What this issue owes
A label extension first (process rule 2 applies: its own commit, and the two-direction token
audit). The shape is per-query per-rank judgements over the served list, not new notes — the
corpus does not need to grow, the labels need to get deeper. Decide the cheapest honest form:
full per-rank relevance, or a per-query "last relevant rank" mark.
D1's prefix requirement binds here exactly as in discovery: per-hit eligibility may not punch
holes in the fused order. The fold is a cut, or it does not ship.
zero labelled relevant notes below the tail fold — no headroom (D2's tripwire, per-hit form)
dense fixture
a single-domain vault's lists are all real matches; a tail rule that truncates them there is disqualified, the same absolute #200 enforced for discovery
What this is
Split out of #201, which shipped the query-level half of D2 (PR #205) and left this half
untouched by design. Step 1 answers does the vault hold evidence for this query at all. This is
the other question: a real query whose list runs out before
limitdoes — the "results 7–10 arefiller" complaint, the per-hit sibling of the same dishonesty.
Why it was not shipped with #201
The evidence exists and is measured; the labels do not.
evals/queries.jsonnames therelevant note for a query — it says nothing about the irrelevance of ranks 5–10, which is exactly
what a tail rule must be judged on. So the provenance is reported and no rule is drawn from it:
HitProvenanceper served row (bm25_rank,vector_rank,distance), reaching a caller throughVault::search_evidence→EvidencedResult.just eval'ssearch evidence bake-offprints thedense_onlyreading — served rows the lexicalhalf never ranked. At the shipped bar: 0 of 410 served positive rows are dense-only, against 20
of 50 negative ones. That separation is the reason to think a tail rule is findable; it is not
evidence that any particular one is right.
What this issue owes
audit). The shape is per-query per-rank judgements over the served list, not new notes — the
corpus does not need to grow, the labels need to get deeper. Decide the cheapest honest form:
full per-rank relevance, or a per-query "last relevant rank" mark.
b2 similar(runs #197 Phase 2, on D1 as redrafted) #200's fold and Phase C: search evidence honesty —b2 searchanswers zero when the vault holds no evidence (implements D2) #201's bar, onboth corpora, with "no tail fold at all" an admissible winner (it won for discovery in Phase B: the discovery disclosure bake-off — a default-view fold for
b2 similar(runs #197 Phase 2, on D1 as redrafted) #200).holes in the fused order. The fold is a cut, or it does not ship.
just calibrate --searchis the bench, and it is the one that caught Phase C: search evidence honesty —b2 searchanswers zero when the vault holds no evidence (implements D2) #201's first rule.Judged on
just calibrate --searchon real vaultsOut of scope
b2 searchanswers zero when the vault holds no evidence (implements D2) #201.References
b2 searchanswers zero when the vault holds no evidence (implements D2) #201 (the parent; Step 1's engine work, PR GH #201: Search evidence bar — query-level rule for hybrid_search #205), Phase D: the search evidence verdict reaches the surfaces — CLI, JSON, desktop, and the exit-gate moves (lands #201; #200 shipped no fold) #202 (the surfaces), Phase B: the discovery disclosure bake-off — a default-view fold forb2 similar(runs #197 Phase 2, on D1 as redrafted) #200 (the fold bake-offwhose method this follows, and whose answer was "no rule"), Discovery floor: the centroid z rule suppresses genuinely-related multi-topic notes outright, and its stated calibration window no longer holds #187 (constants in code, measurements
in the harness).
invariants.mdD2's closing sentence, which names this as deliberately unshipped;docs/evals/README.mdprocess rules 2 and 5.