The ruler · artifacts from the published store
Evaluation evidence
Ten adversarial fixtures, three frozen corpora, and a recorded trace for every article. Pick a run, then follow a story that should have reached a reader and did not.
Algorithm comparison: prod-llm against proto-hybrid-judge-events, same corpus and same k.
- protocol — Neither run recorded a protocol identifier — the harness only began emitting one after these scorecards were written. Protocol equality therefore cannot be verified from the artifacts, only assumed from the runner name.
- eval_revision — Both runs executed at the same revision, so any difference reflects the two pipelines, not a code change over time.
prod-llm · 2026-09-02 · k=12
compared with proto-hybrid-judge-events · 2026-09-02 · k=12
- Estimated model cost, all personas
- $0.265
- Model calls, all personas
- 44
- Peak model calls for one persona
- 7
- Cache misses
- 0
Across 1 fixture: 5 unwanted placements delivered (5 distinct articles), and 2 wanted ones lost before the scorer saw them (2 distinct). A story lost before scoring cannot be rescued by better ranking.
- delivered-wanted2 pairs / 2 articles
- delivered-unwanted5 pairs / 5 articles
- delivered-unlabelled5 pairs / 5 articles
- lost-before-scorer2 pairs / 2 articles
- lost-at-or-after-scorer1 pairs / 1 articles
Two runs, one scale
Fixture Paula- Capped recall@k+20.0 ppCapped recall@k: prod-llm 40.0 per cent, proto-hybrid-judge-events 60.0 per cent, a difference of +20.0 pp, past the fixed materiality cutoff.
- Raw recall@k+20.0 ppRaw recall@k: prod-llm 40.0 per cent, proto-hybrid-judge-events 60.0 per cent, a difference of +20.0 pp, past the fixed materiality cutoff.
- Reached the scorer+40.0 ppReached the scorer: prod-llm 60.0 per cent, proto-hybrid-judge-events 100.0 per cent, a difference of +40.0 pp, past the fixed materiality cutoff.
- Unwanted rate−41.7 ppUnwanted rate: prod-llm 41.7 per cent, proto-hybrid-judge-events 0.0 per cent, a difference of −41.7 pp, past the fixed materiality cutoff.
- Need-to-know recallunchangedNeed-to-know recall: prod-llm 50.0 per cent, proto-hybrid-judge-events 50.0 per cent, a difference of unchanged.
- World-critical delivery+100.0 ppWorld-critical delivery: prod-llm 0.0 per cent, proto-hybrid-judge-events 100.0 per cent, a difference of +100.0 pp, past the fixed materiality cutoff.
- Planted-needle recallunchangedPlanted-needle recall: prod-llm 100.0 per cent, proto-hybrid-judge-events 100.0 per cent, a difference of unchanged.
- Lookalike rateunchangedLookalike rate: prod-llm 0.0 per cent, proto-hybrid-judge-events 0.0 per cent, a difference of unchanged.
One 0–100% scale for every row, ticked at 0, 50 and 100. Fraction metrics only — counts, costs and latencies share no scale with a recall rate. A filled dot marks a difference past the fixed ±0.02 cutoff, which is a chosen threshold, not a significance test.
| Metric | prod-llm | proto-hybrid-judge-events | Difference |
|---|---|---|---|
| Capped recall@kNeeded stories that made the feed, capped at the slots available. | 40.0% | 60.0% | Capped recall@k moved +20.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Raw recall@kSame, divided by every needed story. Always ≤ capped recall. | 40.0% | 60.0% | Raw recall@k moved +20.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Reached the scorerNeeded stories that survived long enough to be judged at all. | 60.0% | 100.0% | Reached the scorer moved +40.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Unwanted rateDelivered stories the reader had labelled never-wanted. | 41.7% | 0.0% | Unwanted rate moved −41.7 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Need-to-know recallRecall over stories the reader genuinely needed to know. | 50.0% | 50.0% | Need-to-know recall is unchanged. |
| World-critical deliveryWorld-critical events that put at least one story in the feed. | 0.0% | 100.0% | World-critical delivery moved +100.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Planted-needle recallRecall on planted articles whose right answer is known by construction. | 100.0% | 100.0% | Planted-needle recall is unchanged. |
| Lookalike rateDecoys that got through: right keyword, wrong thing. | 0.0% | 0.0% | Lookalike rate is unchanged. |
| Major-event delivery | 0.0% | 40.0% | Major-event delivery moved +40.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Judge precision | 66.7% | 100.0% | Judge precision moved +33.3 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Judge recall | 66.7% | 80.0% | Judge recall moved +13.3 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Distinct sources | 8 | 9 | Distinct sources moved +1, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
| Feed size at k | 12 | 12 | Feed size at k is unchanged. |
| Build latency | 0.460s | 0.020s | Build latency moved −0.440s, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test. |
What each metric means, exactly
- Capped recall@khigher is better
- Of the stories this reader had to see, the share that made the feed — with the target capped at the number of slots available.
- Raw recall@khigher is better
- The same count divided by every must-see story, even when there are more must-see stories than feed slots. Always at or below capped recall.
- Reached the scorerhigher is better
- The share of must-see stories that survived long enough to be judged at all. A story never loaded cannot be ranked.
- Unwanted ratelower is better
- The share of the delivered feed that the reader had explicitly labelled as never wanted. Lower is better.
- Need-to-know recallhigher is better
- Recall restricted to stories labelled as ones the reader genuinely needed to know.
- World-critical deliveryhigher is better
- Of the events everyone should have been told about, the share that put at least one story in the feed.
- Planted-needle recallhigher is better
- Recall on stories injected at run time whose correct answer is known by construction rather than by opinion.
- Lookalike ratelower is better
- How often a deliberate decoy got through — the right keyword attached to the wrong thing.
- Follow-up recallhigher is better
- Recall restricted to stories that continue a thread the reader was already following.
- Major-event deliveryhigher is better
- The same measure for events rated major rather than world-critical.
- False-major ratelower is better
- How often the pipeline promoted a story as a major event when no major event existed. Measured on the quiet-day corpus, where the right answer is restraint.
- Judge precisionhigher is better
- When the model accepted a story, how often the labels agreed. Reported only for runners whose judge emits an explicit verdict.
- Judge recallhigher is better
- Of the stories the labels wanted, how many the model accepted once it saw them.
- Distinct sourceshigher is better
- How many different publications the feed drew from. More breadth is generally healthier, but this is a diversity proxy, not a correctness measure.
- Feed size at kno direction
- How many slots the feed actually filled. A size, not a quality score.
- Build latencylower is better
- Wall-clock time to assemble one reader’s feed during the run. Offline replay timings are not production latency.
Every fixture, no averaging
- Anna37.5%17 labelled story placements for fixture Anna: 8 delivered, wanted, 4 delivered, unwanted, 5 lost before the scorer. Capped recall 37.5 percent.
- Cold start8.3%23 labelled story placements for fixture Cold start: 9 delivered, wanted, 3 delivered, unwanted, 4 lost at the scorer, 7 lost before the scorer. Capped recall 8.3 percent.
- Daniel20.0%16 labelled story placements for fixture Daniel: 1 delivered, wanted, 1 delivered, unlabelled, 10 delivered, unwanted, 4 lost before the scorer. Capped recall 20.0 percent.
- Frank18.2%21 labelled story placements for fixture Frank: 2 delivered, wanted, 9 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 8 lost before the scorer. Capped recall 18.2 percent.
- Lena16.7%27 labelled story placements for fixture Lena: 10 delivered, wanted, 2 delivered, unwanted, 15 lost before the scorer. Capped recall 16.7 percent.
- Maya16.7%25 labelled story placements for fixture Maya: 11 delivered, wanted, 1 delivered, unlabelled, 1 lost at the scorer, 12 lost before the scorer. Capped recall 16.7 percent.
- Paula40.0%15 labelled story placements for fixture Paula: 2 delivered, wanted, 5 delivered, unlabelled, 5 delivered, unwanted, 1 lost at the scorer, 2 lost before the scorer. Capped recall 40.0 percent.
- Ray30.0%19 labelled story placements for fixture Ray: 7 delivered, wanted, 5 delivered, unwanted, 7 lost before the scorer. Capped recall 30.0 percent.
- Tom33.3%22 labelled story placements for fixture Tom: 7 delivered, wanted, 1 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 12 lost before the scorer. Capped recall 33.3 percent.
- Will0.0%29 labelled story placements for fixture Will: 5 delivered, wanted, 6 delivered, unlabelled, 1 delivered, unwanted, 1 lost at the scorer, 16 lost before the scorer. Capped recall 0.0 percent.
- Delivered, wanted
- Delivered, unlabelled
- Delivered, unwanted
- Lost at the scorer
- Lost before the scorer
Bar length is labelled story placements; the right column is capped recall at k, mean 22.1%. A fixture at 0.0% received none of the stories its labels said it needed. Select one to filter everything above.
Provenance
What can and cannot be establishedProvenance · prod-llm · 2026-09-02— 4 caveats on this run
Three revisions, kept apart
- Executed the evaluation
- 47edb50not reachable from the default branch
- Stores the scorecard
- 3b11a3c
- Built this artifact
- 3b11a3c
This artifact imports a stored scorecard. It is not a run this pipeline executed, and no current revision is stamped onto its numbers.
Run identity
- Runner
- prod-llm
- CLI argument
- prod
- Protocol
- unknownabsent in the source scorecard
- k
- 12
- Execution mode
- unknownThe scorecard does not record an execution mode. A zero cache-miss total cannot stand in for one, because the harness only counts misses on the live network path.
- Recorded at
- 2026-09-02T07:14:29+00:00
- Needles planted
- true
- Quiet corpus
- false
- Cache keys touched
- 89
Models observed: gpt-4o-mini
Corpus
- Snapshot
- 2026-09-02
- Articles
- 1,358
- Frozen at
- 2026-09-02T05:06:49.958134+00:00
- Content hash
- 73f93e5b8cc02c18…
- Derived from
- original capture
- Clusters removed
- none
second live RSS snapshot, 50+ hours after 2026-08-31
Ground truth
- Label rows
- 3,918persona/article pairs
- Distinct articles
- 1,260labelled at least once
- Fixtures
- 10
- Contested
- 209
- Status
- provisional model and agent
- Label models
- gpt-4.1, gpt-4.1-mini
By source: agent 345 · model 3,573. Labels are model-written with an agent editorial pass; product-owner human review is outstanding, so absolute values are provisional and only run-to-run differences are gate-enforced.
Caveats
- Warning. The revision that executed this run (47edb50) is not reachable from the current default branch, so "this result came from that code" cannot be verified by checking out the revision. Everything under backend/evals/ reached the default branch in a single squash commit.git merge-base --is-ancestor 47edb50 HEAD
- Warning. This scorecard records no protocol identifier. The harness only began emitting one later (backend/evals/run.py writes meta.protocol), so the evaluation protocol cannot be read from the artifact and is reported as unknown rather than inferred from the runner name.backend/evals/run.py
- Caution. This scorecard is missing summary keys that the current harness emits unconditionally (reader_pipeline_calls_max_per_persona, event_slot_contamination_mean). It therefore could not have been produced by the code on the current default branch, and re-running today would not yield a document of the same shape.backend/evals/run.py, backend/evals/metrics.py
- Caution. A reported cache-miss total of zero does not prove offline replay: the harness increments its miss counter only on the live network path, so the counter is structurally zero whenever offline mode is active. Execution mode is therefore reported as unknown unless the scorecard states it.backend/evals/llm_cache.py
- Note. Stored in the repository tree at backend/evals/results/47edb50-prod-llm-2026-09-02.json.backend/evals/results/47edb50-prod-llm-2026-09-02.json
Provenance · proto-hybrid-judge-events · 2026-09-02— 5 caveats on this run
Three revisions, kept apart
- Executed the evaluation
- 47edb50not reachable from the default branch
- Stores the scorecard
- 3b11a3c
- Built this artifact
- 3b11a3c
This artifact imports a stored scorecard. It is not a run this pipeline executed, and no current revision is stamped onto its numbers.
Run identity
- Runner
- proto-hybrid-judge-events
- CLI argument
- unknown
- Protocol
- unknownabsent in the source scorecard
- k
- 12
- Execution mode
- unknownThe scorecard does not record an execution mode. A zero cache-miss total cannot stand in for one, because the harness only counts misses on the live network path.
- Recorded at
- 2026-09-02T07:14:42+00:00
- Needles planted
- true
- Quiet corpus
- false
- Cache keys touched
- 282
Corpus
- Snapshot
- 2026-09-02
- Articles
- 1,358
- Frozen at
- 2026-09-02T05:06:49.958134+00:00
- Content hash
- 73f93e5b8cc02c18…
- Derived from
- original capture
- Clusters removed
- none
second live RSS snapshot, 50+ hours after 2026-08-31
Ground truth
- Label rows
- 3,918persona/article pairs
- Distinct articles
- 1,260labelled at least once
- Fixtures
- 10
- Contested
- 209
- Status
- provisional model and agent
- Label models
- gpt-4.1, gpt-4.1-mini
By source: agent 345 · model 3,573. Labels are model-written with an agent editorial pass; product-owner human review is outstanding, so absolute values are provisional and only run-to-run differences are gate-enforced.
Caveats
- Warning. The revision that executed this run (47edb50) is not reachable from the current default branch, so "this result came from that code" cannot be verified by checking out the revision. Everything under backend/evals/ reached the default branch in a single squash commit.git merge-base --is-ancestor 47edb50 HEAD
- Warning. This scorecard records no protocol identifier. The harness only began emitting one later (backend/evals/run.py writes meta.protocol), so the evaluation protocol cannot be read from the artifact and is reported as unknown rather than inferred from the runner name.backend/evals/run.py
- Caution. This scorecard is missing summary keys that the current harness emits unconditionally (reader_pipeline_calls_max_per_persona, event_slot_contamination_mean). It therefore could not have been produced by the code on the current default branch, and re-running today would not yield a document of the same shape.backend/evals/run.py, backend/evals/metrics.py
- Warning. This runner name belongs to the corrected default prototype pipeline, but the repository’s own regression gate maps this runner key to the historical proto-s0-legacy-v1 adapter, whose protocol reproduces known defects (per-reader event discovery, uncapped forced delivery that ignores reader exclusions). Treat this run as evidence about the historical S0 protocol, not as validation of the corrected default pipeline.backend/tests/test_eval_gate.py, backend/evals/legacy_s0.py
- Caution. A reported cache-miss total of zero does not prove offline replay: the harness increments its miss counter only on the live network path, so the counter is structurally zero whenever offline mode is active. Execution mode is therefore reported as unknown unless the scorecard states it.backend/evals/llm_cache.py
- Note. Stored in the repository tree at backend/evals/results/47edb50-proto-hybrid-judge-events-2026-09-02.json.backend/evals/results/47edb50-proto-hybrid-judge-events-2026-09-02.json