Daily

The ruler · artifacts from the published store

Evaluation evidence

Ten adversarial fixtures, three frozen corpora, and a recorded trace for every article. Pick a run, then follow a story that should have reached a reader and did not.

Algorithm comparison: prod-llm against proto-hybrid-judge-events, same corpus and same k.

  • protocol — Neither run recorded a protocol identifier — the harness only began emitting one after these scorecards were written. Protocol equality therefore cannot be verified from the artifacts, only assumed from the runner name.
  • eval_revision — Both runs executed at the same revision, so any difference reflects the two pipelines, not a code change over time.

prod-llm · 2026-09-02 · k=12

compared with proto-hybrid-judge-events · 2026-09-02 · k=12

Estimated model cost, all personas
$0.265
Model calls, all personas
44
Peak model calls for one persona
7
Cache misses
0

Across 1 fixture: 1 unwanted placements delivered (1 distinct articles), and 12 wanted ones lost before the scorer saw them (12 distinct). A story lost before scoring cannot be rescued by better ranking.

  • delivered-wanted7 pairs / 7 articles
  • delivered-unwanted1 pairs / 1 articles
  • delivered-unlabelled1 pairs / 1 articles
  • lost-before-scorer12 pairs / 12 articles
  • lost-at-or-after-scorer1 pairs / 1 articles
01

Two runs, one scale

Fixture Tom
prod-llmproto-hybrid-judge-eventsmoved the wrong way
  • Capped recall@k+8.3 ppCapped recall@k: prod-llm 33.3 per cent, proto-hybrid-judge-events 41.7 per cent, a difference of +8.3 pp, past the fixed materiality cutoff.
  • Raw recall@k+5.9 ppRaw recall@k: prod-llm 23.5 per cent, proto-hybrid-judge-events 29.4 per cent, a difference of +5.9 pp, past the fixed materiality cutoff.
  • Reached the scorer+47.1 ppReached the scorer: prod-llm 29.4 per cent, proto-hybrid-judge-events 76.5 per cent, a difference of +47.1 pp, past the fixed materiality cutoff.
  • Unwanted rate−11.1 ppUnwanted rate: prod-llm 11.1 per cent, proto-hybrid-judge-events 0.0 per cent, a difference of −11.1 pp, past the fixed materiality cutoff.
  • Need-to-know recall−11.1 ppNeed-to-know recall: prod-llm 22.2 per cent, proto-hybrid-judge-events 11.1 per cent, a difference of −11.1 pp, past the fixed materiality cutoff.
  • World-critical delivery+100.0 ppWorld-critical delivery: prod-llm 0.0 per cent, proto-hybrid-judge-events 100.0 per cent, a difference of +100.0 pp, past the fixed materiality cutoff.
  • Planted-needle recallunchangedPlanted-needle recall: prod-llm 50.0 per cent, proto-hybrid-judge-events 50.0 per cent, a difference of unchanged.
  • Lookalike rateunchangedLookalike rate: prod-llm 0.0 per cent, proto-hybrid-judge-events 0.0 per cent, a difference of unchanged.

One 0–100% scale for every row, ticked at 0, 50 and 100. Fraction metrics only — counts, costs and latencies share no scale with a recall rate. A filled dot marks a difference past the fixed ±0.02 cutoff, which is a chosen threshold, not a significance test.

Evaluation metrics for reader fixture Tom, with the comparison run’s difference
Metricprod-llmproto-hybrid-judge-eventsDifference
Capped recall@kNeeded stories that made the feed, capped at the slots available.33.3%41.7%Capped recall@k moved +8.3 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Raw recall@kSame, divided by every needed story. Always ≤ capped recall.23.5%29.4%Raw recall@k moved +5.9 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Reached the scorerNeeded stories that survived long enough to be judged at all.29.4%76.5%Reached the scorer moved +47.1 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Unwanted rateDelivered stories the reader had labelled never-wanted.11.1%0.0%Unwanted rate moved −11.1 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Need-to-know recallRecall over stories the reader genuinely needed to know.22.2%11.1%Need-to-know recall moved −11.1 pp, which is worse for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
World-critical deliveryWorld-critical events that put at least one story in the feed.0.0%100.0%World-critical delivery moved +100.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Planted-needle recallRecall on planted articles whose right answer is known by construction.50.0%50.0%Planted-needle recall is unchanged.
Lookalike rateDecoys that got through: right keyword, wrong thing.0.0%0.0%Lookalike rate is unchanged.
Follow-up recall0.0%33.3%Follow-up recall moved +33.3 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Major-event delivery20.0%20.0%Major-event delivery is unchanged.
Judge precision80.0%100.0%Judge precision moved +20.0 pp, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Judge recall80.0%76.9%Judge recall moved −3.1 pp, which is worse for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Distinct sources710Distinct sources moved +3, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
Feed size at k912Feed size at k changed by +3. This metric has no better-or-worse direction.
Build latency0.450s0.020sBuild latency moved −0.430s, which is better for this metric, past compare.py’s fixed ±0.02 materiality cutoff. This is a fixed cutoff, not a significance test.
What each metric means, exactly
Capped recall@khigher is better
Of the stories this reader had to see, the share that made the feed — with the target capped at the number of slots available.
Raw recall@khigher is better
The same count divided by every must-see story, even when there are more must-see stories than feed slots. Always at or below capped recall.
Reached the scorerhigher is better
The share of must-see stories that survived long enough to be judged at all. A story never loaded cannot be ranked.
Unwanted ratelower is better
The share of the delivered feed that the reader had explicitly labelled as never wanted. Lower is better.
Need-to-know recallhigher is better
Recall restricted to stories labelled as ones the reader genuinely needed to know.
World-critical deliveryhigher is better
Of the events everyone should have been told about, the share that put at least one story in the feed.
Planted-needle recallhigher is better
Recall on stories injected at run time whose correct answer is known by construction rather than by opinion.
Lookalike ratelower is better
How often a deliberate decoy got through — the right keyword attached to the wrong thing.
Follow-up recallhigher is better
Recall restricted to stories that continue a thread the reader was already following.
Major-event deliveryhigher is better
The same measure for events rated major rather than world-critical.
False-major ratelower is better
How often the pipeline promoted a story as a major event when no major event existed. Measured on the quiet-day corpus, where the right answer is restraint.
Judge precisionhigher is better
When the model accepted a story, how often the labels agreed. Reported only for runners whose judge emits an explicit verdict.
Judge recallhigher is better
Of the stories the labels wanted, how many the model accepted once it saw them.
Distinct sourceshigher is better
How many different publications the feed drew from. More breadth is generally healthier, but this is a diversity proxy, not a correctness measure.
Feed size at kno direction
How many slots the feed actually filled. A size, not a quality score.
Build latencylower is better
Wall-clock time to assemble one reader’s feed during the run. Offline replay timings are not production latency.
02

Every fixture, no averaging

  • Delivered, wanted
  • Delivered, unlabelled
  • Delivered, unwanted
  • Lost at the scorer
  • Lost before the scorer

Bar length is labelled story placements; the right column is capped recall at k, mean 22.1%. A fixture at 0.0% received none of the stories its labels said it needed. Select one to filter everything above.

—

Provenance

What can and cannot be established
Provenance · prod-llm · 2026-09-02— 4 caveats on this run

Three revisions, kept apart

Executed the evaluation
47edb50not reachable from the default branch
Stores the scorecard
3b11a3c
Built this artifact
3b11a3c

This artifact imports a stored scorecard. It is not a run this pipeline executed, and no current revision is stamped onto its numbers.

Run identity

Runner
prod-llm
CLI argument
prod
Protocol
unknownabsent in the source scorecard
k
12
Execution mode
unknownThe scorecard does not record an execution mode. A zero cache-miss total cannot stand in for one, because the harness only counts misses on the live network path.
Recorded at
2026-09-02T07:14:29+00:00
Needles planted
true
Quiet corpus
false
Cache keys touched
89

Models observed: gpt-4o-mini

Corpus

Snapshot
2026-09-02
Articles
1,358
Frozen at
2026-09-02T05:06:49.958134+00:00
Content hash
73f93e5b8cc02c18…
Derived from
original capture
Clusters removed
none

second live RSS snapshot, 50+ hours after 2026-08-31

Ground truth

Label rows
3,918persona/article pairs
Distinct articles
1,260labelled at least once
Fixtures
10
Contested
209
Status
provisional model and agent
Label models
gpt-4.1, gpt-4.1-mini

By source: agent 345 · model 3,573. Labels are model-written with an agent editorial pass; product-owner human review is outstanding, so absolute values are provisional and only run-to-run differences are gate-enforced.

Caveats

  • Warning. The revision that executed this run (47edb50) is not reachable from the current default branch, so "this result came from that code" cannot be verified by checking out the revision. Everything under backend/evals/ reached the default branch in a single squash commit.git merge-base --is-ancestor 47edb50 HEAD
  • Warning. This scorecard records no protocol identifier. The harness only began emitting one later (backend/evals/run.py writes meta.protocol), so the evaluation protocol cannot be read from the artifact and is reported as unknown rather than inferred from the runner name.backend/evals/run.py
  • Caution. This scorecard is missing summary keys that the current harness emits unconditionally (reader_pipeline_calls_max_per_persona, event_slot_contamination_mean). It therefore could not have been produced by the code on the current default branch, and re-running today would not yield a document of the same shape.backend/evals/run.py, backend/evals/metrics.py
  • Caution. A reported cache-miss total of zero does not prove offline replay: the harness increments its miss counter only on the live network path, so the counter is structurally zero whenever offline mode is active. Execution mode is therefore reported as unknown unless the scorecard states it.backend/evals/llm_cache.py
  • Note. Stored in the repository tree at backend/evals/results/47edb50-prod-llm-2026-09-02.json.backend/evals/results/47edb50-prod-llm-2026-09-02.json
Provenance · proto-hybrid-judge-events · 2026-09-02— 5 caveats on this run

Three revisions, kept apart

Executed the evaluation
47edb50not reachable from the default branch
Stores the scorecard
3b11a3c
Built this artifact
3b11a3c

This artifact imports a stored scorecard. It is not a run this pipeline executed, and no current revision is stamped onto its numbers.

Run identity

Runner
proto-hybrid-judge-events
CLI argument
unknown
Protocol
unknownabsent in the source scorecard
k
12
Execution mode
unknownThe scorecard does not record an execution mode. A zero cache-miss total cannot stand in for one, because the harness only counts misses on the live network path.
Recorded at
2026-09-02T07:14:42+00:00
Needles planted
true
Quiet corpus
false
Cache keys touched
282

Corpus

Snapshot
2026-09-02
Articles
1,358
Frozen at
2026-09-02T05:06:49.958134+00:00
Content hash
73f93e5b8cc02c18…
Derived from
original capture
Clusters removed
none

second live RSS snapshot, 50+ hours after 2026-08-31

Ground truth

Label rows
3,918persona/article pairs
Distinct articles
1,260labelled at least once
Fixtures
10
Contested
209
Status
provisional model and agent
Label models
gpt-4.1, gpt-4.1-mini

By source: agent 345 · model 3,573. Labels are model-written with an agent editorial pass; product-owner human review is outstanding, so absolute values are provisional and only run-to-run differences are gate-enforced.

Caveats

  • Warning. The revision that executed this run (47edb50) is not reachable from the current default branch, so "this result came from that code" cannot be verified by checking out the revision. Everything under backend/evals/ reached the default branch in a single squash commit.git merge-base --is-ancestor 47edb50 HEAD
  • Warning. This scorecard records no protocol identifier. The harness only began emitting one later (backend/evals/run.py writes meta.protocol), so the evaluation protocol cannot be read from the artifact and is reported as unknown rather than inferred from the runner name.backend/evals/run.py
  • Caution. This scorecard is missing summary keys that the current harness emits unconditionally (reader_pipeline_calls_max_per_persona, event_slot_contamination_mean). It therefore could not have been produced by the code on the current default branch, and re-running today would not yield a document of the same shape.backend/evals/run.py, backend/evals/metrics.py
  • Warning. This runner name belongs to the corrected default prototype pipeline, but the repository’s own regression gate maps this runner key to the historical proto-s0-legacy-v1 adapter, whose protocol reproduces known defects (per-reader event discovery, uncapped forced delivery that ignores reader exclusions). Treat this run as evidence about the historical S0 protocol, not as validation of the corrected default pipeline.backend/tests/test_eval_gate.py, backend/evals/legacy_s0.py
  • Caution. A reported cache-miss total of zero does not prove offline replay: the harness increments its miss counter only on the live network path, so the counter is structurally zero whenever offline mode is active. Execution mode is therefore reported as unknown unless the scorecard states it.backend/evals/llm_cache.py
  • Note. Stored in the repository tree at backend/evals/results/47edb50-proto-hybrid-judge-events-2026-09-02.json.backend/evals/results/47edb50-proto-hybrid-judge-events-2026-09-02.json