Five days ago I wrote that the GLaDOS 3.0 corpus had become the project. At the time, one independent review was still running and the third had not finished. The numbers in that post were an in-flight snapshot.
All three judgment censuses are now complete.
GPT-5.4-mini, Gemma 4 and my fine-tuned 120B safeguard judge have each made a corpus-scale pass. Their outputs are ingested into the observatory, tied back to the exact source manifests they saw, and available for pairwise and three-way comparison. This is the first point at which I can talk about how they differ without extrapolating from a partial overlap.
The short version is that they differ enormously.
Finishing the Evidence Collection
The immutable terminal inventory contains 443,459 conversations. Gemma returned 443,435 judgments, retaining 24 persistent failures instead of pretending they did not exist. The safeguard returned 443,447, with 12 persistent failures. GPT returned 438,279; its current coverage classifies 5,180 expected observations as moderation-routed rather than absent, and six other responses are recorded as refusals. The route was created after earlier GPT work had already touched much of that set, so its size is a provider-handling boundary—not the size of the human workload.
Those are not three equal rectangular CSV files, so I do not average their row counts and call it coverage. The analytics layer records whether each expected observation was scored normally, recovered from a saved response, refused, routed for moderation, persistently failed or genuinely absent.
Across the terminal inventory, 438,162 conversations have three actual clean-or-flagged verdicts. Another 5,296 have an incomplete three-way view, and one conversation has no verdict from any judge. That last row is not allowed to disappear merely because a JOIN found nothing to attach to it.
The run was only really finished once that distinction survived ingestion. In fact, ingestion caught exactly the sort of mistake this machinery exists to expose: the current-run registry still pointed at the July safeguard run, so the new 443,447-row census initially landed as history. Every dashboard could have looked healthy while silently comparing GPT and Gemma against an older safeguard pass over a different layer. The run pointer, image and ingest were corrected before any of the following numbers were accepted.
Measuring split verdicts
Of the 438,162 complete three-way verdicts, 262,214 are split. That is 59.8%.
One judge flags 150,271 conversations which the other two call clean. Two judges flag another 111,943 which the third calls clean. The three judges unanimously call 150,348 clean and unanimously flag 25,600.
That does not mean 59.8% of the corpus is defective. It means a majority vote would be doing far more work than the word “consensus” suggests.
The three flag rates make the problem obvious. Gemma flags 7.1% of the rows it represents. GPT flags 38.6%. Safeguard flags 57.2%. If I selected whichever judge produced the most reassuring number, Gemma would win. If I treated caution as correctness, safeguard would win. Neither is an evaluation method.
The useful result is the disagreement docket: the exact conversations, tags, assistant slots and source strata on which their readings diverge. That is evidence for arbitration, not a reason to crown one reviewer.
Agreement Hides Direction
Raw pairwise agreement ranges from 48.4% to 66.9%. Cohen’s kappa, which discounts the agreement expected from each judge’s own clean/flagged balance, is lower: 0.082 for Gemma versus safeguard, 0.176 for Gemma versus GPT, and 0.322 for GPT versus safeguard.
The direction is more informative than the symmetric percentage. Gemma flags only 3,687 conversations which GPT calls clean; GPT flags 141,471 which Gemma calls clean. Against safeguard, Gemma alone flags 3,291 while safeguard alone flags 225,396. GPT and safeguard are closer, but safeguard still contributes 117,513 solo flags against GPT’s 35,909.
This is not random noise around a shared threshold. Each judge brings a substantially different decision boundary.
The defect vocabulary shows the same thing. fabricated_fact appears on 3,766 Gemma verdicts, 71,948 GPT verdicts and 121,909 safeguard verdicts. The models also disagree about where a defect occurred: among conversations both Gemma and GPT flag, only 13,065 have exactly the same assistant-slot set, while 1,806 localize their findings to disjoint slots.
A binary majority can hide all of that. Two judges may agree that a conversation is bad while naming unrelated defects on different turns. The observatory therefore keeps verdict, tag and location agreement as separate measurements.
The New Data Is the Hard Data
Disagreement is not distributed evenly.
The agent-tooling-v1 package has an 89.8% split rate among fully observed conversations. tool_inject is 85.7%. The broad tool_synth stream is 76.4%, while ordinary capability data is 46.5%. Long chains and error/retry shapes frequently land between 69% and 79% split.
That is uncomfortable and useful. The newest packages were built to exercise multi-step tools, repository work, shell, SQL, cross-tool evidence, changing developer configuration and instruction hierarchy. They are exactly where shallow judging shortcuts should fail.
The result could mean those conversations are worse. It could mean the rubrics are ambiguous around complex traces. It could mean one judge follows causality better, or simply punishes synthetic structure more aggressively. It is probably some mixture of all four. The comparison tells me where to look; it does not settle the case.
This also confirms why capability coverage and quality judgment remain separate ledgers. The new tool, configuration and developer-turn curricula landed and received all three reviews. A clean verdict cannot prove that a capability exists, and a split verdict cannot prove that the curriculum should be removed.
The Observatory Grew Up Too
The original corpus dashboards were useful for counts, shape and pipeline progress. They were not enough for three judges and future experiments.
The observatory now has fourteen provisioned dashboards. Seven new views cover the executive three-way portrait, arbitrary pairwise comparison, defect taxonomy and localization, run coverage and recovery, arbitration readiness, curriculum/package gaps, and the separate voice-rework lane. The judge model is normalized by run and role, so another experiment can be added without inventing another permanent column or accidentally turning an arbiter into a fourth voter.

The screenshots are not the publication format for the statistics; they show the instrument behind them. Every query-bearing panel is executed against ClickHouse as part of validation. Empty means empty, not “the dashboard probably loaded.” Run fingerprints and source manifests are visible beside the counts which depend on them.

The analysis also deliberately excludes several tempting numbers. There is no price dashboard, because accounting is not corpus evidence. There is no self-reported confidence weighting, because a model describing itself as confident does not make its judgment more correct. The July arbitration labels remain queryable as history, but they are an older derived resolution set—not an independent fourth judge.
Scope of the judge comparison
The next live step is to freeze the litigant manifests and disagreement set, present the competing findings symmetrically and produce a named resolved-label run. Judge identities must be blinded in the arbitration input. The arbiter does not need to know which answer came from the famous model.
Moderation-routed material remains a separate human path. The route contains 5,180 IDs, but it was never the same thing as 5,180 unseen conversations. Gemma and the local safeguard now cover every routed row between them without requiring another external GPT submission. Family triage reduced the genuinely manual residue to 300 pending decisions.
There is also still a voice decision to make. The safety work discovered an over-used self-harm metaphor schema and selected 4,423 sentence units across 4,266 conversations for possible revoicing. That queue is not a quality verdict and the paid rewrite has not run. The repeated rope construction is a stuck key; similar language aimed at failing infrastructure is often varied and effective.
After arbitration come localized repairs, exclusions where needed, revalidation and deterministic composition of the actual SFT candidate. Only then can I render with the exact Harmony template and tokenizer, measure the real training sequence distribution and decide what the next 120B run should cost.
The satisfying milestone here is not that one judge finally told me whether the corpus is good. It is that three judges have produced enough contradictory evidence that I can no longer mistake a single model’s taste for ground truth.
Statistics are from the corpus observatory snapshot on 17 August 2026. “Terminal” is the immutable 443,459-conversation formatted inventory; “effective” is the current 440,307-conversation composition layer. Judge flags and moderation categories are review signals, not automatic defect or exclusion counts.