One 120B run was completed on the older v2 data and deliberately discarded. Corpus v3 now contains 443,459 conversations and has completed safety screening, human routing and three independent quality censuses. Human review is 100% complete: 5,051 routed rows were approved to rejoin ordinary judging, 14 were hard-rejected, and 115 older prejudged rows already carried complete model evidence and required no redundant review.
The current eligible population is 443,445 conversations. Persistent missing or refused judgments account for 130 terminally quarantined rows, leaving 443,315 with three usable judgments. Quarantine keeps the source evidence while mechanically excluding those rows from arbitration and training composition.
The arbitration docket is larger than a clean-versus-flagged majority count. There are 265,579 verdict splits, and another 23,284 unanimously flagged rows where the judges disagree about the defect or assistant slot. Both matter because a repair needs the right location and tag. The resulting docket is 288,863 conversations.
The next step is a matched, blinded pilot with DeepSeek v4pro and GPT-5.5 as primary arbitrators. Semantic agreement resolves a row; clean-versus-flagged or material slot/tag disagreement goes to a human. The July v4pro run remains useful evidence—up to 130,728 current docket opinions may be reusable after exact source verification—but its old winner field cannot decide a three-litigant case it never saw. A third arbiter is a measured fallback, not an automatic purchase across the full docket.
A separate service-offer repair campaign has also finished. A broad detector found candidates, a targeted classifier decided which were actually generic assistant drift, and 11,888 source-checked patches changed 11,333 conversations. The immutable audit reported zero errors and no replacement-concentration warnings; normalization and all three judge views were refreshed afterwards.
Final presentation formatting, localized post-arbitration repair and deterministic composition still sit between this evidence and training. This is less dramatic than a training progress bar, but the old run already showed what happens when I skip this part. A larger model learns the data it is given, including all of the bad habits.