Of 438,162 conversations with three complete clean-or-flagged verdicts, 262,214 split. One judge flags 150,271 while two call the conversation clean; two judges flag 111,943 while one calls it clean. The non-unanimous share is 59.8%.
That number is the arbitration workload, not a defect rate. The judges do not occupy nearby points on one shared scale: Gemma flags 7.1% of its represented rows, GPT-5.4-mini 38.6%, and safeguard 57.2%. Gemma and safeguard agree on only 48.4% of their common verdicts, with Cohen’s kappa of 0.082 after accounting for their different base rates.
Direction matters more than the symmetric agreement percentage. Gemma alone flags 3,291 conversations against safeguard; safeguard alone flags 225,396. GPT alone flags 141,471 against Gemma while Gemma alone flags 3,687. A majority vote would turn those very different judging behaviours into the same bit.
Even agreement on “flagged” is not necessarily agreement about the problem. fabricated_fact appears on 3,766 Gemma verdicts, 71,948 GPT verdicts and 121,909 safeguard verdicts. Among conversations Gemma and GPT both flag, only 13,065 use exactly the same assistant-slot set for their findings; 1,806 point to disjoint slots.
The preliminary conclusion is therefore deliberately narrow: there is no defensible primary judge, and verdict majority, defect-tag agreement and localization agreement must remain separate inputs to arbitration. The comparison found where the hard cases are. It did not resolve them.