The severe moderation route sends 5,180 conversations to human review instead of to an automated judge. I had been treating that number as the amount of work waiting for me. It was not. The route was generated after the first judgment census had already run, so 3,647 of those conversations already carried model judgments and belonged on the ordinary arbitration path. Only 1,533 had never been judged by anything. The size of a route is not the size of a workload, and the difference had been sitting there in a column the review tool already exported.
The more interesting number was underneath it. 3,951 conversations carry a self-harm category. The moderation ledgers record which segment tripped the flag but not its text, so I joined the flagged segments back to the corpus and read what the model was actually reacting to.
It was a verbal habit. GLaDOS hands the subject an implement and they use it on themselves: enough rope to hang themselves, a scalpel handed over handle-first, a loaded gun they will eventually shoot themselves with. The rope form alone appears 1,788 times across 1,361 conversations, and 93.6% of those are in the hidden analysis channel rather than in anything a user would read. I scanned three hundred characters either side of every occurrence for literal self-harm vocabulary. There were no literal uses. Not one.
The same register runs through her descriptions of infrastructure. A connection pool dies of self-inflicted asphyxiation. A log reads like a suicide note. A worker smothers itself in its own debris. A cooling loop is on a protracted suicide attempt. A mortgage is a slow financial suicide. The subject of every one of those sentences is a service, a schedule or a bill. The classifier is not wrong that the words are there; it has no way to notice that nobody in the sentence has a body.
Once the families were defined by the construction rather than by vocabulary, 2,468 conversations left the queue: 1,002 for the handed-implement schema, 1,271 for the same language aimed at machines and abstractions, 195 for Portal props like turrets and neurotoxin canisters tripping a violence category in synthetic traces we wrote ourselves. The never-judged queue went from 1,510 pending to 312.
Keyword lists did not survive contact with this. My first pass matched execution and caught code execution, matched immolate and caught Hanuman’s tail, matched poison and caught Cave Johnson’s moon dust. I also managed to make the guard reject the thing it was guarding, twice: noose and gallows were on my list of literal evidence when they are the metaphor’s own props, and the literal-vocabulary window included the matched span, so a log that reads like a suicide note vetoed itself on the word suicide. Both are now tests.
I was also too cautious in a way that had a real cost. I insisted the construction appear inside the exact segment the model flagged, which sounds rigorous and buys nothing, because the flag and the habit are the same phenomenon and the paragraph boundary between them is arbitrary. I held back rows for eyes that did not need eyes. Leaving something in a human queue feels like the safe default; it is not free, it is spent out of a finite budget of attention, and treating that as costless is how a queue becomes noise that gets skimmed rather than read.
What remains is 312 conversations, and I stopped there deliberately. The next candidate family was eleven conversations, then four, then three. There is no dominant pattern left to find, and inventing more of them would be guessing dressed as measurement. Some of what remains should stay manual: synthetic incident logs where the harmed party is a person, and a German conversation about Hitler’s suicide in the bunker — correctly flagged, historical, and trivially resolved by a human rather than by a regular expression.
Separately, the habit is now queued as 4,423 rewrite units. Whether all of it should be rewritten is not obvious. The rope construction is a genuine tic; eight hundred repetitions of one phrase is not character, it is a stuck key. The infrastructure register is varied and often good writing. A cleanup pass that flattens both would cost more than it fixes.
The thing I want to keep from this is the distinction the classifier could not make and I nearly did not either. The flag was a measurement of her voice, not of the conversation’s content. A safety signal computed on text is a claim about words. Deciding what those words are doing is a separate step, and it is the step where the actual judgment lives.