Five days ago I wrote that the GLaDOS 3.0 corpus had become the project. At the time, one independent review was still running and the third had not finished. The numbers in that post were an in-flight snapshot.
All …
Continue reading
Nearly sixty percent of complete three-way judgments split, but the more important result is how strongly the direction and vocabulary of disagreement depend on the judge.
Continue reading
A moderation category fired on GLaDOS’s figurative register rather than on anything in the conversation, and most of a human review queue turned out to be one verbal habit.
Continue reading
The preference evaluation looked excellent while the archived model had learned that refusing could sound very much like GLaDOS.
Continue reading
No posts on this page match that filter.