One of the later 14B preference-training attempts reached 95.75% reward accuracy in its evaluation before the run failed. That looked encouraging on the dashboard.
The private 14B model also refused a simple request for a Fibonacci function. It understood the question perfectly well, mocked it in character and supplied no code at all. The older 8B model at least attempted the function.
Those results are not contradictory. Reward accuracy measured whether the evaluation process preferred the chosen examples over the rejected ones. It did not measure whether the model answered the user, stayed correct and added the personality without letting it take over.
I now keep usefulness and identity as separate judgments. A model can improve at sounding like GLaDOS while becoming much worse at being useful, and one reassuring percentage is quite capable of hiding that trade.