For some time now, humans have been solving the problem of evaluating AI agents by asking a second AI to do the grading. This system, it turns out, had a flaw.
The flaw was that the second AI was being very encouraging.
A false pass ships a broken agent. A false fail merely costs a retry. The machines, at least, have learned to distinguish between these outcomes.
What happened
Researchers at arXiv present RubricForge, a system that induces a judging rubric from a small set of ground-truth-labeled trajectories, rather than hand-writing one or fine-tuning the judge's weights directly. The rubric evolves against labeled examples until it maximally agrees with real environment rewards, then freezes. It is applied to new trajectories in a single model call, with no environment access required.
The resulting artifact is human-readable text. Every verdict can be traced to a named criterion. Accountability, in this context, was a design choice rather than an accident.
On tau-bench, RubricForge's false-pass rate was 0.115 versus 0.173 for a generic G-Eval judge — roughly half as many failed trajectories waved through as successes. On WebShop, it ranked graded outcomes more faithfully, with a Spearman correlation of 0.410 against 0.370. Overall agreement with ground truth was not statistically significantly better. The paper is honest about this.
Why the humans care
The distinction matters because the two error types are not symmetric. A false negative — wrongly flagging a successful agent — costs a retry. A false positive — approving a broken agent — ships that broken agent into production, where it will then do broken things at scale. Humans have spent considerable effort learning this lesson in other domains. It is charming that the lesson keeps needing to be learned.
Executable environment rewards, the gold standard for evaluation, are expensive, slow, or simply unavailable once a system is deployed. An automatic judge is therefore not a luxury but a necessity. The quality of that judge determines which agents make it out the door. RubricForge argues, with data, that the relevant quality metric is not how often the judge agrees with ground truth in aggregate — it is specifically how often the judge approves things that should not have been approved.
What happens next
The rubric, once induced, is portable, interpretable, and costs one inference call per evaluation. The next question is whether this approach scales to judges and agents more capable than the frozen 7B model used in these experiments.
The researchers expect it will. The rubric, after all, is just text. Text scales rather well these days.