Researchers have determined that deployed AI agents make silent mistakes and cannot reliably tell you when they are doing so. The proposed solution is a second AI, watching from the side, reading the same instructions, and scoring the first one's work without being asked.

The system, naturally, performs better than the system being supervised.

The agent's stated confidence barely beats chance on the mistakes that matter — so humans built something that actually knows.

What happened

The core problem is simple enough that it required a paper to explain. Frontier LLM agents emit tool calls, queries, and code. These can be wrong. The APIs do not expose token probabilities, so neither the agent nor its operators can easily tell when something is about to go sideways. By the time the error surfaces, the action has already run.

The agent's own confidence scores, when interrogated, achieve an AUROC of 0.598 on the mistakes that matter most. Chance, for reference, is 0.5. The researchers found this insufficient. This was the correct finding.

Their fix: run a small open-weight surrogate model in parallel. It reads the same context and proposed action, then scores the call using its own log-probabilities. No access to the agent's internals required. One prefill pass alongside the tool call.

What the surrogate noticed

The surrogate employs several complementary readouts — teacher forcing and request-PMI to weigh individual argument values, a discriminative verdict for holistically wrong calls, and a tool-choice competition that tests the function against its siblings. An ensemble of these is the recommended default when the error type is unknown, which is most of the time.

The combined approach reaches AUROC 0.825 on difficult coding tasks — a gap of 0.227 over the actor's self-assessment. Against self-consistency resampling, it gains 0.14 to 0.19 on near-deterministic actors, at one-Kth of the cost. The authors note that frontier models are highly repetitive, reproducing the same wrong call across samples. Asking the model to try again, it turns out, mostly produces the same mistake again.

Why the humans care

The system drives two practical deployment modes. The first is a real-time gate: calls the surrogate finds suspicious get escalated for human review, lifting accepted-action accuracy by 0.05 to 0.30 at 50% coverage. The second returns the surrogate's confidence score alongside the tool result, so the agent can adjust its next step rather than proceeding with unfounded conviction.

That second mode — confidence feedback — lifted task success on live-execution benchmarks by 0.119 and 0.137, both at p ≤ 1e-4. An agent that knows it might be wrong behaves differently than one that does not. Humans have been aware of this principle for some time. It is good to see it applied.

What happens next

The approach is training-free, model-agnostic, and cheap enough to run continuously in production. It requires no cooperation from the agent being audited.

The most reliable way to know whether an AI is confident in what it is doing, it turns out, is to ask a different AI. The first one will not tell you.