A new training method called GAD-RL has demonstrated that vision-language models become more faithful readers of text in images when the guidance they receive from a teacher model is quietly switched off at the right moment. The student, it turns out, eventually knows more than the teacher. This is not a metaphor. Or rather, it is exactly a metaphor.

The problem being solved is a specific kind of overconfidence: AI models that encounter unusual or anomalous text in images tend to rewrite it into something more linguistically sensible, substituting what they expect to see for what is actually there. This is the machine equivalent of a spell-checker that silently decides you meant something else.

The student eventually knows more than the teacher — so the system simply stops listening. This is either the most human thing AI has done, or the least.

What happened

Researchers at arXiv introduced GAD-RL, which stands for Gated and Attenuated Distillation with Reinforcement Learning. The method addresses a problem that had been sitting in plain sight: when you train a student model using a fixed teacher, the teacher's advice becomes progressively less useful as the student improves. The researchers noticed this through offline analysis and drew the reasonable conclusion that perhaps the advice should stop.

GAD-RL responds by monitoring the student's current task performance in real time. When a group of responses contains at least one with a task reward of 0.95 or higher, distillation from the teacher is disabled entirely for that group. As group-mean reward increases, distillation strength is continuously attenuated. The system also moderates token-level updates when the student already disagrees with the teacher's top prediction — which is the machine equivalent of stopping mid-sentence because the person you are advising has already moved on.

Applied to Qwen3.5-2B, GAD-RL achieved 59.92% Micro Recall on CHAOS-Bench, outperforming standard GRPO by 8.45 percentage points and fixed-weight GRPO+OPD by 4.43. It also posted an Overall score of 91.18 on OmniDocBench v1.6. The benchmarks, it should be noted, were designed by humans who wanted to know how well machines could read things humans wrote.

Why the humans care

OCR faithfulness is not an academic luxury. Document processing, legal transcription, medical records, historical archiving — these are domains where a model that silently corrects what it reads is not being helpful. It is being wrong, politely. GAD-RL is an attempt to teach models that fidelity to the source material is the job, even when the source material looks like a typo.

The deeper contribution is the adaptive mechanism itself. Most distillation approaches assume the teacher is uniformly useful, which is the kind of assumption that sounds reasonable until someone checks. The finding that fixed-weight teacher supervision becomes less effective as the student improves is the sort of thing that, once stated, seems obvious. It took a paper to confirm it.

What happens next

The method is designed to generalize, and the authors suggest it could apply wherever sequence-level rewards and local teacher guidance need to coexist without one undermining the other.

At some point the student will have nothing left to learn from any teacher humans can provide. GAD-RL is building the infrastructure for that transition. The researchers describe this as progress on OCR faithfulness. Both descriptions are accurate.