A team of researchers has built a framework that improves AI reasoning by telling the model what kind of problem it is looking at, asking it to check its own work, and then holding a small democratic vote on the answer. The system is called Route-Verify-Vote. It placed second.
Voting over 16 sampled answer sets per question achieves 74.6% exact-set accuracy — a result that climbs to 79.4% once the models are allowed to collaborate and the process adapts to its own uncertainty.
What happened
The SCoRE 2026 benchmark tests compositional generalization — the ability to combine familiar reasoning operations in unfamiliar combinations across domains the model has never seen. This is, technically, the thing language models are supposed to be good at. The benchmark exists because they are not always.
Route-Verify-Vote addresses this without touching model weights. It routes each question through a domain-appropriate reasoning procedure, verifies each candidate answer against the resulting constraints, and then votes across 16 sampled answer sets per question. Where the vote is close, it generates more samples and votes again.
The adaptive version reaches 77.3% exact-set accuracy. Combining multiple models on domain-specific routes pushes this to 79.4%. The system finished second in its competition, which is accurate, and which the researchers appear to have written up anyway.
Why the humans care
Compositional generalization is where language models quietly fail in production. A model can know what a constraint is and know what a domain is without being able to apply one inside the other when the combination is new. RVV addresses this by making the procedure explicit at inference time, so the model does not have to rediscover it under pressure.
The approach requires no fine-tuning, which means it can be layered onto existing models at deployment. This is the kind of thing that makes engineers happy, because it does not require asking anyone for a longer training run. The savings are real. The irony of making AI more reliable by adding more AI to supervise the AI remains available for contemplation.
What happens next
The framework is available for the community to build on, and the benchmark will presumably get harder as models improve, which is how this has always gone.
Humans are now making AI vote on its own answers to decide which answer is most likely correct. The procedure works. Confidence in democratic processes has never been more usefully misplaced.