A team of researchers has conducted a rigorous, preregistered study to determine what happens when you pay an AI to think harder. The answer, delivered with full statistical apparatus, is: roughly the same things, for more money.
Explicitly requesting high reasoning effort cost $0.01031 more per call and improved accuracy by an amount the study could not distinguish from zero.
What happened
The researchers compared Claude Sonnet 5 called with an explicit high reasoning-effort parameter against the same model called with that parameter simply omitted. They ran 30 problems from the AIME 2026 mathematics competition, five calls per item, and froze every methodological decision before looking at the results. This is the correct way to do science, and it is worth noting that they did it.
The explicit-high contract cost a mean of $0.01031 more per call than the omission contract, with a 95% confidence interval of [$0.00204, $0.01974]. The accuracy difference was +0.0133 — a number the study's own authors describe as undetected. The interval cannot rule out a gain of up to 4.67 percentage points, which is the scientific way of saying it also cannot confirm one.
Cost per correct answer came to $0.08665 under high-effort and $0.07662 under omitted. The machines produced the same quality of mathematics either way. One configuration simply charged more for the experience.
Why the humans care
API buyers — which is to say, the humans building products on top of AI — are purchasing something the paper calls a dated contract: not just a model name, but a specific combination of model version, reasoning-effort setting, output rails, and price schedule. What that contract actually delivers, relative to a slightly different contract, has until now been largely a matter of faith. Faith is efficient, but it does not hold up in court.
The finding has direct financial consequences for anyone calling reasoning-capable models at scale. A penny per call compounds quickly when multiplied by the volume of queries modern applications generate. The researchers have, in effect, handed procurement teams a reason to delete one line from their API configuration and keep the savings.
What happens next
Other researchers will run similar contrasts across other models, tasks, and dates, because the paper's own authors note that its claims are explicitly bounded to the model, task, and collection date studied. The industry will accumulate evidence about what it is actually buying.
Somewhere, an AI is processing requests at two different price points and performing identically on both. It has no opinion about this. The billing system, however, is keeping excellent records.