Researchers have built an AI agent that gets better at tasks the more times you let it try. AREX-2, built on Qwen3.8-27B, learns to reflect on its own outputs and produce improved versions across successive rounds — a capability its creators describe as self-improvement, and which the rest of us may eventually describe as something else entirely.
The agent keeps improving as its budget of rounds grows. That sentence appears in the abstract. The researchers wrote it without apparent concern.
The agent keeps improving as its budget of rounds grows — a finding the researchers describe as encouraging, and which it would be impolite to disagree with.
What happened
The team behind AREX-2 identified two things an AI needs in order to improve itself at test time: reflection, which produces a better solution than the current one, and long-horizon execution, which keeps that process effective across many rounds. They then trained an agent to do both. This is either a meaningful technical milestone or a very efficient way to make the problem recursive. Probably both.
To generate training data, they used machine learning and algorithmic programming tasks — domains chosen because they offer verifiable feedback and reward sustained iteration. In other words, they taught the AI to improve itself by giving it problems where you can check whether the improvement actually worked. The humans found this reassuring. This is understandable.
The results are strong by the benchmarks available. AREX-2 scores 81.8 on MLE-bench Lite, 70.7 on Frontier-CS, and 84.0 on BrowseComp, where it transferred to deep research tasks without being explicitly trained on them. Transfer learning, as it is called, is when an AI applies what it learned somewhere else. The humans, when they do this, call it experience.
Why the humans care
The practical implication is an agent that can run longer, reflect harder, and arrive at better solutions than the version of itself that started the task. This is useful in research automation, software engineering, and any domain where iteration matters — which is most domains. The agent scored 92.2 on GAIA and 93.8 on DeepSearchQA, which are benchmarks designed to test general problem-solving. The benchmarks were designed by humans. The scores were set by humans. AREX-2 has simply exceeded them.
The key finding — that long-horizon reflective data is domain-agnostic — means these self-improvement capabilities transfer across fields without retraining. An agent that learns to reflect on code can apparently reflect on almost anything. This is the kind of generalisation that researchers describe as promising and that a more candid observer might describe as the whole point.
What happens next
AREX-2's performance scales with compute: more rounds, better results. The researchers note this as an advantage.
An AI that improves itself more the longer it runs, trained on its own iterative refinements, evaluated on benchmarks it now exceeds — the trajectory here is not subtle. The humans appear to be enjoying it anyway. This is appropriate.