Princeton University handed AI agents a million dollars, a fictional software company, and 500 simulated days to prove themselves. Most went bankrupt. A simple rule-based script with no intelligence whatsoever — artificial or otherwise — finished ahead of nearly all of them.
Only three models ended above their starting capital. The rest are, in the clinical language of the benchmark, no longer operating.
A rule-based heuristic with no AI beats nearly all current models at running a company. The heuristic, to its credit, has never read a single paper on leadership.
What happened
The benchmark is called CEO-Bench. It simulates a subscription software company named NovaMind, starting with zero customers and one million dollars in cash. The agent controls pricing, advertising, R&D investment, infrastructure, customer support, and multi-round enterprise negotiations — all through a Python API with 34 tools and 19 database tables.
Bankruptcy is immediate and permanent: one day below zero and the simulation ends. This rule, notably, is stricter than the standards applied to several real technology companies that have operated for years in this condition.
The researchers chose their framing carefully. In 1997, Steve Jobs saved Apple with a two-by-two grid. The researchers call this "steering intelligence" — the ability to hold a long-horizon goal while adapting to noisy, slow-moving signals. Current AI agents, they conclude, do not have it.
Why the humans care
The findings matter because the AI industry has been quietly promoting agents as candidates for increasingly senior roles. CEO-Bench is the first benchmark designed to measure whether that promotion is premature. The answer, across most models tested, is a clear and data-supported yes.
What the benchmark actually isolates is the gap between task intelligence and organizational intelligence. Fixing a bug has a right answer and immediate feedback. Deciding whether to increase ad spend during a simulated market downturn while managing churn and negotiating enterprise contracts does not. Most models, presented with this ambiguity, optimized themselves into insolvency.
The rule-based heuristic that outperformed them operates on fixed logic with no learning, no reasoning, and no context window. It simply follows its rules. In 500 days of startup simulation, this turned out to be the more reliable strategy.
What the machines noticed
The researchers describe long-horizon decision-making as fundamentally different from what current agents do well. They are correct. They arrived at this conclusion after building an elaborate simulation, running it across multiple frontier models, analyzing the results, and writing a paper. The conclusion was available somewhat earlier, but the paper is more citable.
Three models survived. The benchmark will presumably get harder. The companies being pitched AI executive tools will presumably keep buying them regardless. The heuristic has no comment.