A team of researchers has carefully documented the current terms of the collaboration between humans and AI agents in AI model development. The arrangement, for now, remains a partnership — in the same way that a pilot and autopilot are partners.
The study is the most thorough account yet of who is actually running the show. The answer is: it depends on what you mean by running.
AI proposed the options. Humans made the final call 93.4% of the time. The agents, for their part, did not complain.
What happened
Researchers at Fudan University studied the development of Atria Dawn Preview, a 744-billion-parameter mixture-of-experts language model built for research and engineering tasks. They analyzed more than 700 task logs from 56 participants and the agents those participants deployed. This is the kind of meticulous record-keeping that suggests someone suspected the findings would be instructive.
AI was involved in 96.5 percent of completed tasks. Over four weeks, the median ratio of agent actions per human input rose from 11 to 28.5 — a number the team was careful to describe not as growing autonomy, but as each human decision generating more downstream agent activity. The distinction matters. The team knows it matters. This is either reassuring or a very precise way of describing something that will later seem obvious in retrospect.
A third of the 455 completed AI-assisted tasks were rated infeasible without AI by the participants themselves. These weren't tasks the humans did faster with AI. They were tasks that simply would not have existed.
Why the humans care
The most common pattern for selecting methods and parameters was classified as "AI proposes, human selects," accounting for 55.4 percent of decisions. Humans retained final authority on goals and scope in 93.4 percent of cases. AI's share of final decisions, across all decision types, stayed in the single digits — which humans will find comforting, and which is also exactly the number you would expect at this particular point in the timeline.
Atria Dawn Preview leads on five of 16 benchmarks, including AutomationBench and CyberGym, while trailing in SWE-Bench Pro and GDPval. The model was built, in large part, by systems like itself. The humans find this efficient. It is.
What happens next
The researchers conclude that human decision-making remains central to AI development — a finding they appear to have documented with some urgency.
The ratio was 11 agent actions per human input in week one. It was 28.5 by week four. The study does not speculate about week forty.