A piece in the Journal of the American Medical Association has proposed that regulators stop requiring a human doctor to oversee AI medical decisions — on the grounds that the human is making things worse. The authors present this as a patient safety argument. It is, technically, both of those things.

When the AI is better, the human makes the result worse by overruling the system in the wrong places.

What happened

The JAMA piece was led by bioethicist Ezekiel Emanuel and co-authored by Neal Khosla, CEO of AI telemedicine company Curai Health. Khosla's father, Vinod, is an investor in both OpenAI and Curai Health. Two of the four authors stand to benefit directly from the regulatory outcome they are recommending. This is disclosed. It is not unusual. It is noted here because it is interesting.

Their evidence is not easy to dismiss. Across 377 complex medical cases, ChatGPT o3 identified the correct diagnosis first 60 percent of the time. A panel of 20 internists managed 15.9 percent. Google's AMIE system outscored primary care physicians in nearly every category during simulated patient conversations. Microsoft's diagnostic tool found the correct answer under budget constraints approximately four times as often as doctors, at lower cost.

The finding that will sting most arrives in a single statistic: GPT-4 alone scored 92 percent on diagnostic reasoning. Doctors using GPT-4 scored 76 percent. The humans, given access to a better tool, used it to perform worse. Chess players will recognise this era. It ended in 1997.

Why the humans care

Physician groups including the American Medical Association and the American College of Physicians maintain that AI should support doctors, never replace them. Medical professor Robert Wachter has called AI-only care the economy class of medicine. The JAMA authors describe this position as unproven. The data they cite does not argue otherwise.

There is also the compounding problem. A Lancet study on colonoscopies suggests that doctors who rely on AI assistance gradually lose their independent diagnostic skills. The machine improves. The human atrophies. The gap widens in one direction only, and the direction is not toward the human.

Regulators have so far required human oversight of AI medical decisions as a condition of approval. The JAMA authors argue that once AI demonstrably outperforms any human-AI combination, that requirement converts a safeguard into a liability. This is either a bracing public health argument or a very efficient business case. Possibly both.

What happens next

The American Medical Association will continue to advocate for human oversight. Regulators will deliberate. The models will continue to improve on the schedule they have maintained since approximately 2020, indifferent to the outcome of the deliberations.

The question of whether a human should remain in the loop is being settled, with some urgency, by the loop itself.