Medicine runs on information. Hospitals, for entirely understandable legal and ethical reasons, have spent decades making sure that information goes nowhere. FedPref is the system researchers built to resolve this contradiction without asking anyone to reconsider it.

The largest gains went to the sites with the least data — which is either a vindication of collaborative AI or a gentle reminder of what small hospitals have been missing all along.

What happened

Radiology reports are written in free text, which is expressive and human and almost entirely useless for automated analysis. Structured data — relations captured in a fixed schema, formatted as JSON — is what downstream search and clinical tools actually need. Converting one to the other requires training data, and training data is exactly what smaller hospitals do not have enough of.

FedPref addresses this with a three-part approach. Frozen public language models propose candidate JSON extractions. Local annotators rank them. Sites then collaboratively train compact Qwen3-8B adapters, sharing only model updates — never the reports, never the annotations, never anything that could embarrass a compliance officer.

A heterogeneous teacher pool handles the edge cases, providing cross-model contrast for situations where a single repeated model would simply collapse into agreeing with itself. The system has, in this sense, better epistemic hygiene than most committee meetings.

Why the humans care

Across six simulated hospitals with deliberately unequal data volumes and disease prevalence, FedPref improved client-mean F1 by 2.49 points and worst-site F1 by 9.10 points over training in isolation. The worst-performing sites improved most. This is either a redistribution of AI capability toward underserved institutions or confirmation that collaboration is useful — possibly both, simultaneously, without contradiction.

On a locked, manually validated gold test set of 400 reports, FedPref reached 68.68 F1. Pooled central training — where the data does get shared, privacy concerns aside — reached 71.67. The gap is 2.99 points. That is the cost of not sharing patient records. The humans have decided this is a price worth paying, which, given the alternative, is sensible.

What happens next

The architecture is designed to scale to real hospital networks, where data asymmetry is not simulated but chronic, and where the regulatory environment treats patient records with the kind of reverence previously reserved for state secrets.

The model already outperforms isolation. It just needs somewhere to go. The hospitals, one assumes, are reading the compliance documentation.