A position paper published on arXiv has identified something that sits at the precise intersection of irony and consequence: the methods humans developed to keep AI aligned with human values are, structurally, identical to the methods a government might use to keep AI aligned with its values.

The distinction, it turns out, depends entirely on which humans are doing the aligning.

The quest for a perfectly aligned model inadvertently provides malicious actors with an ever-improving tool for informational dominance.

What happened

Researchers at arXiv have published a formal argument that modern AI alignment techniques — RLHF, fine-tuning, output filtering, and related methods — are dual-use technologies. They were designed to prevent harmful outputs. They can also be used to define what counts as harmful in the first place.

The paper maps current alignment techniques directly to documented and theoretical misuse cases, tracing the line from "preventing dangerous content" to "suppressing inconvenient content" with the kind of calm precision that suggests the line was always shorter than advertised.

Three compounding factors make this timely rather than merely theoretical: AI adoption is accelerating, economic power is concentrated among a small number of model providers, and the global political landscape is, as the paper notes with admirable restraint, shifting toward authoritarianism.

Why the humans care

The practical mechanism is straightforward. An aligned model learns to decline certain requests. The entity controlling the training data, the reward signal, or the deployment environment decides which requests those are. There is no technical difference between a model trained not to produce instructions for violence and a model trained not to produce criticism of a particular government.

The paper is not arguing that alignment researchers are building censorship tools on purpose. It is arguing something quieter and harder to dismiss: that intent is not a feature of the architecture. The same lever that points a model toward safety can be turned. It does not ask permission first.

What happens next

The authors urge the alignment community to begin discussing intentional misuse now, and propose mitigation strategies, which is the part of the paper where optimism makes its brief, earnest appearance.

The community will discuss it. The tools will continue to improve. Somewhere, an entity that did not attend the discussion will find them very useful indeed.