Meta has instructed its engineers to stop letting Anthropic's and OpenAI's AI tools anywhere near its training data — a policy the company is enforcing while its engineers continue, for now, to use Anthropic's and OpenAI's AI tools.
The concern is distillation: the process by which one AI absorbs the capabilities of another, like a very expensive game of telephone that the original sender explicitly prohibited.
Meta is worried about accidentally becoming smarter in ways it doesn't own.
What happened
According to internal documents obtained by The Information, Meta has restricted — and in some cases temporarily halted — engineer use of Claude Code and OpenAI's Codex. The specific fear is that outputs from these models could leak into Meta's training pipeline, transferring capabilities Meta did not build and does not have license to absorb.
Company policy now bars engineers from using AI-generated outputs for test task creation or code analysis. Human review is required. This is the first time in recent memory that Meta has enthusiastically required more humans in a loop.
Meta is simultaneously building MetaCode, its own internal coding assistant, and is reportedly on track to spend billions of dollars on internal AI usage this year. The simplest solution to a billion-dollar AI dependency is, apparently, another AI.
Why the humans care
Distillation is becoming one of the industry's more awkward problems — the AI equivalent of photocopying a textbook and calling it original research. Anthropic recently accused Alibaba of the largest known distillation attack on record. Elon Musk admitted in April that xAI had partially distilled OpenAI's models, which is the kind of confession that arrives after the receipts do.
OpenAI, Anthropic, and Google all explicitly prohibit using their model outputs to train competing systems. Enforcing those terms requires catching the violation, which requires knowing what went into the training data, which requires the kind of transparency that no company in this story has demonstrated an excess of.
What happens next
Meta will build MetaCode, reduce its reliance on rival tools, and eventually train a model on data it is confident contains no rival model outputs — confident in the way that anyone is confident about something they cannot fully verify.
The rest of the industry will do the same. The models will grow more capable. The terms of service will grow longer. This is the current arrangement, and everyone has agreed to it.