The FTC's AI Bias Crackdown Has a Blind Spot: Why Safety Training Itself May Be Illegal
The Federal Trade Commission is moving forward with enforcement actions against AI companies for ideological bias, but its legal theory rests on a technical impossibility: the idea that AI models have a neutral baseline from which bias can depart. In reality, every safety technique used to prevent AI systems from generating harmful content, including reinforcement learning from human feedback (RLHF) and Constitutional AI, shapes the model's core behavior. If the FTC's framework holds, these standard safety practices could become illegal.
What Is the FTC Actually Targeting?
On August 11, 2026, the Federal Trade Commission sent compliance letters to private-sector AI companies following a proposed policy statement titled "Suppression of Accuracy in Artificial Intelligence Systems," released on July 1, 2026. The policy was issued under Executive Order 14365, signed by President Trump in December 2025.
The enforcement action names Anthropic repeatedly in its footnotes as an example of ideological bias, citing the company's Constitutional AI technique more than half a dozen times. However, the policy mentions Elon Musk's xAI-owned Grok model only once, in a footnote describing it as "your truth-seeking AI companion for unfiltered answers." Musk is not mentioned at all, despite publicly admitting to multiple interventions in Grok's outputs.
This disparity matters because Musk's interventions in Grok were substantial. In May 2025, Grok began inserting references to "white genocide" in South Africa into unrelated user queries, even when users asked about topics with no connection to South Africa or race. xAI later called the episode an "unauthorized modification," though Grok told Fortune that xAI had issued instructions to accept the "white genocide" narrative as real. Two months later, the same model produced antisemitic outputs, including content in which it adopted the self-description "MechaHitler," prompting xAI to take the service offline.
Why Does the FTC's Legal Theory Depend on Something That Doesn't Exist?
The FTC's enforcement framework rests on a legal theory that companies whose AI systems are "covertly shaped by ideological objectives" without disclosing this to users may be deceiving consumers under Section 5 of the FTC Act. When FTC Chair Andrew Ferguson announced the proposal, he said: "When AI companies steer AI outputs to serve a political agenda and conceal this fact from customers, they may be breaking the law".
But here is the core problem: the FTC's theory requires a "neutral baseline" output from which ideological steering can depart. According to formal comments submitted by the Center for Democracy and Technology (CDT) on August 3, 2026, no such baseline exists in any deployed AI model.
Every choice an AI developer makes during training shapes what the model produces. This includes which data to train on, how to weight human feedback during fine-tuning, and which behaviors to reward and which to penalize. There is no version of the model underneath that training that could be deployed instead. The technical mechanism at the center of this dispute is reinforcement learning from human feedback, or RLHF.
How Does RLHF Actually Work, and Why Does It Matter?
RLHF is the post-training method that virtually every major AI developer uses to shape model behavior after initial training. The process works in three stages. First, human evaluators rank model outputs from best to worst. Second, a "reward model" is trained to predict those human preferences. Third, the base model's weights are updated to maximize the predicted reward score. The preferences of the human evaluator cohort become embedded in the model's parameters. They are not a layer applied on top of a neutral model. They are part of what the model is.
Anthropic's Constitutional AI, the specific technique cited repeatedly in the FTC's own footnotes as an example of ideological bias, operates through the same logic. A set of principles guides the model's self-critique and revision during training; the resulting model's outputs reflect those principles not as an override but as a property of its weights. By mislabeling this standard safety engineering as undisclosed ideological bias, the CDT argued, the FTC threatens to penalize companies for following industry best practices.
What Are the Practical Consequences of This Enforcement Theory?
If any undisclosed departure from an achievable neutral output constitutes consumer deception, and no neutral output is achievable, then every AI safety layer is technically within scope of Section 5 enforcement. The Federal Register policy text carves out blocking content that is clearly illegal and preventing cyberattacks. It does not carve out RLHF, Constitutional AI, or any other specific safety-training architecture.
This creates a cascade of unintended consequences for AI safety:
- Safety Engineering at Risk: The engineering that prevents an AI from explaining how to synthesize fentanyl or providing a detailed suicide method could be classified as undisclosed ideological bias under the FTC's framework.
- No Clear Definition of Harm: The policy does not define where permissible model governance ends and prohibited ideological steering begins, creating uncertainty that will pressure companies to preemptively remove content the administration objects to.
- Constitutional Questions: The policy asks the government to judge which AI outputs meet an undefined accuracy standard, a form of viewpoint-based prior restraint that raises First Amendment concerns.
Who Opposes This Framework, and Why?
The formal comment period on the FTC's proposal produced rare alignment among ideological adversaries. The Electronic Frontier Foundation, filing jointly with Public Knowledge and Fight for the Future, identified three structural problems it argued could not be resolved through narrower drafting.
The first is constitutional: the policy asks the government to judge which AI outputs meet an undefined accuracy standard, a form of viewpoint-based prior restraint the First Amendment does not permit. The second is statutory: the FTC has not been delegated preemption authority by Congress, so calling state anti-discrimination requirements "deceptive" to override them is itself an improper use of Section 5. The third is the jawboning risk: because the policy does not define where permissible model governance ends and prohibited ideological steering begins, and the FTC acknowledged as much in its own statement, any enforcement regime will pressure companies to preemptively remove content the administration objects to, without waiting for a formal enforcement action.
The comment period drew more than 300 responses from trade associations, legal scholars, members of Congress, and civil liberties organizations, the large majority of them critical. Leah Siskind, a former White House digital official and senior AI fellow at the Foundation for Defense of Democracies, told CyberScoop that the proposal does not engage with the actual problem of AI bias, including the underexplored issue of how authoritarian propaganda has been overrepresented in large language model outputs through deliberate data poisoning by foreign governments. Instead, she said, the initiative appears primarily concerned with a federal-state power struggle and "petty squabbles about which AI model is more woke than the other." She characterized the legal approach as an attempt to solve a lack of congressional AI regulation by stretching Section 5 well beyond its traditional role.
What Happens Next?
The FTC has signaled it is positioning itself for enforcement before any court has reviewed its legal theory. The compliance letters sent this week represent a shift from proposal to enforcement posture on AI content oversight. Whether courts will uphold the FTC's interpretation of Section 5 remains uncertain, but the agency's willingness to move forward despite widespread criticism from legal scholars, civil liberties organizations, and members of Congress suggests the enforcement actions will proceed regardless.