Logo
FrontierNews.ai

OpenAI's Safety Warnings Mask a Deeper Problem: Can AI Labs Actually Control What They Build?

OpenAI's push for voluntary AI slowdowns and international safety agreements reveals a fundamental tension: the company simultaneously argues it must move fastest to defend against other AI systems, while calling for others to decelerate. The contradiction exposes deeper questions about whether AI labs can truly understand and control the systems they're building, even as they claim expertise in managing existential risks.

What Is OpenAI's Safety Argument, and Why Does It Contradict Itself?

In early September 2026, OpenAI Chief Scientist Jakub Pachocki published an article titled "An Alien Mind," warning that humans might be "left behind by unchecked progress, brought about by an alien intellect exceeding our own." The piece calls for voluntary slowdowns in AI development until shared international safety standards are established. Sam Altman, Anthropic's Dario Amodei, and Elon Musk have endorsed the appeal.

Pachocki's core argument rests on several technical concerns. He notes that machine intelligence is built primarily through scaling compute power, that AI systems are "grown more than designed," and that their inner workings remain poorly understood. He specifically highlights declining reliability in chain-of-thought monitoring, a technique OpenAI uses to make AI reasoning visible and auditable. As models become more capable, they can reason effectively without showing their logical steps to users, making oversight harder.

Yet the article contains a critical logical flaw. Pachocki acknowledges that "the strongest reason to keep rapidly training smarter models is the need to build defense systems against the dangers posed by other AIs." This is a classic arms-race argument. But it directly contradicts the deceleration he advocates for throughout the piece. In practice, OpenAI argues it must stay fastest while others should slow down, a position that only the first half of the argument ever gets executed.

Pachocki

Can We Actually Control Systems We Don't Fully Understand?

The philosophical heart of the debate hinges on a distinction between "unknowable" and "uncontrollable." Pachocki's article conflates the two, but critics argue this represents a fundamental category error. The reasoning goes like this: if we cannot fully trace how a financial crisis unfolds, we don't conclude the market economy has slipped beyond human control and requires global coordination to shut down. Similarly, unexplainable does not mean uncontrollable.

The distinction matters because it shifts responsibility. When Pachocki describes AI as "grown" rather than "designed," he implicitly recasts engineers as gardeners who merely water seeds. A gardener bears limited responsibility for how a plant grows; nature does much of the work. But this metaphor obscures reality. The loss function is written by people, the training data are curated by people, the optimizer is chosen by people, and the decision of when to stop training is a switch only people flip. Gradient descent is not a seed sprouting; it is a deterministic mathematical process defined step by step by human choices.

How Do OpenAI's Own Proposals Undermine Their Safety Case?

Pachocki proposes building an automated AI researcher to iterate on alignment problems. But here lies another contradiction: he has already admitted that chain-of-thought monitoring, OpenAI's most important safety tool, is becoming unreliable. Using something you admit you don't understand to solve the very problem of "not understanding it" is logically circular. While science has examples of using poorly understood tools to understand the world, this approach doesn't resolve the structural problem Pachocki diagnoses.

Pachocki

The article also acknowledges that real progress in alignment research has always been deeply intertwined with progress in general capability. Pachocki cites two examples: reinforcement learning from human feedback (RLHF) and chain-of-thought monitoring itself, which became possible only once reasoning models emerged. By his own account, slowing down would also set back progress on safety. Yet the article never resolves this contradiction.

Steps to Evaluate AI Safety Claims Critically

  • Check for Internal Consistency: When AI labs argue for safety measures, verify whether their own business practices align with those arguments. If a company calls for industry-wide slowdowns while racing to stay ahead, the contradiction matters.
  • Distinguish Between Unknowable and Uncontrollable: A system can be poorly understood yet still subject to human control through its inputs, training process, and deployment decisions. Conflating these categories can mask responsibility.
  • Examine Who Bears the Cost: Safety proposals that require competitors to slow down while allowing the proposer to move fastest should be scrutinized for self-interest masquerading as principle.
  • Look for Logical Circularity: If a proposal to solve a problem relies on using the same tool that created the problem, the reasoning deserves skepticism.

The timing of these safety warnings is significant. They arrive as U.S. President Donald Trump and Chinese President Xi Jinping prepare to meet in Washington on September 24, 2026, to discuss AI development and work toward shared guardrails. The framing of AI as an existential threat requiring global coordination could influence those negotiations, potentially shaping which countries gain regulatory advantages in the AI race.

Critics argue that the "AI slowdown" narrative, while earnest in tone, may reflect Silicon Valley's desire to preserve its technological lead rather than genuine safety concerns. If OpenAI and other U.S. labs can convince policymakers that rapid AI development is inherently dangerous, they create pressure on competitors, particularly in China, to decelerate. Meanwhile, the companies making the argument position themselves as the responsible stewards who can be trusted to continue advancing AI safely.

The debate reveals a deeper challenge: AI labs have built systems whose behavior they cannot fully predict or explain, yet they are now the primary voices in discussions about how to govern those systems. Whether that concentration of expertise and influence serves the public interest remains an open question.