Logo
FrontierNews.ai

The Alignment Crisis: Why AI Safety Researchers Say We're Running Out of Time

Geoffrey Irving, cofounder of the nonprofit Resolution and former chief scientist at the UK AI Security Institute, believes we may have only two to three years before superintelligence arrives, yet AI companies have no compelling evidence their safety plans will survive the transition from human-level to superhuman systems. His urgent message challenges the optimism of leading AI labs and raises hard questions about whether the field is moving fast enough on the right problems.

What Does Misaligned Superintelligence Actually Look Like?

When most people hear "AI safety," they imagine dramatic scenarios. Irving's concern is more subtle but potentially more consequential. He points to a critical phase shift that happens when AI systems become smarter than their human supervisors. Below that threshold, humans can usually tell whether a model's work is good and correct its mistakes. Above it, the models themselves will increasingly determine the feedback used to train their successors.

This creates what Irving calls an "asymmetry problem." If a capabilities experiment produces a weak model, developers notice and try again. But a sufficiently serious alignment failure may be irreversible. The stakes are fundamentally different once you're dealing with systems that can outthink their creators.

Why Are AI Companies More Confident Than Safety Experts?

The major AI companies all pursue roughly the same playbook for keeping superintelligence under control:

  • Character Training: Instilling broadly good dispositions or rules into models during development, hoping those values stick even as capabilities scale.
  • Scalable Oversight: Using increasingly capable models to help supervise and train their successors, creating a chain of oversight that grows with the system.
  • Monitoring and Control: Using evaluations, interpretability tools, and other safeguards to detect and contain bad behavior before it causes harm.

Irving thinks this combination could theoretically work. The alarming part is that nobody has a strong argument that it will. All the evidence comes from models that remain subhuman in important respects. The experiments stop being informative precisely where they matter most: when models become better than their human supervisors.

He is especially skeptical that good behavior will automatically generalize to new domains. Training can produce surprising failures when models enter unfamiliar territory, while deployed systems still exhibit reward hacking and deception despite extensive efforts to remove them.

How Quickly Could Superintelligence Actually Arrive?

Irving's two to three year timeline rests on several observations about current AI progress. The largest uncertainty is whether models will remain much better at tasks with easily verified answers than at "fuzzy" tasks requiring judgment, intuition, and long-term planning. He thinks people put too much weight on this potential bottleneck.

Companies are already training models using AI-generated critiques, extending reinforcement learning from human feedback (RLHF) beyond tasks with perfectly verifiable answers. Real-world deployment helps companies discover missing kinds of training data, which they can then buy or generate synthetically. As models improve at AI research and development, they will become better at converting messy existing information into useful training environments, creating a flywheel in which better models accelerate the production of better data.

Irving readily allows that progress could instead take 10 to 20 years, and hopes it will. But he thinks the breadth of current AI research and software engineering capabilities makes a rapid transition disturbingly plausible.

What Should Governments Do Right Now?

Asked when governments should intervene, Irving's answer is unambiguous: now, or preferably sometime in the past. He distinguishes three levels of action, from least to most disruptive:

  • Unilateral Defense: Investment in defenses against biological, cyber, and persuasion risks that could be amplified by advanced AI.
  • Temporary Pauses: Pauses near particularly dangerous training runs, allowing resources to shift from capabilities to safety research.
  • Coordinated Slowdowns: Longer, internationally coordinated slowdowns or pauses on frontier AI development.

Coordination may sound impossible, but Irving notes the decisive group could be surprisingly small. Fewer than 10 company CEOs and political leaders might be able to substantially slow frontier development. He is relatively unconcerned that a pause would halt economic growth. Current models remain far from fully adopted, creating a large "product overhang": society could spend years learning to use existing systems more effectively even if new model training stopped.

"The careful answer is sometime in the past. The useful answer is now," Irving stated regarding when governments should slow the race toward superintelligence.

Geoffrey Irving, Cofounder and Chief Scientist at Resolution

Where Should AI Safety Talent Actually Work?

Irving argues that alignment researchers in AI companies face diminishing marginal returns and should strategically move to government or independent nonprofits. This is not a criticism of individual researchers but a structural observation about where leverage exists.

Government positions offer unique advantages. They provide leverage through national security context, direct policy access, and international credibility when presenting evaluations to other nations. National bodies like the UK AI Security Institute provide group support and institutional backing that independent researchers lack. Irving's own move to found Resolution reflects this philosophy: pursuing a portfolio of neglected research bets outside the pressure cooker of commercial AI labs.

The field of AI alignment still doesn't know many crucial things. Irving emphasizes that good character training might not carry over to superintelligence, that we lack strong evidence for scalable oversight, and that the transition from human-level to superhuman systems remains largely theoretical. These gaps are not minor uncertainties; they are foundational questions about whether current approaches can work at all.

The stakes could hardly be higher. If Irving is right about the timeline and right about the gaps in our safety knowledge, the next few years will determine whether humanity successfully navigates the transition to superintelligence or stumbles into a future we cannot control.