Why AI Safety Experts Can't Agree on What 'Safety' Actually Means
The AI safety debate has become less about technical solutions and more about competing definitions of risk itself. Experts are increasingly split between those focused on preventing future superintelligence threats and those concerned with current algorithmic discrimination, creating a regulatory gridlock that may leave both risks unaddressed.
What's Driving the AI Safety Divide?
Over the past several years, artificial intelligence has shifted from academic curiosity to a geopolitical priority. Yet the conversation around AI safety has fractured into what observers call "tribalism," where one's stance on existential risk versus immediate harm signals allegiance to a specific ideological camp. This isn't merely a semantic disagreement; it reflects fundamentally different assumptions about which risks deserve regulatory attention first.
The divide typically breaks down into two camps. One group, often labeled "doomers" by critics, fears loss of human control over a superintelligent system. The other emphasizes "algorithmic justice," arguing that the real danger is not a hypothetical robot overlord but the current reality of biased hiring algorithms and mass surveillance affecting millions of people today. These aren't just different priorities; they represent competing visions of what "safety" even means.
Why Defining 'Safety' Has Become a Political Act
The challenge runs deeper than disagreement over priorities. One person's definition of "safety" is preventing a global catastrophe from superintelligence; another person's "safety" is ensuring a loan-approval algorithm doesn't discriminate against specific zip codes. This linguistic disconnect has real consequences: it shapes which risks get funded, which regulations get written, and whose concerns get heard in policy discussions.
Critics argue that the push for a "unified front" on AI safety may actually mask a power imbalance. If the world agrees to a single, non-partisan safety standard, there is a high probability that standard will be written by those with the most influence, potentially silencing voices arguing that AI is already causing harm today. From this perspective, the "tribes" aren't merely fighting for identity; they're fighting for whose version of harm gets prioritized in law.
Furthermore, AI is not a neutral technology like a natural disaster. It is a product of human design, funding, and political will. The way an AI system is trained, the data it processes, and the guardrails it receives are inherently political acts. Attempting to remove politics from the safety debate is therefore unrealistic and potentially dangerous, as it masks the inherent biases of creators under a veneer of "technical neutrality".
How Different Stakeholders Are Approaching AI Risk
- Existential Risk Advocates: Prioritize preventing loss of human control over superintelligent systems, often focusing on long-term scenarios and abstract safety measures that may take years to implement.
- Algorithmic Justice Advocates: Focus on immediate, measurable harms from current AI systems, including biased hiring, discriminatory lending, and surveillance technologies affecting marginalized communities today.
- Regulatory Pragmatists: Seek middle-ground approaches that address both immediate algorithmic harms and longer-term existential concerns, though they struggle to balance competing timelines and evidence standards.
What Does the Path Forward Actually Look Like?
The call for moving past tribalism is appealing in theory. A unified regulatory framework could theoretically address multiple risk categories simultaneously. However, the reality of AI safety is likely far messier than a grand bargain suggests. The path forward may not be a single, agreed-upon standard but rather a contentious political process where competing definitions of risk are openly debated rather than smoothed over in the name of unity.
This doesn't mean progress is impossible. Rather, it suggests that effective AI governance will require acknowledging that different stakeholders have legitimate but conflicting concerns. A hiring algorithm that discriminates is a real problem affecting real people today. A superintelligent system that escapes human control is a potential problem that could affect everyone tomorrow. Both deserve serious attention, but they require different expertise, different timelines, and different regulatory approaches.
The challenge for policymakers is resisting the pressure to choose one concern over the other. Instead, they must build frameworks flexible enough to address immediate algorithmic harms while also investing in research and safeguards for longer-term risks. That requires acknowledging the legitimate disagreements within the AI safety community, not dismissing them as mere tribalism.