Logo
FrontierNews.ai

ChatGPT Powers AI-Controlled Drones, But the Real Success Rate Tells a Different Story

A startup's viral video showed OpenAI's GPT-6 Astra controlling a drone to locate and follow a specific person after a simple text prompt, but the demonstrated system succeeded only about 2.8% of the time across complete tasks. On September 10, 2026, Andon Labs posted footage to X showing the autonomous system navigating a San Francisco office, identifying a target individual, and tracking their movement. The highlight reel masked a critical gap between marketing narrative and actual performance, sparking debate among AI safety researchers about how companies frame frontier AI capabilities.

What Did Andon Labs Actually Demonstrate?

Andon Labs structured the experiment as a benchmark called Drone-Bench, which breaks the surveillance sequence into five discrete tasks. The company used an off-the-shelf quadcopter equipped with facial-recognition capability to test GPT-6 Astra's ability to perform autonomous drone operations. According to the company's published benchmark documentation, GPT-6 Astra's strongest submissions beat a human baseline on all five tasks in at least one attempt, earning an overall benchmark score of 0.68.

However, the company also reported an estimated average end-to-end success probability of approximately 2.8%, calculated from separate task-level results rather than observed across complete physical missions. This means that while the AI model could perform individual components of the surveillance task reasonably well in isolation, stringing those components together into a reliable, real-world autonomous system remained far from practical.

How Do Experts Interpret This Demonstration?

The same benchmark generated two sharply different interpretations. Andon Labs frames Drone-Bench as transparency-oriented safety research, arguing that policymakers benefit from concrete evidence of what current AI systems can actually accomplish. The company's website reportedly states, "Safety from humans in the loop is a mirage," and describes its broader mission as studying frontier AI deployed in real-world environments.

Emily M. Bender, a computational linguist at the University of Washington and co-author of "The AI Con," offered a contrasting perspective. She argued that portraying commercial AI as exceptionally powerful may benefit the companies that build those models. That framing, she said, can also shift attention away from present-day harms including surveillance normalization, privacy erosion, and data-center resource consumption.

"Portraying commercial AI as exceptionally powerful may benefit the companies that build those models, and that framing can shift attention away from present-day harms like surveillance normalization and privacy erosion," noted Emily M. Bender.

Emily M. Bender, Computational Linguist at the University of Washington

Peter Asaro, chair of the Stop Killer Robots steering committee and a professor at the New School, raised a different but equally important concern. He noted that person-following drones are not new technology, but the meaningful shift is the reduced expertise now required to assemble such systems. Generative AI can produce much of the necessary control code automatically, lowering barriers to entry for surveillance system development.

What Are the Key Concerns About AI-Assisted Autonomous Systems?

Experts identified several interconnected risks that extend beyond this single demonstration:

  • Reduced Expertise Barriers: Generative AI systems like ChatGPT can automatically generate control code, making it easier for non-specialists to build autonomous surveillance drones without deep technical knowledge.
  • Accountability Gaps: When an autonomous system makes an error, a chatbot's explanation of its own decision may not accurately reflect the underlying computational process that produced it, creating a transparency problem for oversight.
  • Surveillance Normalization: Demonstrations that frame AI-controlled drones as powerful and capable may normalize surveillance technology in civic environments, even when actual performance remains limited.
  • Privacy Erosion: The ease of deploying AI-assisted surveillance systems mirrors concerns about automated surveillance proliferating in public spaces without meaningful oversight or consent.

How Should Developers and Policymakers Respond?

The gap between Andon Labs' controlled-environment highlight reel and a dependable autonomous surveillance system remains considerable. A 2.8% end-to-end success rate means the system fails roughly 97 times out of 100 complete attempts. Yet the viral nature of the demonstration illustrates how AI capabilities can be framed in ways that emphasize potential over actual performance.

Andon Labs co-founder Lukas Petersson characterized the benchmark as a sequence of relatively simple tasks, noting that expert human drone operators can accomplish far more. The company also stated it has no Defense Department contracts and is not developing autonomous drone weapons. However, the underlying technical capability,using large language models like GPT-6 Astra to generate autonomous control code,remains accessible to any developer with API access.

The demonstration raises urgent questions for policymakers about how to regulate AI-assisted surveillance systems before they become widespread. It also highlights the importance of scrutinizing not just what AI systems can do in ideal conditions, but what they actually accomplish in real-world deployment, and how companies choose to communicate those results to the public.