Logo
FrontierNews.ai

OpenAI's Sora Struggles as Contractors Caught Gaming the System

OpenAI has terminated contractors hired to improve its AI models after discovering they were using artificial intelligence to complete their work, according to internal documents and contractor interviews. The firings reveal a fundamental tension in how AI companies scale human oversight, and they come as OpenAI faces mounting competitive pressure in the broader AI market.

Why Is OpenAI Firing Contractors for Using AI?

OpenAI employs thousands of contractors through its "Project Lily" initiative to read real ChatGPT user prompts and conversations, then rate and critique the AI's responses. The goal is straightforward: inject human judgment into model training to prevent problems like "model collapse," a phenomenon where AI models trained repeatedly on AI-generated text become progressively worse.

But some contractors have been circumventing this process by using AI tools themselves to complete their assignments. According to internal documents obtained by 404 Media, contractors are explicitly prohibited from using AI detection tools, Grammarly, AI translation services, or any large language model (LLM) to review, write feedback, or compose comments. One contractor described the enforcement as strict: "People are let go for it all the time, it's pretty much the one thing that will get you kicked off ASAP," they said. In a group of thousands of contractors, "tons have been caught".

The irony is sharp. OpenAI's entire business model centers on encouraging people to use AI at work. Yet the company is simultaneously firing people for doing exactly that in their contractor roles. This contradiction underscores a real problem: if AI-generated responses are good enough to fool human reviewers and slip into training data, the resulting model degradation could compound over time.

What Signs Do Reviewers Look For to Catch AI Use?

OpenAI's quality assurance teams are trained to spot telltale patterns of AI-generated work. Reviewers watch for repetitive word choices, overzealous punctuation (particularly excessive em dashes), and suspiciously fast completion times. In internal Slack channels, contractors frequently post examples asking colleagues, "Is this AI?" to calibrate their detection skills.

One contractor who was terminated shared what they described as their dismissal letter, which cited issues with the "authenticity" of their work. "I'm not a bad person or worker," the contractor reflected. "I just needed a little boost and turned to AI to help me which eventually led to my downfall. I felt no joy in the work or that I was contributing to society in any way".

Mercor, the AI-training company that hires many of these contractors on behalf of OpenAI, stated that it takes the issue seriously. A Mercor spokesperson explained: "Our experts are hired for their expertise and judgement, which is essential to the ongoing advancement of AI. Our contracts strictly prohibit the use of LLMs to complete projects and we enforce that. We invest heavily in our tools and systems to detect misuse and ensure our experts comply with project rules and contract terms. When we confirm an expert has used AI to complete a task, we immediately remove them from the project".

How to Spot and Prevent AI Misuse in Training Data

  • Pattern Recognition: Train reviewers to identify repetitive vocabulary, unusual punctuation patterns, and unnaturally fast work completion as potential indicators of AI-generated content.
  • Blind Auditing: Implement secondary review processes where auditors evaluate work without knowing which pieces are flagged as suspicious, reducing confirmation bias.
  • Contractor Rotation: Regularly rotate contractors between projects and teams to prevent long-term patterns of undetected AI use from accumulating in training datasets.
  • Transparency Training: Educate contractors about the real consequences of model collapse and explain why human authenticity matters to the quality of downstream AI systems.

What Does This Mean for OpenAI's Competitive Position?

The contractor firings arrive at a precarious moment for OpenAI. According to Comscore's Q2 2026 AI Intelligence Report, ChatGPT's market share has fallen 20 percent over the past six months, dropping to 50 percent of total prompt volume. Google Gemini has surged to 30 percent market share, up 13 percentage points, while Anthropic's Claude holds 11 percent, up 9 points.

One major factor cited in ChatGPT's decline is OpenAI's decision to shut down Sora, its video generation model, as Google's competing Veo model gains traction. The company has also faced criticism over documented biases in ChatGPT's responses. Meanwhile, Google Gemini has benefited from search integration and partnerships with influential creators like MrBeast.

The contractor issue compounds these challenges. If OpenAI's training data has been contaminated by AI-generated responses, the quality degradation could accelerate the shift of users toward competitors. Model collapse is not a theoretical concern; it represents a real erosion of model performance that users will eventually notice.

Beyond ChatGPT, Meta's new Muse AI app has also disrupted the landscape. Muse logged 2.5 million downloads in its first 12 days on the market, surpassing ChatGPT's download rate during the same window after launch. Muse recorded 1.8 million iOS downloads in the US and Canada compared to ChatGPT's 1.3 million during equivalent launch periods. Daily active users for Muse reached 642,000 across iOS and Android in the US, compared to ChatGPT's 231,000 at the same stage.

OpenAI declined to comment on the contractor firings, leaving questions unanswered about the scale of the problem and whether contaminated training data has already affected ChatGPT's performance. As competition intensifies and user growth slows, the company's ability to maintain training data integrity will become increasingly critical to its market position.