OpenAI's ChatGPT Trainers Are Getting Fired for Using AI,Here's the Irony
OpenAI has fired multiple contractors caught using artificial intelligence to do the very work they were hired to perform as human judges for ChatGPT. The contractors, employed to review and improve AI-generated responses, violated explicit rules prohibiting the use of large language models (LLMs) to complete their assignments, according to a report from 404 Media.
Why Would an AI Company Ban AI Use Among Its Own Trainers?
The rule might seem counterintuitive at first, but it reflects a fundamental reality of modern AI development: human judgment remains irreplaceable. These contractors are hired specifically for their ability to evaluate AI responses and decide which answers are better, a task that loses its value if another AI is doing the evaluating instead. OpenAI explicitly states that information provided by human trainers and researchers is one of the core sources used to develop its foundation models.
The quality of human feedback directly influences how AI models learn and behave. OpenAI has acknowledged in its own research that models trained using human feedback can be shaped by the people labeling that data, making the authenticity of that feedback critical to the model's development.
How Did OpenAI Catch Contractors Using AI?
Interestingly, OpenAI did not rely on AI detection software to catch the violators. According to internal documents obtained by 404 Media, reviewers are explicitly instructed not to use tools like GPTZero because they are considered unreliable. Instead, the company trains human reviewers to spot suspicious patterns in contractor submissions.
The detection method relies on human observation rather than automated tools. Reviewers look for telltale signs of AI-generated work, including repetitive wording, unusually fast completion times, and writing patterns commonly associated with AI outputs. Even excessive use of em dashes can raise suspicion.
Steps to Identify Suspected AI Use in Human Feedback Work
- Writing Pattern Analysis: Reviewers examine submissions for repetitive phrasing and linguistic patterns that match known AI outputs rather than natural human variation.
- Completion Speed Monitoring: Unusually fast submission times relative to the complexity of the task can indicate AI assistance rather than genuine human deliberation.
- Stylistic Inconsistencies: Sudden changes in writing style, vocabulary, or tone across a worker's submissions may suggest different sources of work.
- Punctuation Anomalies: Excessive or unusual use of specific punctuation marks, such as em dashes, can be a red flag for AI-generated content.
Notably, reviewers are instructed not to reveal exactly what triggered their suspicion of AI use. This deliberate opacity prevents workers from learning how to better hide AI assistance in future submissions.
What Do the Contractors Say About the Enforcement?
One contractor told 404 Media that they see people using AI "all the time" and that workers are regularly removed for doing so. Another contractor shared what they described as a termination letter citing problems with the "authenticity" of their work, highlighting how seriously OpenAI takes violations of this policy.
The contractors work across various OpenAI projects, some of which involve thousands of workers. Their responsibilities include reading AI responses, rating them, and deciding which answers are better or more helpful. Mercor, an AI-training company that employs contractors working on OpenAI projects, confirmed to 404 Media that its contracts explicitly prohibit workers from using large language models to complete their assignments.
The internal instructions reportedly prohibit workers from using AI to write feedback or comments and specifically name tools including Grammarly and AI translation services. This broad restriction reflects OpenAI's commitment to ensuring that human judgment remains genuinely human.
Does This Affect ChatGPT's Quality?
There is no evidence in the 404 Media report that the AI-generated contractor work caused measurable damage to any OpenAI model. However, the situation does reveal a deeper truth about AI development: even companies building cutting-edge AI systems still fundamentally depend on humans to tell those systems when they have gotten something right.
OpenAI does use synthetic data, which is information generated with the help of AI models, as part of some training processes. The company argues that synthetic data can improve model performance and fill gaps where other training data is scarce. But there is a critical distinction between intentionally using AI-generated data as a training resource and secretly substituting an AI model for a person hired specifically to provide human feedback.
The irony is sharp: in an industry racing to build more capable AI systems, the bottleneck remains fundamentally human. No matter how advanced ChatGPT or other large language models become, they still need people to evaluate them, judge their outputs, and guide their development. And those people, it turns out, cannot be replaced by the very technology they are training.