Skip to content
AI.info

The Pulse

OpenAI Contractors Removed After Using AI to Grade ChatGPT

Multiple contractors working on OpenAI model-evaluation projects were fired or removed after using AI tools to complete assignments, according to 404 Media. Mercor, one staffing company involved, says its contracts prohibit large language m

OpenAI Contractors Removed After Using AI to Grade ChatGPT

AI.info Team ·

OpenAI contractors hired to judge and improve ChatGPT have been fired or removed from projects after using AI to complete the work they were supposed to perform themselves, according to an investigation published by 404 Media on September 22, 2026.

The finding exposes a direct conflict in the human-feedback system behind modern language models. OpenAI and its contractors rely on people to read prompts, compare model responses and explain which answers are better. When another language model supplies that judgment, the process no longer produces the human assessment the contractor was hired to provide.

OpenAI’s reviewers were hired for human judgment

404 Media reports that OpenAI uses thousands of contractors across projects intended to improve its models. One internal document obtained by the publication describes projects that can involve more than 10,000 contractors. Their assignments include reviewing AI-generated responses, rating outputs and writing comments about quality, tone and accuracy.

The work can involve real ChatGPT prompts and other user data. In a separate investigation published the previous week, 404 Media described Project Lily, an OpenAI effort in which hundreds of contractors reviewed real user conversations, including prompts that could contain personal information.

Internal instructions for contractors who review other workers’ submissions explicitly prohibit them from using AI. The guidance says reviewers must not use GPTZero or other AI-detection tools because they are unreliable, and also bars tools such as Grammarly and AI translation for reviewing work, writing feedback or producing comments.

“Do not use AI detection tools, or AI yourself,” the instructions say. “Do not use GPTZero or any other AI detection tool. They are not reliable.”

Instructions obtained and quoted by 404 Media

Contractors were judged by patterns, not a detector

The instructions tell reviewers to assess the overall pattern of a worker’s submissions rather than rely on one suspicious phrase. Signs of possible AI use can include repetitive wording, punctuation associated with AI-generated prose and unusually fast completion times. Contractors were also told not to disclose the specific clues that led them to suspect AI use.

Three contractors interviewed by 404 Media said reviewers were warned not to use AI in their own work. Two sources said workers had been fired or offboarded after using AI. The contractors spoke anonymously because they were not authorized to speak with the press.

One contractor shared what they presented as a termination letter citing concerns about the authenticity of their work. The person told 404 Media: “I’m not a bad person or worker. I just needed a little boost and turned to AI to help me which eventually led to my downfall.”

Mercor says it removes workers who violate the rule

Two of the contractors interviewed by 404 Media worked for Mercor, an AI-training company that recruits workers for assignments connected to OpenAI. Mercor, rather than OpenAI, employed those contractors directly.

A Mercor spokesperson said the company’s contracts prohibit large language models from completing project work and that it removes workers when it confirms a violation.

“Our experts are hired for their expertise and judgement, which is essential to the ongoing advancement of AI. Our contracts strictly prohibit the use of LLMs to complete projects and we enforce that.”

Mercor spokesperson, quoted by 404 Media

Mercor also said it invests in tools and systems to detect misuse and “immediately remove[s]” an expert from a project when the company confirms that AI was used to complete a task. The company’s public language-model policy allows limited grammar, wording and tone edits, but prohibits using a model to judge another model’s responses or write the reasoning behind an evaluation.

OpenAI declined to comment

OpenAI declined to comment to 404 Media on contractors being fired for using AI. The company’s silence leaves unanswered how many workers have been removed, which OpenAI projects were affected and whether any AI-generated evaluations entered training or quality-control systems before detection.

The reporting does not establish that the incidents damaged a released OpenAI model. It does show how dependent model development remains on a fragile chain of human review: the companies building systems to automate judgment still need workers whose own judgment they can trust.

Source

404 Media

Explore

More articles