Contractors hired to help improve OpenAI's systems have been removed from projects after using artificial-intelligence tools to complete work that required human judgment, according to reporting based on internal documents and interviews with workers. The account highlights a central quality-control problem for the data-labelling industry: companies buying human evaluation must be able to distinguish it from machine-generated output.
404 Media reported that it spoke with three contractors working across OpenAI-related projects and obtained instructions that prohibited reviewers from using AI. The rules reportedly extended beyond chatbots to tools such as AI-assisted translation and writing software. Reviewers were also told not to rely on automated AI-detection products, which the instructions described as unreliable, and to assess patterns in a worker's output rather than disclose individual warning signs.
Two of the sources said contractors had been fired or otherwise offboarded for using AI. One worker shared what was presented as a termination notice referring to problems with the authenticity of submitted work. OpenAI declined to comment to the publication about the reported removals. The account therefore establishes the contractors' and intermediary's descriptions, not an independent measurement of how frequently violations occurred across all projects.
Mercor, an AI-training company that employed two of the contractors interviewed, confirmed the underlying policy in a statement to 404 Media. The company said its experts are hired for their judgment, contracts prohibit using large language models to complete assignments, and confirmed misuse results in immediate removal from the project. Mercor also said it invests in systems intended to detect violations.
The restriction reflects the purpose of many evaluation tasks. Contractors may compare model responses, identify undesirable behavior, or provide critiques intended to make later systems more useful and reliable. If a second model supplies those judgments, the customer no longer receives the independent human signal it commissioned. It also introduces the risk that recurring model habits or errors are fed back into the evaluation process.
The reported enforcement illustrates an unusual boundary in the AI labor market. Technology companies promote assistants as workplace tools, yet some of the human work supporting those assistants depends precisely on workers not delegating their judgment to another model. Speed alone can be a warning sign, but the reported guidance recognized that no single stylistic clue reliably proves AI use.
The evidence does not quantify the number of people removed, identify every contracting firm involved, or show that prohibited output entered a released model. It does show that intermediaries view undisclosed AI assistance as a breach serious enough to end an assignment. As model developers rely on large distributed workforces, maintaining a trustworthy chain between instructions, human reviewers and training data remains an operational challenge rather than a solved detection problem.



