A leading AI evaluation company has issued a directive that may seem paradoxical: contractors evaluating AI models should not use AI to do their work. The policy, which states that AI contractors shouldn't use AI to evaluate AI models, is designed to preserve human judgment in a process often assumed to be fully automated.

What You Need to Know

The company, which builds and audits large language models, has noticed that some contractors were outsourcing their evaluation tasks to AI tools. This practice, the firm warns, can introduce errors and reduce the diversity of feedback needed to improve model performance. The ban applies to all external evaluators working on model quality assessments.

The Core Policy

Says AI company that specializes in model benchmarking, the move is about protecting evaluation integrity. Contractors are required to manually review model outputs and provide graded feedback without assistance from generative AI tools. The company argues that using AI to evaluate AI creates a feedback loop that can mask real-world flaws.

Under the new rules, evaluators must rely on their own expertise and domain knowledge. The policy applies specifically to tasks where models are tested for accuracy, safety and bias. Violations could result in termination of contracts.

Industry Context

The directive highlights a growing tension in the AI industry. Companies spend heavily on human evaluators to fine-tune models, yet many of those same evaluators have begun using AI assistants to speed up their work. This creates a risk that models are graded by other AI systems rather than by human reasoning.

Several major AI labs have started restricting how external contractors use AI tools. The policy from this evaluation firm is among the strictest so far. It reflects a broader recognition that human oversight remains essential for catching subtle errors and cultural nuances that machines may miss.

Why This Matters

The ban has direct consequences for the quality of AI products. If contractors rely on AI to evaluate models, the resulting feedback may reinforce existing biases rather than correcting them. Over time, this could degrade model performance and erode user trust.

For AI developers, the policy means that evaluation will remain a slower and more expensive process. But the trade-off is potentially higher reliability. For the industry as a whole, the move signals that even as AI capabilities advance, human judgment remains the gold standard for measuring those capabilities.

The Human Element

The company’s approach underscores a fundamental paradox: AI systems are often better than humans at certain analytical tasks, but they are not yet good at catching their own blind spots. By insisting that contractors evaluate AI models without AI, the firm is betting that human intuition and contextual awareness can still outperform automation in this specific domain.

  • Evaluation quality: Human graders can identify subtle logic errors and cultural missteps that AI tools overlook.
  • Feedback diversity: A team of human evaluators brings varied perspectives, reducing the risk of homogeneous scoring.
  • Accountability: Requiring manual evaluation makes it easier to trace and correct individual contractor mistakes.

The directive is unlikely to be the last of its kind. As AI models become more embedded in daily life, the demand for rigorous, human-led evaluation will only increase. Says AI company that enforces this policy, the ultimate goal is not to slow progress but to ensure that progress is built on a trustworthy foundation.