A leading AI evaluation company has issued a directive that may seem paradoxical: contractors evaluating AI models should not use AI to do their work. The policy, which states that AI contractors shouldn't use AI to evaluate AI models, is designed to preserve human judgment in a process often assumed to be fully automated.
The Core Policy
Says AI company that specializes in model benchmarking, the move is about protecting evaluation integrity. Contractors are required to manually review model outputs and provide graded feedback without assistance from generative AI tools. The company argues that using AI to evaluate AI creates a feedback loop that can mask real-world flaws.
Under the new rules, evaluators must rely on their own expertise and domain knowledge. The policy applies specifically to tasks where models are tested for accuracy, safety and bias. Violations could result in termination of contracts.
Industry Context
The directive highlights a growing tension in the AI industry. Companies spend heavily on human evaluators to fine-tune models, yet many of those same evaluators have begun using AI assistants to speed up their work. This creates a risk that models are graded by other AI systems rather than by human reasoning.
Several major AI labs have started restricting how external contractors use AI tools. The policy from this evaluation firm is among the strictest so far. It reflects a broader recognition that human oversight remains essential for catching subtle errors and cultural nuances that machines may miss.
Why This Matters
The ban has direct consequences for the quality of AI products. If contractors rely on AI to evaluate models, the resulting feedback may reinforce existing biases rather than correcting them. Over time, this could degrade model performance and erode user trust.
For AI developers, the policy means that evaluation will remain a slower and more expensive process. But the trade-off is potentially higher reliability. For the industry as a whole, the move signals that even as AI capabilities advance, human judgment remains the gold standard for measuring those capabilities.
The Human Element
The company’s approach underscores a fundamental paradox: AI systems are often better than humans at certain analytical tasks, but they are not yet good at catching their own blind spots. By insisting that contractors evaluate AI models without AI, the firm is betting that human intuition and contextual awareness can still outperform automation in this specific domain.
The directive is unlikely to be the last of its kind. As AI models become more embedded in daily life, the demand for rigorous, human-led evaluation will only increase. Says AI company that enforces this policy, the ultimate goal is not to slow progress but to ensure that progress is built on a trustworthy foundation.



