A new analysis reveals a stark disconnect in how websites interact with artificial intelligence systems. Only 8.9% of websites actively block AI crawlers from collecting their data, yet 94.8% of those same sites are never cited as sources in AI-generated answers.

What You Need to Know

The study examined a broad sample of websites and found that while most allow AI systems to access their content, almost none receive attribution in AI responses. This imbalance reflects a growing divide between site accessibility and the credit publishers expect. It also raises questions about the fairness of current AI training and answer generation practices.

The Blocking Disparity

Website owners have several technical options to block AI crawlers, including robots.txt rules and server-side restrictions. The analysis, however, shows that the vast majority choose not to use them. Only 8.9% of sites employ any form of blocking, leaving the remaining 91.1% open to automated data collection.

  • Limited blocking adoption: Most site owners either lack awareness of AI crawlers or choose not to restrict them.
  • Global variation: Blocking rates vary by region and industry, but no category exceeds 15%.
  • Technical barriers: Some website operators find current blocking methods ineffective or difficult to implement.

This low blocking rate stands in contrast to the high level of non-citation in AI answers. The data suggests that site accessibility does not translate into visibility or credit in AI-generated content.

Why Most Sites Stay Uncited

Several factors explain why 94.8% of websites never appear in AI answers. AI models tend to rely heavily on a small set of authoritative sources, often major news outlets, encyclopedic sites and government domains. Smaller or niche websites rarely surface in training data prioritization or retrieval processes.

AI answer systems also use citation mechanisms that differ from traditional web search. They may not reference sources at all, or they may attribute information to aggregate knowledge rather than individual sites. This approach leaves most publishers without recognition even when their content contributes to model outputs.

Why This Matters

This imbalance carries significant consequences for both website owners and the AI industry. If content creators see no benefit from allowing access, more sites may begin blocking AI crawlers. That reduction in training data could degrade model quality and limit the breadth of AI knowledge.

Publishers, on the other hand, face an economic challenge. Many rely on traffic and attribution to generate revenue from their content. When AI systems use that content without citation, those business models weaken. Regulators and industry groups may need to address this attribution gap through new standards or policies.

The 8.9% blocking rate could rise sharply if website owners decide that access without citation is no longer worthwhile. The next few years will determine whether the current open-access environment persists or gives way to a more restricted web.