A new independent analysis of nearly 25,000 real AI conversations shows that usage patterns vary sharply across models and that official company reports omit nearly half of all interactions, including many involving sensitive topics like health, relationships and harassment. The findings, released by the AI Observatory, challenge the completeness of data published by major AI firms.

What You Need to Know

AI companies like Anthropic and OpenAI control the data they release about how people use products like Claude and ChatGPT, leaving researchers without independent verification. The AI Observatory aggregated seven public datasets to create a more complete picture. Its research shows that non-work conversations are frequent and that different models attract very different user behaviors.

The Limits of Corporate Data

Anthropic’s Economic Index is widely cited but focuses exclusively on work-related uses of Claude. When the AI Observatory applied the same filtering methods to its own dataset, 48% of conversations were excluded. Those filtered conversations were more likely to involve health and relationships, adult or illicit topics, harassment and sexual content. OpenAI’s 2025 ChatGPT report similarly found that only 30% of consumer use was work related.

“There is no independent source to corroborate it,” said Anka Reuel, a PhD candidate at Stanford and co-lead of the AI Observatory. Reuel noted that policymakers are making high-stakes decisions based on incomplete, self-serving data.

The AI Observatory’s platform aims to fill that gap by providing an unfiltered view of real interactions collected with user consent. Its analysis spans conversations from 2023 to 2025 and reveals changes in behavior over time. There was a notable increase in small talk and longer exchanges, suggesting a rise in AI companionship, while self-disclosure by chatbots decreased.

Model by Model Differences

The study found that each major AI model attracts distinct use cases. Researchers observed clear specialization across platforms.

  • Grok: Used heavily for news and politics but also a hub for misinformation.
  • Claude: Popular for coding tasks and technical work.
  • Gemini: Frequently used for social and roleplay conversations.
  • ChatGPT: The dominant choice for homework help and general assistance.

Even different versions of the same model showed variation. Sensitive exchanges involving hate speech or sexual content became less frequent over time, which may reflect improved safeguards. However, the concentration of misinformation on Grok aligns with other research and raises concerns about platform accountability.

Why This Matters

The findings have direct consequences for regulation and safety. Without independent data, regulators cannot assess whether AI systems are being used in harmful ways or whether companies are underreporting risks. The AI Observatory’s approach gives researchers and policymakers a tool to verify claims made by firms like Anthropic and OpenAI. As AI use grows, relying solely on corporate reports creates dangerous blind spots. Independent monitoring is not optional; it is essential for informed governance.