AI companies including Anthropic and OpenAI regularly publish reports on how people use products like Claude and ChatGPT, but researchers say these reports present only a partial view of actual usage patterns.
What Happened
Anthropic Economic Index is one of the most widely cited sources for data on AI usage. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations unrelated to those purposes. Researchers at Stanford's Trustworthy AI Research Lab and MIT Media Lab have launched a project called the AI Observatory that aims to provide independent verification of these claims. The platform aggregated 24,521 real conversations across 85,633 conversational turns with models including Claude and Gemini, collected with user consent through seven existing datasets. When the team applied Anthropic's own methodology to their dataset, they found that nearly half of all conversations—48%—would have been filtered out under those criteria. Those excluded conversations were more likely to include health and relationships (44.2% versus 31.2%), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). The datasets covered conversations from 2023 to 2025, during which the researchers observed that interactions grew longer and more elaborate over time, with increasing small talk suggesting rising use of AI for companionship.
Why It Matters
Policymakers are making consequential decisions about AI regulation based on limited data, according to Anka Reuel, co-lead of the AI Observatory and a PhD candidate at Stanford. The research found significant variation across models: Grok and Gemini were used more frequently for information retrieval, with Grok showing concentration of misinformation around news and politics; Anthropic saw heavier use for coding tasks; Gemini attracted more social and roleplay uses; ChatGPT was favored for homework assistance. Different versions of the same model also showed distinct patterns—users had shorter conversations with GPT-3.5 but longer, more iterative ones with GPT-4o. "No single company report tells the whole story," said Shayne Longpre, a recent MIT Media Lab graduate who co-led the research.
The Bottom Line
The AI Observatory team says stakeholders lack independent sources to corroborate vendor-reported data on AI usage. Anthropic has released separate blog posts covering topics such as using Claude for companionship and for generating CSAM, but researchers argue that consolidated cross-platform analysis helps paint a more complete picture of how people actually interact with generative AI systems.