Independent study reveals gaps in corporate AI usage narratives
Research co-led by Stanford’s Anka Reuel finds that corporate disclosures often filter out personal, sensitive, and illicit interactions to emphasise productivity metrics, leaving policymakers with incomplete data.

The AI Observatory, an independent research initiative co-led by Stanford’s Anka Reuel, has released findings that challenge the reliability of AI usage data reported by major technology firms such as Anthropic and OpenAI. By analysing 24,521 conversations across 52 models from seven datasets between 2023 and 2025, the study indicates that corporate disclosures frequently omit personal, sensitive, and illicit interactions to emphasise productivity metrics.
The research highlights distinct user behaviours across platforms, noting that Grok is often utilised for news and political discourse but also disseminates misinformation, whereas Anthropic is predominantly preferred for coding tasks. Researchers caution that the absence of independent verification may lead policymakers and stakeholders to make significant decisions based on incomplete corporate narratives.
Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research Lab, described the AI Observatory as a public platform designed to aggregate and analyse real AI conversations collected with user consent. The project aims to provide independent sources of information for researchers and policymakers to assess how people are using generative AI, filling a gap left by proprietary company reports.
When the AI Observatory team applied Anthropic’s own filtering methods to their dataset, they found that nearly half of the conversations would have been excluded from corporate analyses. These filtered-out conversations were more likely to include health and relationship queries, adult or illicit topics, harassment, hate speech, and sexual content compared to the work-focused data typically published by companies.
The study also revealed significant variations in how users interact with different models. While Grok was popular for information retrieval on news and politics, it was also where misinformation tended to concentrate. Anthropic was the preferred tool for coding, while Gemini was used more for social and roleplay purposes, and ChatGPT for homework assistance.
Conversations within datasets such as WildChat grew longer and more elaborate over time, with increased small talk suggesting a rise in AI companionship. Conversely, exchanges labelled as sensitive use, including sexual harassment and hate speech, dropped, which researchers suggested might indicate more effective platform safeguards.
Despite the breadth of its findings, the AI Observatory noted that its dataset represents a fraction of the data held by major labs and may underrepresent sensitive uses due to the voluntary nature of data submission. The team hopes to expand its datasets over time and encourages AI companies to share data with independent researchers to ensure a more comprehensive understanding of AI adoption.

