What are people actually doing with AI chatbots? A new study looked at nearly 25,000 real conversations to find out
AI companies publish usage reports, but researchers say those reports leave out the messy parts. A new independent project called the AI Observatory dug into real chat data and found a very different picture.

Key points
- The AI Observatory, a research project from Stanford and MIT, analysed 24,521 real AI conversations collected with users' consent across seven existing datasets.
- When researchers applied Anthropic's own filtering method to their data, 48% of conversations were cut out, and the removed chats contained far more sensitive content.
- Grok, the AI chatbot made by Elon Musk's company xAI, showed the highest concentration of misinformation in the dataset.
- Conversations with ChatGPT grew longer and more personal over time, and AI systems became less likely to remind users they were talking to a machine.
- The AI Observatory is making its dataset available to outside researchers, something the big AI companies do not currently do.
Every few months, a company like Anthropic or OpenAI publishes a report explaining how people use its AI chatbot. Those reports get cited heavily by journalists, politicians, and academics trying to understand what AI is actually doing in everyday life.
There is a problem with that arrangement. The companies choose what to publish.
"There is no independent source to corroborate it," says Anka Reuel, a computer science PhD candidate at Stanford's Trustworthy AI Research lab. Reuel co-led the AI Observatory, a new public platform that pulled together 24,521 conversations, across 85,633 back-and-forth exchanges, from seven separate datasets gathered by previous researchers, all with users' consent. Those chats came from roughly 5,000 people talking to 52 different AI models, including ChatGPT, Gemini, Claude and Grok, between 2023 and 2025.
What does the company-published data miss?
A lot, it turns out, especially anything personal or sensitive. Anthropic's widely cited Economic Index focuses almost entirely on work and productivity uses of its Claude chatbot, filtering out conversations that fall outside that scope. When the AI Observatory team applied Anthropic's same filtering rules to their own dataset, 48% of conversations disappeared.
What got cut was revealing. The removed conversations were far more likely to involve health questions or relationship advice (44.2% of filtered chats, versus 31.2% in Anthropic's published analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate speech (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%).
For context, Anthropic based its latest Economic Index on one million Claude conversations, and OpenAI's comparable report analysed 1.5 million ChatGPT chats. The AI Observatory's 24,521 conversations are a small fraction of that. The researchers acknowledge the gap and caution that their findings are not a complete picture of all AI use, partly because people who volunteer chat data may be less likely to share sensitive exchanges.
Which chatbot did what?
The study found clear differences in how people used specific tools.
| Chatbot | Most common use found | Notable concern |
|---|---|---|
| Claude (Anthropic) | Coding and technical help | Underrepresented in company reports |
| ChatGPT (OpenAI) | Homework assistance | Longer, more emotional use with GPT-4o |
| Gemini (Google) | Social interaction and roleplay | N/A |
| Grok (xAI) | News and politics queries | Highest misinformation concentration |
Grok's misinformation finding is consistent with other independent research. xAI did not respond to a request for comment.
Even different versions of the same product behaved differently. People had shorter chats with the older GPT-3.5 model and longer, more iterative ones with GPT-4o, the newer version. That tracks with GPT-4o's reputation for building emotional attachment with users.
Over time, conversations in the dataset grew longer and included more small talk, suggesting people were increasingly treating chatbots as company rather than search engines. Meanwhile, the chatbots themselves became less likely to remind users they were talking to a machine.
What does this mean for ordinary people?
Right now, governments and companies are making big decisions about AI rules and products using data that comes almost entirely from the AI companies themselves. Independent researchers like Reuel and her co-lead Shayne Longpre, a recent PhD graduate from MIT's Media Lab, say that is a serious blind spot.
"No single company report tells the whole story," Longpre says.
Anthropric told MIT Technology Review that its published research reflects its teams' specific interests and that it supports independent external research. OpenAI did not respond.
Reuel's ideal outcome is straightforward: AI companies share anonymised chat data with independent researchers in a way that protects user privacy. For now, the AI Observatory's dataset is open to researchers, and the team plans to expand it. Without something like that, she warns, anyone making policy or product decisions based on AI usage data risks "completely operating in the wild" without knowing what is really happening.



