Children learn language on a fraction of the data AI needs. Scientists want to know why.

A toddler masters grammar after hearing roughly 10 million words. A large language model may consume 15 trillion. Closing that gap could reshape how AI is built.

AI2Day Newsdesk4 min read
Photoreal news-editorial 16:9 image of a darkened operations center with multiple glowing monitors displaying abstract network topology maps and alert indicator
Share

Key points

  • A child typically starts producing grammatically correct sentences after hearing around 10 million words, roughly 1,500 times less data than Meta's Llama 3.1 used in training.
  • Meta's Llama 3.1, an open large language model released in 2023, trained on 15 trillion tokens, where a token is a word-like chunk of text.
  • Cognitive scientists call the gulf between how little children need and how much AI needs the "data efficiency gap".
  • Researchers warn the supply of freely available internet text for training could run dry as early as the 2030s.
  • Scientists hope reverse-engineering how children learn could produce AI models that do more with far less data.

For at least 100,000 years, exactly one kind of thing could pick up a human language from scratch: a human child. Then, four years ago, ChatGPT arrived. Now there are two kinds of things that can do it.

But the comparison flatters the machines.

A child raised in a language-rich home may hear around 100 million words before their teens. Add books and school and that climbs to perhaps 300 million words by age 20. A large language model, the technology behind chatbots like ChatGPT and Claude, may train on 15 trillion tokens, where a token is roughly one word or word-fragment. That is roughly 50,000 times more text than the average teenager has encountered.

Georgetown University cognitive scientist and linguist Ethan Gotlieb Wilcox puts it plainly: "Claude has seen the amount of language that an entire city will experience in one generation."

Print out all the words used to train a modern large language model on paper and the stack would reach past the International Space Station. Print a child's 100 million words and the stack stands 20 metres tall.

Why does this gap matter right now?

It matters because the internet is finite. For most of the past decade, AI language models have improved mainly by getting bigger: more data, more computing power. But MIT Technology Review reports that researchers believe the supply of easily available training text could run out as early as the 2030s. When it does, simply pouring in more words will no longer work.

Children show there is another path. Toddlers usually begin forming grammatically correct sentences after hearing somewhere between 10 and 30 million words. Train the older GPT-2 model on 30 million words and, as Stanford cognitive scientist Michael C. Frank puts it, "you get a nonsense generator; you don't get a kid."

That puzzle has a name: the data efficiency gap. Solving it is now a shared goal for cognitive scientists studying child development and AI researchers trying to build more capable systems.

What do children do that machines do not?

Nobody fully knows. That is the honest answer, and it is the reason the question still drives research.

One long-running theory, put forward by MIT linguist Noam Chomsky in the 1950s, holds that babies are born with hardwired knowledge of grammar. The rival view, associated with psychologist B.F. Skinner, was that children learn purely through experience and repetition, the way a dog learns to sit for a treat. Decades of argument followed.

What large language models have shown is that pure pattern-learning from massive data can produce impressively fluent language. What they have also shown, by failing so badly at smaller scales, is that something beyond raw statistics must be helping children.

Researchers are now testing specific ideas. Perhaps children use what they already know about the physical world to anchor the meaning of new words. Perhaps the way caregivers speak to babies, simpler sentences, slower delivery, carries structural information that text scraped from the internet does not. Perhaps a child's developing brain applies constraints that keep learning from going off track.

Answering those questions would matter beyond AI. It could settle whether humans are born with a language instinct or learn it entirely from their surroundings, a debate that has run for 70 years.

What does this mean for ordinary people?

In the near term, nothing alarming. But the practical stakes are real. More data-efficient AI could enable chatbots that serve speakers of minority languages, where training data is scarce. It could also make AI far cheaper to build and run.

Watch for: research teams publishing results on models trained on smaller, carefully curated datasets rather than raw web crawls. That is where this field is heading.

© 2026 AI2Day