Four years after ChatGPT's debut, large language models can converse fluently enough to pass as human—but they still need dramatically more data than a child does to master language, and researchers are grappling with why.
What Happened
Cognitive scientists at Stanford and Georgetown universities say the gap between how children and machines acquire language is staggering. A preteen raised in a linguistically rich home may hear roughly 100 million words; with literacy exposure, that figure reaches around 300 million by age 20. By contrast, Meta's Llama 3.1 processed 15 trillion tokens during pretraining alone. Anthropic's Claude has seen language equivalent to what an entire city experiences in one generation—researchers estimate printing out its training data would stack beyond the International Space Station, while a child's lifetime word exposure would reach just 20 meters on paper. Toddlers typically begin producing grammatically correct sentences after hearing approximately 10 million words (or up to 30 million at the higher end). Stanford cognitive scientist Michael C. Frank notes that training GPT-2 on 30 million words yields "a nonsense generator; you don't get a kid."
Why It Matters
The so-called data efficiency gap has practical stakes for AI development. Frontier model makers could be pretraining on 10 times more data than Llama 3.1, according to Georgetown cognitive scientist Ethan Gotlieb Wilcox, but the well of easily available internet text may run dry by the 2030s. If researchers can reverse-engineer how children learn language with far less exposure, it could enable training AI effectively on limited or low-resource domains—video data, minority languages, specialized fields—and reduce compute costs significantly. The question also cuts to foundational debates in cognitive science: whether humans possess an innate language instinct and what universal constraints shape how any learner—whether child or machine—can acquire syntax.
The Bottom Line
Despite remarkable recent progress in large language models, children still outperform them in data efficiency by orders of magnitude. How babies infer grammar from limited exposure remains unsolved—a gap that matters both for the future scalability of AI and for understanding human cognition itself.