Children continue to master human language using a fraction of the data required by the most advanced AI systems, a disparity that researchers call the data efficiency gap. Despite rapid progress with models like ChatGPT, Claude and OpenAI's GPT, a toddler's ability to learn from everyday conversation remains unmatched by machines that train on trillions of words. Now cognitive scientists are studying this gap to understand what it reveals about both human development and the future of artificial intelligence.
The Data Efficiency Gap
A typical LLM like Meta's Llama 3.1 was pretrained on 15 trillion tokens, while a preteen raised in a linguistically rich home may hear around 100 million words by age 20. The difference in scale is immense. Michael C. Frank, a cognitive scientist at Stanford University, describes the achievement of child language learning as something LLMs still cannot replicate without enormous resources. "The progress recently has been amazing," Frank says of LLMs. "But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year."
Why This Matters
The data efficiency gap has direct consequences for AI development. As the supply of easily available internet data approaches exhaustion, likely by the 2030s, scaling models larger becomes unsustainable. Cracking how children learn could unlock more data-efficient approaches that reduce energy costs and allow training on new modalities like video. For cognitive science, building computational models that mimic child learning could settle enduring debates about whether language is innate or learned entirely from experience. The outcome may affect how future chatbots serve minority language communities where training data is scarce.
What Researchers Hope to Learn
By reverse-engineering the way kids acquire language, scientists aim to create models that learn with far less data. Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University, notes that the gap is so wide it can only be described in analogy: "Claude has seen the amount of language that an entire city will experience in one generation." If such model efficiency can be improved, it could enable AI to handle video understanding and support languages with limited digital text. The research also touches on theories from Noam Chomsky about innate grammar, though newer findings suggest experience may play a larger role than previously thought.
The fact remains that children still outlearn AI, and researchers have only begun to understand why. Bridging this gap could redefine how machines process human language and what it means to truly communicate.



