People have been talking to each other for at least a hundred thousand years. In all that time, there was only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two.
In the short years since ChatGPT arrived, most of us take it for granted that we can converse naturally with our phones. But peek behind the computational curtain and there is a catch: teaching a computer to use human language still requires an inhuman amount of data — many thousands of times more words than a person encounters while mastering a mother tongue.
Michael C. Frank, a cognitive scientist at Stanford, puts it plainly: the recent progress has been amazing, "but we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year."
This divide between children and machines is called the data efficiency gap. It poses a question for cognitive scientists and a challenge for model architects: how do kids still outperform the most linguistically sophisticated machines ever built?
The scale of the divide
For the past decade, language models have mostly gotten better by getting bigger. Meta's open-weight Llama 3.1 chewed through 15 trillion tokens in pretraining. Frontier models could be pretraining on ten times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown. But there is only so much internet to train on, and the well of easily available data could run dry as early as the 2030s.
Kids show it could be possible to learn more with far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy and you can boost that to maybe 300 million words by age 20.
Wilcox reaches for analogy: "Claude has seen the amount of language that an entire city will experience in one generation." If you printed out all the words used to train a modern language model, the stack would reach past the International Space Station. The preteen's 100 million words would stack up just 20 meters. And we can make do with far less than that.
The comparison lines up like this:
- A toddler — starts producing grammatical sentences after roughly 10 million words.
- A preteen — around 100 million words.
- A reader at twenty — perhaps 300 million words.
- Llama 3.1 — 15 trillion tokens in pretraining.
Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words. As Frank says: "It's just totally miraculous. If you train GPT-2 on 30 million words, you get a nonsense generator; you don't get a kid."
The old argument: innate or learned
Exactly how babies pull this off is a mystery. Syntax includes recursive, nested structures that let us express virtually infinite ideas with a finite lexicon. Babies only splash about in the shallows of a fathomless ocean of language — and yet from a drop, they infer the depths.
In the 1950s the MIT linguist Noam Chomsky proposed a solution: babies are born with hardwired grammar. He was reacting to B. F. Skinner, who held that language is learned through conditioning, the way a dog figures out how to sit for treats.
Chomsky countered by citing the "poverty of the stimulus": language, especially syntax, is too complex and children's exposure to it too impoverished for them to learn entirely from experience. As Richard Futrell, a linguist at the University of California, Irvine, summarizes it: "His signature argument was, essentially, that language cannot be learned on the basis purely of statistics."
That view dominated US linguistics for decades as generative grammar. Researchers tried to teach language by explicitly coding the rules — less immersion, more grammar class. This approach, called symbolic AI, largely failed to handle human language at scale.
The assumption models overturned
In 2018 and 2019, BERT and GPT-2 — built on the transformer and trained on billions of tokens — made it clear that learning from a massive glut of data could work for language. Language models are not brains; they are naive pattern-learning machines without the evolved quirks of the human cortex, exactly the kind of thing a generative linguist two decades ago would have said could not learn language.
Alison Gopnik, a developmental psychologist at the University of California, Berkeley, concedes the point directly: "No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax. I didn't think that was going to turn out to be true."
BabyLM: training at child scale
But what about learning from a small sample of language — a child-size one?
Alex Warstadt, a linguist and data scientist, turned that question into a competition. With Leshem Choshen and others he founded BabyLM, an annual contest to train models on small data sets.
The main track asks for training on a "developmentally plausible" corpus of just 100 million words; the toddler track allows 10 million, drawn from storybooks, dialogue, subtitles, Wikipedia and transcripts of speech directed at children.
Models are evaluated on the grammar benchmarks psycholinguists use with humans, which look for signs of confusion at ungrammatical features. A test might compare "The keys to the cabinet are on the table" with "The keys to the cabinet is on the table." In a human that surprise shows in eye movements; in a model, researchers use a measure called surprisal.
The unexpected finding: curriculum did not help
The competition has already challenged assumptions. One was curriculum learning: starting with simple data and working up to complex inputs, like beginning with baby talk. It was by far the most popular approach in the first round, and it did not work as well as expected.
Aaron Mueller, a computer scientist at Boston University, explains: "The appeal is just kind of hard to resist. It seems to really line up with ways that we believe humans are learning. But it seems like these transformers don't really need to have their data ordered in such a way to learn effectively."
Perhaps ironically, the best BabyLM models are not inspired by babies at all. The 2024 champion, GPT-BERT, is a transformer trained partly to predict the next token and partly to fill in blanks. Pretrained on about 100 million words, it beat Llama 2 70B — pretrained on thousands of times that amount — on one BabyLM benchmark.
Still, BabyLM models are not on the same level as commercial language models. Many cannot produce text at all.
What is missing: bodies, senses, curiosity
Kids are not disembodied programs whose only experience arrives as written text. Some researchers think machines will need to learn through the eyes and ears of children.
Frank and colleagues used headcams to see how babies experience the world: "Kids' experience looks really radically different than we thought. It's much more focused: They've got these little short arms, so the objects are, like, right in front of them. And they live in a forest of knees."
A project called SAYCam recorded two hours a week of three babies' lives. Brenden Lake of Princeton showed that a model trained on 61 hours of that footage learned to identify objects and tie them to words, without the built-in biases many theories consider necessary. But, he adds, "we don't get a two-year-old out of training when we're done."
Children choose their data; models receive it
For Gopnik the missing ingredient lies elsewhere: "Children are actively exploring, which means that they're actively choosing their own data. Kids are constantly experimenting."
Research by her group shows that what looks like child's play is an effective way to learn cause and effect. Kids seek out experiences that maximize their ability to make a predictable impact on the world.
Elizabeth Bonawitz, a developmental cognitive scientist at Harvard, adds another difference: unlike models, children know what they do not know and are driven to fill those gaps. Her research shows they interpret information differently when they know an adult is teaching them: "Children are not only reasoning about the evidence they're being told. They're reasoning about the teacher, about the teacher's knowledge, and about why the teacher is telling them this particular information."
That is very different from how models learn: passively and in isolation. Last year's BabyLM admitted models that learn by interacting with other models. They did not outperform standard ones.
Why it matters
Closing the gap would pay off in three distinct ways.
Minority languages. David Samuel, a machine-learning researcher at the University of Oslo and one of GPT-BERT's architects, has a personal reason: he is Czech and works in Norway, and both languages have far less training data than English. Minority languages like Sami might have only tens of millions of tokens — about a toddler's scale of exposure. His question: "How can we develop language models that are just as capable as the English ones for small languages?"
The question applies to Turkish too. Turkish is not a minority language, but the volume of Turkish text online is small beside English. Closing the gap would mean competent Turkish models without English-scale data.
Democratization. Warstadt wants the gap closed so universities and others without hyperscaling resources can train good models and stay relevant in AI research.
Understanding ourselves. Bonawitz was initially skeptical that models could reveal anything about cognition: brains are embodied, our neurons are living cells rather than tidy code, and brains keep changing while models pretrain once. Even so, she says, "I'm sort of revising my beliefs."
Researchers now use models as a kind of linguistic lab rat, building hypotheses into them — simulating degrees of bilingualism, withholding certain grammatical forms — and testing them in ways impossible with real children.
As Warstadt puts it: "Humans have been the only entities in the universe that use language. Now there's this other linguistic entity. Finally we have a model; not in the sense of a language model, but in the sense of a model organism."
Where things stand
Frontier labs are not racing to borrow tricks from children. Gopnik thinks it will be the next generation of AI — whatever replaces the transformer — that takes lessons from developmental psychology. Interest is growing from another direction, though: the NanoGPT Slowrun benchmark, launched in March 2026, shares BabyLM's goals but drops the human-learning motivation and targets data efficiency directly.
So the gap is not closed and the reason is not known. But it is now being measured — and a thing that gets measured tends, sooner or later, to become a thing that gets worked on.