A Question of Foundations
I spent a career arguing that human judgment is not the gold standard. We are, I believe, demonstrably worse at statistics than we think and demonstrably worse at predicting our own futures than we care to admit. But something stranger is happening now. A new wave of cognitive-science research is revealing that the gap between human and machine thinking is not narrowing along the axis most people assume. It is not that machines are getting slow — they are unfathomably fast. It is that they are getting nowhere near the kind of understanding that a six-month-old already possesses.
Three recent studies illustrate this from different angles. A baby can read intentions that a neural network cannot. A multimodal model cannot solve puzzles a child finds in a newspaper. And a human translator working alongside AI is the one who ends up with a sharper mind. Each finding, individually, is modest. Together, they sketch something important about where the field is wrong.
Infants Outperform Machines at Theory of Mind
The first study comes out of NYU, published in the journal Cognition in early 2023. Moira Dillon and Brenden Lake — joined by Harvard's Tomer Ullman — designed what they called the Baby Intuitions Benchmark. The idea was deceptively simple. Take eleven experiments that developmental psychologists have long used to measure what infants understand about other people's goals and preferences. Then feed the same tasks to machine-learning models trained on large corpora of human activity video. Compare scores.
The result is not close. Human infants achieved 69.7% accuracy on these tasks. They beat three of four AI models, each of which had been trained on more than 12 million video clips of human behavior. The only model that outperformed the infants, called Ego4D, had consumed roughly 3,000 hours of human activity footage. Even that advantage comes with the caveat that these models were optimized on far more data than any infant has ever encountered.
What the study reveals is not a failure of scale. It is a failure of architecture. Infants extract goal-directed structure from behavior with almost no data. They attribute preferences and intentions as a baseline cognitive function. Current models treat those attributions as something to be inferred from a statistical distribution of observed behavior — which is a categorically different task, and one they do poorly at even with millions of examples.
Dillon puts it plainly: adults and even infants can easily make reliable inferences about what drives other people's actions. Current AI finds those same inferences challenging. Lake adds the normative point — if we want machines that think flexibly like adults do, we may need to give them the cognitive starting point that infants already possess. The authors call their framework "reciprocal development of intelligence," which suggests infant cognition is not merely a curiosity but a blueprint.
Abstract Reasoning Stalls at the Visual Layer
The second study comes from USC's Viterbi Information Sciences Institute, published on arXiv in October 2024. The team, led by Babak Salimi, a doctoral student working under USC ISI researchers, put multimodal large language models against abstract visual reasoning puzzles of the kind that appear on human IQ tests. Matrix puzzles. Property-change sequences. Tasks that, if you have ever taken an intelligence assessment, you know feel like pattern recognition with shapes and shading.
The scores were underwhelming. GPT-4V managed 49% on matrix puzzles and 63% on sequences. PaLM-E did 51% on matrices and 38% on sequences. The open-source models did worse still. What interests me is not just that these numbers are low, though they are, but why they are low. The researchers found two distinct failure modes. The models struggled with visual processing (they could not reliably extract the relevant information from the image) and with reasoning (they could see the information but not manipulate it into a coherent solution).
One partial remedy: Chain of Thought prompting, which forces the model to articulate intermediate reasoning steps, improved accuracy by up to 100% for the smaller open-source models. That improvement is real. But I want you to notice the asymmetry. We are improving AI reasoning by asking it to explain itself step by step, that is, by giving it a scaffold that mirrors how human System 2 thinking works. The model does not reason spontaneously. We have to prompt it into something that resembles reasoning.
This is not a minor engineering detail. It points to the same structural observation the infant study makes from a different direction: the cognitive machinery humans deploy effortlessly is not present in these systems unless we explicitly construct it.
The Translator Who Gets Sharper
The third finding reverses the usual script. A team at the University of Surrey, Dr. Marianna Tsergas and Professor Marco Fabbri, published in the International Journal of Language and Linguistics, studied what happens when human language professionals collaborate with AI rather than compete against it.
The practice is called Interlingual Respeaking, or IRSP. A translator listens to spoken language, simultaneously translates it, and speaks the translation into speech-recognition software that produces live subtitles. Add punctuation and content labels orally. This is not a comfortable task. It demands that a single person maintain two languages, monitor speech recognition output, make real-time editorial judgments, and keep the utterance flowing, all in parallel.
Fifty-one language professionals completed a 25-hour upskilling course in IRSP. The researchers measured executive functioning and working memory before and after. Both improved significantly. Not marginally. Significantly.
I find this result the most quietly subversive of the three. The prevailing narrative about AI and human cognition runs in one direction: machines get sharper, humans atrophy. The Surrey team found the opposite in a setting where humans are not replaced by AI but forced to operate alongside it at a level of cognitive complexity neither party would face alone. The human brain, it turns out, treats the collaboration itself as a training load.
The authors frame this as a professional development insight, interpreters and translators should not fear AI but should learn to work with it in real time. I read it as something more general. Cognition, in humans, is adaptive to demand. Put a person in a situation that is hard enough, where the difficulty is genuine and the feedback is immediate, and the machinery gets better.
The Underlying Asymmetry
Pull these findings together and you get a picture with two arms. On one side: machines cannot do the small things. Infants read intentions that twelve-million-clip models cannot. IQ-test puzzles with abstract shapes defeat models that can pass a bar exam. On the other side: humans doing cognitively demanding work alongside machines become measurably more capable. The machine does not make the human dumber. The collaboration makes the human sharper, provided the collaboration is hard enough to constitute training.
This is a consistent picture. Human cognition is a general-purpose adaptation system that bootstraps from almost nothing. A baby gets the world for six months and has functional theory of mind. A translator gets twenty-five hours of IRSP practice and her working memory improves. The cognitive architecture is efficient, flexible, and responsive to challenge. Machine cognition is a statistical engine that does not spontaneously build representations of intention, does not reason without a scaffold, and, critically, does not get better by being made to struggle. It gets better by being made to train.
That asymmetry is the real story. Not whether AI will eventually surpass us on some benchmark. But whether we are building the right kind of intelligence or merely a very fast kind of computation. The infant study suggests we are missing foundational structures that evolution gave us in the first year of life. The abstract-reasoning study suggests that even with language fluency and pattern exposure, the visual and inferential machinery for fluid intelligence is absent. The Surrey study suggests that the gap between us and them is not narrowing, and, more importantly, that the human side of that gap is still getting better.
For those who have followed the debate about whether AI systems are "really thinking," these results should be clarifying. They are not thinking. Not yet. The gap is not a matter of scale. It is a matter of what cognition is, an adaptive system that grows under pressure, versus what computation is: a fixed architecture that requires external pressure to change its weights.
The infant knows. The translator improves. The model waits for its next training run.
Sources:
- Infants Outperform AI in "Commonsense Psychology", Neuroscience News (NYU, 2023)
- Can AI Tackle Abstract Reasoning? Study Tests Cognitive Limits, Neuroscience News (USC Viterbi ISI, 2024)
- Human Cognition Enhanced By AI Use, Neuroscience News (University of Surrey, 2023)