Research notes
Outline: 1) What analogical reasoning measures and why UCLA tested it; cite the original comparison and task variety. 2) Results across different benchmarks, contrasting successful matrices/SAT analogies with story and physical-tool failures; cite Neuroscience News and the original paper. 3) Human-cognition-inspired model: describe the UCLA team's own psychological model and how GPT-3's capability eventually caught or surpassed it per the report. 4) Interpret cautiously: report the authors' emergent-reasoning interpretation alongside the published critique about task variations and possible memorization. Avoid asserting models think as humans do. Evidence notes are attributed source by source above.
Source: https://neurosciencenews.com/chatgpt-ai-student-reasoning-23730/
- Neuroscience News reported July 31, 2023 on UCLA psychologists' comparison of GPT-3 with human participants on analogical reasoning problems.
- GPT-3 correctly solved 80% of one problem set; the human subjects averaged just below 60%, though top human scores overlapped its performance range.
- On SAT analogy questions researchers considered unlikely to have appeared online, GPT-3 scored above the average college applicant result.
- GPT-3 did worse than student volunteers on analogies based on short stories; GPT-4 did better than GPT-3 on those problems.
- The UCLA team developed a computer model inspired by human cognition and compared its ability with commercial AI; professor Keith Holyoak said their psychological model had led on analogy tasks until a GPT-3 upgrade caught up or did better.
- GPT-3 struggled with physical-space/tool-use scenarios, offering implausible plans for transferring gumballs using a cardboard tube, scissors and tape.
Source: https://arxiv.org/abs/2212.09196
- The study by Taylor Webb, Keith J. Holyoak and Hongjing Lu compared text-davinci-003, a GPT-3 variant, with human reasoners across multiple analogy tasks, including a nonvisual matrix reasoning task based on Raven's Standard Progressive Matrices.
- The authors reported GPT-3 matched or surpassed human capabilities in most tested settings on abstract pattern induction; preliminary GPT-4 tests indicated better performance.
- The authors argued large language models had acquired an emergent ability to find zero-shot solutions to a broad range of analogy problems.
Source: https://arxiv.org/abs/2308.16118
- Damian Hodel and Jevin West's 2023 response offered counterexamples using letter-string analogies and reported GPT-3 failed simple task variations while human performance stayed consistently high.
- The response authors argued evidence for zero-shot reasoning must rule out memorization from training data and challenged the strength of the original reasoning claim.
Introduction: The Benchmark of Human Analogical Thought
Analogical reasoning—the cognitive process of transferring information or meaning from a particular subject to another—has long been considered a hallmark of uniquely human intelligence. When faced with novel problems without prior direct instruction, humans readily identify structural similarities between familiar domains and unfamiliar scenarios, applying abstract rules with remarkable flexibility. For decades, cognitive scientists treated this capacity as a crucial dividing line between biological minds and artificial computation.
However, recent advancements in large language models (LLMs) have upended traditional assumptions about machine cognition. In a landmark study published in Nature Human Behaviour, researchers at the University of California, Los Angeles (UCLA)—including Taylor Webb, Keith J. Holyoak, and Hongjing Lu—investigated whether commercial language models can exhibit genuine analogical reasoning capabilities. By comparing models such as GPT-3 and GPT-4 against human participants across a diverse array of reasoning problems, the UCLA team explain both the astonishing capabilities and the stark limitations of contemporary artificial intelligence.
Testing GPT-3 Against College Undergraduates
To evaluate whether generative AI can perform analogical reasoning on par with humans, the UCLA researchers administered a series of rigorous tests traditionally used to assess human intelligence. These included matrix reasoning tasks modeled after Raven’s Standard Progressive Matrices, verbal analogies derived from standardized tests like the SAT, and short story-based reasoning puzzles.
The results surprised many observers. Across one major problem set, the text-davinci-003 variant of GPT-3 correctly solved approximately 80% of the problems. In comparison, human undergraduate subjects averaged just below 60% correct, although top-performing human participants achieved scores overlapping with the upper limits of the AI's range. Furthermore, on custom SAT-style analogy questions specifically designed by researchers to minimize the likelihood of direct memorization from online training corpora, GPT-3 scored higher than the average college applicant.
In nonvisual matrix reasoning tasks—which measure abstract pattern induction without relying on linguistic prompts, GPT-3 demonstrated an unexpectedly strong capacity for discerning underlying rule structures. Preliminary evaluations of GPT-4 indicated even higher performance, suggesting a scaling trend in abstract reasoning proficiency across successive model generations.
The UCLA Human-Cognition-Inspired Model and Comparative Benchmarking
To better contextualize these findings, the UCLA researchers did not merely observe commercial AI; they also developed and benchmarked their own computer model specifically inspired by human cognition. For years, computational models rooted in cognitive psychology, such as those designed by Holyoak and colleagues to simulate human analogical mapping, held the performance lead in solving complex relational matching tasks.
In comparative evaluations, the UCLA team's human-cognition-inspired model served as a baseline for how rule-based cognitive architectures handle structural alignment and abstraction. Interestingly, as commercial language models scaled up and underwent successive architectural and training updates (such as the transition from standard GPT-3 variants to more refined iterations like GPT-4), these massive neural networks eventually caught up to and, in certain zero-shot benchmarks, surpassed the performance of specialized psychological models designed explicitly for analogy.
Yet, this convergence raised profound theoretical questions: Are large language models actually implementing cognitive mechanisms akin to human mental models, or are they executing statistical pattern matching that mimics the outward appearance of reasoning? The tension between surface-level competence and mechanistic difference is exactly what critics describe as a cognitive illusion: performance that looks like understanding without the underlying processes that give human reasoning its robustness.
Striking Strengths and Peculiar Blind Spots
Despite achieving high marks on standardized matrix and verbal analogies, GPT-3 exhibited stark, asymmetrical failures that underscored the fundamental differences between biological reasoning and transformer-based text generation.
While the model excelled at abstract geometric and linguistic analogies, it stumbled significantly when tasked with reasoning about physical-space mechanics and everyday tool use. For instance, when presented with physical problem-solving scenarios, such as figuring out how to transfer gumballs across a room using simple household items like a cardboard tube, scissors, and tape, GPT-3 frequently generated impractical or physically implausible plans. These failures highlight a persistent gap in grounded physical common sense, even among models capable of scoring well on abstract academic tests.
Moreover, model performance varied depending on the semantic domain. On story-based analogies, GPT-3 performed notably worse than human student volunteers, though GPT-4 showed marked improvements in bridging that gap. These uneven performance profiles suggest that high aggregate test scores mask specific vulnerabilities in handling grounded context and multi-step physical interactions.
Emergent Capabilities vs. Training Memorization
The interpretation of these results sparked an intense academic debate within the cognitive science and AI communities. The UCLA researchers argued that large language models have acquired an emergent ability to find zero-shot solutions to a broad range of analogy problems, meaning the capacity arises spontaneously as a byproduct of scale and generalized training, rather than explicit programming for analogical mapping.
However, this conclusion was quickly challenged. In a published scientific response in late 2023, researchers Damian Hodel and Jevin West offered counterexamples utilizing letter-string analogies. Their experiments revealed that when simple variations were introduced to the original tasks, GPT-3's performance dropped sharply, whereas human participants maintained consistently high accuracy across all modified versions.
Hodel and West argued that extraordinary claims, such as attributing genuine zero-shot reasoning to language models, demand rigorous empirical evidence that definitively rules out data memorization. Because large language models ingest vast quantities of internet text, what appears to be novel reasoning during zero-shot evaluation may instead represent sophisticated retrieval of patterns deeply embedded in training data.
Conclusion: Bridging Artificial and Human Intelligence
The UCLA study and the subsequent scholarly debates illuminate the complex landscape of modern artificial intelligence. While commercial AI models like GPT-3 and GPT-4 demonstrate remarkable proficiency in abstract pattern induction, linguistic analogy, and zero-shot problem solving, they operate through mechanisms fundamentally distinct from human cognition.
By comparing commercial systems with both human participants and custom cognitive models, researchers continue to map the boundaries of machine intelligence. Whether current LLMs possess true emergent reasoning or merely mirror human thought through statistical interpolation, the ongoing dialogue between cognitive science and artificial intelligence remains essential for understanding the true nature of machine cognition. This episode sits within a broader set of epistemic fault lines between human and artificial intelligence, where each new benchmark both narrows and reveals the gap.