The Statistical Illusion Behind Dunning-Kruger
For decades, the Dunning-Kruger effect has been shorthand for calling out clueless confidence. The popular version of the story is simple enough: people who know the least overpredict their performance the most, while high performers stay grounded. It's been replicated hundreds of times. It's been weaponized in boardrooms and on social media. Everyone knows it.
Except it appears not to be real.
A new study by Professor Chris Dawson at the University of Bath and Professor David de Meza at LSE, published in Psychological Review, shows that the celebrated Dunning-Kruger pattern is actually a statistical artifact. When you correct for the math, the entire psychological narrative flips. The least competent aren't overconfident because they're clueless. They're overconfident because bad luck makes their test scores look artificially low, which inflates the gap between what they predicted and what they actually achieved.
Dawson put it bluntly: "Our research demonstrates that the original methodology fell victim to a statistical illusion. The popular image of the incompetent person brimming with confidence tells us less about human psychology than it does about the hidden traps in statistical analysis."
The original studies sorted participants strictly by their raw test scores without accounting for volatility—random noise from luck, ambiguous question phrasing, or a bad day. That sorting method creates a mathematical distortion. Low scores often reflect bad luck, not low ability. When you group people by their noisy test scores, you manufacture the exact pattern Dunning and Kruger claimed to have discovered.
How Bad Luck Manufactures Overconfidence
Test performance is naturally volatile. You get lucky with a question you happen to know cold, or you get burned by a poorly worded option that looks plausible. These fluctuations aren't noise to ignore—they're baked into every assessment, every A/B test, every experiment.
Here's what happens when you sort by raw scores. Someone with high actual ability gets unlucky on a test. Their score drops. But their self-assessment—made before seeing the results—stays the same. The gap between their prediction and their actual score looks enormous. They appear wildly overconfident.
Now flip it. A genuinely lower-ability person gets lucky. Their score inflates beyond what their self-assessment predicted. The gap shrinks—or reverses. They look unrealistically accurate.
De Meza explained it clearly: "When a test score is exceptionally low, it's often because the person just got unlucky with the questions—meaning their estimated performance naturally ends up looking way too high. The exact opposite happens to top performers, who benefit from a lucky break."
The researchers applied improved statistical techniques to data from thousands of participants across massive replication datasets. This approach handles randomness in both people's test scores and their predictions about those scores. Once the model respects this reality, the classic Dunning-Kruger story flips. High-ability individuals display the highest levels of overconfidence. The pattern doesn't disappear—it inverts.
Why High Performers Take the Biggest Risks
So if high performers are actually the most overconfident, why? The answer isn't cognitive incompetence. It's strategic.
The researchers propose that overconfidence functions as a social signaling mechanism. High-ability individuals have the strongest incentives to display confidence because they need to communicate unobservable competence to others. Your actual skill might be hidden. Your results might be ambiguous. But your confidence? That's loud, clear, and hard to fake at scale.
This isn't about delusion. It's about signaling. When your ability is genuinely high, the benefit of being perceived as able increases with that ability. The decision error associated with high self-belief is less burdensome for the more capable—they can afford to be wrong sometimes and still deliver. But the less able? Their errors are costlier. They can't afford the gap between confidence and delivery.
Overconfidence, then, isn't a bug. It's a feature. It's a credible signal precisely because the more able can back it up. The research shows that overconfidence is a widespread human tendency across all skill tiers, but its magnitude scales positively with actual competence. The most capable people exhibit it the most—not because they're deluded, but because they have the most to gain from signaling.
The Signaling Game Behind Confidence
The paper's full title—"Talking the Talk, Not Walking the Walk: The Coevolution of Overconfidence and Loss Aversion"—hints at a deeper puzzle. If overconfidence is so adaptive, why does loss aversion exist? Loss aversion curbs initiative. It tells you to hold back. Overconfidence tells you to leap.
The researchers argue these biases are symbiotic. Loss aversion partially ameliorates the decision costs of overconfidence, but because it's usually hidden, it doesn't eliminate the signaling role of confidence. The payoff to agents from an integrated set of biases is higher than would be the case if either were absent.
This has a direct implication for how we design experiments and interpret results. Kahneman's advice that individuals should eliminate both overconfidence and loss aversion is, by this framework, "poorly founded." These aren't errors to correct. They're evolved strategies that work together. The system functions because the biases complement each other.
For anyone running A/B tests or making business decisions, this means something important: confidence and caution aren't opposites. They're interdependent. Suppress one without understanding the other, and you break the equilibrium that keeps decision-making functional.
What This Means for A/B Testing
The practical takeaway for operators is straightforward. When you're evaluating team performance, test results, or experiment outcomes, don't trust raw scores. They're noisy. They're contaminated by luck. Sorting people or campaigns by raw performance without correcting for volatility will give you exactly the wrong picture. As explored in When AI Makes Testing Cheap, Your Experimentation Standards Must Get Harder, scaling test output without scaling decision standards is a direct route to wasted budget—and the same statistical noise that creates false Dunning-Kruger patterns also produces false positives in your A/B test results.
The Dunning-Kruger effect didn't reveal a psychological truth about incompetence. It revealed a mathematical truth about how we measure things. And once you correct for that math, you discover that the people who deliver the best results are also the ones who project the most confidence—not because they're wrong about themselves, but because confidence is how they signal value in a world where competence is often invisible.
That's not a flaw in human judgment. It's a feature of how high performers operate. And it's something worth accounting for when you're designing experiments, evaluating teams, or trying to understand why your best performers sometimes sound like they're overselling.
They're not overselling. They're signaling. There's a difference.
Source: Dawson, C., & de Meza, D. (2026). "Talking the Talk, Not Walking the Walk: The Coevolution of Overconfidence and Loss Aversion." Psychological Review. DOI: 10.1037/rev0000644. Research from the University of Bath and London School of Economics.