ProBackend
ai cognitive offloading
2 weeks ago8 min read

Everyone Overestimates AI Performance: The Reverse Dunning-Kruger Effect

A new study reveals that all AI users, regardless of skill level, overestimate their performance when interacting with tools like ChatGPT, reversing the classic Dunning-Kruger effect where lesser-skilled users typically overestimate more.

Introduction: When Confidence Outpaces Ability

When people sit down with a powerful large language model, a strange thing happens. They start to think they're better at the task than they actually are. A new study reveals that this overconfidence isn't limited to novices — it hits everyone, regardless of skill level. In fact, the classic Dunning-Kruger effect, where less-skilled users overestimate most, disappears entirely when AI enters the conversation. Instead, AI-literate users become the most overconfident of all.

This isn't just a quirky finding. It has real consequences for how people interact with AI tools in work, education, and daily life. When users can't accurately judge their own performance, they stop reflecting, they stop double-checking, and they offload critical thinking to the machine. The study, led by Aalto University and published in Computers in Human Behavior, puts this under the microscope with two large-scale experiments involving hundreds of participants and logical reasoning tasks from the LSAT.

Dr. Ryan Vance · October 28, 2025 · · 6 min read

The Modern Dunning-Kruger Twist

The original Dunning-Kruger effect describes a well-known pattern: people who perform poorly on a task tend to overestimate their ability. The worse they score, the more they think they've aced it. But put AI in the loop, and the rules change. Robin Welsch, professor at Aalto University, and doctoral researcher Daniela da Silva Fernandes designed two experiments to see whether AI use flips the script.

In Study 1, 246 participants used AI to solve 20 logical reasoning problems from the Law School Admission Test. Some used AI, some didn't. After each task, they reported how well they thought they'd done — and they were paid extra if their self-assessment was accurate. The results were striking. Even though participants' actual performance improved by about three points compared to a norm population, they overestimated their performance by four points. Every single subgroup overestimated, but the pattern was different than expected.

Higher AI literacy correlated with lower metacognitive accuracy. In plain language: the more someone knew about how AI works, the more confident they felt — but the less precise they were in judging their own performance. The Dunning-Kruger effect, as traditionally understood, ceased to exist with AI use. What replaced it was a reverse pattern: AI-literate users showed even greater overconfidence than novices.

Study Design: LSAT Tasks and AI Assistance

The researchers didn't pick random tasks. They chose logical reasoning problems from the LSAT, a high-stakes test used for law school admissions. These tasks require careful analysis, and they're cognitively demanding. The study included two experiments. Study 1 had N = 246 participants. Study 2 replicated the design with N = 452 participants, confirming the same findings.

Half the group used AI and half didn't. This setup allowed the team to isolate the effect of AI itself, separate from any practice effect or inherent skill difference. After each task, subjects monitored how well they performed. If they were accurate, they earned extra compensation — giving them a real incentive to self-assess honestly.

What happened next is where the story gets interesting. Most participants rarely prompted ChatGPT more than once per question. Often, they simply copied the question, put it in the AI system, and were happy with the AI's solution without checking or second-guessing. The researchers called this cognitive offloading — when all the processing is done by AI, and the user plays a passive role.

"These tasks take a lot of cognitive effort. Now that people use AI daily, it's typical that you would give something like this to AI to solve, because it's so challenging," says Welsch. The data confirmed the theory. Shallow engagement may have limited the cues needed to calibrate confidence and allow for accurate self-monitoring. It's plausible that encouraging or experimentally requiring multiple prompts could provide better feedback loops, enhancing users' metacognition.

The Reverse Dunning-Kruger Effect

So what exactly is this reverse Dunning-Kruger effect? In the classic version, low performers overestimate themselves. With AI, the pattern flips. Users who consider themselves more AI literate tend to assume their abilities are greater than they really are. They know enough about the technology to trust it, but not enough to critically evaluate their own outputs.

"We found that when it comes to AI, the DKE vanishes. In fact, what's really surprising is that higher AI literacy brings more overconfidence," says Professor Robin Welsch. "We would expect people who are AI literate to not only be a bit better at interacting with AI systems, but also at judging their performance with those systems — but this was not the case."

This creates a dangerous loop. The very people who understand AI best are also the most likely to overestimate their skill with it. They're confident in answers that may be wrong, and they're less likely to spot their own mistakes. The study found that AI literacy alone is not enough. People need platforms that foster metacognition and critical thinking — tools that make reflection visible, not optional.

Cognitive Offloading: The Single-Prompt Habit

One of the most practical findings from the study involves how people actually use AI. The researchers designed their experiments to probe real-world usage patterns. What they found was disappointing, from a metacognition standpoint. Most users relied on single prompts. They put in one request, accepted the answer, and moved on. Rarely did they ask follow-up questions, challenge the AI's reasoning, or ask it to explain its thought process.

"We looked at whether they truly reflected with the AI system and found that people just thought the AI would solve things for them. Usually there was just one single interaction to get the results, which means that users blindly trusted the system. It's what we call cognitive offloading, when all the processing is done by AI," Welsch explains.

This shallow level of engagement has consequences. When a user only interacts with AI once, they miss the feedback loops that could help them calibrate their confidence. They don't see where the AI struggled, where it might have hallucinated, or where their own reasoning diverged from the machine's output. It's a one-way street: the AI gives an answer, the user accepts it, and no metacognitive learning occurs.

"AI could ask the users if they can explain their reasoning further. This would force the user to engage more with AI, to face their illusion of knowledge, and to promote critical thinking," says doctoral researcher Daniela da Silva Fernandes. The study suggests that design interventions — requiring explanation, prompting reflection, or building in checkpoints — could counteract the offloading tendency and help users maintain a more accurate sense of their own skill.

The Metacognition Gap

Why do current AI tools fail at helping users self-assess? The study points to a fundamental gap. Generative AI systems are optimized for producing plausible answers, not for calibrating user confidence. They don't nudge users to reflect. They don't show uncertainty estimates. They don't provide a mirror back to the user about how well they're doing.

"Current AI tools are not enough. They are not fostering metacognition [awareness of one's own thought processes] and we are not learning about our mistakes," adds da Silva Fernandes. "We need to create platforms that encourage our reflection process."

The implications go beyond individual users. If people consistently overestimate their AI-assisted performance, it can lead to workforce de-skilling, poor decision-making in high-stakes domains, and a general erosion of critical thinking. The study warns that blind trust in AI output comes with risks like "dumbing down" people's ability to source reliable information.

What Can Be Done: Platform Solutions

The study doesn't just identify problems — it suggests design directions. The most repeated recommendation is that AI platforms should foster reflection. This could take several forms:

  • Built-in prompts that ask users to explain their reasoning after receiving an answer
  • Uncertainty indicators that show when the AI is less confident, forcing the user to pause
  • Multiple-interaction workflows that require a follow-up check before finalizing an answer
  • Self-evaluation tools that ask the user to rate their confidence, then compare that rating to actual accuracy

These aren't just nice-to-have features. The study's data shows that without intentional design, users will default to the path of least resistance: one prompt, accept the answer, move on. Platforms that build in metacognitive checkpoints can interrupt this default and help users develop a more accurate sense of what they can and can't do.

Conclusion: Toward Wiser AI Interaction

The Aalto University study offers a clear-eyed look at what happens when people meet AI. The results are sobering: everyone overestimates their performance, AI literacy amplifies rather than reduces this effect, and most users engage in shallow, single-prompt interactions that short-circuit reflection. The classic Dunning-Kruger effect goes out the window, replaced by a reverse pattern where the most knowledgeable users become the most overconfident.

But the study also points the way forward. By designing AI platforms that encourage reflection, that require users to explain their thinking, and that make metacognition visible, we can counteract the natural tendency to overtrust. The goal isn't to stop people from using AI — it's to help them use it wisely, with a realistic sense of their own skill and a habit of critical self-assessment.

The reverse Dunning-Kruger effect is a reminder that technology isn't neutral. How we design it, and how we use it, shapes whether it lifts our thinking or dulls it. The choice, as the researchers emphasize, is ours.


Source: Neuroscience News article (Oct 28, 2025) covering Aalto University study published in Computers in Human Behavior. Two experiments: Study 1 N=246, Study 2 N=452. Participants used AI to solve LSAT logical reasoning tasks.

confidence outpaces ability

More blogs