ProBackend
ai human perception cognition
5 hours ago8 min read

The Deep AI Impact on Human Psychology: When Flawed Video Summaries Implant False Memories

Expanded research report on how misleading AI video summaries dropped eyewitness recall from 83.6% to 44.8% in a Georgetown and University of Washington study — and what the result reveals about the AI impact on human psychology, what AI in psychology means, and whether AI can understand human psychology at all.

As modern institutions increasingly delegate the compression of complex video and other evidence to AI, a new question is emerging: what happens when a summary does more than omit details, and instead actively rewrites them? New research out of Georgetown University and the University of Washington, reported by Neuroscience News in September 2026, delivers an uncomfortable answer. People who watched a traffic accident and then read a misleading AI-generated summary of it recalled the scene they had witnessed with their own eyes at barely half the accuracy of people who read an accurate summary — 44.8 percent versus 83.6 percent. The errors were not merely forgotten; they were absorbed, incorporated, and reported as memory.

This is among the cleanest controlled demonstrations yet of the AI impact on human psychology: not through chatbots that argue with us or deepfakes that deceive us, but through quiet, plausible condensation — an ordinary summary with an ordinary mistake, landing in an ordinary human memory that cannot tell the difference.

Memory Was Never a Video Recording

To understand why this finding matters, it helps to discard the folk model of memory as a filing cabinet or a video file. Cognitive psychologists have long established that human episodic memory is a malleable, reconstructive process. Each time we recall an event, we rebuild it from fragments — and that rebuilding is open to contamination from anything we encounter afterward, whether misleading or accurate. Decades of misinformation-effect research show that post-event information can alter what people believe they saw.

Large language models have now inserted themselves into that vulnerable window at industrial scale. From corporate meeting transcripts and clinical case notes to law enforcement body-worn camera logs, institutions are accelerating the deployment of LLMs to condense long video and audio streams into concise narrative digests. The unstated assumption behind nearly all of these deployments is that a human reader will catch what the model got wrong. The Georgetown–UW work tests that assumption directly, and it fails.

The Experiment: Summaries That Erased the Crash

In the study, published in the Journal of Experimental Psychology: General, participants first watched footage of a traffic accident in which a car hit a pedestrian. They were then asked to recall what they had seen and to read an AI-generated summary of the video. Crucially, the researchers first audited real summaries produced by leading commercial multimodal models — including ChatGPT and Gemini — and characterized their failure modes:

  • Severe omission rates. Across consumer multimodal AI models, automated video summaries omitted an average of 51.6 percent of the central events in the footage.
  • Missing the core event entirely. In 95 percent of the summaries, the single most important fact — that a car hit a pedestrian — went unmentioned. When the summary of an accident does not contain the accident, the reader's mental model of the event is being built from its absence.
  • Misinformation, not just silence. A significant proportion of summaries contained false details, not merely gaps. In one example, a summary described "anxiety in the onlookers' faces", language that invited an emotional inference the footage did not support.

Participants were then randomly assigned to read either an accurate or an inaccurate AI summary before taking a memory recognition test for the original event. Those who read the accurate summary correctly recalled key scene details 83.6 percent of the time. Those exposed to the misleading summary managed only 44.8 percent. Nearly half of what accurate readers remembered was gone, or replaced, for the misleading-summary readers, despite the fact that every participant had watched the video themselves.

The Human-in-the-Loop Fallacy: Labels Did Not Protect Memory

The most policy-relevant result is what happened when participants knew the summary came from a machine. Labeling a summary as AI-produced failed to insulate observers from memory contamination. Participants internalized the errors regardless of their stated familiarity with AI or their declared distrust of it. Experience with the technology offered no immunity; skepticism offered no immunity.

Lead author Kyle Whyne, a Ph.D. student at the University of Washington, frames the implication bluntly: "Though humans-in-the-loop are often expected to correct for AI's mistakes, our work suggests human memory can instead be distorted by these mistakes. AI has the potential to generate misinformation, even absent any adversarial intent, which can meaningfully impact human memory."

That last clause deserves emphasis. Nothing adversarial was required here. No bad actor prompted the model to lie. The misinformation emerged from the models' ordinary, mundane unreliability, the same dropout and confabulation behavior users see every day, and it was still enough to halve eyewitness accuracy. Disclosure norms, media-literacy training, and "verify before relying" warnings, the standard toolkit for AI misinformation policy, are aimed at a threat model of deliberate deception. This study shows the quieter failure mode operates underneath all of those defenses, because it attacks the memory of the verifier.

What Is AI in Psychology? A Field With Two Directions

The coverage of this study raises a definitional question readers often ask: what is AI in psychology? The phrase covers two distinct projects, and the Georgetown–UW work sits at the intersection.

The first project is AI for psychology: using machine-learning tools to study, model, and support the mind, computational models of memory and decision-making, diagnostic screening tools, therapy chatbots, and experimental platforms that measure human cognition at scale. This study belongs here: it uses AI not to treat or model the mind therapeutically, but as an experimental variable, probing how a specifically artificial kind of post-event information reshapes a core human function.

The second project is psychology for AI: applying what we know about human cognition, perception, and trust to the design and governance of AI systems, understanding why people over-attribute competence to fluent text, why automation bias takes hold, and where human oversight can and cannot be load-bearing. The finding that an "AI-generated" label changes nothing is a psychology-for-AI result: it tells regulators that disclosure-based frameworks rest on an assumption about human vigilance that this experiment contradicts.

Can AI Understand Human Psychology? Not From the Inside

Closely related is the question readers ask most often: can AI understand human psychology? The honest answer, on current evidence, is that systems like ChatGPT and Gemini process and generate language about psychological states without any of the machinery that gives the real thing its stakes. They have no episodic memory to distort, no social standing to lose when they are wrong, and no phenomenal experience of the events they summarize. They model correlations in text that describes minds; they do not model minds.

The study is instructive precisely because it inverts the usual framing. The interesting cognition in the experiment is not in the model at all. It is in the participant, whose reconstructive memory accepts a plausible narrative as raw material. The AI's "understanding" is irrelevant to the causal chain, a fact that should deflate both hype and panic. A system does not need to understand psychology to affect it. It only needs to produce fluent text in front of a brain that is built to believe fluent narratives about events, especially when it is told those narratives summarize something that happened.

The asymmetry is worth naming plainly: humans are deeply vulnerable to AI outputs, while AI remains entirely indifferent to the humans it is altering. Any serious account of AI and human psychology has to start from that asymmetry rather than from the science-fiction image of machines that read us while we read them.

Where the Damage Could Land

The experiment used a brief accident video, but the deployment contexts flagged by the researchers are anything but small. Eyewitness testimony already carries known reliability problems; a misleading digest layered on top of it in a police review workflow could harden a false account that a full re-watch might have corrected. Clinical case notes summarized by an LLM could implant false recollections of a therapy session's content in a patient or clinician. Meeting recaps could rewrite what attendees believe was decided. Media consumers who read an AI digest of an event they watched may find their memory quietly edited, 51.6 percent of the central events at a stroke.

The researchers are careful about generalization: the effect was measured in a controlled lab task with short videos and immediate testing, and longer exposure histories or higher-stakes settings could strengthen or weaken it. What is not speculative is the direction. The very institutions that stand to gain most from AI condensation, courts, clinics, newsrooms, law enforcement, are the ones whose decisions most depend on human memory staying clean.

Practical Guardrails While the Evidence Catches Up

Until replication studies map the boundaries of the effect, several low-cost practices follow directly from the findings:

  1. Treat AI summaries as leads, not records. In high-stakes review, evidence, clinical notes, incident footage, open the underlying media before forming or recording a judgment. The cost of watching again is almost always lower than the cost of a contaminated account.
  2. Don't rely on labels. The experiment shows provenance disclosure does not inoculate memory. Design review processes so that verification happens against the source artifact, not against the reader's confidence.
  3. Summarize with the core event pinned. Since 95 percent of summaries dropped the central event, systems that digest evidence should structurally guarantee the core facts appear and are checked first.
  4. Record impressions before reading the digest. If a witness, clinician, or attendee writes down their own recall before encountering the AI summary, the summary has less raw material to overwrite.
  5. Fund the misinformation-effect studies this field still needs. The authors call their work the first rigorous experimental test of AI-generated misinformation's effect on memory. One experiment, however elegant, should be the opening of a research program, not the last word.

The Bottom Line

Georgetown and University of Washington researchers showed that reading a misleading AI summary of a video people had witnessed firsthand slashed their recall accuracy from 83.6 percent to 44.8 percent, and that knowing the summary was AI-made changed nothing. The result reframes the AI impact on human psychology as something more mundane and more intimate than manipulation by intelligent machines: a reconstructive memory system colliding with a compression system that loses half the story and sometimes invents the rest. Understanding AI in psychology today means taking that collision seriously on both sides of the label, building AI that respects how memory actually works, and educating humans about a vulnerability that no amount of distrust, apparently, switches off.

Sources

  • Neuroscience News: "Flawed AI Summaries Distort Eyewitness Memory" (September 27, 2026), summary of research from Georgetown University and the University of Washington published in the Journal of Experimental Psychology: General. https://neurosciencenews.com/ai-false-memory-llm-31261/

memory was never a video recording

More blogs