Skip to content
AI.info

The Pulse

Misleading AI Summaries Cut Memory Accuracy to 44.8% in Study

A study by Mattea Sim, Yael Eiger and Tadayoshi Kohno found that misleading AI-generated summaries reduced participants’ accuracy on a memory question about a traffic sign. The researchers also found that all 20 summaries they tested contai

Misleading AI Summaries Cut Memory Accuracy to 44.8% in Study

AI.info Team ·

Only 44.8% of participants who read a misleading AI summary correctly remembered a traffic sign in a video, compared with 83.6% of those given an accurate summary. The results come from a study posted to arXiv on September 23 by Mattea Sim, Yael Eiger and Tadayoshi Kohno, which tested whether errors in AI-written summaries can change what people remember.

The experiment involved a short animated car-pedestrian accident. The researchers say the findings raise a concern for settings where AI summaries may influence later judgments: human review may not reliably undo a misleading account once someone has read it.

A traffic sign became the test of memory

Researchers recruited U.S. adults online to watch one of two 25-second videos. Both showed a red car approaching an intersection, with either a stop sign or a yield sign, before it collided with a pedestrian. After a delay of 24 to 48 hours, participants read a summary and answered questions about the original video.

The study analyzed 328 participants who completed both stages and passed the required attention checks. Researchers varied whether the summary described the sign accurately or misleadingly, and whether participants were told it came from an AI or a human transcriber. In fact, the summary was generated by ChatGPT in both source-label conditions; the researchers changed the sign detail to create the accurate and misleading versions.

“People who read a misleading AI summary were significantly less likely to accurately recall the original event, compared to people who read an accurate AI summary.”

Mattea Sim, Yael Eiger and Tadayoshi Kohno, authors of “AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory”

Every tested summary contained errors

In a separate analysis, the researchers prompted ChatGPT-5.5 and Gemini 2.5 Flash-Lite to summarize the same two videos. They generated 20 summaries in total. Every one contained at least one error, and the number of errors ranged from seven to 21 per summary.

Omissions appeared in all 20 summaries. Across the prompts and models, summaries left out 51.6% of the details the researchers had identified as central to the events. In 19 of the 20 summaries, the car’s collision with the pedestrian—the central event—was missing. The researchers also recorded inaccurate descriptions and added details that did not appear in the videos.

Labeling the summary did not prevent the effect

Calling the summary AI-written rather than human-written did not significantly change participants’ accuracy. The researchers also found no significant relationship between memory accuracy and participants’ reported trust in AI or frequency of AI use. The label did not protect participants from the misleading detail.

The authors connect the result to the established misinformation effect: people can incorporate information they encounter after an event into their later recollection of it. Their experiment tested a specific version of that problem, using a traffic sign in an animated video and a summary whose wording was deliberately adjusted for the study.

The test was narrow, but the setting matters

The researchers caution that their analysis of summary errors was limited to two brief videos, one standardized prompt and 20 outputs. The footage was an animated accident scene, not real-world police or body-camera video. The authors say those limits mean the results cannot establish how often similar memory effects occur in other contexts or with other systems.

Even so, the experiment suggests that an inaccurate summary can do more than introduce a wrong detail: it can affect a person’s later answer about what they saw. Whether that effect holds for longer, realistic footage—and how people using AI-generated reports can check the original evidence without relying on altered recollection—remains untested by this study.

Source

Explore

More articles