Study finds LLM-generated autobiographies inaccurate 96.7% of the time, with fabricated and distorted events
A new study finds that large language models fabricate events in nearly all generated autobiographical accounts when compared with documented life records.
A recent study titled "Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes," published on arXiv in August 2026, presents the first quantitative examination of confabulation by Large Language Models (LLMs) when writing a person's autobiography.
The study's author, Heather Renze, tasked an LLM with drafting 366 daily autobiographical entries using only a template, two sample days, and a quote of the day as initial input. Each entry was then rigorously checked against a ground-truth corpus using a predefined four-level evaluation framework.
The study found a verification-failure rate of 96.7%, with 354 of the 366 entries failing verification. Only 12 entries described events consistent with the documented record. Another 19 entries, or 5.2%, directly contradicted reality.
The most common failure mode was "grounded drift," in which real people, places, or employers were woven into events invented by the model. The study concludes that using current AI models to construct life stories carries a high risk of inaccuracy without appropriate data grounding.
The findings highlight a major limitation of AI in handling personal memories and biographies, warning users to watch for inaccuracies and fabricated information when applying these models to work that demands high factual accuracy.