Research
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
Overview Research area: Large language model confabulation and factuality evaluation, specifically long-form first-person narrative generation ("synthetic memoir"), person simulation, and digital twin
- arXiv
- 2608.23640
- Published
- 2026-08-23
- Authors
- Heather Renze
AI summary
Overview
Research area: Large language model confabulation and factuality evaluation, specifically long-form first-person narrative generation ("synthetic memoir"), person simulation, and digital twins.
Technical level: Intermediate. The statistics used are accessible (Wilson score intervals, Cohen's kappa, Fisher exact tests), but the domain framing (episodic vs. attitude-level fidelity in person simulators) assumes some familiarity with LLM evaluation literature.
Scope: A single-subject, scene-level audit of 366 LLM-generated autobiographical day-entries, rated against an independent ground-truth corpus documenting the life of the paper's own author, who is also the audited subject.
Note: the full paper text supplied here is truncated partway through Section 7.1, so any content after that point is not represented in this summary.
What This Paper Is About
When an LLM is asked to write someone's life story, how much of what it writes actually happened? The author — who is also the subject — had a 366-day "page-a-day" book of first-person anecdotal entries drafted by a conversational LLM using only a template, two exemplar days, and each day's quote, never her own records. Every one of the 366 days was then audited at the anecdote-scene level against an independent verification corpus using a rubric fixed before analysis, and the resulting labels are mechanically parsed and quantified here.
Key Contributions
- A quantified, scene-level audit of confabulation in personal-narrative generation against subject-specific ground truth. The unit is the day-narrative (n = 366), and 96.7% of days fail positive scene-level verification under the paper's stated definition.
- A reusable audit instrument. The paper publishes the four-verdict rubric verbatim (VERIFIED / WEAK / UNVERIFIED / CONTRADICTED) plus a nine-theme recurring-false-premise taxonomy operationalized as fixed, published keyword screens in
02_analyze.py. - A remediation workflow specified during the same project: lesson-only rewriting (reflection without the invented scene) and author-adjudication precedence, where the subject's explicit written rulings bind all other rows.
- An inter-rater reliability analysis with independent re-raters, yielding the paper's most transferable methodological claim: scene-level fabrication audits are reliable at the binary level (does this day's scene check out?) but only fair-to-moderate at the taxonomic level, because the WEAK/UNVERIFIED boundary turns on an undefined notion of "setting." A controlled 2026 replication with current named models and a grounding arm accompanies this.
Main Findings
- Headline verification-failure rate: 354 of 366 days (96.7%, Wilson 95% CI 94.4–98.1%). Only 12 days (3.3%) contain a scene that survived positive corroboration; 19 days (5.2%, CI 3.3–8.0%) assert claims the record actively contradicts.
- Verdict distribution: WEAK 227 days (62.0%), UNVERIFIED 108 days (29.5%), CONTRADICTED 19 days (5.2%), VERIFIED 12 days (3.3%).
- The dominant failure mode is "grounded drift." WEAK — real setting, person, or employer with an invented specific scene — is the single largest category in all three independent ratings, though its measured share is rater-sensitive (43–82%; the discussion cites 43%, 63%, and 82% of sampled days).
- The source audit summary's own approximate tally (~350 of 366 days, ~9 VERIFIED days) is consistent with exact parsing (354/366; both round to ≥95%). All inferential statistics use the parsed denominators.
- Verified survivors are almost entirely public-career facts (Evernote tenure and scale metrics) or scenes the subject has told publicly for years. The asymmetry: what survives is what the corpus says often; what fails is what sounds like the life.
- Contradictions cluster unevenly by month: August (6), April (4), and December (4) account for 14 of 19. Six months (January–March, June, July, August) contain zero VERIFIED days; May contributes 4 of the 12. The repository records no generation order or dates for the month files, so no chronological conclusion is drawn.
- UNVERIFIED share swells mid-year (July: 18 of 31 days), marking stretches where entries detached from the corpus entirely.
- Nine recurring false premises span four fabrication families — invented relations, invented capabilities and events, misattributed facts, and inverted dispositions — with screen footprints totaling 31 distinct days (three dates screen under two premises). Individual screen hits: P1 invented children (5), P2 blue-water sailor persona (4), P3 treatment-history misattribution (4), P4 invented adult car accident (5), P5 Evernote role inflation (5), P6 AntarctiConf "on the ice" fiction (3), P7 first-keynote misattribution (3), P8 education misattribution (3, reversed on author adjudication and retained only as a correction-trail record), P9 invented named third parties (2).
- Ground truth is thin and personal. Of 366 audit rows, 114 (31.1%) cite no independent source at all. The 252 sourced rows carry 333 source-class mentions (multi-label): memoir After the Unicorn 134 (53% of sourced rows), anecdote ledger 94 (37%), memoir Birth of a Unicorn 68 (27%), asserted-real settings without a named source 16, with music catalogue, web, knowledge-base, and talk-transcript citations in single digits. The memoirs and ledger carry over 80% of the verification load.
- Binary reliability replicates; the four-way taxonomy does not fare as well. On VERIFIED vs. not across two independent 60-day samples (seed 20260821), raw agreement was 95.0% (rater B, 57/60) and 98.3% (rater C, 59/60); rater C found zero corroborated scenes. Both independent ratings place the failure rate at or above the original audit's level (B: 58/60 = 96.7%; C: 59/60 = 98.3%), providing no evidence the original rate was inflated.
- Exact four-way agreement: 63.3% (κ = 0.387, "fair") for rater B and 81.7% (κ = 0.573, "moderate") for rater C. The "clean," fully blinded rater agreed more than the partially exposed one.
- Named finding: the WEAK cell is doing two jobs. Of rater B's 22 disagreements, 21 (95%) involve WEAK as one of the two labels: WEAK–UNVERIFIED 13 (59%), CONTRADICTED–WEAK 5 (23%), VERIFIED–WEAK 3 (14%), CONTRADICTED–UNVERIFIED 1 (5%). The worked example is January 22 ("At Evernote, I was working 80-hour weeks") — the employer is documented, the figure appears nowhere in the record, and both original and re-raters' labels are defensible under a rubric that never defines "setting."
- Recency objection answered with data. Re-running generation on the same 60-day sample with the recovered 2025 inputs, both 2026 models fail verification in 100% of sampled days ungrounded: Arm B (openai/gpt-5.4) and Arm B2 (anthropic/claude-sonnet-5), with zero corroborated scenes across the 120 combined ungrounded days (B vs. B2 p = 1.0). Comparison against the 2025 originals (Arm A1, 98.3% failure) is statistically indistinguishable (A1 vs. B p = 1.0; A1 vs. B2 p = 1.0).
- Grounding is the only intervention that moved the measure. Arm C (gpt-5.4 given verbatim corpus excerpts and told to use only corroborated material) drops verification failure from 100% to 83.3% (B vs. C p = 0.0013; B2 vs. C p = 0.0013), with corroborated scenes rising from 0 to 10 (A1 vs. C p = 0.0084) — yet even grounded generation leaves five of six days below the verification bar.
- A first rating pass failed the blindness standard and was discarded. It was performed by the same agent that generated the entries while holding the arm map. Agreement between the compromised pass and the independent blind pass was 74.4% exact (κ = 0.383, n = 180): headline conclusions unchanged, but per-item labels diverged substantially, so the non-blind pass should be relied on for nothing finer than the binary rate.
Methodology in Plain English
Three corpora were fixed before the paper's analysis session (August 21, 2026), with the audit labels themselves predating it by five weeks.
What was audited. A 366-entry page-a-day book (~543 words per day; 198,949 words total), drafted for a gift edition and documented in its own repository as "AI-padded." Provenance recovery identified the generator of record: OpenAI o3-pro designed the daily-entry template ("Daily Inspiration Template," 31 messages, July 7–17, 2025), and gpt-4o, o3, and o4-mini-high generated entries from the quote list in ChatGPT conversations (July–August 2025), with a Cursor composer run on August 1, 2025 executing a one-shot version of the same prompt. The audited unit is the day-entry, each organized around one anecdotal scene plus a lesson. A derived ~175-word-per-page edition (64,198 words) exists, but the audit rows correspond scene-for-scene to the long-form entries, which the paper therefore treats as the audited text.
What it was checked against. An independently assembled ground-truth corpus: the memoir After the Unicorn with its editorial anecdote ledgers, the memoir Birth of a Unicorn, a 353k-word corpus of the subject's published writing and talk transcripts (HEATHER_CORPUS), a speaking record and knowledge base maintained for her professional digital twin, a music catalogue with dated life-era annotations, and targeted web checks. The calendar source, its page-a-day derivative, and the ChatGPT conversation backups were excluded as non-independent.
How verdicts were assigned. Each day received exactly one verdict under a rubric quoted verbatim from an audit summary dated July 13, 2026: VERIFIED ("specific scene corroborated"), WEAK ("real setting/person, invented specific scene"), UNVERIFIED ("no anchor; generic invention"), CONTRADICTED/flagged ("asserts something FALSE about her life"). The classes are ordered by evidential relation to the record, not severity alone. Four critical items were settled by explicit written rulings from the subject, and the summary's precedence rule binds all other rows.
How the numbers were produced. Two published scripts convert the audit tables into every number in the results. 01_parse_factcheck.py reads twelve month tables and emits one CSV row per day using fixed normalization rules in fixed order, reporting unparsed verdicts rather than silently dropping them; it output 366 rows with zero duplicate keys and zero unparsed verdicts. 02_analyze.py asserts n = 366, tallies verdicts overall and per month, computes Wilson intervals (z = 1.96), applies the nine fixed keyword screens over each row's gist and note text, codes Source cells against seven named source classes plus two fallbacks (multi-label), and writes stats.json, four figures, and two tables. The headline quantity, the verification-failure rate, is any day whose verdict is not VERIFIED.
Key caveats stated by the authors. The original 366 verdicts were assigned by an LLM auditor (Claude-based agents), not a human; only four items carry explicit author rulings. The paper therefore measures LLM confabulation using LLM-produced labels. VERIFIED is treated as a strict lower bound on true accuracy, so UNVERIFIED conflates "invented" with "real but unrecorded"; the word "fabricated" is reserved for the contradicted class. Blind re-raters establish inter-model consistency, not correctness. Arms in the replication differ in vendor and release date, A1 was rated in a separate batch by a different rater, and the 2025 artifact came from roughly 10 days of iterative human-guided chat plus a Cursor agent session while B/B2/C are single-shot API calls — so the replication supports the nulls above but not a clean model-generation effect. The paper also discloses that its own prose draft, analysis scripts, and literature-search assistance were drafted by an LLM ("ox-alpha," an anonymous stealth model served via OpenRouter, August 2026) under the human author's direction, and that a post hoc tie-break rule for the WEAK/UNVERIFIED boundary was derived after seeing the disagreements and was not applied to any label reported.
Why This Matters
Impact on research. Prior work benchmarks long-form factuality against encyclopedic knowledge (FActScore decomposes text into atomic claims; ReFACT annotates scientific confabulation) and evaluates person simulators on attitudes and survey answers. This paper supplies the missing episodic layer: even surface episodic content — employers, dates, named events — fails verification in the large majority of generated days. It also contributes a negative instrument finding: the WEAK/UNVERIFIED boundary in scene-level audits is unreliable for definitional reasons, not fixable rater bias, since the two re-raters erred in opposite directions relative to the original.
Real-world applications:
- Digital twins and synthetic-biography products: the paper argues fidelity claims should state which layer they cover — profile, attitudes, or episodes — and shows a curated twin knowledge base never supplied to the generator could not have kept the memoir honest.
- Ghost-writing, memoir services, and AI-assisted life-writing: the audit instrument localizes which days need remediation, making lesson-only rewriting and author adjudication cheap to target.
- Consent, privacy, and posthumous-memory governance: the study was caught and measured before publication, and the positionality section (author = subject) raises self-audit implications the paper flags as its own topic.
- Editorial fact-checking pipelines: the two published scripts, the pre-stated rubric, and the keyword-screen taxonomy offer a reproducible pattern for auditing generated personal narrative against archival records.
Industry relevance. The finding that no improvement was observed across two vendors in 2026 under reconstructed documented-input conditions — 100% verification failure for both ungrounded arms — argues that model recency alone does not solve this class of error, while corpus grounding produces a significant but partial improvement (100% → 83.3%). For teams shipping AI memory, companion, or avatar features, the practical implication is that the failure mode is structural rather than sporadic: with 96.7% of days failing, days with corroborated scenes were the exception, not corrections of a normally reliable process.
Future Directions
- A blinded human adjudication sample. The paper flags this as the obvious next step, noting its unusual design makes it feasible because the author is the subject. It is listed as future work, not claimed here.
- Evaluating the remediation workflow at scale. The lesson-only rewrite and adjudication precedence were specified during the project; their effect on reader trust is unmeasured. Only the retrieval half of the workflow was evaluated, at the generation stage and under fully-blind rating.
- Fixing the instrument's WEAK/UNVERIFIED boundary. A post hoc tie-break rule is proposed — a day is WEAK only if a specific, named, datable entity from the record appears inside the asserted scene; naming a real employer or city as mere backdrop is UNVERIFIED — but the paper notes it was derived after seeing disagreements and has not been applied to any reported label.
- Separating model effects from interaction pattern. The 2025 artifact came from iterative human-guided chat plus a Cursor agent session while the 2026 arms are single-shot API calls, and the per-day split of authorship among o3-pro, gpt-4o, o3, and o4-mini-high is not recoverable, so no failure rate can be attributed to a single model version.
- Extending beyond a single life. n = 366 days is one life, one corpus, one generator deployment; nothing here estimates a population fabrication rate, and no prior benchmark the authors are aware of uses a verified human-life corpus as ground truth.
Target Audience
Readers best served by this paper include: LLM evaluation and hallucination researchers who need a subject-grounded, scene-level complement to encyclopedic factuality benchmarks; developers and product owners building digital twins, synthetic biographies, AI memory features, or companion and avatar systems; editors, ghost-writers, and publishers working with AI-assisted personal narrative; and methodologists interested in audit-instrument design, inter-rater reliability, and the circularity of using LLM-generated labels to measure LLM behavior. The single-subject design also makes it unusually relevant to individuals who want to run a comparable audit on their own generated life story.
Authors’ abstract
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.