Skip to content
AI.info

Research

Humans Introduce, Models Elaborate: Asymmetric Narrative Agency in Human-LLM Co-Writing

Overview Research area: Human–Computer Interaction and computational linguistics, specifically the interactional dynamics of turn-based collaborative storytelling and creative writing with large langu

arXiv
2609.07920
Published
2026-09-07
Authors
Halfdan Nordahl Fundal, Yuri Bizzoni, Charlotte Gjørup Bilde, Ida Bække Johannesen, Rebekah Baglini

AI summary

Overview

Research area: Human–Computer Interaction and computational linguistics, specifically the interactional dynamics of turn-based collaborative storytelling and creative writing with large language models.

Technical level: Advanced. The paper uses mixed-effects models, bootstrap confidence intervals, nonparametric tests, concept vector projection for sentiment, and information-theoretic surprisal measures (novelty, transience, resonance) computed with a separate base language model.

Scope: A matched comparison of Human–Human (HH), Human–LLM (HA), and LLM–LLM (AA) turn-based co-writing, analyzing how affect, semantic distance, and narrative influence are distributed across turns.

What This Paper Is About

Most research on human–LLM co-writing evaluates final text quality, productivity, or user experience, leaving open the question of how a co-written story came to be. The authors compare three dyadic conditions on an identical storytelling task to test whether Human–LLM collaboration is an intermediate case between Human–Human and LLM–LLM collaboration, or a distinct interactional regime. They measure, turn by turn, who introduces new material, whose contributions persist, and how strongly each partner adapts to the other.

Key Contributions

  1. A matched comparative dataset. The authors build a corpus of turn-based stories under three conditions — HH (36 stories, 360 exchanges, 720 turns), HA (97 stories, 873 exchanges, 1746 turns), and AA (80 stories, 800 exchanges, 1600 turns) — using identical instructions across conditions, so any role specialization in HA reflects the pairing rather than the prompt.

  2. A turn-level measurement framework. They combine concept-vector-projection valence scores, directional valence alignment, adjacent semantic distance, and surprisal-based Novelty, Transience, and Resonance to quantify affective adaptation and directional narrative influence in co-writing.

  3. Evidence that HA is a distinct regime, not a midpoint. Metric-level slot asymmetry in HA exceeds both same-type baselines for novelty, transience, and resonance, and all planned HA-versus-baseline contrasts are significant (p_BH < .001).

  4. A directional account of agency. Within HA, human turns are more novel and persist more strongly into the following narrative, while LLM turns elaborate and stabilize the existing context.

Main Findings

  • HA shows the largest baseline affective gap. Baseline valence asymmetry was higher in HA (mean A_baseline = 0.060, 95% CI [0.051, 0.069]) than HH (mean = 0.041, 95% CI [0.032, 0.052]) and AA (mean = 0.028, 95% CI [0.024, 0.033]); the planned contrast was significant (Δ = 0.025, 95% CI [0.014, 0.036], p < .001).

  • Directional alignment asymmetry is weak across conditions. Unsigned per-story alignment asymmetry was highest in HA (mean A_valence = 0.608, 95% CI [0.518, 0.705]) versus HH (mean = 0.521, 95% CI [0.399, 0.646]) and AA (mean = 0.496, 95% CI [0.418, 0.578]), but cross-condition tests were not reliable (ANOVA p = .190; Kruskal–Wallis p = .365; planned contrast p = .106).

  • Within HA, the LLM adapts more than the human. The mixed-model directional slope gap showed stronger LLM alignment than human alignment (Δ = 0.143, p = .008), and a signed per-story diagnostic showed the same direction (mean signed Δz = 0.194, p = .012). Corresponding signed slot tests were not reliable for HH or AA.

  • HA turns diverge most semantically. Adjacent turn-to-turn semantic distance was highest in HA (mean = 0.316), followed by AA (mean = 0.296) and HH (mean = 0.287). HA distances exceeded HH (Δ = 0.0299, p < .0001) and AA (Δ = 0.0198, p < .0001); the planned contrast was significant (Δ = 0.0248, p < .0001).

  • Narrative-role asymmetry is largest in HA. Slot asymmetry means were novelty 0.828 (HA) vs 0.355 (HH) and 0.209 (AA); transience 0.484 (HA) vs 0.300 (HH) and 0.197 (AA); resonance 1.290 (HA) vs 0.605 (HH) and 0.313 (AA). All planned HA-versus-baseline contrasts were significant (p_BH < .001).

  • Novelty couples most steeply with resonance in HA. The novelty-to-resonance slope was 1.217 in HA (95% CI [1.184, 1.250]), 1.066 in HH (95% CI [1.003, 1.128]), and 0.956 in AA (95% CI [0.913, 0.999]). Both HH and AA slopes were flatter than HA.

  • Humans introduce, models elaborate. Within HA, 30.2% of human turn-information was explained by prior context (mean Novelty = −1.82) versus 52.8% for LLM turns (mean Novelty = −2.65, p < .0001). Human turns produced a 30.0% reduction in predictive surprisal of the following turn versus 12.0% for LLM turns (p < .0001), giving higher mean net resonance for humans (−0.796 vs −2.07).

  • Partner identity changes how much a turn is taken up. Keeping the speaker constant, human turns gained 12.8 percentage points more uptake with an LLM partner than with a human partner, while LLM turns dropped by 26.4 percentage points when writing with a human instead of another LLM. Turns followed by human writers showed low uptake (12.0%–17.2%), whereas turns followed by LLMs yielded roughly double (30.0%–38.4%).

  • The agent × novelty interaction was not significant. Novelty predicted resonance for both agents, descriptively more strongly for humans, but the interaction term was β₃ = −0.058, p = .083.

  • Surface and lexical explanations were ruled out. Lexical overlap showed no significant asymmetry (0.052 vs 0.054, p = .53); LLM turns contained a higher rate of unused content words and were longer, meaning human novelty emerges despite opposite surface trends. Surface metrics (word count, sentence length, lexical diversity, readability, first-order coherence) were significantly larger in HA than HH and AA (PERMANOVA, R² = 0.179, p < .001), identifying asymmetry but not its direction.

  • Shuffled-context control preserved the gap. Recomputing turns against a random story's context collapsed predictability (G_shuffled, human = 0.140 bits/token; G_shuffled, AI = 0.318 bits/token, against true context gains of 1.82 and 2.65 bits/token), preserving the HA gap (G_content, AI − G_content, human = +0.669 bits/token).

  • Scorer familiarity does not explain the dynamics. Stratifying AA by model family, unconditional base surprisal ranged from 4.27 bits/token (Qwen) to 5.91 bits/token (Claude), but proportional context gain stayed between 50% and 64%, and all generator families exceeded both human baseline conditions (28%–33%).

Methodology in Plain English

The researchers ran the same turn-based storytelling task under three pairings: two humans, a human and an LLM, or two instances of the same LLM. Every agent received identical instructions to continue a story from where their partner left off, with 10 exchanges per story. The HA data came from an existing corpus (Fundal and Bizzoni, 2026); the HH and AA conditions were newly constructed to match its structure. The four models used across HA and AA were gpt-4.1-2025-04-14, claude-sonnet-4-5-20250929, Llama-3.3-70B-Instruct, and Qwen2.5-72B-Instruct.

Each turn was converted into a sentiment score by embedding it with paraphrase-multilingual-mpnet-base-v2 and projecting the embedding onto a pretrained sentiment concept vector, which places the turn on a negative–positive valence axis. Alignment was estimated as the within-story correlation between one agent's turn valence and the partner's next turn valence, computed in both directions, and asymmetry was the absolute difference between the two directions. Semantic distance was the cosine distance between adjacent turns, embedded with Kingsoft-LLM/QZhou-Embedding.

For narrative influence, the team used a separate base model, google/gemma-4-31B — chosen because it belongs to a different model family from all four generation models — to compute how surprised it was by each turn. Novelty compared a turn's surprisal given the full prior story against its surprisal from the beginning-of-sequence token. Transience measured how much the turn reduced the surprise of the immediately following turn beyond what was already established. Resonance was novelty minus transience, capturing whether a surprising contribution stuck. Statistical comparisons used mixed-effects models with story-level random intercepts, bootstrap confidence intervals, nonparametric tests, and Cliff's delta, since the conditions differ in sample size (HH is the smallest cell).

Why This Matters

Impact on research. The paper reframes human–LLM co-writing from an output-quality question to an interactional one, and shows that HA collaboration is not simply a halfway point between HH and AA. It provides a reusable turn-level framework for measuring how narrative agency is distributed, and it cautions against treating "human–AI" and "human–human" collaboration as interchangeable baselines.

Real-world applications:

  • Writing-assistant design. Knowing that models predominantly elaborate rather than originate suggests interfaces could deliberately surface or protect human-introduced material that is at risk of being diluted.
  • Creative-writing pedagogy. Instructors using LLM partners can anticipate that students retain authorship of novel plot material while the model consolidates direction, which matters for assessing student contribution.
  • Collaborative storytelling platforms and games. Systems that pair a player with an AI narrator can be tuned knowing that AI turns are taken up roughly twice as often as human turns when they follow.
  • Evaluation of generative writing tools. Productivity and surface-quality metrics (word count, lexical diversity, readability) are shown to flag asymmetry in HA without indicating its direction, so richer interaction metrics are needed.

Industry relevance. Teams building co-writing products, educational writing tools, or narrative-generation systems have a concrete reason to measure turn-level influence rather than only final text quality. The finding that LLMs adapt more than humans within HA, and that LLM contributions are consumed more strongly when they follow, has direct implications for how much narrative steering an AI co-writer silently exerts.

Future Directions

  1. Test longer and more varied interactions. The paradigm is capped at ten exchanges, participants were mostly students at a single university, and future work should examine longer horizons, more diverse participant populations, and different writing genres.

  2. Extend to other dyadic text settings. Applying the framework to argument, negotiation, or collaborative explanation would establish how general these interactional asymmetries are.

  3. Investigate what LLMs do when they elaborate. The authors ask whether models steer narratives toward latent structural or archetypal patterns encoded in their representations.

  4. Expand the HH sample and separate model-specific effects. The HH condition has the smallest cell, and the HA and AA conditions pool stories across four LLM families, which increases ecological validity but may obscure model-specific dynamics.

Target Audience

Researchers in human–computer interaction, computational linguistics, and digital humanities who study human–AI collaboration, creativity, or dialogue dynamics; developers and designers of co-writing and narrative-generation tools; and scholars in writing studies or education interested in how authorship and agency are redistributed when an LLM joins the writing process. Readers without a background in mixed-effects modeling or information-theoretic surprisal will find the framing accessible but the methods demanding.

Authors’ abstract

Human-LLM co-writing is increasingly used for open-ended text generation, but much prior work focuses on final outputs rather than the interactional dynamics through which stories are produced. We study turn-based collaborative storytelling across three matched conditions: Human-Human (HH), Human-LLM (HA), and LLM-LLM (AA). Using a shared storytelling paradigm, we measure how agents align, introduce novel material, and influence narrative development through turn-level measures of valence adaptation, semantic novelty, transience, and resonance. Our results show that HA co-writing is not intermediate between HH and AA collaboration. Instead, it displays a distinctive asymmetry where humans tend to introduce more novel and persistent narrative material, while LLMs tend to elaborate and stabilize the existing context. These findings suggest that, in this setting, LLMs function less as human co-authors and more as adaptive narrative amplifiers that reshape how agency is distributed in collaborative writing.

Read the original paper