Skip to content
AI.info

Research

Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning

Overview Research area: Human-Computer Interaction, specifically persuasive technology, serious games, conversational agents, and sustainability education. Technical level: Intermediate. The statistic

Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning
arXiv
2602.17905
Published
2026-02-20
Authors
Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, Vered Shwartz

AI summary

Overview

Research area: Human-Computer Interaction, specifically persuasive technology, serious games, conversational agents, and sustainability education.

Technical level: Intermediate. The statistical methods (Kruskal–Wallis, Mann–Whitney U, ordinal logistic regression) are applied rather than derived, and the study design is described in accessible detail, but readers should be comfortable with nonparametric tests and self-report scales.

Scope in one sentence: A controlled between-subjects user study (43 participants, two sustainability topics) comparing a static essay, an LLM chatbot, and a narrative text-based game built from identical arguments and facts, to see how delivery format alone shapes subjective experience, perceived persuasion, and 24-hour delayed knowledge retention.

What This Paper Is About

Interactive systems such as chatbots and games are increasingly used to teach and persuade people about sustainability, but most evaluations rely on what users say they learned or felt rather than what they actually retained. Because content and interaction are usually confounded, it has been hard to tell whether a format's persuasive power comes from its information or from its interface. The authors isolate format by holding the arguments and facts constant across all three conditions, then measure subjective experience, perceived attitude change, and objective knowledge after a 24-hour delay.

Key Contributions

  1. PersuLab, a shared experimental system that delivers identical argument–fact content through three formats: a pre-generated essay, a free-form chatbot, and a narrative text-based game. It enforces complete content exposure in every condition (all five predefined facts must be presented before the session can end in the interactive conditions) and logs all interactions. The code is available at https://github.com/salavi/persulab.

  2. A controlled between-subjects comparison with a single-factor three-level design (essay, chat, game), two topics (recycling, public transit), and a fixed set of five argument–fact pairs per topic, so that only delivery format varies.

  3. Evidence of a dissociation between felt and actual learning: the chatbot dominates subjective measures and raises perceived importance, while the text-based game is rated lowest in subjective learning yet produces the highest delayed quiz scores.

  4. Exploratory analysis of interaction logs (turn counts, word counts, word ratios, session duration, reaction time) suggesting that common engagement proxies track subjective impressions more closely than they track actual knowledge.

Main Findings

  • Chatbot wins on subjective measures. Across subjective dimensions, the chatbot condition was consistently rated highest and increased perceived importance of the topic. The game condition tended to receive lower ratings, with the essay generally falling between the two interactive conditions depending on the measure.

  • Ease of following differed significantly by mode. Essay and chatbot tied at the top (both Mean = 4.64), while the game was lower (Mean = 4.20). A Kruskal–Wallis test found a significant effect of delivery mode (p = 0.0497). The provided content is truncated immediately after this, before the pairwise Mann–Whitney U results and the remaining subjective measures are reported.

  • Perceived learning diverged from objective retention. Participants in the text-based game condition reported learning less than those who read essays, yet scored higher on the delayed (24-hour) knowledge quiz. This is the paper's headline finding.

  • Engagement proxies track feelings, not learning. Exploratory analyses suggest that verbosity and interaction length are more closely related to subjective experience than to actual learning. As these analyses were not pre-registered and involved multiple comparisons, the authors report them descriptively and without multiple-comparison correction.

  • Topics behaved similarly on persuasiveness. Perceived convincingness of arguments was comparable between recycling (Mean = 3.82) and public transit (Mean = 3.76), which justified pooling topics for the mode-level comparisons.

  • Subjective ratings were high overall. Mean ratings were generally above the midpoint of the 5-point scale (3) for most measures in all three conditions.

  • Detailed inferential results for perceived change, the knowledge quiz, and the interaction logs (Sections 6.2–6.4) are not included in the provided content, so their specific effect sizes, p-values, and directional patterns across all measures cannot be reported here.

Methodology in Plain English

Design. A between-subjects study where each person experienced exactly one delivery mode and one topic. Topic assignment was balanced across conditions.

Content held constant. Two topics were chosen: recycling and public transit. Each had five persuasive arguments, each paired with one supporting fact. For example, the recycling condition included facts such as "Recycling one ton of paper saves approximately 17 trees" and "Recycling aluminum saves up to 95% of the energy compared to producing new aluminum." The public transit condition included facts such as "commuters can save $10,000+ per year by using public transit instead of driving" and "Every $1 invested in public transit generates approximately $4 in community benefits."

The three conditions.

  • Essay: GPT-4.1 generated persuasive essays from the fixed facts and arguments. Twenty essays were generated per topic and one was randomly sampled per participant. Argument coverage was verified automatically by an LLM and manually on a random sample. Because the essay is non-interactive, no runtime fact tracking was needed. The paper states a minimum exposure period of 60 seconds (Section 3.3); the Figure 1 caption describes the End Session button activating after three minutes of reading.
  • Chatbot: Participants held a free-form conversation with an LLM that delivered the same facts conversationally in response to user input, controlling pacing and order themselves.
  • Text-based game: A narrative-driven game in which the participant plays a protagonist, reads narrative text, and responds either by choosing one of three predefined options or typing a custom action. The same facts were woven into the story.

Keeping the comparison fair. In both interactive conditions, a moderator module built prompts and called the language model to generate each turn, and an automated fact-checking module tracked which target facts remained uncovered and fed that back into later prompts. The End Session button only activated once all facts had been presented. Participants could continue interacting afterward, up to a maximum interaction time of 25 minutes.

Participants. 45 were recruited via university advertisements and word of mouth; 43 were analyzed (one did not complete all steps, one took part in a dry run). Most were aged 18–34 (28 participants aged 18–24, 11 aged 25–34, 3 aged 35–44, 1 aged 45–54). Condition counts were chat 14, essay 14, game 15; topic counts were recycling 22 and public transit 21.

Measures. A pre-study questionnaire captured baseline importance, behavioral intention (recoded to a 1–5 ordinal scale), epistemic confidence, and demographics. A post-study questionnaire captured perceived change in importance, behavioral intention, and belief in effectiveness (Less / Same / More / Not sure), plus 5-point Likert ratings of ease of following, engagement, enjoyment, trust, self-reported learning, satisfaction, convincingness, increased motivation to act, influence on thinking, willingness to recommend, and desire to re-encounter similar experiences. Twenty-four hours later, participants took a multiple-choice quiz with five content-covered questions and two control questions about information never presented; only the five content questions counted toward the score. Confidence ratings were collected but not analyzed. Interaction logs recorded turns, words per turn, word ratios, session duration, and reaction time.

Analysis. Kruskal–Wallis tests for omnibus differences across the three modes, followed by Mann–Whitney U tests for pairwise comparisons (chosen because Likert data are ordinal and may violate normality assumptions). Perceived change was modeled with ordinal logistic regression using a logit link and BFGS optimizer, with Wald z-tests for significance. Robustness checks re-ran importance and behavioral intention models with pre-study scores as covariates, which did not qualitatively change the pattern (Appendix H). Interaction log analyses used Spearman rank-order correlations and were treated as exploratory. All tests were two-sided with α = .05.

Why This Matters

Impact on research. The study is a clean demonstration that self-reported learning and engagement are unreliable stand-ins for knowledge retention. It reinforces prior findings of a perceived-versus-actual learning gap (e.g., Persky et al.'s study of 277 students and Aperapar Singh and Anthonysamy's study of 382 participants, both cited) and extends it specifically to comparisons between interaction modalities with content held fixed. For HCI and learning-science researchers, it argues that evaluations resting only on engagement or perceived learning risk drawing the wrong conclusion about which system teaches better.

Real-world applications.

  • Designing sustainability education campaigns where the goal is durable knowledge, not just a pleasant session.
  • Building serious games and games-for-change, where the paper warns that low felt learning does not mean low actual learning.
  • Developing LLM-based conversational agents for public health, civic, or environmental outreach, where chatbots excel at subjective persuasion.
  • Selecting formats for corporate or community training on recycling, transit, energy, and similar actionable behaviors.

Industry relevance. Organizations deploying LLM-powered chatbots and interactive games at scale have a direct stake in knowing that interactivity and persuasion are separable from retention. The PersuLab architecture, with its automated fact-coverage tracking that guarantees every intended message is delivered before a session can end, is a practical pattern for anyone who needs to guarantee message exposure in an open-ended generative interaction.

Future Directions

  • Larger, better-powered samples. The authors note that splitting data by both topic and mode would yield cells of roughly 6–7 participants, limiting statistical stability; a larger study could test topic-by-mode interactions directly.
  • Pre-registered interaction analysis. The log-based correlations were exploratory, involved multiple comparisons, and received no formal correction; the authors frame them as hypothesis-generating and call for confirmatory follow-up.
  • Isolating the mechanisms behind the paradox. The paper raises trade-offs between interactivity, realism, trust, and cognitive load but does not identify which of these explains why game players felt they learned less while retaining more.
  • Broadening the design space. Testing other modalities, additional sustainability topics, longer retention delays than 24 hours, and designs that aim to raise both perceived and actual learning at once.

Target Audience

Researchers and practitioners in HCI, persuasive technology, and serious games; designers of LLM-based chatbots and interactive learning experiences; sustainability and climate educators; and learning scientists or methodologists interested in the gap between subjective self-reports and objective knowledge assessment. The paper is also useful to product teams deciding whether to invest in conversational or game-based formats for educational content.

Authors’ abstract

Interactive systems such as chatbots and games are increasingly used to persuade and educate on sustainability-related topics, yet it remains unclear how different delivery formats shape learning and persuasive outcomes when content is held constant. Grounding on identical arguments and factual content across conditions, we present a controlled user study comparing three modes of information delivery: static essays, conversational chatbots, and narrative text-based games. Across subjective measures, the chatbot condition consistently outperformed the other modes and increased perceived importance of the topic. However, perceived learning did not reliably align with objective outcomes: participants in the text-based game condition reported learning less than those reading essays, yet achieved higher scores on a delayed (24-hour) knowledge quiz. Additional exploratory analyses further suggest that common engagement proxies, such as verbosity and interaction length, are more closely related to subjective experience than to actual learning. These findings highlight a dissociation between how persuasive experiences feel and what participants retain, and point to important design trade-offs between interactivity, realism, and learning in persuasive systems and serious games.

Read the original paper