Skip to content
AI.info

Research

Re:Member: Emotional Question Generation from Personal Memories

Overview Research area: Natural Language Processing, Human–Computer Interaction, affective computing, and second-language (L2) learning technology. Technical level: Intermediate. Scope: The paper pres

Re:Member: Emotional Question Generation from Personal Memories
arXiv
2510.19030
Published
2025-10-21
Authors
Zackary Rackauckas, Nobuaki Minematsu, Julia Hirschberg

AI summary

Overview

Research area: Natural Language Processing, Human–Computer Interaction, affective computing, and second-language (L2) learning technology. Technical level: Intermediate. Scope: The paper presents Re:Member, an open-source system that turns short user-uploaded videos of personal memories into emotionally styled, spoken Japanese questions for language practice.

What This Paper Is About

Most language-learning tools rely on generic or de-contextualized content, which limits their ability to use the emotional and mnemonic power of a learner's own lived experiences. The authors build Re:Member to test whether emotionally expressive, memory-grounded interaction can support more engaging L2 learning by generating stylized spoken questions in the target language from a learner's personal videos. The system is offered as a "stylized interaction probe"—a design and technical demonstration rather than a validated learning intervention.

Key Contributions

  1. An open-source system (available at github.com/zackrack/Re-Member) that converts personal memory videos into emotionally voiced, interactive prompts for language learning.
  2. A modular pipeline combining Silero VAD-based segmentation, WhisperX transcription with word-level timing, 3-frame visual sampling via OpenCV, GPT-4o question generation, and Style-BERT-VITS2 emotional speech synthesis.
  3. An emotion-style selection mechanism over five fixed Japanese styles (cheerful, silent whisper, voiced whisper, neutral, late-night relaxed) with a re-query scheme designed to prevent repetition.
  4. Illustrative demonstrations on two videos recorded with Meta Ray-Ban Glasses, showing grounded questions and matched emotion styles, plus an explicit discussion of limitations around consent, privacy, and emotional safety.

Main Findings

  • Pipeline performance by design, not evaluation: The paper reports no user study or quantitative benchmark; the system "has not yet been evaluated with users" and is presented as a design and technical demonstration.
  • Example video (1): A 1 minute 31 second walk along Tokyo's Sumida River was segmented and analyzed into 13 moments, each receiving an emotion, a student question, and text-to-speech output.
  • Emotion variety in video (1): Styles were well-matched to the riverfront setting, with a majority in the gentle voiced whisper style interspersed with cheerful and late-night relaxed tones; all five available emotion styles appeared at least once.
  • Single illustrative moment: For a frame where the user points at Tokyo Sky Tree, the system generated the question 東京スカイツリーの高さはどのくらいですか? ("About how tall is Tokyo Sky Tree?") with the cheerful emotion (るんるん).
  • Example video (2): A 31 second clip of a user boarding a train in Japan yielded three distinct segments. Emotion styles balanced voiced whisper, silent whisper, and neutral speech; the short duration limited style range, but the variation mechanism avoided repetition.
  • Video (2) questions: Silent whisper — 試合が行われている場所での電車の利用はどのように便利ですか? ("How is using the train convenient near where the event is being held?"); Neutral — この電車の車内はどのように見えますか? ("What does the inside of this train look like?"); Voiced whisper — この電車の車両には特別な座席やスペースがありますか? ("Does this train car have special seats or areas?").
  • Design claim: Emotionally salient, personally meaningful content voiced in stylized speaking styles (e.g., a whisper from a friend, an excited exclamation) is proposed to create deeper engagement than synthetic neutrality.

Methodology in Plain English

Given a video, the system separates the audio and uses Silero VAD to find speech segments, merging segments when the silence between them is shorter than 0.7 seconds. Each segment is transcribed with WhisperX, which also supplies word-level timing so transcript and visuals stay aligned. For visual grounding, three frames are extracted per segment—one before, one during, and one after the segment midpoint—using OpenCV, and resized to a consistent format. GPT-4o then receives the transcript segment plus the frames and is prompted to act as a curious, friendly student asking one short, relevant question to the person who filmed the video, with instructions not to guess identity, age, gender, or names. A separate prompt asks the same model to act as an emotion label classifier choosing exactly one of five Japanese style labels, with temperature set to 1 and a short history of recent labels; if the chosen style matches either of the last two used, the model is re-queried up to five times. The question and style label are passed to a local Style-BERT-VITS2 model—trained from the Ami Koharune UTAU voicebank—to produce emotionally expressive Japanese speech. A Gradio interface lets users upload videos and browse three representative frames per segment alongside the question, emotion text, and styled audio playback.

Why This Matters

Impact on research: The paper connects HCI work on agent-assisted creativity and NLP work on context-sensitive and empathetic question generation, arguing for grounding LLM-generated questions in video-based emotion cues rather than generic content. It positions affect and personal media as central rather than incidental to learner-centered educational technology.

Real-world applications:

  • Second-language conversation practice built from a learner's own travel, family, or everyday recordings.
  • Memory-centered reflection tools that prompt users to revisit and talk about personally meaningful moments.
  • Affect-aware educational interfaces that adapt spoken output style to the mood of a scene.
  • A reusable modular pipeline that the authors state generalizes to other language settings beyond the current Japanese target.

Industry relevance: The work sits at the intersection of LLM-based content generation, expressive text-to-speech, and multimodal video understanding—capabilities relevant to consumer language-learning products, wearable camera ecosystems (the demos used Meta Ray-Ban Glasses), and personal-media applications that need to handle sensitive autobiographical data responsibly.

Future Directions

  • Adaptive selection of emotion styles rather than prompt-based classification alone.
  • More nuanced alignment between visual and emotional cues, including validation of emotion–scene alignment.
  • Interactive learner control over style and content during question generation.
  • Longitudinal deployments to measure how learners interact with memory-grounded prompts over time and whether affectively voiced questioning measurably improves learning.
  • Addressing the stated limitations: robustness to overlapping dialogue, background noise, and multilingual utterances, plus consent, emotional safety, and data privacy measures such as local-only processing and explicit opt-in.

Target Audience

Researchers and practitioners in NLP, HCI, affective computing, and educational technology, particularly those working on question generation, empathetic or context-sensitive dialogue, and expressive speech synthesis. It is also useful for language-learning product designers interested in personalized, multimodal content, and for readers seeking a clearly scoped design demonstration with openly stated limitations.

Authors’ abstract

We present Re:Member, a system that explores how emotionally expressive, memory-grounded interaction can support more engaging second language (L2) learning. By drawing on users' personal videos and generating stylized spoken questions in the target language, Re:Member is designed to encourage affective recall and conversational engagement. The system aligns emotional tone with visual context, using expressive speech styles such as whispers or late-night tones to evoke specific moods. It combines WhisperX-based transcript alignment, 3-frame visual sampling, and Style-BERT-VITS2 for emotional synthesis within a modular generation pipeline. Designed as a stylized interaction probe, Re:Member highlights the role of affect and personal media in learner-centered educational technologies.

Read the original paper