Skip to content
AI.info

Research

PREFAB: PREFerence-based Affective Modeling for Low-Budget Self-Annotation

Overview Research area: Affective computing and human-computer interaction (HCI), specifically retrospective self-annotation of emotion, with connections to preference learning, user modeling, and per

PREFAB: PREFerence-based Affective Modeling for Low-Budget Self-Annotation
arXiv
2601.13904
Published
2026-01-20
Authors
Jaeyoung Moon, Youjin Choi, Yucheon Park, David Melhart, Georgios N. Yannakakis, Kyung-Joong Kim

AI summary

Overview

Research area: Affective computing and human-computer interaction (HCI), specifically retrospective self-annotation of emotion, with connections to preference learning, user modeling, and personalization.

Technical level: Advanced. The paper combines a transformer-based Siamese network, Feature-wise Linear Modulation (FiLM), ordinal cross-entropy, dynamic time warping clustering, and a controlled user study with qualitative interviews.

Scope: The paper proposes PREFAB, a low-budget retrospective self-annotation method that uses an ordinal preference-learning model to identify affective inflection regions so that participants annotate only selected clips instead of an entire session, then evaluates it with a model-performance study and a 25-participant user study using the AGAIN dataset and the PAGAN annotation tool.

What This Paper Is About

Self-annotation — people labeling their own emotional states, usually after a task — is the gold standard for collecting affect labels, but full annotation is slow, cognitively demanding, and prone to fatigue and memory errors. The authors' goal is to keep annotation quality high while asking people to label much less of a session, by having a model predict where the affectively important moments are and asking users to annotate only those.

Key Contributions

  1. A new low-budget self-annotation method. PREFAB (PREFerence-based Affective modeling for low-Budget self-annotation) performs preference-learning-based selective annotation focused on peak-end inflection points, grounded in the peak-end rule and ordinal representations of emotion.
  2. A model-performance demonstration. Through model performance evaluation, the authors show that the ordinal approach outperforms naïve sampling (random, uniform), a heuristic (rule-based event-driven) sampling method, and cardinal modeling (regression) for modeling affective inflections.
  3. A 25-participant user study. The study shows PREFAB significantly reduces mental and physical workload, conditionally reduces temporal cost, and increases annotation confidence compared to full annotation, while preserving annotation quality.
  4. A complementary qualitative interview study. The authors provide in-depth insight into how participants perceived PREFAB and derive design implications for future affective self-annotation systems.

Main Findings

  • Ordinal modeling wins on inflection detection (RQ1). Across all nine games in the evaluation, PREFAB provided the best combination of region-level F1 score and time-efficiency alignment (ΔTE) compared with random sampling, uniform sampling, rule-based event-driven heuristic sampling, and cardinal modeling (regression). The specific F1 and ΔTE values are not reported in the provided excerpt.
  • Workload is reduced (RQ2). The user study shows PREFAB mitigates workload relative to full annotation, and the abstract specifies this as both mental and physical workload.
  • Temporal burden is reduced only conditionally (RQ2). Both the abstract and the results summary describe the reduction of temporal burden as conditional rather than universal.
  • Confidence improves without quality loss (RQ4). PREFAB improves annotator confidence and does not degrade annotation quality, even though the remainder of the affect trace is interpolated rather than annotated.
  • Interpolation is a viable substitute for annotation. Because the peak-end rule holds that retrospective evaluations are disproportionately shaped by the most intense moment and the ending, the authors assume non-annotated periods can be approximated by interpolation from surrounding inflection regions; the user study supports this without degrading quality.
  • Two design elements shaped the method. The model conditions its latent space on participant biographical data via FiLM, and trains a secondary task that classifies segments into four arousal-trend clusters derived from dynamic time warping clustering; the paper states the effects of FiLM and the auxiliary task are detailed in Appendix A, which is not included in the provided content.
  • Exact statistical results are not reported here. Numeric F1 scores, ΔTE values, and the significance tests behind the workload and confidence claims are not given in the provided excerpt.

Methodology in Plain English

The method has four stages.

Step 1: Predict the affective trajectory. Data recorded during a main task is fed into a trained model. The dataset is the AGAIN dataset, which contains play logs, gameplay videos, self-annotated arousal, and biographical information from over 120 participants who played nine different games; game logs were cleaned to 4 timesteps per second, and arousal was collected with the PAGAN tool. Biographical data (age, gender, nationality, dominant hand, gaming frequency, gamer type, preferred platform, favorite games) is quantized and used as conditional input — 86 favorite game titles were converted into 32 game genres using MetaCritic genre codes. For learning, the authors build input segments from a 3-second window with a 1-second temporal gap, each consisting of 12 frames of images, game log features, and biography. Paired labels are three-class based on arousal change between time i and j = i+4: increase, no change, or decrease. The model is a Siamese transformer: CNN and MLP extractors compress images and game logs, the concatenated vectors pass through a transformer encoder with input dimension 200 and two stacked layers, and FiLM rearranges the latent space using personal information. Two MLP heads then predict the relative affect value (main task) and the arousal trend cluster (auxiliary task), with the auxiliary loss scaled by α = 0.001. Training uses ordinal cross-entropy with cut-points set to [-1, 1]. At inference, the model is applied to single segments rather than pairs, and the predicted values are aggregated to reconstruct a full arousal-change trajectory.

Step 2: Detect inflection regions. The reconstructed arousal graph is passed to the find_peaks function of the scipy library with default hyperparameters; local minima are found by applying it to the inverted graph. Because find_peaks can miss subtle changes in long flat or gently sloped regions, the authors add a rule that includes segments with significant gradient change. Each inflection point becomes the center of a 5-second annotation clip, extending 2.5 seconds before and after. The 5-second window is justified by a reported reaction lag of about 1–3 seconds between perception and labeling, and evidence that humans generally need at least 3 seconds of exposure to comprehend video content. Overlapping regions from adjacent inflection points are merged into one longer region.

Step 3: Selective annotation. Users label only inside the selected clips. A preview mechanism that provides brief contextual cues was also introduced to assist annotation, and the user study tested PREFAB with and without preview against a full-annotation baseline.

Step 4: Interpolation. Unannotated segments are filled in by linear interpolation, formalized as Algorithm 1 (cumulative slope propagation): the slope of the final half of an annotated clip is computed and propagated forward until the next annotated clip, and the propagated value becomes the offset for the next region's annotated values, repeating until the end of the task.

Evaluation. A technical evaluation compared region-level F1 and ΔTE for PREFAB against random and uniform sampling, a rule-based event-driven heuristic, and cardinal regression. A user study compared three conditions — Baseline (full-annotation), PREFAB without preview, and PREFAB with preview — across four research questions.

Why This Matters

Affect labels are the bottleneck for training and validating any emotion-aware system, and the current gold standard costs roughly twice the duration of the main task plus substantial cognitive effort. PREFAB shows that a model can point to a small number of important moments and that labels collected only there, with interpolation for the rest, preserve quality while lowering workload and raising confidence. This shifts self-annotation from "label everything" toward "label what matters," and does so without the intrusive physiological sensors and unhandled remainder that limited earlier selective-annotation work.

Real-world applications:

  • Game user research and playtesting: Studios recording player arousal or engagement from gameplay logs could cut annotation sessions while still capturing the emotional beats that matter, since the evaluation used nine commercial games.
  • Crowdsourced affect labeling platforms: Tools like PAGAN could integrate inflection detection to pay annotators for fewer, more decisive segments, reducing cost per labeled hour.
  • Training, education, and exercise applications: Retrospective annotation suits active, goal-directed activities where in-situ prompting disrupts immersion, so lower-budget annotation makes affect modeling practical in these settings.
  • VR and multimedia experience evaluation: Trajectory-based retrospective tools for reviewing a recorded experience could adopt the peak-end-driven region selection to shorten review time.

Industry relevance: The method targets the data-collection economics of affective computing — less annotator time per session, less fatigue-driven error, and therefore cheaper, more scalable label sets for products such as adaptive games, personalization systems, UX analytics, and wellness applications that depend on emotional ground truth.

Future Directions

  • Generalize beyond games and arousal. The evaluation used the AGAIN dataset of nine games and self-annotated arousal; whether the approach holds for other affects (valence, frustration, immersion) and other task types such as training or education is not established in the reported work.
  • Reduce dependence on the model's trajectory estimate. Because unannotated regions are interpolated from the model's reconstruction, errors in the predicted trajectory may propagate; the paper notes that prior selective-annotation approaches left non-annotated segments unhandled, and the robustness of interpolation under poor predictions remains open.
  • Tune the annotation window and preview design. The 5-second clip is described as an assumed "lower bound," and the preview variant was included to compensate for individual differences in recognition and comprehension ability, leaving room to study window length and preview content systematically.
  • Remove reliance on quantitative biographical input. The model conditions on quantized biography, including a mapping of 86 favorite game titles into 32 MetaCritic genres; how much personalization this adds, and whether it is necessary, is a design question the paper defers to Appendix A and future work.

Target Audience

Researchers and practitioners in affective computing and HCI who collect or use self-reported emotion data — especially those building annotation tools, designing user studies with retrospective labeling, or applying preference learning and ordinal models to human affect. It is also relevant to game user researchers, crowdsourcing platform designers, and machine-learning engineers who need to cut the cost of obtaining reliable emotional labels without sacrificing data quality.

Authors’ abstract

Self-annotation is the gold standard for collecting affective state labels in affective computing. Existing methods typically rely on full annotation, requiring users to continuously label affective states across entire sessions. While this process yields fine-grained data, it is time-consuming, cognitively demanding, and prone to fatigue and errors. To address these issues, we present PREFAB, a low-budget retrospective self-annotation method that targets affective inflection regions rather than full annotation. Grounded in the peak-end rule and ordinal representations of emotion, PREFAB employs a preference-learning model to detect relative affective changes, directing annotators to label only selected segments while interpolating the remainder of the stimulus. We further introduce a preview mechanism that provides brief contextual cues to assist annotation. We evaluate PREFAB through a technical performance study and a 25-participant user study. Results show that PREFAB outperforms baselines in modeling affective inflections while mitigating workload (and conditionally mitigating temporal burden). Importantly PREFAB improves annotator confidence without degrading annotation quality.

Read the original paper