Research
The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends
Overview Research area: Cognitive neuroscience meets multimodal machine learning — specifically brain-tuning (fine-tuning a model to predict fMRI brain activity) of an audio-video transformer toward a

- arXiv
- 2511.07988
- Published
- 2025-11-11
- Authors
- Nico Policzer, Cameron Braunstein, Mariya Toneva
AI summary
Overview
Research area: Cognitive neuroscience meets multimodal machine learning — specifically brain-tuning (fine-tuning a model to predict fMRI brain activity) of an audio-video transformer toward a social cognition brain region, evaluated on social cognition tasks.
Technical level: Intermediate. The reader needs comfort with fMRI encoding models, ROI-based analysis, and multimodal transformer fine-tuning, but the paper is short and its method is described step by step.
Scope: A proof-of-concept study that brain-tunes the TVLT audio-video model to the Superior Temporal Sulcus (STS) using n=6 subjects watching Friends, and tests whether this improves brain alignment and downstream social cognition performance.
What This Paper Is About
Prior brain-tuning work had only been done on audio-only models tuned to language and auditory regions. This paper asks whether the same technique can be extended to a multimodal audio-video model, and whether targeting a social cognition region — the STS — can improve performance on tasks that require interpreting other people's intentions and mental states. The core problem is that multimodal AI models still lag behind humans in social perception, so the authors test brain activity as a supervision signal for closing that gap.
Key Contributions
- Extends the brain-tuning methodology from audio-only models to a multi-modal audio-video domain, tuning the joint audio-video transformer TVLT to fMRI data.
- Shows, for the first time, that brain-tuning a model to an ROI involved in social cognition (the STS) can increase alignment both to that ROI and to an adjacent lateral-stream ROI.
- Demonstrates that this brain-tuned model improves on a related social cognition task — sarcasm detection in sitcoms (MUStARD) — including when all Friends clips are removed from the evaluation.
- Reports a boundary condition: the gains do not generalize to a social cognition task in a markedly different context (sentiment and emotion prediction on CMU-MOSEI).
Main Findings
- Targeted alignment increases: Compared to both a pretrained baseline and a stimulus-tuned baseline, the n=6 brain-tuned models showed significantly increased alignment (p<0.05) to both subdivisions of the STS — the anterior STS (aSTS) and posterior STS (pSTS) — and to one of two neighboring lateral-stream ROIs (LOC). No significant change was found for EBA.
- Gains are not just stimulus exposure: No significant alignment changes were observed between the pretrained and stimulus-tuned baselines, confirming the improvement comes from the fMRI training objective rather than simply training on Friends.
- Downstream sarcasm detection improves: Brain-tuned models significantly outperformed baselines on MUStARD both on the full dataset (p<0.05) and on the subset with all Friends clips removed (p<0.01).
- No transfer to a different context: No improvements (and some decreased performance) were observed on CMU-MOSEI sentiment and emotion prediction.
- Emotion breakdown: Sadness was the only emotion with a significantly improved F1 score, but it also showed decreased accuracy (A2). The authors note sadness occurs in Friends but is not the show's dominant emotion.
- Voxel threshold constraint: The authors attempted to use the 0.4 cross-subject prediction accuracy threshold from prior brain-tuning work but found that beyond 0.25, all STS voxels were removed for some subjects. They set the threshold to 0.25, leaving subjects with 100–700 STS brain-tuning target voxels. Subject-05 had no remaining voxels above 0.25.
- Normalization caveat: Because cross-subject prediction accuracy underestimates the true noise ceiling, some normalized brain alignment scores exceed 1.0, but the authors state relative performance is unaffected by this scaling.
Methodology in Plain English
The researchers took TVLT, a pretrained audio-video transformer (12 encoder layers, embedding size 768, roughly 90M parameters, pretrained on around 130K hours of audio-video), and fine-tuned it to predict fMRI activity in the STS.
- Data: A subset of the preprocessed fMRI data from the 2022-alpha release of the Courtois Neuromod Dataset — n=6 subjects watching seasons 1–4 of Friends, with seasons 1–3 used for training and season 4 for evaluation.
- Noise handling: They estimated cross-subject prediction accuracy per voxel and filtered out unreliable voxels, tuning only on voxels reliably related to the stimulus. This measure was also used to normalize alignment scores.
- Training objective: For each fMRI time point, the model receives the previous 8 TR-lengths of audio-video stimulus (T = 11.92 seconds), matching the approximate hemodynamic response cycle. Video patches (16×16) and log-mel spectrograms are encoded jointly; the output tokens are mean-pooled and passed through a linear projection layer to predict the masked STS voxel vector. The L2 loss between predicted and true voxel activations is backpropagated through both the projection layer and the TVLT transformer layers.
- Training details: 68,063 TRs from seasons 1–3, TR = 1.49s, 10 epochs, 8 evenly sampled frames per clip, audio sampled at 44,100 Hz, Adam optimizer with constant learning rate 1.0×10⁻⁶. One model is tuned per subject.
- Compute: Each brain-tuning run uses 1 H100 GPU and 16 AMD EPYC 9654 CPUs on 244 GB of RAM, taking approximately 70 hours; each evaluation takes approximately 90 minutes.
- Encoding evaluation: Voxel-wise ridge regression models predict fMRI activations from concatenated [CLS] and mean-pooled tokens, using 8298 TRs from season 4 for training and 2630 for testing. Alignment is measured as voxel-wise Pearson correlation divided by cross-subject prediction accuracy and averaged within each ROI. Significance is tested with a Wilcoxon signed rank test (p<0.05).
- Downstream evaluation: A linear binary classifier is trained on concatenated [CLS] and mean-pooled tokens from the last layer. MUStARD is evaluated with mean performance across 10-fold cross validation; CMU-MOSEI uses the original 15,288/4,830 train-test split from the TVLT paper. Results are averaged across n=6 subject models, with one-sided one-sample t-tests (p<0.05 marked *, p<0.01 marked **) and error bars reporting SEM.
Why This Matters
Impact on research: The paper provides evidence that brain-tuning can be extended beyond audio-only models and beyond language/auditory regions to a multimodal model and a social cognition region. It also adds an important negative result — the benefits appear context-dependent rather than universally transferable — which sharpens the open question of when brain-derived supervision helps and when it does not.
Real-world applications:
- AI systems that better interpret social cues, such as sarcasm and intent, in conversational or media settings.
- AI-assisted therapy or communication tools that need to model a person's internal state.
- In-silico models of human social cognition that could help researchers study how social processing works in the brain.
- Socially aware media analysis, for example detecting sarcasm or emotional subtext in video content.
Industry relevance: The work connects to any product built on multimodal understanding of human behavior — assistants, content moderation, social robotics, and mental health technology — by suggesting that brain-derived objectives can shape model representations toward socially meaningful features. The authors also flag the dual-use risk that improved social cognition models could enhance AI-driven manipulation, and urge future researchers to weigh these pros and cons.
Future Directions
- Test the approach on larger, LLM-based multi-modal architectures rather than the roughly 90M-parameter TVLT.
- Evaluate on more diverse training and evaluation datasets to determine what governs whether brain-tuning gains transfer across contexts.
- Investigate the emotion breakdown on CMU-MOSEI more deeply — particularly why sadness showed improved F1 but decreased accuracy — which the authors explicitly leave to future work.
- Explore larger-scale validation, since the current study is limited to a single model and a small number of evaluations.
Target Audience
Researchers working at the intersection of computational neuroscience and multimodal machine learning, particularly those interested in brain alignment, brain-tuning, and encoding models. It is also useful for practitioners exploring whether neural signals can serve as a training objective for socially intelligent AI, and for readers looking for a concise example of extending a brain-tuning pipeline to a new modality and a new functional target region.
Authors’ abstract
Recent studies on audio models show brain-tuning - fine-tuning models to better predict corresponding fMRI activity - improves brain alignment and increases performance on downstream semantic and audio tasks. We extend this approach to a multimodal audio-video model to enhance social cognition, targeting the Superior Temporal Sulcus (STS), a key region for social processing, while subjects watch Friends. We find significant increases in brain alignment to the STS and an adjacent ROI, as well as improvements to a social cognition task related to the training data - sarcasm detection in sitcoms. In summary, our study extends brain-tuning to the multi-modal domain, demonstrating improvements to a downstream task after tuning to a relevant functional region.