Research
PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research Communication
Overview Research area: Human-Computer Interaction (HCI) / Human-centered computing — specifically generative AI for science communication, with a focus on short-form video as a research dissemination

- arXiv
- 2601.18218
- Published
- 2026-01-26
- Authors
- Meziah Ruby Cristobal, Hyeonjeong Byeon, Tze-Yu Chen, Ruoxi Shang, Donghoon Shin, Ruican Zhong, Tony Zhou, Gary Hsieh
AI summary
Overview
Research area: Human-Computer Interaction (HCI) / Human-centered computing — specifically generative AI for science communication, with a focus on short-form video as a research dissemination medium.
Technical level: Intermediate. The paper assumes familiarity with LLM prompting pipelines and text-to-video models, but its core argument is about workflow design and user perceptions rather than model architecture.
Scope: A formative study, system design, and two-part evaluation (a researcher user study and a crowdsourced audience survey) of PaperTok, a human-in-the-loop generative AI system that turns academic PDFs into short-form science communication videos.
What This Paper Is About
Academic papers are written in ways that create an "ivory tower" gap between researchers and the public, while short-form video platforms (TikTok, Instagram Reels, YouTube Shorts) have become dominant channels for information consumption. Researchers often lack the time, skills, and resources to produce public-facing audio and visual content, so this paper asks how generative AI can be responsibly used to transform papers into accurate, engaging, and trustworthy short-form videos. The authors build PaperTok, a system that generates hook options, scripts, visuals, voiceovers, and a credit screen from an uploaded paper, and then let the researcher act as a creative director over the AI output.
Key Contributions
- A formative study (N = 8) with science communicators and content creators that surfaces insights for transforming academic papers into engaging short-form videos, organized as four cross-cutting design challenges (content selection, capturing immediate attention, maintaining engagement, and communicating credibility).
- The design and implementation of PaperTok, described as a novel human-AI collaborative system for authoring short-form scientific videos from academic papers, including a hook-and-script step, a storyboarding step, and a credit screen step.
- Empirical insights from a user study with N = 18 researchers who had published at least one academic paper, plus a survey involving those researchers and N = 100 crowdsourced audience participants, comparing PaperTok videos against videos from existing PDF-to-video platforms (SciSpace and PDFtoBrainrot).
- A set of design implications for future generative AI tools that aim to scaffold rather than automate complex creative tasks, emphasizing user control, iterative refinement, and integration of AI suggestions with human expertise.
Main Findings
- PaperTok videos were rated more engaging and entertaining: In the comparison against videos generated by existing PDF-to-video platforms (SciSpace and PDFtoBrainrot), the authors report PaperTok was significantly more engaging and entertaining while providing similar levels of informational value.
- Lowered barriers to visual and audio production: Researcher participants described PaperTok as useful for lowering the barrier to generating engaging visuals and voiceovers, while still allowing some personalization to suit their communication style.
- Static visuals were rejected by experts: Formative participants universally dismissed SciSpace's primarily static screenshots of paper figures; one participant stated that just having the image of the paper means "no one's gonna watch that."
- AI-sounding narration damaged credibility: Artificial-sounding narration triggered instant rejection in the formative study, with one participant saying the AI voice "immediately lowers my interest," and another saying hearing an AI voice made them think no one had worked on the video.
- Hooks and narrative closure are tightly coupled: Participants identified the opening 2–5 seconds as decisive, wanted short, punchy hooks with enough context, and expected conclusions to resolve the hook's opening question with actionable takeaways.
- Duration matters for retention: One participant stated that 15 to 30 seconds is what people will pay attention to, with others extending to under 60 seconds or around 2 minutes as an absolute maximum; PaperTok's generated scripts target a typical short-form duration of around 45 seconds across 8 scenes.
- Human presence is the primary credibility signal: Participants pointed to creators who appear briefly on camera and to human-read narration as establishing trust, and asked for explicit citation, authorship, and institutional branding as academic authority markers.
- Desire for more fine-grained control: The authors report identifying the need for more fine-grained controls in the creation process, and describe nuanced insights about researchers' desired level of control over the workflow and the need to signal the human-in-the-loop in AI-facilitated science communication.
- Contextual statistics from prior work: A 2024 Pew report cited in the paper states more than 50% of people surveyed at least sometimes get news from social media, and a Pew 2025 study on TikTok usage found 17% of adults in the US report they regularly get news from TikTok.
Methodology in Plain English
The researchers started with a formative study: 60-minute semi-structured video-call interviews with 8 science communicators and content creators recruited from Instagram, RedNote, TikTok, YouTube, and university communications roles, each compensated with $30 USD in a gift card of their choice. To ground the discussion, they converted three CHI 2025 Best Paper awardees — representing artifact, empirical, and methodological contributions — into videos using three methods: a pipeline generating scripts with Google's Gemini 2.5 Flash and video clips with Veo 2 (veo-2.0-generate-001), PDFtoBrainrot, and SciSpace's PDF-to-video service. Participants reviewed these probes, critiqued hooks and scripts, and co-designed improvements; sessions were recorded, transcribed, and analyzed thematically.
Those findings shaped PaperTok's design as a three-step workflow. Users upload a paper PDF, then in Step 1 receive four hook options with matching scripts plus AI-recommended voiceover tones they can preview and edit by prompt. Hooks are constrained to a maximum of 15 words, must be conversational and jargon-free, and definitive statements are converted into questions to reduce oversimplification risk. Step 2 presents the selected script as segmented scenes, each paired with a visual to generate, letting users either generate all scenes at once or work scene by scene, iterating until satisfied. Scripts are structured into 8 scenes — Hook (Scene 1), Body (Scenes 2–7, covering context, findings, and relevance), and Closing (Scene 8) — with each scene limited to 18–22 words, roughly 6–7 seconds. Step 3 appends a credit screen with an auto-filled author attribution that users can add their own name to, signaling credibility and the human-in-the-loop process. The default voiceover tone is "an influencer vibe with fast speech," with the LLM suggesting a content-aware stylistic modifier based on the script.
Evaluation combined a mixed-methods user study with N = 18 researchers who had previously published at least one academic paper, who created short video summaries of their own work with PaperTok, with a survey taken by those researchers and N = 100 crowdsourced audience participants comparing PaperTok output against the existing PDF-to-video platforms.
Why This Matters
Impact on research: The paper argues that no prior work had investigated the design and perception of AI-generated short-form videos specifically for research communication, and that existing commercial PDF-to-video tools do not allow collaborative input from users during video creation and have unclear reception. It contributes both a working human-AI workflow and evidence about how researchers want to control such tools, addressing concerns about LLM hallucination, bias, and the growing "infodemic" of AI-generated content.
Real-world applications:
- Researchers producing a video companion to a newly published paper, complementing social media posts or university press releases.
- University communications and news teams translating faculty research for public audiences.
- Science communicators and independent creators on TikTok, YouTube Shorts, Instagram Reels, and RedNote who currently source and adapt academic work manually.
- Educational and outreach contexts where abstract findings need visual translation into animations, metaphors, or relatable examples.
Industry relevance: The paper benchmarks against commercial PDF-to-video services SciSpace and PDFtoBrainrot, suggesting a competitive landscape for research-communication tooling. Its design implications — user control, iterative refinement, credibility signaling, and scaffolding rather than full automation — are directly relevant to builders of generative media tools for expert users, and its findings suggest that audiences are becoming more sensitive to clearly AI-generated content.
Future Directions
- Designing more fine-grained controls in the creation process, which the authors explicitly identify as a need uncovered in their evaluation.
- Better mechanisms for signaling the human-in-the-loop in AI-facilitated science communication, so audiences can assess credibility.
- Extending beyond one-to-one translation of a paper into a video, since PaperTok was built to complement rather than replace existing research communication practices.
- Continued inquiry into accuracy, trust, and responsible use of generative models for science communication in high-stakes domains where precision matters, given known hallucination, logical fallacy, and bias concerns.
Target Audience
HCI researchers and designers working on human-AI collaboration, generative AI tooling, and creativity support tools; science communication scholars and practitioners; university communications staff; developers of research dissemination and PDF-to-video products; and researchers themselves who want to understand what makes short-form science video credible and engaging, and how much control they should expect to retain when AI produces the content.
Authors’ abstract
The dissemination of scholarly research is critical, yet researchers often lack the time and skills to create engaging content for popular media such as short-form videos. To address this gap, we explore the use of generative AI to help researchers transform their academic papers into accessible video content. Informed by a formative study with science communicators and content creators (N=8), we designed PaperTok, an end-to-end system that automates the initial creative labor by generating script options and corresponding audiovisual content from a source paper. Researchers can then refine based on their preferences with further prompting. A mixed-methods user study (N=18) and crowdsourced evaluation (N=100) demonstrate that PaperTok's workflow can help researchers create engaging and informative short-form videos. We also identified the need for more fine-grained controls in the creation process. To this end, we offer implications for future generative tools that support science outreach.