Skip to content
AI.info

Research

ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision Creators

Overview Research area: Human-Computer Interaction, specifically accessibility systems and tools (CCS category: "Human-centered computing — Accessibility systems and tools"), at the intersection of au

ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision Creators
arXiv
2602.07266
Published
2026-02-06
Authors
Franklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. Kane

AI summary

Overview

Research area: Human-Computer Interaction, specifically accessibility systems and tools (CCS category: "Human-centered computing — Accessibility systems and tools"), at the intersection of audio description (AD) production, screen reader accessibility, and multimodal large language models.

Technical level: Intermediate. The paper is written for HCI and accessibility researchers, but the system itself (a web app built with HTML, JavaScript, and CSS that calls Gemini 2.5 models) is approachable for readers familiar with screen readers and basic web accessibility concepts.

Scope (one sentence): The paper presents ADCanvas, a screen reader-accessible, conversational authoring tool that lets blind and low vision (BLV) creators independently generate, edit, and preview audio description scripts for video, evaluated through a qualitative study with 12 BLV video creators.

What This Paper Is About

Audio Description makes visual media accessible to blind and low vision audiences, and some of its most skilled practitioners are themselves BLV. Yet the mainstream AD authoring tools — Digital Audio Workstations, non-linear video editors, and tools like Ooona or Subtitle Edit Pro — depend on visual metaphors such as timelines, waveform editors, drag-and-drop interfaces, and frame-accurate scrubbing, none of which translate into a linear auditory stream through a screen reader. This forces BLV creators into inefficient workarounds or dependence on sighted collaborators, who may not prioritize the same visual details the creator would.

The paper's goal is to reimagine AD authoring so that BLV creators drive the process with an AI agent as a supporting partner. ADCanvas combines a conversational multimodal LLM agent (for visual question answering, script generation, and script editing), keyboard-and-screen-reader-based media controls, and a plain-text WebVTT editor.

Key Contributions

  1. The ADCanvas system itself: a novel, screen reader-accessible multimodal AD authoring tool that lets BLV creators generate and revise AD scripts through contextual conversational interaction, visual question answering (VQA), keyboard navigation, and real-time in-line narration.

  2. Empirical findings from a qualitative user study with 12 BLV participants, covering human-AI co-creation practices, how participants negotiated trust and verification with the agent, and the breakdowns they encountered in conversational workflows.

  3. Design implications for future accessible creative tools, emphasizing agent configurability and fine-grained authoring control for BLV AD creators — including precise editing controls, accessibility support for creative ideation, and configurable rules for human-AI collaboration.

The paper also frames three design goals that guided the system: enabling independent authoring by BLV creators, facilitating video understanding, and shifting cognitive load from visual information seeking to creative synthesis.

Main Findings

  • The conversational agent served as an informational aide and drafting assistant. Participants adopted it both to ask visual questions about video content and to produce first drafts of AD text, while they retained the authoring role themselves.

  • Participants acted as curators, not passive recipients. They saw themselves as receiving information from the model and filtering it down for their audience, exercising creative judgment over what the model produced. The paper's example is a creator synthesizing the agent's factual list of face-like objects (a bathroom sink, trouser pockets, coffee foam, fried eggs) into a single evocative line for a fast-paced montage.

  • Participants maintained a supervisory stance balancing trust with verification. They treated the agent's output as a starting point that required checking rather than accepting uncritically.

  • Conversational interaction made the AD-creation process accessible. The paper's first stated finding is that the conversational paradigm supported an accessible AD-creation process for participants.

  • Participants guided the AI through distinct creative interaction patterns. These included using the AI to create first drafts, and drafting text manually with VQA support from the AI.

  • Usability breakdowns clustered around agency, precision, and interaction fluidity. Participants' experiences surfaced specific problems in these areas, which the authors used to derive design implications.

  • Participants wanted to keep using the system. The paper reports that participants enjoyed collaborating with the AI agent and expressed a strong desire to continue using ADCanvas after the study, describing it as a responsive co-author that enabled independent, high-quality creative judgment with appropriate accessibility support.

  • No quantitative performance measures are reported. The study is qualitative, and the paper does not report task-completion times, accuracy rates, or comparative benchmarks against other tools.

Methodology in Plain English

The researchers first reviewed prior work on professional media creation tools, collaborative and community AD platforms (including YouDescribe, Rescribe, and DescribePro), and instruction-based LLM agents, using that review to identify where existing tools exclude BLV creators.

They then built ADCanvas as a web application in HTML, JavaScript, and CSS, tested across Google Chrome, Apple Safari, and Mozilla Firefox and made compatible with JAWS, NVDA, and Apple VoiceOver. The system's generative features use the Gemini 2.5 series: Gemini 2.5 Pro for initial AD generation, Gemini 2.5 Flash for the low-latency conversational agent, and Gemini 2.5 Flash text-to-speech for the line-narration preview. Model temperature was set to 0.3 for all prompts. Client-side IndexedDB storage preserves work in progress against accidental tab closes or refreshes.

Key interface choices came from pilot testing: timestamps are shown as "xx min xx sec" rather than the traditional 00:00:08 format because novice creators did not read the latter fluently with a screen reader; hotkeys were mapped to avoid conflicts with system-level and screen reader shortcuts; and state persistence returns users to the exact script line they were editing after a detour to the agent.

The authors positioned ADCanvas as a technology probe — an exploratory artifact rather than a full replacement for commercial DAWs — and ran a qualitative user study with 12 BLV participants who were professional AD practitioners and/or video creators. Participants used the system to author AD scripts for short videos, and the paper reports findings organized around three research questions: how collaboration with the embedded multimodal agent shapes creators' practices (RQ1), what interaction design challenges and opportunities arise for non-visual conversational workflows on complex creative tasks (RQ2), and how BLV creators perceive human-AI collaborative workflows for AD authoring (RQ3).

The paper's participant table lists demographics including age bands, gender, screen reader used (VoiceOver, NVDA, JAWS), visual condition (legally blind, totally blind, low vision), professional AD experience (writing, QC, narration, ranging from under 1 year to more than 5 years), and content creator experience. The version of the table included in the paper content provided here is truncated partway through, so the full set of rows is not available.

Why This Matters

Impact on research: The paper reframes AD authoring as a site for human-AI co-creation rather than automation, arguing that BLV users should be creators and directors of the process rather than consumers of finished descriptions. Prior HCI work on AD (such as Rescribe and DescribePro) largely targets sighted users, and fully automated description systems prioritize rudimentary coverage over narrative coherence. ADCanvas contributes both an artifact and empirical evidence about how trust, verification, and creative agency play out when a multimodal agent is embedded in a professional accessibility workflow.

Real-world applications:

  • Independent AD production by BLV professionals. The paper's example scenario shows a blind creator drafting, refining, timing, and exporting an AD track for a short film without relying on a sighted collaborator for visual information.
  • Creator platforms and social media video. Personal video creators — including the study participants who run YouTube and TikTok channels and manage social media for businesses — could describe their own content.
  • Community and volunteer AD platforms. The authors describe their work as complementary to platforms like YouDescribe, positioned as a core authoring engine that could be integrated to support both sighted and BLV contributors.
  • Video question answering during authoring. The conversational VQA capability (asking what makes slippers resemble smiling faces, or what visuals show a man looking distressed) applies to any workflow where a non-visual user needs to interrogate visual media.

Industry relevance: The paper speaks directly to the tooling gap in professional post-production, where accessibility is treated as an add-on handled by sighted operators. It demonstrates that a screen reader-first, keyboard-first, plain-text (WebVTT) workflow can support creative work that currently requires timeline-based editors. It also illustrates a concrete pattern for embedding multimodal models in accessibility tools — using a stronger model for generation, a faster model for interaction, and a low-latency TTS model for preview — while the authors note their contributions are independent of any specific model and should remain relevant as generative models advance.

Future Directions

  • Agent configurability. The paper argues for agent behavior that users can progressively shape by instilling their own expert knowledge and stylistic preferences through iterative feedback, rather than a fixed set of behaviors.

  • Precise, fine-grained editing controls. The authors explicitly note that ADCanvas does not replace professional DAWs: advanced operations such as fine-grained timing alignment and waveform-based gap detection remain outside its current scope, leaving open how to bring these to non-visual users.

  • Design rules for human-AI collaboration. The paper calls for configurable rules governing how much the AI does versus how much the human does, addressing the agency and precision breakdowns participants reported.

  • Accessibility support for creative ideation. Beyond removing mechanical labor, the paper identifies a need for tooling that actively supports the ideation side of AD — narrative construction, word choice, and tonal consistency — for BLV creators.

Target Audience

This paper is most useful for HCI and accessibility researchers studying audio description, screen reader interaction, and human-AI co-creation; designers and engineers building accessible media or creative tools; AD practitioners and trainers interested in non-visual workflows; and platform developers considering how to make AD authoring or community description tools usable by BLV contributors. Readers who are experts in multimodal model architecture will find the model use deliberately lightweight and implementation-agnostic, while readers new to accessibility research will find the framing of barriers and design goals accessible.

Authors’ abstract

Audio Description (AD) provides essential access to visual media for blind and low vision (BLV) audiences. Yet current AD production tools remain largely inaccessible to BLV video creators, who possess valuable expertise but face barriers due to visually-driven interfaces. We present ADCanvas, a multimodal authoring system that supports non-visual control over audio description (AD) creation. ADCanvas combines conversational interaction with keyboard-based playback control and a plain-text, screen reader-accessible editor to support end-to-end AD authoring and visual question answering (VQA). Combining screen-reader-friendly controls with a multimodal LLM agent, ADCanvas supports live VQA, script generation, and AD modification. Through a user study with 12 BLV video creators, we find that users adopt the conversational agent as an informational aide and drafting assistant, while maintaining agency through verification and editing. For example, participants saw themselves as curators who received information from the model and filtered it down for their audience. Our findings offer design implications for accessible media tools, including precise editing controls, accessibility support for creative ideation, and configurable rules for human-AI collaboration.

Read the original paper