Skip to content
AI.info

The Pulse

Google Research Shows a 10-Minute AI-Generated Video

Google Research describes four systems for planning, generating and refining long-form video, including a ten-minute A²RD demonstration.

Google Research Shows a 10-Minute AI-Generated Video

AI.info Team ·

Ten minutes is the length of the AI-generated film Google Research uses to demonstrate a system intended to keep characters and settings consistent across a long story. In a post published September 24, researchers Yale Song and Yiwen Song describe four related systems for planning, generating and refining longer video sequences. The work aims at a familiar weakness in AI video: details can drift between shots, while errors in one stage can carry into the next.

A ten-minute film built from segments

The longest example comes from A²RD, an autoregressive system that generates video segment by segment and keeps a multimodal memory of what has appeared. For each segment, it retrieves context, generates new material, refines it and updates that memory. Google says A²RD switches between extrapolating the story into new events and interpolating around characters and settings that have already appeared.

“To test this system, the ten-minute movie below showcases a long-form generation, which demands consistent narrative progression across minutes-long temporal gaps.”

Yale Song and Yiwen Song, research scientists at Google

The demonstration is a ten-minute film, not evidence that a single prompt produces a finished ten-minute video in one pass. A²RD’s staged approach is designed to manage continuity across the gaps between segments. Google says the system helps retain details such as character identity, clothing and the layout of locations.

Four systems tackle four sources of drift

Google’s AI video co-director plans a production across several agents. An orchestrator selects a creative direction, then pre-production and production agents build a storyboard and generate keyframes, video and audio. A multimodal model reviews the assembled cut and sends feedback into another round of planning.

CANVAS focuses on continuity between shots by keeping structured records of characters, locations and object states. VQQA, or Video Quality Question Answering, takes a different role: it asks questions about a generated clip, uses a vision-language model’s critique to revise the prompt, and selects the strongest result across iterations. Google describes the systems as orchestration layers, with the co-director built on Gemini and Veo rather than presented as a new foundation video model.

The 81.4 score comes from Google’s benchmark

Google reports a peak quality score of 81.4 for the AI video co-director on GenAD-Bench. The benchmark tests 400 scenarios involving 50 fictional brands, each with four products, against detailed marketing constraints. The blog also describes HardContinuityBench for spatial continuity and LVBench-C, which contains 120 scenarios testing changes in characters, objects and environments.

The post says CANVAS improves continuity on HardContinuityBench and that A²RD improves consistency on long-video tests, but it does not give comparable numerical results for those claims in the article. Google directs readers to the individual papers for model configurations and baseline evaluations. That makes the 81.4 figure a result on a Google-described benchmark, not a universal measure of video coherence.

Google presents research, not a finished editing product

The announcement frames the four systems as tools for automating repetitive work such as prompt coordination, shot chaining and visual review. The co-director and CANVAS papers are listed as forthcoming at COLM 2026 and EMNLP 2026, respectively. The post does not announce a consumer release or provide a deployment date.

For creators, the demonstration points to a workflow in which models help maintain story continuity while people retain control of direction and narrative choices. The evidence Google publishes here is a ten-minute sample and benchmark claims; detailed comparisons and configuration information remain in the linked research papers.

Source

Explore

More articles