Skip to content
AI.info

The Pulse

Alibaba Releases Qwen3.8-Omni-Flash With 1M-Token Context

Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a multimodal model that accepts text, images, audio and video. The model supports a 1-million-token context window, tool calling, web search and audio input across 113 languages and diale

Alibaba Releases Qwen3.8-Omni-Flash With 1M-Token Context

AI.info Team ·

Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a multimodal model built to process text, images, audio and video through a single API. The model supports a context window of up to 1 million tokens, giving developers room to submit long recordings, large collections of images or extended video alongside ordinary text prompts.

Qwen’s announcement is dated September 18, 2026. Alibaba Cloud documentation lists Qwen3.8-Omni-Flash for Chat Completions and Responses APIs, with text-only output. The model is positioned for audio and video understanding, meeting summaries and content analysis rather than spoken replies.

No named spokesperson quotation accompanies the release materials reviewed for this article.

Qwen puts audio, video and text in one request

Qwen3.8-Omni-Flash accepts text, images, audio and video as input. Developers can combine modalities in a request and ask the model to summarize a recording, analyze visual content or answer questions about material spread across different files.

Alibaba Cloud’s documentation lists support for files of up to 2GB when developers provide public URLs. The documented duration limit reaches two hours for video and three hours for audio, while the API accepts up to 64 files in one video-oriented request and up to 2,048 files in an audio-oriented request when public URLs are used.

Those limits make the model suitable for workloads that involve more than a short clip or a single image. A meeting archive, a training library or a set of recorded interviews can be handled as one analysis job, subject to the service’s token and file constraints.

The 1-million-token window changes the unit of analysis

The model’s 1-million-token context window is its clearest specification. A context window determines how much material the model can consider within a request and conversation. For multimodal systems, the practical limit also depends on how audio, video and images are converted into tokens for billing and processing.

Alibaba’s documentation says Qwen3.8-Omni-Flash converts audio at seven tokens per second. Images are counted at one token for each 32-by-32-pixel block, with a default maximum of 1,280 tokens per image and a high-resolution mode that raises the maximum to 16,384 tokens. Video processing uses sampled frames, with up to 2,048 frames supported for the model.

Those conversion rules mean that a 1-million-token window does not translate into a fixed number of hours of video or audio. Resolution, frame sampling and the mix of modalities determine how much source material fits into a request.

Tool calling and web search are built into the API

Qwen3.8-Omni-Flash supports function calling and web search. Alibaba Cloud says the Responses API currently exposes web search as its only built-in Responses tool, while developers can also use function calling for applications that connect the model to external services.

The model supports thinking controls, including a setting that allows developers to disable reasoning when lower latency or reduced processing is preferred. Documentation examples show the model being called through OpenAI-compatible Chat Completions interfaces as well as the Responses API.

Supported deployment regions listed by Alibaba Cloud include Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia. Access requires an API key for the selected region, and regional endpoints and credentials differ.

Qwen separates understanding from speech generation

Qwen3.8-Omni-Flash produces text, not audio. Alibaba Cloud directs developers who need generated speech responses to the Qwen3.5-Omni family, which supports audio output and broader voice interaction features.

The distinction separates two jobs that are often presented as one multimodal capability: interpreting audio and video, and responding with synthesized speech. Qwen3.8-Omni-Flash focuses on the first task, with support for 113 audio languages and dialects according to Alibaba Cloud’s model documentation.

The release also marks a transition away from older Qwen Omni-Turbo models. Alibaba Cloud says that series is no longer being updated and recommends Qwen3.8-Omni-Flash for text analysis and Qwen3.5-Omni for audio output.

Where the model fits

Qwen3.8-Omni-Flash targets developers building systems around long recordings, video archives and mixed-media documents. Its combination of a large context window, multimodal input and tool access lets one model inspect source material before producing summaries, extracting information or initiating an external action.

The model is available through Alibaba Cloud’s Model Studio interfaces and Qwen’s own platform. The release does not make Qwen3.8-Omni-Flash an open-weight model; the materials reviewed describe hosted API access rather than a downloadable checkpoint.

For developers choosing among Alibaba’s current Omni models, the split is direct: Qwen3.8-Omni-Flash handles long-form multimodal understanding and text responses, while Qwen3.5-Omni supplies generated speech and interactive audio output.

Read Qwen’s announcement and Alibaba Cloud’s model documentation for the release details and API limits.

Source

Qwen

Explore

More articles