← News

Official Qwen3.8-Omni-Flash release graphic showing the Qwen model name and layered text, audio, video, and document interfaces

18 Sep 2026

Qwen / Alibaba

Alibaba's Qwen launches Qwen3.8-Omni-Flash for multimodal agents

Qwen announced Qwen3.8-Omni-Flash, an API model that natively handles text, images, audio, and video, with a 1-million-token context window, tool calling, and web search. Qwen says it is designed to turn long media into completed work, such as video editing, translation, and recaps.

SOFTWARE desk — this pushes multimodal AI toward “watch/listen, then do” instead of merely describing media. A 1-million-token context and tool calling could let one model handle long recordings and videos without stitching together separate transcription, vision, and workflow systems, while the cost claims make that workflow more practical if they hold up in use.

What the company and the official docs say the model does: it accepts text, images, audio, and video, and it returns text. Qwen describes native multimodal understanding, reasoning, and tool use for agentic workflows — meaning the model can read several kinds of media, think across them, and ask connected software to take a next step. File that watch / listen / reason / use-tools picture as Qwen’s. This desk did not sit on an API call.

Official Alibaba Cloud docs list a 1-million-token context window, function calling, web search, context caching, and support for 113 audio languages and dialects. A context window is how much the model can hold at once; 1 million tokens is enough room for long recordings and videos instead of short clips. Function calling is the same idea as tool calling: the model can request an action from software you connect. Context caching means already-processed media can be reused so you do not pay to re-read it. File those limits and features as Alibaba Cloud’s. This desk did not meter a token counter.

Qwen says the model can coordinate tools for jobs such as searching a long video, editing workflows, translation, and recaps. Official docs also name audio/video analysis, meeting summaries, and subtitle generation. File those as company-described use cases — not as outcomes this desk watched. Do not file them as independently verified product tests.

Cost and score claims, still company: Qwen says video-input costs are about 89% lower than Qwen3.5-Omni-Plus, the previous omni model, and reports audio-video performance near Google’s Gemini 3.8 Flash. File the 89% cut and the Gemini comparison as Qwen / Alibaba’s. This desk did not rerun a benchmark and is not treating those lines as an independent test or a desk ranking.

Qwen says it is open-sourcing Qwen-MM-Plugins, a tool layer for multimodal agents, and Qwen-Live Harness, a harness built around a realtime API. The model itself is API-only in the cited documentation. Weights — the downloaded numbers that would let you run the model on your own machines — are not claimed as released here. File the plugin / harness open-source line as Qwen’s. This desk did not treat a missing GitHub page as proof the harness already shipped.

Plain English for the rest of the card: omni-modal = one model that can work across several media types, not a separate speech tool plus a vision tool plus a chat tool. tool calling / function calling = the model can ask connected software to perform an action. context window = how much it can hold at once. web search = it can look things up on the live web. context caching = reuse already-processed media so you do not pay to ingest it again. API-only = you call it over the internet; the weights are not published here. Qwen-MM-Plugins = Qwen’s open tool pack for multimodal agents. Qwen-Live Harness = the company-stated realtime harness. Gemini 3.8 Flash = Google’s same-generation speed model Qwen compares itself to — a company comparison, not a desk ranking.

PRIMARY here: Alibaba Cloud Model Studio’s qwen3.8-omni-flash documentation and omni-modal overview, plus Qwen’s 18 Sep 2026 company blog — Tier A PRIMARY company sources, the original record. @Alibaba_Qwen is the same-day company social announce, not a substitute primary. GIGAZINE is same-day discovery corroboration, not a second originating Qwen newsroom. The 18 Sep announce, text/image/audio/video-in and text-out behavior, native understanding / reasoning / tool use, the 1-million-token window, function calling, web search, context caching, 113 audio languages and dialects, the long-video / editing / translation / recap use cases, the about-89% video-input cost cut versus Qwen3.5-Omni-Plus, the near-Gemini-3.8-Flash performance line, and the Qwen-MM-Plugins / Qwen-Live Harness open-source line are Qwen / Alibaba-attributed. Cost, benchmark, and Gemini-comparison figures stay company-attributed — not independently tested here. NOT claimed: AGI, independent benchmark leadership, that the model is open-weight, that this desk ran the API, a stock tip, or investment advice. Distinct from the already-filed gemini-3-8-live, gemini-agentic-video-understanding, gemini-3-8-flash-cyber, and nvidia-vera-rubin-nvl72-mlperf.

RELATED

ONLINE

article thread

guidelines

warming…

warming…

On 18 Sep 2026 Qwen announced Qwen3.8-Omni-Flash. The company PRIMARIES are Alibaba Cloud Model Studio’s qwen3.8-omni-flash model page, the same studio’s omni-modal overview, and Qwen’s company blog at https://qwen.ai/blog?id=qwen3.8-omni-flash. Qwen also posted the launch on @Alibaba_Qwen. Those company pages are the filing event. These are Qwen / Alibaba words. This desk did not run the model.

Sources