Alibaba's New Model Watches and Listens, and Charges Almost Nothing for It
Qwen3.8-Omni-Flash handles text, images, audio and video in one model with a million-token context. Alibaba reports hourly audio ingestion costs falling by more than 98 percent. No open weights at launch.
Alibaba’s Qwen team released Qwen3.8-Omni-Flash on 18 September, a model that takes text, images, audio and video together rather than through separate specialist systems. “Omnimodal” is the marketing word, and what it means in practice is that the model hears a recording as sound rather than being handed a transcript that some other tool produced first. It holds a million tokens of context, enough for hours of material at once, and Alibaba reports an average improvement of more than 26 percent across roughly 30 evaluations compared with the previous Qwen3.5-Omni-Plus.
The prices are the part worth writing down. Alibaba lists 0.8 RMB per million input tokens and 2.7 RMB per million output tokens on Alibaba Cloud Model Studio, and says the cost of ingesting an hour of audio has fallen by more than 98 percent, with mixed audio and video input down more than 93 percent. On agent-style video tasks it reports 45.7 percent fewer tokens used. The model is API only, available through QwenCloud, Alibaba Cloud Model Studio and Qwen Studio, with no downloadable weights announced at launch. That last point is a departure for a team that has built much of its reputation on open releases.
What is behind this
The older way to make an AI understand a video was a relay race: one model transcribes the speech, another describes the frames, a third reads both descriptions and reasons about them. Every handoff loses information, and the things that get lost are exactly the ones that carry meaning. Tone of voice, whether someone sounded unsure, what was on screen at the moment a sentence was said, none of that survives a transcript. A single model taking the raw streams keeps those details in play, which is why “native” is the word the labs keep reaching for.
The cost collapse is what makes it more than a research curiosity. When processing an hour of audio was expensive, you only did it for material you already knew mattered. At these prices it becomes reasonable to point a model at everything and search afterwards, which is a different kind of tool entirely. Worth keeping expectations grounded, though: the 26 percent figure is Alibaba’s own, measured against Alibaba’s previous model, and independent evaluation of omnimodal systems is still thin.
What this means for you: If you have a pile of recordings, meetings, interviews, lectures, support calls, that has been too expensive or too tedious to do anything with, this class of model is the reason that is changing. You will meet it through products rather than the API. For anyone building something, the honest note is the missing open weights: an API-only model means your audio and video go to Alibaba’s servers, which is a straightforward question of whether the material is yours to send.
Sources
Source: https://technode.com/2026/09/18/alibabas-qwen-releases-qwen3-8-omni-flash-with-1m-token-context/
Xiaomi Now Has the Strongest Open Model in the World, and You Can Just Download It
The phone and car maker released MiMo-V2.6-Pro and Flash under an MIT licence on Monday. Third-party benchmarker Artificial Analysis scored Pro at 46, the highest any downloadable model has reached.