Qwen3.8-Omni-Flash Costs Five Times Less Than Gemini Flash
Alibaba released Qwen3.8-Omni-Flash, its first multimodal Qwen model for agents. It handles audio and video and costs five times less than Gemini 3.8 Flash on input tokens.
Alibaba released Qwen3.8-Omni-Flash, its first multimodal model for agents. It processes audio and video at the same time and calls tools on its own. The model costs $0.15 per million input tokens and $0.47 per million output tokens. Gemini 3.8 Flash costs $0.75 and $3.75, respectively. The context window is 1 million tokens. The model is available via API and has plugins for Claude Code and Gemini CLI.
On audio-video benchmarks, Qwen scores close to Gemini Flash. These are Alibaba's own benchmarks, not independent ones. I would use this model for pipelines that need to process a lot of video or audio. The price difference there is real, and tasks like translation or video summarization do not require frontier-level quality.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles