GitHub radar
Qwen Just Open-Sourced Weights at Its Flagship Qwen-Max Level
This isn't a browser demo — Qwen has released open weights matching its paid flagship, Qwen-Max, for the first time, so you can run it through any API provider instead of being locked to Qwen's own service.
It's a massive open-weight language model for coding and AI agents — the first time Qwen has published weights at this quality tier.
When it helps
Useful if you want flagship-level coding/agent model access through a third-party API provider, without depending on Qwen's own service. Useless for a laptop or home PC — it needs a server.
Pros
- First time Qwen opens flagship-level model weights
- Strong results on coding and agent benchmarks
- Context window up to 1M tokens
- Reachable through third-party API providers
Cons
- Even 4-bit quantization needs ~1400GB VRAM — server only
- Reasoning mode is always on, can't be disabled
- chat.qwen.ai actually runs a different model, Qwen3.8-Max
How to set it up — step by step
- 1Don't try to run this on your own machine — it needs a server with roughly 1400GB VRAM.
- 2Open your AI agent (Claude Code, Codex) in your project folder.
- 3Send the agent this exact request: "Find me an API provider with access to Qwen/Qwen3.8-2.4T-A95B: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B — show me how to make a test call to this model through Together AI, Fireworks, or another provider that supports it."
- 4Check the provider's listing for Qwen/Qwen3.8-2.4T-A95B and confirm access terms — pricing isn't given in the facts.
- 5Make a test API call and compare the output on a coding or agent task you care about.
- 6If you just want browser chat, note that chat.qwen.ai actually runs Qwen3.8-Max, a different model.
Text for your agent
Copy this and send it to your agent — Claude Code, Codex, any of them:
Find me an API provider with access to Qwen/Qwen3.8-2.4T-A95B: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B — show me how to make a test request via Together AI, Fireworks, or another provider that supports it.
Server-grade model — you can't run this at home. Even the 4-bit version needs around 1400 GB of VRAM. Think of it as news, not a laptop tool.
Open on Hugging Face▌ More finds