GitHub radar

Qwen Just Open-Sourced Weights at Its Flagship Qwen-Max Level

This isn't a browser demo — Qwen has released open weights matching its paid flagship, Qwen-Max, for the first time, so you can run it through any API provider instead of being locked to Qwen's own service.

01Qwen/Qwen3.8-2.4T-A95B 1.1k18k downloads/mo2.4T (активных 95B) paramstext-generation

It's a massive open-weight language model for coding and AI agents — the first time Qwen has published weights at this quality tier.

When it helps

Useful if you want flagship-level coding/agent model access through a third-party API provider, without depending on Qwen's own service. Useless for a laptop or home PC — it needs a server.

Pros

  • First time Qwen opens flagship-level model weights
  • Strong results on coding and agent benchmarks
  • Context window up to 1M tokens
  • Reachable through third-party API providers

Cons

  • Even 4-bit quantization needs ~1400GB VRAM — server only
  • Reasoning mode is always on, can't be disabled
  • chat.qwen.ai actually runs a different model, Qwen3.8-Max

How to set it up — step by step

  1. 1Don't try to run this on your own machine — it needs a server with roughly 1400GB VRAM.
  2. 2Open your AI agent (Claude Code, Codex) in your project folder.
  3. 3Send the agent this exact request: "Find me an API provider with access to Qwen/Qwen3.8-2.4T-A95B: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B — show me how to make a test call to this model through Together AI, Fireworks, or another provider that supports it."
  4. 4Check the provider's listing for Qwen/Qwen3.8-2.4T-A95B and confirm access terms — pricing isn't given in the facts.
  5. 5Make a test API call and compare the output on a coding or agent task you care about.
  6. 6If you just want browser chat, note that chat.qwen.ai actually runs Qwen3.8-Max, a different model.

Text for your agent

Copy this and send it to your agent — Claude Code, Codex, any of them:

Find me an API provider with access to Qwen/Qwen3.8-2.4T-A95B: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B — show me how to make a test request via Together AI, Fireworks, or another provider that supports it.

Server-grade model — you can't run this at home. Even the 4-bit version needs around 1400 GB of VRAM. Think of it as news, not a laptop tool.

Open on Hugging Face