Qwen3.8-Max: a 2.4-trillion MoE and the benchmarks it claims

Qwen3.8-Max is a Mixture of Experts with 2.4 trillion parameters and context up to one million tokens, accessible from 3 August 2026 through the Model Studio API and QwenWork. Open weights are announced for next week on Hugging Face and ModelScope, with no exact day and with the licence not yet defined. The benchmarks Qwen published, where the picture does not run one way, and the Max series so far.

Open SourceAIOpen SourceQwenAlibabaLLMAIOpen WeightsMoELocal InferenceBenchmarksDigital Sovereignty
Four technical figures on Qwen3.8-Max as of 3 August 2026: 2.4 trillion total parameters, a Mixture of Experts architecture and context up to one million tokens, accessible through the Model Studio APIs and the QwenWork platform; weights announced for next week on Hugging Face and ModelScope, with no exact day and with the licence not yet defined, alongside a Qwen3.8-27B sized for on-premise GPUs; jumps over Qwen3.7-Max with FrontierSWE from 40.7 to 73.5, DeepSWE 1.1 from 21.6 to 56.6, JobBench from 31.3 to 53.4 and GPQA Diamond from 92.4 to 92.6, so the largest jumps are where the model was weakest; benchmarks stated by Qwen with PaperBench at 93.0, GPQA Diamond at 92.6, OmniDocBench 1.5 at 92.1 and IFBench at 82.8, with no independent checks. A note at the bottom says the declared picture does not run one way, because on Terminal-Bench 2.1 Qwen gives itself 86.6 against 88.8 for GPT-5.6 Sol and on SWE-bench Pro 67.7 against 80.0 for Fable 5
The state of the release as of 3 August. Sources at the end.

Today Alibaba made Qwen3.8-Max accessible to developers worldwide through the Model Studio APIs and the QwenWork platform, which entered public beta today on web and desktop. The open weight release is announced for next week on Hugging Face and ModelScope.

Trying it is possible today; hosting it yourself depends on the weights, which have not been published yet.

What is available from today

The model was presented on 19 July 2026 at the World AI Conference in Shanghai as a preview, reachable only through Alibaba’s Token Plan and the Qoder and QoderWork platforms, at a price sources put at around 10% of list. From today access goes through the standard Alibaba Cloud APIs and is open to everyone.

The figures are the ones stated by the vendor, with no independent confirmation: 2.4 trillion total parameters, a Mixture of Experts architecture, context up to one million tokens and maximum output of 131,072 tokens. QwenWork enters an already crowded field, alongside Tencent WorkBuddy, Moonshot’s Kimi Work, Claude Cowork and ChatGPT Work.

The stated numbers

At launch Qwen published a broad comparison table. These are the most cited values, all measured by Qwen:

BenchmarkQwen3.8-MaxStated comparison
PaperBench93.0GPT-5.6 Sol 90.5 · Fable 5 88.8
GPQA Diamond92.6GPT-5.6 Sol 94.1 · Qwen3.7-Max 92.4
OmniDocBench 1.592.1
Terminal-Bench 2.186.6GPT-5.6 Sol 88.8 · Claude Opus 4.8 84.6
OSWorld-Verified86.1
IFBench82.8GPT-5.6 Sol 72.7 · Fable 5 63.5
FrontierSWE73.5Fable 5 88.8
SWE-bench Pro67.7Fable 5 80.0

The declared picture does not run one way, and that is the most useful thing about the table. Qwen leads on IFBench, which measures instruction following, and on PaperBench. It sits behind GPT-5.6 Sol on Terminal-Bench and behind Fable 5 by twelve points on SWE-bench Pro and fifteen on FrontierSWE, which is to say precisely on software engineering tasks over real repositories.

The jumps over the previous generation are largest where the model was weak: DeepSWE 1.1 goes from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5 and JobBench from 31.3 to 53.4.

Three cautions on reading them. The scores are all Qwen’s and no independent reproductions appear to exist. In the multimodal comparisons the internal reference point is Qwen3.7-Plus rather than Max, so the improvement on that part should be read carefully. And the reinforcement learning curve shown in the documentation peaks at 0.725 around four thousand training environments and then declines, which is a more interesting detail than most of the scores.

On activated parameters, the figure of 95 billion circulates and has been picked up by several outlets, but Alibaba has not disclosed it. It is best treated as a third-party estimate.

The four Max models before it

Qwen’s Max line was born closed and has stayed that way in every release so far: Qwen3-Max-Preview in September 2025, Qwen3.6-Max-Preview on 20 April 2026, Qwen3.7-Max on 20 May and Qwen3.8-Max-Preview on 19 July.

Until today the downloadable part of the family was of a very different size, from the Qwen3.6 cycle onwards: models that fit on a single GPU, while one with 2.4 trillion parameters needs infrastructure of an entirely different order.

If the weights do arrive, it would be the first time a Max-class model leaves the API. A Qwen3.8-27B is announced along with them, described as a checkpoint sized for ordinary on-premise GPU hardware, which is the part of the release most likely to end up on other people’s machines.

Promises and dates

In the same weeks other labs did the same thing verifiably. Moonshot paired Kimi K3 with a stated date for the weights. Poolside published Laguna S 2.1 on 21 July, putting the weights on Hugging Face the same day under the OpenMDW-1.1 licence. DeepSeek released V4 Flash 0731 with weights on Hugging Face under MIT, downloadable and self-hostable.

Every lab has constraints of its own that are invisible from outside. A promise of openness does become something you can plan against when three things are there: a day, a licence and a model card. Qwen3.8-Max currently has a generic “next week”.

What we think

For anyone weighing adoption today the situation reads simply. The model can be tried through the API, and if the task is working out whether it holds up on your own use cases the API is enough. What cannot be done is designing on top of it a system that has to keep data inside the perimeter, because that choice depends on a file that does not exist yet and a licence nobody has seen.

The difference counts, as in Open Intelligence, Secure Governance: an open-weight model can be run on your own hardware, as DwarfStar 4, colibri and KTransformers show, and it makes the choice of supplier reversible. A model behind an API, however good, leaves that lever in someone else’s hands.

At 2.4 trillion parameters, moreover, opening the weights changes less than it seems for most companies: few have the hardware to host it. It matters for those building services on top of these models, and it matters as a signal that an open-weight frontier keeps existing, as in open weights and American AI leadership.

We will come back to it when there are a file, a licence and an independent reproduction of the numbers.

Sources

Need support?Under attack?Service Status
Need support?Under attack?Service Status