- Start
- Aug 3, 202690% CONFIDENCEfrom the source
Qwen3.8-Max Released
- Alibaba launched Qwen3.8-Max on 3 August 2026, "the most powerful model in its Qwen series to date", on its Model Studio API for global developers and in its QwenWork agent platform, after previewing it in July[1][3]
- It ranked fifth in Text Arena, second in Vision Arena and fourth in Frontend Code Arena, and in one internal test coded autonomously for 16 days to build and open-source the "oh-my-cli" agent framework[1][6]
- Alibaba's New York-listed shares climbed 4.5% on the day and its Hong Kong shares 7%[6]
- The weights followed on 12 August as Qwen3.8-2.4T-A95B, the first Max-tier Qwen released openly, under a licence that makes providers earning more than US$50 million sign a commercial licence; the open model omits image input and the non-thinking mode, which the Apache 2.0 Qwen3.8-27B (14 August) keeps[3][5]
- On 2 September the
qwen3.8-maxAPI moved to a new snapshot, qwen3.8-max-0902, and on 23 September Alibaba added Qwen3.8 Max Prime, a higher-throughput tier of the same model at twice the price[2][4]
Notable features
- A sparse mixture-of-experts model with a hybrid attention mechanism, built on Qwen 3.5: 2.4T total parameters with 95B activated, 92 layers and 512 experts (10 routed plus 1 shared per token)[1][5]
- Specifications: a 1M-token context window with up to 131,072 output tokens on the cloud API (the open weights run 262,144 tokens natively, extensible to 1,010,000), text, image and video input, function calling, structured outputs, web search and context caching[2][5]
- Natively multimodal: Alibaba says it can turn hundred-page documents, full television series or 100-hour livestreams into searchable knowledge bases and rebuild a frontend project from one screenshot[1]
- Prices per million tokens: CNY 12 input and CNY 36 output on Model Studio in Beijing; Prime costs $4 input and $12 output[2][4]
- Alibaba's benchmark table puts it at 86.6 on Terminal Bench 2.1 against Fable 5's 84.6 and GPT-5.6 Sol's 88.8, and 67.7 on SWE-bench Pro against Fable 5's 80.0[5]
References 781% CONFIDENCEOverall confidence: 81%How well the pin's source and references back up its dates.Weighted average of how firmly 7 references, the source included, support the pin's start and end times; a reference counts half as much for every 180 days older than the newestShow all pins at 75% confidence or better
The first entry is always the pin's source. Overall confidence is a weighted average of how firmly each reference supports the start and end times used above; a reference counts half as much for every 180 days older than the newest.
- [1]90%alibabagroup.com/en-US/document-2021044032125272064alibabagroup.com· Posted Sep 28, 2026· Starts Aug 3, 2026 ✓· 16% of score
Alibaba Group's release is dated "August 3, 2026" and says "The model is now accessible via APIs on the Alibaba Cloud Model Studio for global developers, with model weights scheduled for release next week"; CNBC[6] reports Alibaba "unveiled" it "on Monday" 3 August.
- [2]90%qwen3.8-max Model Info - Alibaba Cloud Model Studiohelp.aliyun.com· Added Sep 28, 2026· 16% of score
Alibaba Cloud's model page: model ID qwen3.8-max, now the snapshot qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02); image, text and video input; a 1,000,000-token context window with 131,072 max output; function calling, structured outputs, web search and context caching; CNY 12 input and CNY 36 output per million tokens in Beijing.[1]
- [3]75%Qwen - Wikipediaen.wikipedia.org· Added Sep 28, 2026· 16% of score
Dates the cloud release to 3 August 2026, the open weights (Qwen3.8-2.4T-A95B, without image input or non-thinking mode) to 12 August, with a licence requiring providers earning more than US$50 million to get a commercial licence, and the Apache 2.0 Qwen3.8-27B to 14 August.[1]
- [4]70%Qwen3.8 Max Prime: Is it actually better? - DataNorth AIdatanorth.ai· Published Sep 24, 2026· 15% of score
Reports that Alibaba's Qwen team released Qwen3.8 Max Prime on 23 September 2026 as a higher-throughput version of Qwen3.8 Max at $4 per million input and $12 per million output tokens, "exactly twice the price of the standard model", and that it was not measurably faster in OpenRouter's first measurements.[1]
- [5]90%Qwen/Qwen3.8-2.4T-A95B · Hugging Facehuggingface.co· Published Aug 12, 2026· 13% of score
The open-weights model card: "For the first time, Qwen3.8 brings a Qwen-Max-class model to open release"; 2.4T total and 95B activated parameters, 92 layers, 512 experts (10 routed + 1 shared), 262,144 tokens natively extensible to 1,010,000, and benchmarks against Opus 4.8, Fable 5 and GPT-5.6 Sol.[1]
Suggest a correction
Something missing or wrong? Say it in your own words: a link that backs this pin up, a different start or end date and why, or a fact it lacks or gets wrong. The AI checks it against this pin's sources, searches for better ones, and adds any page that backs you up. The pin's own sources still count most. A picture that shows something else, or shows it badly, is looked at too, and moved down or replaced.