Skip to content
AI ReviewsQwen 3.8Qwen 3.8 MaxAlibabaChinese AI modelsAI benchmarksagentic codingopen weight AI

Qwen 3.8 Max Review: Alibaba's 2.4 Trillion Parameter Model Beats GPT-5.6 Sol and Claude Opus 4.8 on Key Coding Benchmarks

Alibaba published the full Qwen 3.8 Max benchmark report on qwen.ai. 2.4 trillion parameters, a leading score on PaperBench and IFBench, and a genuine SWE-bench Pro gap versus Claude Fable 5. Here is the complete, honest breakdown, methodology caveats included.

By Soufiane B.15 min read
Qwen 3.8 Max benchmark comparison chart against Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Pro, GPT-5.6 Sol, and Qwen 3.7 Max across coding, agentic, and general capability benchmarks

TL;DR

What actually shipped:

Alibaba's Qwen team published the full Qwen 3.8 Max technical benchmark report on qwen.ai, replacing the bare 'second only to Fable 5' teaser from the July 19 WAIC preview with a genuine, methodology-documented benchmark table against Opus 4.8, Fable 5, Gemini 3.1 Pro, GPT-5.6 Sol, and its own predecessor Qwen 3.7 Max.

The strongest confirmed wins:

Qwen 3.8 Max leads the published table on PaperBench (93.0), IFBench (82.8), and its own internal QwenReactBench and QwenSVGBench frontend generation tests, ahead of every model in the comparison including GPT-5.6 Sol and Claude Fable 5.

The honest gaps that remain:

On SWE-bench Pro, Qwen 3.8 Max scores 67.7 against Claude Fable 5's 80.0, a 12.3 point gap. On HLE it scores 43.6 versus Fable 5's 53.3. GPQA Diamond sits at 92.6, essentially level with Opus 4.8 and Fable 5 but behind GPT-5.6 Sol's 94.1.

The methodology caveat that matters most:

A meaningful share of the table, including QwenSWEBench, QwenQoderBench, CoWorkBench, SkillsBench, and AndroidBench, are Alibaba's own internal benchmarks, not independently run public evaluations. Treat those rows as directional, not verified against third-party evaluators like Artificial Analysis or LMArena.

Still no open weights, still no standard API price:

Active parameter count remains undisclosed for the sparse Mixture-of-Experts architecture. Open weights are still promised as 'coming soon' with no date, and access continues to run through Alibaba's credit-based Token Plan rather than a published per-token API rate.

The competitive timing:

This full benchmark release lands weeks after Moonshot's Kimi K3 open-weight launch and Alibaba's own bare WAIC teaser, continuing the pattern of Chinese labs racing to publish evidence, then more evidence, in response to each other within the same news cycle.

Qwen 3.8 Max: The Full Benchmark Report, Read Honestly

When Alibaba's Qwen team walked on stage at the World AI Conference in Shanghai on July 19, 2026, two days after Moonshot AI's Kimi K3 open-weight launch rattled global chip stocks, they brought a big number and not much else. 2.4 trillion parameters. A claim of being "second only to Fable 5." No benchmark table, no model card, no disclosed active parameter count. It read like a direct competitive response, timed to reclaim a news cycle rather than to present finished evidence.

That gap has now closed, at least partially. Alibaba has published a full technical benchmark report for Qwen 3.8 Max on qwen.ai, comparing it directly against Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Pro, GPT-5.6 Sol, and its own predecessor Qwen 3.7 Max across dozens of coding, agentic, and general capability evaluations. This is a real step forward from the July teaser. It is also, on close reading, still a document with real limits worth understanding before you route any production decision around it.

Here is the complete breakdown, with every claim traced back to what Alibaba's own methodology notes actually say.


Coding Agent Benchmarks: Where Qwen 3.8 Max Actually Leads, and Where It Doesn't

This is the category Alibaba leans on hardest, and the results are genuinely mixed rather than a clean sweep in either direction.

Benchmark Opus 4.8 Fable 5 GPT-5.6 Sol (max) Qwen 3.7 Max Qwen 3.8 Max
Terminal-Bench 2.1 84.6 84.6 88.8 74.5 86.6
SWE-bench Pro 69.2 80.0 64.6 60.6 67.7
DeepSWE 1.1 59.0 70.0 73.0 21.6 56.6
FrontierSWE 70.0 88.8 74.0 40.7 73.5
MLS-Bench-Lite 42.8 49.9 46.2 31.7 41.0
PaperBench 80.3 88.8 90.5 64.8 93.0 (leads)
AndroidBench 69.8 84.5 74.0 56.5 75.1

Two results here deserve to be read closely rather than skimmed.

PaperBench is the standout win. Qwen 3.8 Max scores 93.0, ahead of every other model including GPT-5.6 Sol's 90.5 and Claude Fable 5's 88.8. PaperBench evaluates research reproduction, whether a model can take a published paper's methodology and correctly reproduce the described results. This is a real, meaningful lead, and it is also the largest single jump from Qwen 3.7 Max in the entire table, up from 64.8 to 93.0.

SWE-bench Pro is the honest counterweight. Claude Fable 5 leads decisively at 80.0, with Qwen 3.8 Max at 67.7, a 12.3 point gap on one of the field's most cited real-world software engineering benchmarks. Claude Opus 4.8 also sits ahead of Qwen 3.8 Max here at 69.2. On DeepSWE and FrontierSWE, the pattern repeats: Qwen 3.8 Max improves substantially over Qwen 3.7 Max but still trails Fable 5 and, on FrontierSWE specifically, GPT-5.6 Sol as well.

The methodology footnotes matter here too. Alibaba's own documentation states that Terminal-Bench 2.1, SWE-bench Pro, and DeepSWE were evaluated using the Claude Code harness for Qwen's own models, while Opus 4.8 and Fable 5's scores on Terminal-Bench come from Artificial Analysis's published evaluation and GPT-5.6 Sol's score comes from OpenAI's own Codex-based preview materials. This is not a single controlled run across every model on identical infrastructure. It is a composite of Alibaba's own testing plus each competitor's best previously published number, which is a common but meaningfully weaker standard of evidence than a true head-to-head evaluation.


Alibaba's Own Internal Benchmarks: Real Signal, Unverified Externally

A significant share of the coding table is not built on public, third-party-auditable benchmarks at all.

QwenSWEBench, QwenQoderBench, QwenReactBench, and QwenSVGBench are described in Alibaba's own footnotes as internal evaluation suites, built and scored entirely inside Alibaba. On these, Qwen 3.8 Max performs strongly: 80.7 on QwenSWEBench and 58.4 on QwenQoderBench, both ahead of Qwen 3.7 Max and GPT-5.6 Sol. On QwenReactBench, a frontend generation test, Qwen 3.8 Max scores 1724, ahead of Opus 4.8's 1694 and GPT-5.6 Sol's 1564, though behind Fable 5's 1770. On QwenSVGBench, a dual English and Chinese SVG generation test, Qwen 3.8 Max scores 1713 against GPT-5.6 Sol's leading 1758 and Fable 5's 1690.

None of these four benchmarks have been independently replicated by an outside evaluator. That does not make them meaningless: internal benchmarks built around real product use cases can be genuinely informative about how a model performs on the specific workflows a company cares about. It does mean you should treat these numbers as Alibaba's own account of its model's strengths rather than a neutral, third-party-verified ranking, and weight them accordingly when making a purchasing or infrastructure decision.


General Agent Benchmarks

Benchmark Opus 4.8 Fable 5 GPT-5.6 Sol (max) Qwen 3.7 Max Qwen 3.8 Max
CoWorkBench 72.3 75.9 71.5 64.6 74.8
WorkSpaceBench 66.8 68.7 65.6 61.4 67.7
JobBench 48.4 57.4 45.4 31.3 53.4
SkillsBench 65.1 70.9 73.5 61.2 70.2
Agents' Last Exam (Pass / Score) 27.0 / 45.1 -- / -- 30.6 / 53.6 11.8 / 31.1 27.0 / 52.4
Toolathlon Verified (Pass@1) 76.2 77.9 74.9 49.7 72.5
WideSearch 72.9 81.2 -- 75.2 81.9 (leads)
HLE w/ tools 57.9 64.5 58.0 53.5 56.2

Qwen 3.8 Max sits in a genuinely competitive middle tier here. It leads on WideSearch (81.9, ahead of even Fable 5's 81.2) and comes second only to Fable 5 on JobBench (53.4 versus 57.4), but trails on Toolathlon Verified, SkillsBench, and HLE with tools. On Agents' Last Exam, a benchmark specifically designed to test long-horizon, high-value professional tasks, Qwen 3.8 Max scores 52.4, just behind GPT-5.6 Sol's leading 53.6 and clearly ahead of both Qwen 3.7 Max's 31.1 and Opus 4.8's 45.1. Claude Fable 5's cell is blank on this benchmark, one of several gaps in the table Alibaba's own footnote attributes to that model's results not being fully disclosed.


General Capabilities: The Closest Competition in the Whole Table

Benchmark Opus 4.8 Fable 5 GPT-5.6 Sol (max) Qwen 3.7 Max Qwen 3.8 Max
GPQA Diamond 92.0 92.6 94.1 92.4 92.6
HLE 45.7 53.3 47.2 41.4 43.6
IFBench 62.2 63.5 72.7 79.1 82.8 (leads)
HealthBench 52.4 -- 55.3 54.5 60.2
PLawBench 69.6 70.2 72.3 58.9 73.2 (leads)
MRCR v2 256K (8-needle) 83.2 -- 93.8 86.7 92.9
LongBench v2 69.1 -- 67.1 65.3 66.3

This is where the results get genuinely interesting, because the gaps compress and in some cases invert.

GPQA Diamond, a graduate-level science reasoning benchmark, is essentially a three-way tie. Opus 4.8 at 92.0, Fable 5 at 92.6, and Qwen 3.8 Max also at 92.6 sit within half a point of each other, with only GPT-5.6 Sol clearly ahead at 94.1.

IFBench is Qwen's clearest and most consistent win. Instruction-following precision is where Qwen 3.8 Max leads outright at 82.8, and notably, Qwen 3.7 Max was already ahead of every closed model here at 79.1 before this release. This is a two-generation pattern of genuine strength, not a one-off result.

PLawBench, a legal reasoning benchmark, also goes to Qwen 3.8 Max, at 73.2, narrowly ahead of GPT-5.6 Sol's 72.3 and Fable 5's 70.2.

HLE is where the gap with Fable 5 is largest in this section. Humanity's Last Exam, one of the hardest general knowledge and reasoning benchmarks currently in use, shows Fable 5 at 53.3 against Qwen 3.8 Max's 43.6, a 9.7 point gap that closely tracks the same pattern seen on SWE-bench Pro. Fable 5 remains the strongest model in this comparison set on the hardest, most demanding general reasoning tasks.

Long-context performance is strong but not the leader. On MRCR v2 at 256K tokens with an 8-needle retrieval test, Qwen 3.8 Max scores 92.9, just behind GPT-5.6 Sol's 93.8 and a meaningful jump over Qwen 3.7 Max's 86.7.


Multimodal and Visual Capabilities

Alibaba's report extends into a separate, larger table covering multimodal reasoning, visual agent and coding tasks, document intelligence, spatial understanding, visual perception, and video intelligence, comparing Qwen 3.8 Max against Opus 4.8, Fable 5, Gemini 3.1 Pro, and GPT-5.6 Sol, alongside Qwen 3.7 Plus specifically rather than Qwen 3.7 Max, since Alibaba's multimodal work has historically run through the Plus line.

The headline pattern holds across this section: Qwen 3.8 Max posts the strongest reported scores on MMMU-Pro, a demanding multimodal reasoning benchmark, ahead of Opus 4.8, Fable 5, Gemini 3.1 Pro, and GPT-5.6 Sol in Alibaba's own comparison. Alibaba also reports leading results across several video intelligence and visual agent and computer-use evaluations in the same table, including OSWorld-Verified.

Given the density of this particular table, which pairs a large number of tool-assisted and non-tool-assisted values in the same cells, we are treating the specific multimodal figures with more caution than the cleaner coding and general capability tables above, and recommend checking Alibaba's own published report directly for exact multimodal numbers before citing them in a technical comparison. The directional signal, that Qwen 3.8 Max is a genuinely strong multimodal system and not just a text model with vision bolted on, is clear regardless of the exact digits.


What Alibaba Still Has Not Told You

Three things are missing from this report that would matter for any real infrastructure decision.

The active parameter count. Qwen 3.8 Max is a sparse Mixture-of-Experts model with 2.4 trillion total parameters, but Alibaba has not disclosed how many of those parameters activate per token. For an MoE architecture, that number is the single figure that actually determines real-world inference cost and latency. Without it, the 2.4 trillion headline is closer to a marketing number than a complete engineering spec.

Open weights. Alibaba has said open weights are "coming soon" since the original July 19 preview, with no date and no license named at any point since. Every previous Qwen Max-tier release, including Qwen 3.7 Max, has stayed closed while a separate line, Qwen 3.6, carried Apache 2.0 open releases. Whether Qwen 3.8 Max actually breaks that pattern remains genuinely uncertain.

A standard per-token API price. Access currently runs through Alibaba's Token Plan, a credit-based subscription system rather than a conventional dollar-per-million-tokens rate card. Reported tiers range from roughly $6 for a smaller weekly credit allotment up to $68 for a higher tier with more concurrent agent capacity, but this is a promotional access structure, not the kind of stable API pricing you would build a production cost model around.


Qwen 3.8 Max vs Qwen 3.7 Max: The Clearest Story in the Report

Setting aside the competitive comparison against Claude and OpenAI's models for a moment, the improvement within Alibaba's own lineup is the most unambiguous result in this entire release.

Terminal-Bench 2.1 jumps from 74.5 to 86.6. SWE-bench Pro moves from 60.6 to 67.7. PaperBench goes from 64.8 to 93.0, the largest single jump in the whole table. DeepSWE improves from 21.6 to 56.6, more than doubling. IFBench, already Qwen 3.7 Max's strongest category relative to the competition, climbs further from 79.1 to 82.8.

This is a genuine generational upgrade, not a parameter-count increase with marginal returns. If you are currently running production workloads on Qwen 3.7 Max, testing Qwen 3.8 Max against your specific tasks is worth the effort based on Alibaba's own numbers alone, independent of how it stacks up against Claude or GPT-5.6.


The Honest Verdict

Qwen 3.8 Max is a real step up from both Qwen 3.7 Max and from the bare teaser Alibaba shipped on July 19. The benchmark report is more substantial than most Chinese lab launches this year, with documented methodology, explicit footnotes about evaluation frameworks, and enough internal consistency across the tables that it reads as a genuine technical report rather than pure marketing.

It is not, however, independently verified, and it does not represent a clean win over the current frontier. Claude Fable 5 remains ahead on SWE-bench Pro, FrontierSWE, and HLE by margins that are too large to dismiss as noise. GPT-5.6 Sol leads GPQA Diamond and long-context retrieval. Where Qwen 3.8 Max genuinely stands out, PaperBench, IFBench, PLawBench, and several of its own internal coding benchmarks, those wins are real but concentrated in specific categories rather than spread evenly across the board.

For teams evaluating Qwen 3.8 Max today: test it directly against your own workload rather than relying on any single benchmark table, including this one and including Alibaba's own. If your use case leans toward instruction-following precision, research reproduction, or frontend code generation, the published numbers suggest genuine strength worth investigating. If your use case is heavy, real-world software engineering closer to what SWE-bench Pro measures, Claude Fable 5 and Claude Opus 4.8 remain the better-evidenced choice based on what is public today.


Compare Qwen 3.8 Max Against the Full Field

Want to see how Qwen 3.8 Max stacks up against Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, Kimi K3, and DeepSeek V4 across pricing, context windows, and benchmark scores as independent verification arrives?

Compare AI models on Renovate QR

The /tools directory is updated as new benchmark data lands, including independent evaluation of Qwen 3.8 Max as Artificial Analysis and LMArena publish their own results.


Last updated August 3, 2026, based on Alibaba's official Qwen 3.8 Max technical benchmark report published on qwen.ai. We will update this article with independent third-party verification, active parameter disclosure, and standard API pricing as they become available.

Frequently Asked Questions

What is Qwen 3.8 Max and is it fully released yet?

Qwen 3.8 Max is Alibaba's flagship large language model, a 2.4 trillion parameter sparse Mixture-of-Experts system first previewed on July 19, 2026 at the World AI Conference in Shanghai. That initial preview shipped with no benchmark table, no model card, and no disclosed active parameter count, just a headline claim that the model was "second only to Fable 5." Alibaba has since published a full technical benchmark report on qwen.ai comparing Qwen 3.8 Max against Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Pro, GPT-5.6 Sol, and Qwen 3.7 Max across dozens of coding, agentic, and general capability evaluations. The model remains accessible as a preview through Alibaba's Token Plan; open weights and a standard per-token API price have still not been published.

Does Qwen 3.8 Max actually beat Claude Fable 5 and GPT-5.6 Sol?

On some benchmarks, yes. On others, no, by a meaningful margin. Qwen 3.8 Max leads the published table on PaperBench (93.0 versus GPT-5.6 Sol's 90.5 and Fable 5's 88.8), IFBench (82.8, ahead of both), and several of Alibaba's own internal coding benchmarks like QwenReactBench and QwenSVGBench. It trails clearly on SWE-bench Pro, where Claude Fable 5 scores 80.0 against Qwen 3.8 Max's 67.7, and on HLE, where Fable 5 leads 53.3 to 43.6. On GPQA Diamond, all three models, Opus 4.8, Fable 5, and Qwen 3.8 Max, land within half a point of each other, while GPT-5.6 Sol leads that specific benchmark outright at 94.1. There is no single winner across the board.

How much of the Qwen 3.8 Max benchmark table is independently verified?

Less than half. Alibaba's own footnotes disclose that benchmarks like Terminal-Bench 2.1, SWE-bench Pro, and DeepSWE were run using the Claude Code evaluation harness, with other labs' scores pulled from their own previously published best results rather than a single controlled run across all models. A significant portion of the table, including QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, CoWorkBench, SkillsBench, AndroidBench, and Automation-Bench, are Alibaba's own internal evaluation suites with no external replication. As of this report, no independent evaluator such as Artificial Analysis or LMArena has published its own Qwen 3.8 Max score, so the ranking claims in Alibaba's table should be read as vendor-reported evidence, not confirmed fact.

Is Qwen 3.8 Max open source?

Not yet. Alibaba has stated open weights are "coming soon" since the original July 19 preview, but has not published a release date or license. Every previous Qwen Max-tier flagship, including Qwen 3.7 Max, has stayed API-only while a separate line, Qwen 3.6, carried the open-weight Apache 2.0 releases. Whether Qwen 3.8 Max actually breaks that pattern remains an open question rather than a confirmed roadmap item.

How does Qwen 3.8 Max compare to Qwen 3.7 Max?

The gap is substantial and consistent across nearly every benchmark in Alibaba's own table. On Terminal-Bench 2.1, Qwen 3.8 Max scores 86.6 versus Qwen 3.7 Max's 74.5. On SWE-bench Pro, 67.7 versus 60.6. On PaperBench, 93.0 versus 64.8, the single largest jump in the entire report. On IFBench, 82.8 versus 79.1, a smaller but still real improvement where Qwen 3.7 Max was already ahead of Claude and GPT-5.6 Sol. This is a genuine generational upgrade within Alibaba's own lineup, not just a larger parameter count with marginal gains.

What is the active parameter count and pricing for Qwen 3.8 Max?

Both remain undisclosed. Qwen 3.8 Max is a 2.4 trillion total parameter sparse Mixture-of-Experts model, but Alibaba has not published how many parameters activate per token, which is the figure that actually determines real inference cost for an MoE architecture. Access currently runs through Alibaba's credit-based Token Plan, with tiers reported from roughly $6 for a weekly allotment up to $68 for a higher tier with more concurrent agents, rather than a published dollar-per-million-token API rate. Treat any specific per-token price you see circulating as an estimate until Alibaba publishes an official rate card.

Published

Related Articles