Qwen3.8-Max vs GPT-5.6 Sol: Computer Use (Aug 2026)
The Short Answer
As of August 2026, Alibaba’s Qwen3.8-Max (released Aug 2, 2026) makes a bold claim: it beats GPT-5.6 Sol Max and Claude Fable 5 on agentic computer use. The benchmarks confirm it — on desktop computer use (OSWorld), Qwen3.8-Max wins; on terminal coding (Terminal-Bench), GPT-5.6 Sol still leads.
What Is Qwen3.8-Max?
Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model with 95B active parameters, a 1M-token context window, and native text, image, and video input. Alibaba published a full benchmark table claiming wins over GPT-5.6 Sol and Fable 5 across 7 coding/general tasks and 36 multimodal benchmarks.
Benchmark Head-to-Head
| Benchmark | Qwen3.8-Max | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|
| OSWorld-Verified (computer use) | 86.1 | 83.2 (Sol Max) | 85.0 |
| Terminal-Bench 2.1 (agentic coding) | 86.6 | 88.8 | 84.6 |
| SWE-bench Pro | 67.7% | — | 80.0 |
| GPQA Diamond | 92.6% | — | — |
| PaperBench | 93.0 | 90.5 | — |
| FrontierSWE | 73.5 | — | — |
Where Qwen3.8-Max Wins
- OSWorld-Verified (86.1) — desktop computer-use agents interacting with real apps. Qwen leads GPT-5.6 Sol Max (83.2) and Fable 5 (85.0).
- PaperBench (93.0 vs 90.5) and AndroidBench — multimodal and mobile agentic tasks.
- FrontierSWE (73.5) — beats the field on this harder SWE variant.
Where GPT-5.6 Sol Wins
- Terminal-Bench 2.1 (88.8 vs 86.6) — GPT-5.6 Sol still leads agentic terminal coding.
- SWE-bench Pro — Claude Fable 5 leads decisively at 80.0; Qwen sits at 67.7%, so for classic repo-fix SWE work, Fable 5 (or GPT-5.6 Sol) is stronger.
The Reality Check
Qwen3.8-Max’s headline “beats GPT-5.6 Sol” is true on computer-use and multimodal benchmarks, but GPT-5.6 Sol and Claude Fable 5 still lead on the core coding benchmarks (Terminal-Bench, SWE-bench Pro). These are Alibaba-reported numbers — independent verification is ongoing.
Which Should You Pick?
- Desktop/computer-use agents + multimodal → Qwen3.8-Max.
- Terminal/agentic coding → GPT-5.6 Sol.
- Classic repo-fix SWE work → Claude Fable 5 (or Opus 5).
Sources
- VentureBeat — Qwen3.8-Max beats GPT-5.6 Sol Max and Fable 5 on agentic computer use: venturebeat.com
- DataCamp — Qwen3.8-Max features, benchmarks, pricing: datacamp.com
- Neowin — Alibaba releases Qwen3.8-Max: neowin.net