Quick Answer
Qwen3.8-27B vs GLM-5.2 vs Kimi K3 (Aug 2026)
The Short Answer
For self-hosting on your own hardware in August 2026: GLM-5.2 is the proven, easy-to-quantize workhorse available today; Kimi K3 (open weights July 27, 2026) leads long-context coding but is heavier; Qwen3.8-27B (open weights due week of Aug 10, 2026) is the freshest single-GPU option with Alibaba’s agentic tuning. Choose by VRAM and task.
Side-by-Side
| Qwen3.8-27B | GLM-5.2 | Kimi K3 | |
|---|---|---|---|
| Vendor | Alibaba | Zhipu | Moonshot |
| Open weights | Due week of Aug 10, 2026 | Available | Available (Jul 27, 2026) |
| Local fit | Single high-end GPU (24-48GB) | 24GB with 4-bit quant | Heavier; multi-GPU or aggressive quant |
| Best at | Agentic tasks, multimodal, fresh tuning | General coding, reliability, quant support | Long-context coding |
| License | TBD at launch | Permissive | Open weights |
How To Choose
- Best long-context coding you can host → Kimi K3 — strong on large codebases, but budget the VRAM (K3 API is $3/$15 if you’d rather not host).
- Easiest, most reliable single-GPU workhorse → GLM-5.2 — mature quantization, broad tooling support, permissive license. The safe default.
- Freshest, agentic, multimodal → Qwen3.8-27B — wait for the week-of-Aug-10 drop and the license terms, then test. Pairs conceptually with the Qwen3.8-Max flagship.
The Trade-Offs
- Availability: GLM-5.2 and Kimi K3 are downloadable now; Qwen3.8-27B is imminent, not shipped, as of Aug 6, 2026.
- License clarity: GLM-5.2 and Kimi K3 have known terms. Qwen3.8-27B’s license was undisclosed at the Aug 3 Max launch — verify before commercial use.
- VRAM reality: “27B” and “K3” mean very different memory footprints. Match the quant to your card before downloading.
Verdict
- Ship today, single GPU → GLM-5.2
- Long-context coding → Kimi K3
- Freshest agentic pick (pending drop + license) → Qwen3.8-27B
Sources
- Medium — Qwen3.8-27B local release preview: medium.com
- MarkTechPost — Qwen3.8 family release: marktechpost.com
- Latent Space — Qwen3.8-Max 2.4T and 27B analysis: latent.space