AI agents · OpenClaw · self-hosting · automation

Quick Answer

Qwen3.8-27B vs GLM-5.2 vs Kimi K3 (Aug 2026)

Published:

The Short Answer

For self-hosting on your own hardware in August 2026: GLM-5.2 is the proven, easy-to-quantize workhorse available today; Kimi K3 (open weights July 27, 2026) leads long-context coding but is heavier; Qwen3.8-27B (open weights due week of Aug 10, 2026) is the freshest single-GPU option with Alibaba’s agentic tuning. Choose by VRAM and task.

Side-by-Side

Qwen3.8-27BGLM-5.2Kimi K3
VendorAlibabaZhipuMoonshot
Open weightsDue week of Aug 10, 2026AvailableAvailable (Jul 27, 2026)
Local fitSingle high-end GPU (24-48GB)24GB with 4-bit quantHeavier; multi-GPU or aggressive quant
Best atAgentic tasks, multimodal, fresh tuningGeneral coding, reliability, quant supportLong-context coding
LicenseTBD at launchPermissiveOpen weights

How To Choose

  • Best long-context coding you can host → Kimi K3 — strong on large codebases, but budget the VRAM (K3 API is $3/$15 if you’d rather not host).
  • Easiest, most reliable single-GPU workhorse → GLM-5.2 — mature quantization, broad tooling support, permissive license. The safe default.
  • Freshest, agentic, multimodal → Qwen3.8-27B — wait for the week-of-Aug-10 drop and the license terms, then test. Pairs conceptually with the Qwen3.8-Max flagship.

The Trade-Offs

  • Availability: GLM-5.2 and Kimi K3 are downloadable now; Qwen3.8-27B is imminent, not shipped, as of Aug 6, 2026.
  • License clarity: GLM-5.2 and Kimi K3 have known terms. Qwen3.8-27B’s license was undisclosed at the Aug 3 Max launch — verify before commercial use.
  • VRAM reality: “27B” and “K3” mean very different memory footprints. Match the quant to your card before downloading.

Verdict

  • Ship today, single GPU → GLM-5.2
  • Long-context coding → Kimi K3
  • Freshest agentic pick (pending drop + license) → Qwen3.8-27B

Sources