GLM-5.3-Flash vs Gemini 3.8 Flash vs Luna: Cheapest
The Short Answer
Three cheap-tier models, as of September 5, 2026:
| Model | Input | Output | Released | Watch out for |
|---|---|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.50 | Aug 26, 2026 | Promo expired Sep 9 |
| GPT-5.6 Luna | $0.20 | $1.20 | Cut 80% Jul 30, 2026 | — |
| Gemini 3.8 Flash | $0.75 | $3.75 | Sep 2, 2026 | ⚠️ Doubles Jan 1, 2027 |
The list price is the wrong number. On measured cost per completed task, the order inverts — Luna is the cheapest at $0.29, Gemini 3.8 Flash the most expensive at $2.04, despite a 3.75x gap in list input price running the other way.
GLM-5.3-Flash — Cheapest List Price
$0.15 / $0.50 per MTok · cached input $0.03 · released August 26, 2026
Z.ai’s natively multimodal small model in the GLM-5 family — and the one that spent late August and early September running anonymously in coding tools as “Ox Alpha” before the reveal. A 50% launch promotion cut it to roughly $0.075 / $0.25 / $0.015 through September 9, 2026 at 24:00 UTC+8; that window has closed and list rates now apply.
On the Artificial Analysis Intelligence Index, GLM 5.3 scores 59.4 — ahead of Gemini 3.8 Flash’s 58.7 and Qwen 3.8 Max’s 57.7. That is genuinely competitive, not a discount-tier caveat.
Strengths: cheapest list price here, natively multimodal rather than vision bolted onto a text model, and cached input at $0.03 is aggressive for repeated-context workloads.
⚠️ Limits: open weights have not shipped — Z.ai has staged GLM-5 weight releases rather than launching with them, so do not plan a self-hosting roadmap around it. Chinese API endpoints may be excluded by your data-residency rules. And tracker figures near $0.071 / $0.238 circulate alongside Z.ai’s own $0.15 / $0.50 listing, reflecting currency conversion and reseller spreads — cite whichever source your invoice reflects.
GPT-5.6 Luna — Best Cost Per Task
$0.20 / $1.20 per MTok · price cut 80% on July 30, 2026
Luna’s list price sits in the middle of this group and its output price is more than double GLM-5.3-Flash’s. It still wins the metric that matters.
On the Artificial Analysis coding-agent harness, Luna at max effort completes a task for $0.29 at index 57.1 — the cheapest measured per-task cost of any model tracked, roughly seven times cheaper per task than Gemini 3.8 Flash despite costing about a quarter as much per token in the other direction.
The reason is token efficiency. Luna does not burn tens of thousands of output tokens reasoning its way around a task.
Strengths: lowest measured cost per completed task, OpenAI ecosystem and tooling, no pending price change.
Limits: index 57.1 is the lowest quality tier here — fine for bounded, well-specified work, wrong for open-ended agent tasks.
Gemini 3.8 Flash — Most Capable, Most Expensive Per Task
$0.75 / $3.75 per MTok · released September 2, 2026 · 1M context, 64K max output
The strongest model of the three on capability breadth, and the only one accepting text, image, audio and video input. It scores 58.7 on the Intelligence Index, roughly 89–91% on Terminal-Bench 2.1 (up from Gemini 3.7 Flash’s 81.6%), and 61.4% on Vals Finance Agent v2 — ahead of every model on Google’s own launch table for that row.
⚠️ Two hard warnings.
The price doubles. $0.75 / $3.75 is introductory through December 31, 2026. On January 1, 2027 it becomes $1.50 / $7.50. Anything still running next year should be budgeted at the doubled rate, which takes it from roughly 5x GLM-5.3-Flash’s input price to roughly 10x.
It collapses on long-horizon work. On Terminal-Bench 4.0, the harder long-running agent suite, Gemini 3.8 Flash scores 19.1% against Claude Opus 5’s 51.8%. On OSWorld-2.0 computer use it manages 59.0% against Opus 5’s 70-plus. It is an excellent scoped-task model and a poor autonomous agent.
Strengths: widest input modality, strongest coding scores of the three, batch and flex tiers at half price.
Limits: the January doubling, ~48,000 output tokens per task, and the long-horizon collapse.
The Cost-Per-Task Table That Should Decide It
Measured on the Artificial Analysis coding-agent harness, September 2026 — these are measured figures, not price-sheet arithmetic:
| Model · effort | Index | Cost per task |
|---|---|---|
| GPT-5.6 Luna (max) | 57.1 | $0.29 |
| Gemini 3.7 Flash (high) | 59.6 | $1.27 |
| GPT-6 Astra (low) | 62.6 | $1.41 |
| Muse Spark 1.3 (xhigh) | 64.2 | $1.72 |
| Gemini 3.8 Flash (high) | 61.1 | $2.04 |
| GPT-6 Astra (xhigh) | 67.0 | $3.27 |
Two uncomfortable observations for anyone shopping on the price sheet:
- Gemini 3.8 Flash costs more per task than Gemini 3.7 Flash ($2.04 vs $1.27) at an identical list price, because 3.8 reasons longer.
- GPT-6 Astra at low effort ($1.41) is cheaper per task than Gemini 3.8 Flash ($2.04) while scoring higher — despite a 13x higher list price. Astra emits roughly 2,200–14,000 output tokens per task against Flash’s ~48,000.
The general rule for 2026: list price stopped predicting spend. Running the full Intelligence Index took GPT-6 Astra about 16M output tokens against roughly 123M for Gemini 3.8 Flash — a 7.7x efficiency gap that swamps token-price differences entirely.
Which Should You Pick?
| If you need… | Use |
|---|---|
| Lowest measured cost per task | GPT-5.6 Luna |
| Lowest list price with multimodal input | GLM-5.3-Flash |
| Text, image, audio and video input | Gemini 3.8 Flash |
| Best coding scores in this tier | Gemini 3.8 Flash |
| Anything running into 2027 | Not Gemini 3.8 Flash without re-budgeting |
| Long-horizon autonomous agents | None of these — use a frontier model |
The method that beats all of this: run 50–100 tasks from your own backlog through each candidate and measure what they actually cost you. Token consumption is a property of your workload as much as the model, and no leaderboard knows your workload.
Last verified: September 5, 2026.