OpenAI Decisions API vs Microsoft-Decision-1 vs Liquid d1
The short answer
Three decision-model options shipped in the same week of October 2026, and they split cleanly by where the decision has to run. As of October 11, 2026:
- Microsoft-Decision-1 — cheapest hosted decision API: $0.042 per million input tokens, output free, on Microsoft Foundry and OpenRouter. Text or JSON in.
- OpenAI Decisions API — the hosted option that also reads images: $0.10 per million input tokens on
gpt-6-luna, no output or cache charges. Public beta since October 6, 2026. - Liquid AI d1-3B — open weights you run yourself: 3.12B parameters, text and images, about 8 ms per decision on an RTX 4090 and 50 ms on a Jetson Orin Nano.
All three answer the same three kinds of question — is this true?, which option?, how much? — and return probabilities your code can threshold.
The comparison
| OpenAI Decisions API | Microsoft-Decision-1 | Liquid d1-3B | |
|---|---|---|---|
| Released | Public beta October 6, 2026 | October 2026 (Foundry docs dated October 7) | Open weights October 7, 2026 |
| Model | gpt-6-luna (only option) | Post-trained from Qwen3.5-9B | Built on LFM2.5-VL-3B, 3.12B params |
| Inputs | Text and images (inline base64 only) | Text or JSON state | Text, JSON and images; 32,768-token context |
| Question types | predicate, choice, score | noul, choice, score | Yes/no, choice, score |
| Price | $0.10 / 1M input; output, cache read and write free | $0.042 / 1M input; output free | Free weights; you pay for the hardware |
| Where it runs | OpenAI API (POST /v1/decisions) | Microsoft Foundry (GlobalStandard, DataZoneStandard), OpenRouter | Your GPU, Mac or Jetson; llama.cpp and Transformers |
| Vendor speed claim | About 10x faster than GPT-6 Luna via Responses | P50 about 35x faster than GPT-6 Sol | 8 ms on RTX 4090, 30 ms on Apple M5 Pro |
| Status | Beta; GA “in the coming weeks” | Generally available in Foundry | Research release; d1-omni-600M experimental |
Picks by job
Routing support tickets or agent requests inside Azure: Microsoft-Decision-1. It is less than half the price of OpenAI’s option on input and Microsoft publishes robustness numbers: the decision flips on 1.3% of paraphrased or reordered requests on average, and never when options are shuffled. DataZoneStandard deployments keep inference inside one data zone, which matters for EU data. Microsoft says it will rebase the model on its own MAI and OpenAI models later.
Checking photos or screenshots: OpenAI Decisions API. It is the only hosted option here that takes images — product-damage checks, UI-state checks for computer-use agents, document-type classification from a scan. Two limits: images must be inline base64 data URLs (no hosted URLs or file_id), and decisions that depend on an earlier answer need separate requests.
On-device, edge or no-cloud decisions: Liquid d1-3B. Liquid reports 48.57 on the Decision Index v0.2.1, ahead of every model under 10B and level with the 35B Decider model. It answers three questions over one state in 21 ms on an RTX 4090. The smaller d1-omni-600M adds audio input but scores 15.95 on the same index, so treat it as an experiment. Liquid’s announcement does not name the licence text, so read the model card before shipping commercially.
What decision models are good for
- Model and tool routing — pick the cheapest model that can handle a request (see our LLM router and gateway comparison).
- Agent guardrails — “does this tool call touch production data?” as a probability, with a threshold that sends borderline calls to a human.
- LLM-as-a-judge — score an answer against a rubric without paying for a written critique.
- Moderation and prompt-injection screening before content reaches the main model.
- Labelling — Microsoft’s Xbox research team sorted 10,000+ pieces of player feedback into fixed themes and reports quality competitive with GPT-6 Sol at more than 14 times the speed.
They are the wrong tool for extracting fields or writing explanations; use structured JSON output from an LLM for that.
Cost at scale
A routing step that reads a 1,500-token request costs about $0.00015 on the OpenAI Decisions API and $0.000063 on Microsoft-Decision-1. At 10 million decisions a month that is roughly $1,500 versus $630, before any LLM call. A self-hosted d1-3B on one RTX 4090 handles about 475 packed states per second by Liquid’s measurement, so a single workstation-class GPU covers most routing volumes. Compare with LLM prices on our current API prices page.
How to choose in three questions
- Does the decision need to see an image? Yes → OpenAI Decisions API or d1-3B.
- Can the data leave your network? No → d1-3B on your hardware, or Microsoft-Decision-1 with
DataZoneStandard. - Is cost per decision the main constraint? Yes → Microsoft-Decision-1 hosted, d1-3B self-hosted.
Java teams can already call OpenAI’s endpoint: LangChain4j 1.22.0 added OpenAiDecisionModel.
Last verified: October 11, 2026.