TL;DR
Kev (jaredpalmer/kev) is a family of four open-weight decision models from Jared Palmer, Vercel’s VP of AI. Each model reads a text plus typed questions and returns calibrated probabilities over the answers you supply: yes/no (noul), multiple choice (choice) or a rating (score), in one forward pass, with no generated text. The server speaks TypeSafe’s System One API, so code written for Jev can switch by changing the base URL and model name.
As of October 7, 2026 the repository has 8,626 stars, 567 forks and 19 open issues. The latest release is Kev 1.0 (October 1, 2026); code and weights are Apache-2.0. The repo was created on September 17, 2026.
- Four sizes: Kev-0.8B, 4B and 9B are LoRA adapters plus a small pointer head on Qwen3.5 base models; Kev-27B is a full fine-tune of Qwen3.8-27B (51 GB of bf16 weights).
- The project’s own numbers: on datasets no Kev trained on, Kev-27B scores 0.851 accuracy against Jev’s 0.857 (development split), and 52.3 against Jev’s 54.0 on a chance-corrected index across 14 public datasets.
- Our lab result: a clean install took 56 seconds on a CPU-only runner, and Kev-0.8B served the README’s example ticket on CPU, though it routed it differently from the README’s Kev-4B output.
Who should try it: anyone paying per call for routing, triage or guardrail decisions who wants weights they can host and fine-tune. Who can skip it: teams without a GPU or a 32 GB Apple Silicon Mac, and anyone needing knowledge-heavy answers rather than classification.
What problem Kev solves
Much of what agents and back-office pipelines do is deciding: which team gets this ticket, is this a prompt injection, should this tool call run. The usual approach is to prompt a chat model for JSON, parse it, retry when it breaks, and get a label with no useful confidence.
TypeSafe’s Jev popularised a different interface in September 2026: send a state and typed questions, get a probability for every option in one pass. It is faster and easier to threshold than a chat completion, but Jev is hosted-only. Palmer’s launch post puts the motivation in one line: “I wanted that interface with weights I could run and fine-tune myself.”
Laya puts the same API on small encoders for raw speed. Kev uses 0.8B–27B decoder backbones and chases Jev’s accuracy instead.
How it works
Each Kev checkpoint is a Qwen backbone plus a pointer head. The state goes in first, then each question with its options wrapped in marker tokens and a final <decide> token:
<state> …state…
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <decide>
The head scores each option’s </opt> hidden state against the question’s <decide> hidden state, and a softmax turns those scores into probabilities. Nothing is decoded, so there is nothing to parse.
Three design choices matter in practice:
- Questions cannot see each other. Qwen3.5 and 3.8 mix attention with recurrent Gated DeltaNet layers, which ignore attention masks, so the server runs each question as its own row after a shared, cached state. The README reports that asking questions together or separately gives probabilities within 4e-6 in fp32 tests.
- The document is cached. Asking more questions about text you already sent only pays for the questions. On an Apple M5, Kev-4B takes 721 ms for five questions on new text and 136 ms when the text is cached, according to the README.
- Calibration ships with the weights. Every checkpoint carries a fitted temperature (2.35 for Kev-0.8B, 1.32 for Kev-27B). It changes confidence, never which answer wins. The project reports that it takes Kev-9B’s share of confident errors (wrong answers at 0.9 probability or higher) on new sources from 8.2% to 2.4%, below Jev’s 3.7%.
The three small models are LoRA rank-16 adapters trained on a 12,576-example decision-v7 set (ten public datasets plus generated policy and rule examples), then briefly fine-tuned on documents and skills. Kev-27B fine-tunes every weight on a 145,840-record corpus; the launch post puts one run at about 16 hours on eight H200s, roughly $650. No Jev outputs were used for training, per the README.
The models
| Model | Base | Runs on (per README) | Validated context | Held-out index |
|---|---|---|---|---|
| Kev-0.8B | Qwen3.5-0.8B-Base | Any 4 GB GPU, any Apple Silicon Mac | 8,192 | 23.3 |
| Kev-4B | Qwen3.5-4B-Base | L40S/H100, 32 GB Mac | 8,192 | 38.0 |
| Kev-9B | Qwen3.5-9B-Base | L40S/H100 | 8,192 | 41.0 |
| Kev-27B | Qwen3.8-27B (post-trained) | B200, H200, H100 80 GB | 65,536 | 52.3 |
| Jev (hosted) | — | TypeSafe API | — | 54.0 |
“Validated context” is the longest document for which accuracy on real contracts (CUAD) stays within 3 points of the same model at 8k tokens. The server accepts 65,536 tokens for every size and refuses longer input with a 422 rather than truncating silently. Palmer recommends starting with Kev-4B and using 0.8B only when size matters more than accuracy.
Quick start
You need Python 3.12 or 3.13 and uv:
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009
The server picks CUDA, then Apple’s MPS/MLX, then CPU. Send it a ticket:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"model": "kev-latest",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
}}'
If you already use the TypeSafe SDK, only the client constructor changes:
from typesafe_sdk import Noul, Choice, TypeSafeClient
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8009", model="kev-latest")
r = client.system_one(
state="I was charged twice. Please fix this ASAP.",
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None}),
},
)
print(r.nouls["billing"].noul, r.choices["tone"].choice)
For a hosted endpoint without cloning anything, the repo ships a Modal script that serves Kev-4B on an L40S behind a bearer key and scales to zero; the README warns the first request after idle waits about 35 seconds. There is also a Hugging Face Space running Kev-4B and Kev-0.8B in the browser.
Hands-on test (October 7, 2026)
We ran Kev in a fresh python:3.12-bookworm container on a GitHub-hosted runner (x86_64, 2 CPUs, 8 GB RAM, no GPU) against commit 5e42a7a03f28. That machine can only realistically hold the smallest model, so we tested Kev-0.8B; nothing here says anything about 4B, 9B or 27B quality. The run took 211.5 seconds.
| Step | Result | Time |
|---|---|---|
pip install uv, clone, uv sync --extra serve | Passed; resolved torch 2.8.0, transformers 5.17.0, typesafe-sdk 0.6.0 | 56.4 s |
| Import check | torch 2.8.0+cu128 ; transformers 5.17.0 ; cuda False | 6.6 s |
Start kev.serve --run jaredpalmer/kev-0.8b, send the README ticket | Passed; served “on cpu via torch (float32)” and answered | 38.4 s |
pytest tests/test_unit.py (allowed to fail) | Did not finish: our 110-second cap stopped it at 57%, with no failures before that point | 110.1 s |
What worked. The install was uneventful, although uv pulled a CUDA build of torch on a machine with no GPU. The server downloaded the base and adapter, logged that it was serving “on cpu via torch (float32)”, and printed its 65,536-token limit. Download, load and first answer took 38.4 seconds together.
What it answered. Kev-0.8B on CPU returned:
department: shipping at 0.458, with billing at 0.281 and returns at 0.261 (confidence 0.187)escalate: 0.554 probability of yesfrustration: score 1.23 on the 0–2 scale
The README’s output for the same ticket comes from Kev-4B on an Apple M5: returns at 0.47 and escalate at 0.93. The 0.8B model picked a different department and was near a coin flip on escalation. That is not a bug: the ticket mentions all three departments, and the README puts Kev-0.8B at 0.648 new-source accuracy against 0.817 for Kev-4B. Both models kept confidence low on the ambiguous question, so code acting only above a threshold would have sent this ticket to a person either way. The 0.8B model is a narrow-task model that needs fine-tuning.
What did not finish. The unit tests did not complete within our limit on 2 CPUs. The recorded exit code of 0 comes from the tail we piped into, not from pytest, so it is not evidence of a full pass.
What the run does not tell you. Our output cut off before the response’s latency_ms field, so we have no CPU latency figure. Every other accuracy and speed number here is the project’s, measured on GPUs and Apple Silicon.
Benchmarks (as published by the project)
The README is unusually careful about what its numbers mean. Each cell below is development / test accuracy on new sources, datasets and rule types Kev never trained on:
| Model | Accuracy, new sources | Brier, new sources (lower is better) |
|---|---|---|
| Kev-0.8B | 0.648 / 0.697 | 0.481 / 0.416 |
| Kev-4B | 0.817 / 0.838 | 0.269 / 0.242 |
| Kev-9B | 0.820 / 0.852 | 0.289 / 0.217 |
| Kev-27B | 0.851 / 0.889 | 0.225 / 0.156 |
| Jev | 0.857 / – | 0.211 / – |
Jev was only run on development splits, so the fair comparison is the first number in each cell: Kev-27B within a point of Jev, Kev-4B and 9B within four. Palmer adds that “we don’t know what Jev was trained on, so this isn’t a controlled comparison.”
Where Jev still wins clearly:
- Knowledge. On MMLU-Pro, Kev-27B scores 0.675 against Jev’s 0.840. The base model sets the ceiling.
- Automation rate. At a 5% error budget, Kev-4B, 9B and 27B can automate 0.52–0.69 of new-source decisions against Jev’s 0.70. Kev-0.8B manages 0.14.
- Dates. Day-precision date arithmetic is weak below 27B.
KEV_DATE_FACTS=1appends explicit day counts and takes Kev-9B from 0.80 to 0.90 on deadline questions.
On speed, the README reports model time for six questions on new short text of 22.7 ms for Kev-0.8B on an L4, 18.1 ms for Kev-4B on an H100 and 67.2 ms for Kev-27B on an H200. The launch post estimates GPU cost per million requests at $3.54, $10.54 and $44.09 respectively, assuming a busy GPU (Kev-4B’s figure is for an L40S).
The project flags its own caveats: results come from Palmer’s harness, Kev-27B’s selection “was not fully blind,” and Qwen’s pre-training data is unknown.
Fine-tuning is the real product
The headline benchmarks matter less than the fine-tuning path. The README reports that one epoch on 5,219 labelled consumer-finance complaints took Kev-4B from 0.804 to 0.904 accuracy on unseen complaints, and that about 1,000 generated support records took it from 67.7% to 73.6%, raising the share of decisions automated at a 5% error budget from 34% to 48%. With 400 records, the gain “was inside the noise.”
Training data is the API request shape plus a label on each question:
{"state": {"subject": "Charged twice", "body": "I see two charges for order #4411."},
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments", "shipping": "Delivery", "access": "Login"}, "label": "billing"}}}
uv run python -m kev.train --data train.jsonl --base Qwen/Qwen3.5-4B-Base \
--init_from jaredpalmer/kev-4b --epochs 2 --lr 2e-5 --batch 1 --accum 8 --dtype bf16 --out runs/mine
uv run python -m kev.benchmark --run runs/mine --data heldout.jsonl --out runs/mine-eval
--init_from is the important flag. The README cites one user’s test where fine-tuning from the bare base on 836 decisions dropped Kev’s own eval score to 0.33, while starting from the released checkpoint kept 0.83 and reached 0.88 on the new domain.
An agent skill (npx skills add jaredpalmer/kev@kev-finetune) runs the whole loop on Modal: finds the questions your code sends to Jev, generates labels if needed, trains, fits the temperature and compares against the untouched model. The README prices a Kev-4B run at about $1 of H100 time.
Community reactions
Kev reached the Hacker News front page in late September. One skeptical commenter reported 95% accuracy on emails with their own approach, trained on 50–100 examples in under five minutes on a CPU. That is fair for single, stable labels; Kev’s case is many questions about one document with calibrated probabilities and no retraining when you add a question.
On r/LocalLLaMA, a user comparing Jev and Kev side by side forked llama.cpp to serve Kev behind the TypeSafe API. On r/buildinpublic, commenters flagged a naming collision with an unrelated “Kev” model; the official weights live under jaredpalmer/ on Hugging Face. Early coverage, such as AlphaSignal’s, describes the older Qwen3 generation (0.6B/4B/8B), which is no longer developed.
Honest limitations
- It only scores your options. Kev does not explain, retrieve facts, or say “none of the above” unless you add that option. Supply the policy text in the state.
- Option order can move answers. Question isolation does not cover order within a question. A
/v1/systemone/permuteendpoint exists to measure it. - The small models are short-context models. Trust 0.8B, 4B and 9B to 8,192 tokens, whatever the server accepts.
- Kev-27B is expensive to host (80 GB GPU), less accurate than its v1 on long contracts, and overconfident there (CUAD ECE 0.053 against 0.007). The release notes suggest
@v1-lorafor contract review. - Kev-0.8B should not route tool calls. The release notes say its When2Call accuracy fell below chance (0.133).
- The 27B training corpus is private, and model selection was not fully blind, by the project’s own account.
When to use Kev, and what to pick instead
- Use Kev-4B or 9B if you already call Jev for routing or triage, have labelled examples, and want to own the weights and the threshold.
- Use Kev-27B if you need accuracy close to Jev on new kinds of questions and have an H100/H200.
- Use Laya if latency on cheap hardware matters more than zero-shot accuracy and you are going to fine-tune anyway.
- Use hosted Jev if you need knowledge-heavy decisions (the MMLU-Pro gap is large) or do not want to run GPUs.
FAQ
What is Kev?
Kev is an Apache-2.0 family of four decision models (0.8B, 4B, 9B and 27B parameters) by Jared Palmer, built on Qwen3.5 and Qwen3.8. They answer yes/no, multiple-choice and rating questions about a text with calibrated probabilities, in one forward pass, through a TypeSafe System One-compatible API.
Is Kev a drop-in replacement for Jev?
At the API level, yes: the README says the TypeSafe Python SDK works against a Kev server unchanged. At the quality level, Kev-27B is within a point of Jev on the project’s out-of-domain development set, but behind on knowledge questions and on how many decisions can be automated at a fixed error budget.
Can Kev run without a GPU?
The server falls back to CPU in fp32. In our lab, Kev-0.8B loaded and answered a three-question request on a 2-CPU, 8 GB runner within 38.4 seconds of starting the server. The project publishes no CPU latency figures; its targets are CUDA GPUs and Apple Silicon.
Which Kev model should I start with?
The README recommends Kev-4B: it fits a 32 GB Mac or an L40S and scores 0.817 on new-source development data against Kev-27B’s 0.851. Use 0.8B only for narrow, fine-tuned tasks.
How much does it cost to fine-tune Kev?
The README estimates about $1 of H100 time for a Kev-4B fine-tuning run through the kev-finetune skill on Modal, excluding data generation and evaluation. It reports gains starting from roughly 1,000 records. With 400 records, the gain was within noise.
Verdict
Kev is the most carefully documented open take on the Jev pattern so far: every headline number comes with its split, its caveats and a list of what Jev still does better. API compatibility makes it cheap to try against an existing Jev integration, and the --init_from fine-tuning path is where it earns its place. Our CPU run confirms it installs cleanly and serves without a GPU, and that 0.8B is no shortcut to 4B’s answers. Start with Kev-4B, measure it on your own data, and set thresholds from those measurements.
Sources
- jaredpalmer/kev on GitHub: README, API, benchmarks, limitations
- Kev 1.0 release notes
- Kev 1.0 GitHub release
- Introducing Kev, Jared Palmer (October 1, 2026)
- Kev model collection on Hugging Face
- Kev demo Space on Hugging Face
- TypeSafe System One API documentation
- Jev’s Architecture Unmasked, Archer Hume
- Hacker News discussion
- r/LocalLLaMA: Jev vs. Kev tested side by side