TL;DR
fast-jev-compaction (tamaratran/fast-jev-compaction) is an MIT-licensed Claude Code plugin, plus an npm library, that replaces the /compact summary with something structurally different. Instead of asking Claude to rewrite your session into a shorter paraphrase, it sends the whole conversation to TypeSafe’s Jev decision model and asks two yes/no questions about every old tool call: does the call still matter? and does its result still need to be there verbatim? Calls that fail both tests are deleted; results that fail only the second are cut to a 300-character head. Everything that survives stays word for word. No text is ever rewritten.
The demo that made it go viral was Alex Volkov’s: a session of nearly 1M tokens down to 86K in about a second. Twelve days after the repo appeared it sits at 7,169 stars, 454 forks, 96 open issues and 59 open pull requests, and the last upstream commit is from 2026-09-17, the day after creation.
Key facts as of 2026-09-29:
- Stars / forks: 7,169 / 454; 96 open issues; last commit 2026-09-17
- Form: Claude Code function-hook plugin (
session.compact,turn.complete) and an npm package (fast-jev-compaction, Node 18+, TypeScript) - Requires: Claude Code 2.1.274+ with
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1, and aTYPESAFE_API_KEY - Speed: one contributor measured a median 743 ms per compaction against 30,375 ms for the native summary (n=6)
- Fallback: on any Jev error, or if it frees less than 25% of the history, Claude Code’s built-in summary runs instead
- Price: the plugin is free; Jev is a paid cloud API (TypeSafe quotes $0.042 per million input tokens, output free)
Verdict: the idea is right and the seam it exposes (JevAsker) is clean enough that Hermes Agent already runs it as a production context engine. But the current release has real, well-documented problems in long sessions: a keep threshold that almost never keeps anything, dropped calls that leave the assistant narrating work it can no longer see, a Cloudflare WAF that blocks typical coding transcripts, and compaction that is undone on --resume. Try it on fresh sessions with a key you have; do not put it on a headless box yet.
Quick reference
| Repo | github.com/tamaratran/fast-jev-compaction |
| Stars / forks | 7,169 / 454 (2026-09-29) |
| License | MIT |
| Language | TypeScript; vitest test suite with a fake Jev, one live npm run demo |
| Plugin version | 0.3.0 (.claude-plugin/plugin.json); library package.json still says 0.2.0 |
| Hooks used | session.compact (replaces the summary), turn.complete (triggers at compactAtPercent, default 60) |
| Jev endpoint | https://api.typesafe.ai/v1/systemone, model jev-latest (resolves to jev-1.13.0 at the time of writing) |
| Install | claude plugin marketplace add tamaratran/fast-jev-compaction then claude plugin install fast-jev-compaction@fast-jev-compaction |
| Notable fork | deadczarvc/hermes-jev-compaction, a context-engine plugin for Hermes Agent |
What it is
Every long Claude Code session hits the context ceiling, and when it does the CLI runs compaction: it asks the model to write a summary of what happened so far, throws the original messages away, and continues from the summary. That summary is lossy by construction. An exact file path, an error string, a constraint the user gave in turn three, a command that must never be re-run: any of these can vanish, and you usually find out when the agent does the thing it was told not to do.
Tamara Tran’s plugin starts from a different question. Most of the bulk in a coding session is not conversation, it is tool traffic: the 4,000-character Read, the 30-line Bash output, the failed grep. The user’s prompts and the assistant’s prose are a small share of the tokens (8% to 10% in the first compaction round of a real session, per issue #70). So instead of summarizing everything, the plugin leaves all text alone and only decides, per tool call, whether that call and its result are still worth their tokens.
The decision is not made by an LLM. It is made by Jev, TypeSafe AI’s “System One Model”, which we covered in the Jev Ultrafast review: a model that takes a state and typed questions (choice, score, or noul, a yes/no probability) and answers all of them in one parallel forward pass with no text generation. A few hundred yes/no questions in under a second is exactly what compaction needs, and fast-jev-compaction is, so far, the most-starred demonstration of that fit.
How it works
The README documents the pipeline in detail, and the source in src/compact.ts matches it:
- Pair and pin. Every
tool_useis matched to itstool_resultbytool_use_id. Calls in the first message and in the newestpreserveRecentMessages(default 6) are pinned and never touched. - Build the state. The whole conversation, oldest first, is sent to Jev with every tool result replaced by a note such as
ok, 4213 chars (omitted). Tool inputs stay in. Text stays in. Nothing is summarized. - Fit the state into
maxStateTokens(25K) in stages, each applied only if the previous one was not enough: tool inputs truncated to 1000, then 200, then 60 characters; long texts abridged to head plus tail, oldest first; old messages collapsed to[… N chars omitted …]; old tool calls reduced to one line each (t12 Read file_path=src/a.ts → ok 480ch); old call-less messages dropped; runs of call-only messages folded. If it still does not fit, compaction throws. - Ask two questions per call. For every non-pinned call Jev gets
keepCall(knowing the call was made, with its input, still matters) andkeepResult(the result’s contents are still needed and re-running the tool would not do). - Batch. Questions are split into as many requests as needed to keep state plus questions under
maxRequestTokens(30K, below Jev’s 32K request limit). The full state is resent every time; requests run concurrently. - Decide against
keepThreshold(0.5):keepResult ≥ 0.5keeps call and result; elsekeepCall ≥ 0.5keeps the call and truncates the result totruncateHeadChars(300) plus a one-line[fast-jev-compaction truncated N chars …; re-run the tool if needed]note; else the call and its result are removed together. - Rebuild. Messages that lose all content are removed; untouched messages are returned as the same objects; no result is ever left without its call.
The Claude Code adapter in hooks/fast-jev.ts is thin. It reads the plugin’s userConfig, finds the key, hands session.compact transcripts to the library, maps the result back onto SessionMessage objects (unchanged messages keep their engine handles, rebuilt ones come back fresh), and shows a toast: fast-jev-compaction: kept N/M messages, no summary (…) on success, fallback to built-in summary (…) when Jev fails or the estimated reduction is under minReductionRatio.
As a library
The plugin is a wrapper around a package you can use in your own agent:
import { compactMessages, reductionRatio, type Message } from 'fast-jev-compaction';
const transcript: Message[] = [
{ role: 'user', text: 'Fix the failing test. Never edit src/generated.', toolUses: [] },
{
role: 'assistant',
text: '',
toolUses: [{ tool_use_id: 'toolu_1', tool: 'Read', input: { file_path: 'src/a.ts' } }],
},
{ role: 'user', text: '', toolUses: [], toolResults: [{ tool_use_id: 'toolu_1', text: '…file…' }] },
// …
];
const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.messages, result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
// not worth it: keep the original transcript, or summarize instead
}
Message is a subset of Claude Code’s SessionMessage, so a session JSONL can go in as is. To bring your own transport (or your own decision model), implement JevAsker, one ask(state, questions) method, and call compact(messages, asker, options). buildJevRequest and parseJevResponse are exported for the HTTP side, and the building blocks (collectToolCalls, fitState, batchCalls, decideCall, applyDecisions) are exported individually. That seam is why the Hermes port below took one adapter file.
Getting started
Function hooks are early-access in Claude Code 2.1.274+, so you need the opt-in flag wherever Claude Code runs. In ~/.claude/settings.json:
{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1", "TYPESAFE_API_KEY": "<your key>" } }
Then add the repo as a marketplace and install:
claude plugin marketplace add tamaratran/fast-jev-compaction
claude plugin install fast-jev-compaction@fast-jev-compaction
The installer prompts for the userConfig options (key, thresholds, truncateHeadChars, model); leave them at defaults to use the environment variable. Restart Claude Code or run /reload-plugins. From then on /compact and auto-compaction go through Jev. To run from a checkout without installing:
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .
The options that matter in practice:
| Option | Default | What it does |
|---|---|---|
keepThreshold | 0.5 | Minimum probability for a call or result to stay (see limitations before trusting this number) |
preserveRecentMessages | 6 | Newest messages pinned |
compactAtPercent | 60 | Context percentage at which turn.complete requests compaction |
minReductionRatio | 0.25 | Below this estimated reduction, fall back to the built-in summary |
maxStateTokens / maxRequestTokens | 25000 / 30000 | Budget for the state and for state plus one batch of questions |
truncateHeadChars | 300 | Head of a dropped result kept before the note |
One practical snag before you start: several users could not find where to get a TYPESAFE_API_KEY at all. Issue #54 documents signing up at console.akka.io and finding only Akka license and repo tokens, none of which work against api.typesafe.ai. The README does not say, and as of writing the issue has no answer. If you already have a key from the Jev Ultrafast wave, it works here.
Community reaction
The r/ClaudeCode thread “Instant Claude Code compaction is my favorite use of Jev so far” passed 300 upvotes and 80 comments. The poster’s summary is the most honest one-liner about the tool: “it’s not compacting text, only removing tool calls. It’s a very basic use of Jev but saves legitimate tokens. As far as Claude is concerned, all the actual context is still there.” Much of that thread, as explainx.ai’s write-up noted, spent its first dozen comments describing a different project, a PostToolUse hook that trims individual Bash outputs before they enter context. The two compose well (issue #18 proposes adding exactly that as a second hook here) but they are different jobs: one is input hygiene, the other is retroactive pruning.
Sébastien Dubois framed it as the clearest demonstration yet of where System One models belong: “‘Is this tool result still relevant?’ is a bounded question asked hundreds of times per session; it doesn’t need a writer, it needs a fast judge.”
The most interesting reaction is the fork. Issue #60, from deadczarvc, reports that the library “ports cleanly and is now the production compaction backend” of Hermes Agent: the fork ships a context-engine plugin that registers as the agent’s compaction backend, an OpenAI-chat ↔ Message[] adapter, a --dry-run CLI, and a trigger derived from the model’s real output budget rather than a bare percentage. It also surfaced a detail worth knowing: jev-latest resolves to jev-1.13.0, the pin jev-1.13.0 works, and jev-1.13 is rejected.
The other reaction is the issue tracker itself, which is unusually rigorous. Several users replayed their own sessions through the library with the real API and posted score histograms, token counts and repro scripts. That is where the limitations below come from.
Limitations
These are not speculative. Each is a filed issue with numbers, and none has an upstream fix merged.
- The keep threshold barely keeps anything. Issue #26 replayed 16 sessions (277 tool calls, 256 scored) through the real API. Zero of 256 results scored ≥ 0.5; the
keepResulthistogram tops out below 0.3. The plugin’s decisions matched a fake asker that answers 0 to everything to within one percentage point of character reduction (87.7% vs 88.5%). Issue #56 found the cause:keepResultandkeepCallcome back on different scales, Jev’s ranking is stable and correct, butdecideCallcompares both against the same 0.5. The practical effect is that the compactor truncates the file the agent is about to edit. LowerkeepThresholdor rescale per question before trusting the defaults. - Dropped calls leave the assistant narrating ghosts. Issue #65 is the most alarming:
drop_callremoves atool_useblock but leaves the assistant’s text (“opened the PR”, “review came back REQUEST_CHANGES”) with no marker. In one live two-day session onclaude-opus-5, after the third/compactthe model produced nine consecutive tool-free “work done” reports over 22 minutes. None of the artifacts existed.drop_resultleaves a note;drop_calldoes not. A fix (PR #69) is open, unmerged. - Repeated compaction runs out of candidates. Issue #70 tracked five rounds in one session: retained context grew 36K → 52K → 59K → 74K → 87K because each round can only score calls made since the previous round, and the text share climbs from 8% to 44%. From round two on, Jev is answering hundreds of questions about a state in which most text is collapsed to
[… N chars omitted …]. Thegoalsent to Jev on all nine rounds studied was the/compactcommand echo, not the task. - Cloudflare blocks real transcripts. Issue #97:
api.typesafe.aisits behind a WAF, and astatecontaining shell commands, SQL, or../../etc/shadowreturns a 403 HTML page. Coding transcripts contain such strings almost by definition, so on those sessions the plugin silently falls back to the summary every time. Fix belongs to TypeSafe, not the plugin. - Compaction is undone on
--resume. Issue #89 measured it: a plugin-compacted session reloads at 166K to 168K tokens (from 173K pre-compaction), while a native one reloads at 28K. The hook’scompact_boundaryrecord lacks thepreservedSegmentanchors Claude Code 2.1.278’s loader branches on. This is a Claude Code function-hook gap (anthropics/claude-code#95328), and it means the in-session win evaporates the moment you restart. - No durable record on headless installs. Toasts and
ui.logare TUI-only, so a scheduler-launched session cannot tell afterwards whether Jev ran or the summary did (#72, closed, apparently without a merged change). - Maintenance has stalled. Everything above was filed after the last commit. The repo went from creation to 7K stars on a two-day burst of Devin-authored PRs, then stopped; 59 PRs wait, including a Laya provider (#87) and fail-closed hardening for malformed Jev answers (#28). Read the fork network before you depend on it.
- Cloud-only by design. Your entire transcript text (with tool outputs stubbed) goes to TypeSafe on every compaction. Issue #86 asks for a local provider; the Laya PR is the obvious answer but is unmerged.
Who should use it
Use it if you run long, tool-heavy Claude Code sessions in one sitting, already have a TypeSafe key, and want to see for yourself what verbatim compaction feels like versus the summary. The first compaction round on a fresh session is where it shines, and the fallback means the worst case is the behaviour you had before. It is also worth reading as a design: if you are building an agent with its own compaction, the two-question-per-call pattern and the JevAsker seam are a good recipe, and the Hermes fork shows it generalises.
Skip it if you rely on --resume, run sessions headless, work in a transcript that trips WAF rules (most real coding work), or cannot send session text to a third party. And skip the defaults regardless: with keepThreshold at 0.5, the evidence says you are running a “drop every non-pinned result” rule with an API call in the middle.
Comparison with alternatives
| fast-jev-compaction | Claude Code built-in /compact | Context Mode | Headroom | |
|---|---|---|---|---|
| When it acts | At compaction time | At compaction time | Before output enters context | Proxy layer, every request |
| What it changes | Deletes/truncates tool calls; text untouched | Rewrites everything into a summary | Sandboxes tool output in SQLite, returns slices | Compresses tool output, AST-aware, reversible |
| Decision maker | Jev (cloud, ~1 s) | Claude (30 s+) | Rules | Rules + ML |
Survives --resume | No (#89) | Yes | n/a | n/a |
| Local option | Not merged (Laya PR #87) | n/a | Yes | Yes |
| Best at | Fast, lossless-for-text first compaction | Cross-session continuity | Keeping raw logs out entirely | Multi-agent, multi-provider token bills |
They stack: an input-side trimmer keeps junk out; fast-jev-compaction prunes what still got in. None of them fixes the fact that only summaries carry across a restart today.
FAQ
Does fast-jev-compaction ever rewrite my messages? No. User and assistant text is never removed or shortened in the output; it is only abridged in the state Jev sees when fitting the 25K budget. Only tool calls and tool results are candidates, and a dropped result keeps its first 300 characters plus a note by default.
What happens if Jev is down or my key is missing?
The hook logs a fallback and delegates to Claude Code’s built-in summary. The same happens if the estimated reduction is under minReductionRatio (25%) or the history cannot be fitted into the state budget. You never end up with no compaction at all.
Is this the same as Claude Code’s hidden background compaction? No. Reddit users inspecting the binary reported experimental Anthropic flags that precompute a summary in the background so hitting the limit causes no delay. That is still a summary, computed earlier; this plugin replaces the summary with pruned originals. Neither depends on the other.
Can I use a local model instead of the TypeSafe API?
Not with what is merged. JevAsker is a one-method interface, so you can implement it against anything that returns probabilities; PR #87 wires Laya (a 421M-parameter ModernBERT decision model, Apache-2.0) as a provider, and issue #115 asks for the same. Accuracy of local substitutes versus jev-1.13.0 is unverified.
Why did my compacted session come back huge after --resume?
Because the plugin’s compact_boundary record has no preservedSegment anchors, and Claude Code’s loader treats it as if compaction never happened (#89). Native compaction writes those anchors; function hooks currently cannot. Until Anthropic exposes them, the saving is per session only.
How much does a compaction cost? The plugin is free. Each round sends the full fitted state (up to ~25K estimated tokens) once per batch of questions; issue #26’s replay used 134K input tokens over 8 requests for 256 calls. At TypeSafe’s published $0.042 per million input tokens, output free, that is well under a cent per compaction, but the exact bill depends on how many batches your session needs.
Sources
- tamaratran/fast-jev-compaction on GitHub: README,
hooks/README.md,.claude-plugin/plugin.json,hooks/fast-jev.ts - GitHub issues #26, #54, #56, #60, #65, #70, #89, #97
- r/ClaudeCode: Instant Claude Code compaction is my favorite use of Jev so far
- explainx.ai: fast-jev-compaction — Claude Code Compaction Without a Lossy Summary, 2026-09-20
- Sébastien Dubois: fast-jev-compaction, 2026-09-23
- deadczarvc/hermes-jev-compaction, the Hermes Agent fork
- Claude Code plugins reference and hooks