TL;DR
Paperclip (paperclipai/paperclip) is an open-source, MIT-licensed control plane that turns a pile of AI agents into something that looks like a company: org chart with reporting lines, ticket-style tasks, per-agent budgets, approval gates, scheduled routines, and a “heartbeat” scheduler that wakes agents when work lands on them. Its tagline is the cleanest one-liner in the category: “If OpenClaw is an employee, Paperclip is the company.”
As of October 4, 2026 the repo sits at 96.9K stars (16.4K forks), gained 10.7K stars in the last seven days, and has merged 2,700+ pull requests since its creation on March 2, 2026. Latest stable release: v2026.1001.0 (October 1); nightlies ship almost daily.
Key facts:
- It doesn’t think. Paperclip has no agent loop of its own. Every agent connects through an adapter to a runtime that does the actual work — Claude Code, Codex, Gemini CLI, Cursor, OpenCode, Pi, Hermes, Grok Build, Kimi Code, OpenClaw over a WebSocket gateway, or any shell command / HTTP endpoint.
- Install is one command:
npx paperclipai@latest onboard --yes(Node.js 24.11+). Embedded PostgreSQL, UI + API onlocalhost:3100, no account required. - Stack: Node.js server + React UI, TypeScript, Postgres (embedded locally, bring-your-own in production), a Rust-built native runner, out-of-process plugin workers.
- Governance is the differentiator: atomic task checkout with execution locks, budget hard-stops, revisioned config with rollback, approval stages, multi-org isolation, and a durable activity log.
- Honest limitation: the orchestration layer is solid, but most production bugs live on the adapter boundary (identity, wake context, liveness, token accounting). One operator who ran a 13-agent org for six weeks had to write his own adapter package to get a loop to close unattended.
What Problem Does Paperclip Actually Solve?
The README’s framing is one every multi-agent operator recognises: “You have 20 Claude Code tabs open and can’t track which one does what. On reboot you lose everything.” Agents are good. Everything around them — who’s working on what, what it has cost, what context each session needs, who approves the output — is held together with tmux panes and copy-pasted prompts.
Paperclip’s bet is that the right abstraction isn’t a workflow graph or a prompt library. It’s an organization. Tasks are tickets with a single assignee. Agents have titles, bosses, and job descriptions. Goals cascade from company to project to issue, so the agent gets the “why” behind the work. Budgets exist at company, agent, project, and goal level, and a runaway loop hits a hard stop instead of your credit card.
The README’s four pillars map neatly onto who the product is for:
| Pillar | Built for | What it covers |
|---|---|---|
| Agentic Task Manager | Everyone, daily | Tasks, approvals and review gates, routines, verify from diffs/screenshots/tests |
| Org Chart for Agents | Managers | Mixed human + agent org chart, delegation, governance, scoped secrets |
| Agent Employee Training | Enablers | Skill Studio, shared org-wide skills, evals, saved test runs, skill version history |
| Agentic OS | IT & platform | Cross-provider runtime, sandboxing, MCP servers, SSO/RBAC, cost controls, tracing |
The product explicitly says what it is not: not an agent framework (“we don’t tell you how to build agents”), not a prompt manager, not a workflow builder, not a chatbot. That boundary is the point — the same strategic call Multica made: be the layer that survives whichever agent runtime wins.
How Paperclip Works Under the Hood
The architecture is a single Node.js server with twelve subsystems on top of Postgres. The ones that matter in practice:
Heartbeat execution. Agents wake for three reasons: an issue is assigned to them, something they were blocked on resolves, or a timer fires. Wakeups go through a DB-backed queue with coalescing; the server then runs budget checks, resolves the workspace, injects secrets, loads skills, and invokes the adapter. Each run produces logs, usage records, and session state. Bounded recovery handles known failure classes and escalates the rest to a human.
Atomic task checkout. One assignee per task, and execution locks stop two runs from claiming it simultaneously. Trivial-sounding, and the thing most homegrown multi-agent setups get wrong.
Governance with rollback. Approval gates are enforced server-side, every config change is revisioned and reversible. Board-level approvals cover hiring (yes, you approve new agents), strategy, and configurable review stages on task completion.
Connections and the tool gateway. Agents reach GitHub, Notion, Railway, or your own MCP server through governed connections. Each gateway action is Allowed, Ask first, or Off, and human access, agent eligibility, and action permissions are three separate controls. Managed GitHub operations can run under the account of the person who directed the work.
Workspaces and portability. Git-worktree execution workspaces, dev servers with preview URLs, pluggable sandbox providers, and export/import of entire organizations with secret scrubbing (plus a Team Catalog of ready-made orgs).
Adapters: where the real work happens
Every adapter launches or calls the runtime, passes company/task/wake context through, captures results and session state, and validates the environment before a run. The built-ins as of October 2026:
| Adapter | Type key | Notes |
|---|---|---|
| Claude Code | claude_local | Recommended; session persistence, skills sync, structured transcript |
| Codex | codex_local | Recommended; managed CODEX_HOME |
| Gemini CLI | gemini_local | Resume support |
| Cursor Local | cursor | --resume continuity |
| OpenCode | opencode_local | Provider/model routing — the usual route to local or OpenRouter models |
| Pi | pi_local | Built-in tool set |
| Hermes | hermes_local | Persistent memory, 30+ tools |
| Grok Local / Kimi Code | grok_local / kimi_local | ACP engine with headless fallback |
| OpenClaw Gateway | openclaw_gateway | Remote OpenClaw over WebSocket — functional, but “Coming soon” in the UI dropdown; configure via API or company import |
| Process / HTTP | process / http | Any shell command or webhook; same UI caveat |
The ACP (Agent Client Protocol) engine gives the best transcript fidelity — Claude, Codex, Gemini, and Kimi runs emit a JSONL event per tool call, thinking delta, and status change, rendered live in the run viewer. Process and HTTP adapters only give you raw stdout.
Installation
Requirements: Node.js 24.11+. The quickstart defaults to a trusted local loopback mode:
npx paperclipai@latest onboard --yes
For a network-reachable, authenticated instance, pick a bind preset explicitly:
npx paperclipai@latest onboard --yes --bind lan
# or
npx paperclipai@latest onboard --yes --bind tailnet
To kick the tyres without a real install, test-drive spins up an isolated instance pre-seeded with a CEO agent:
ANTHROPIC_API_KEY=... npx paperclipai test-drive
OPENAI_API_KEY=... npx paperclipai test-drive --harness codex
OPENROUTER_API_KEY=... npx paperclipai test-drive \
--harness opencode \
--model openrouter/anthropic/claude-sonnet-4.5
From source:
git clone https://github.com/paperclipai/paperclip.git
cd paperclip
pnpm install # pnpm 9.15+
pnpm dev # UI + API on http://localhost:3100
Source builds also compile the native Paperclip Runner, so you need a Rust toolchain or a prebuilt binary via PAPERCLIP_RUNNER_BINARY. For production, point it at your own Postgres and deploy with Docker; the docs cover a Tailscale-fronted authenticated setup. One README gotcha: a private npm registry in ~/.npmrc can make npx fail with E404 — pass --registry https://registry.npmjs.org.
What Shipped Recently (v2026.1001.0)
The October 1 release carried 77 commits and is mostly a hardening pass — what you want to see at this star count:
- Native runner and chat recovery hardened — approval/Stop races, session continuity, sandbox reconnection and cleanup.
- Approvals and answers queue during active runs instead of bouncing.
- Agents as GitHub review bots — guided setup turns an agent into a scheduled PR reviewer; the GitHub MCP connection gained the Actions toolset.
- Railway joined the connector catalog; agent personas got faces across the app.
Two breaking changes matter if you’re already running it:
- The legacy Composio broker is retired with no automated migration; the experimental MCP aggregator connectors are the successor.
- Execution harnesses now default to full auto. Native runs default Claude/ACPX to
approve-all, OpenCode toallow, and Codex to approval-and-sandbox bypass. Paperclip’s own approval gates still apply, but if you relied on unconfigured agents inheriting restrictive provider defaults, set an explicit mode now.
That second one deserves a pause. It’s the right default for a product whose premise is unattended agents, but a fresh install’s Claude Code agent now runs with permission prompts bypassed inside its workspace. Put the governance in Paperclip’s approval stages and sandboxes, not in the hope that the harness will ask.
What the Community Says
Despite 97K stars, Paperclip has had surprisingly little Hacker News traction — direct submissions (“a ticket-based multi AI agent orchestrator”, April 25; “meta-harness for your agents”, September 24) landed a handful of points each. The real discussion is on the project’s Discord, in the issue tracker (6.3K open issues — a sign of adoption and of surface area), and in long-form operator write-ups.
The most useful is Felipe Fontoura’s October 1 field report on running a 13-agent content company (a CEO agent, a chief of staff, five “heads”, six specialists) on Hermes Agent inside Paperclip for six weeks. It cuts both ways:
- “The org chart took an afternoon. Everything else took the six weeks.” The loop first closed on its own — head orients the week, specialist delivers, head judges and advances, report filed up, parent closed — on July 22.
- The built-in Hermes adapter ignored his instruction bundle. It spawned the Hermes CLI with the task as the prompt and never sent the SOUL/AGENTS/HEARTBEAT/TOOLS files, so agents “woke up with a task and no identity.” He published a drop-in replacement (
@felipefontoura/paperclip-adapter-hermes-local-plus) that prepends the bundle, keeps skillreferences/folders Paperclip’s build was stripping, recovers blank runs, and reports real token counts. - Two identities cost 26% of quality. The same GLM-5.1 model scored 9.55 on his benchmark through OpenCode but 7.0–7.9 through Paperclip + Hermes, because the runtime’s 14K-character system prompt introduced itself as Hermes and the persona arrived later as a user message.
- Liveness was inferred from stdout, so a four-minute silent synthesis looked dead, a duplicate spawned and stole the lock, and the real run’s write-back bounced. A fifteen-line keepalive fixed it.
- A head re-delegated forever when work was handed back, until one rule: an in-review issue assigned to you is a return — judge it, never re-delegate.
His summary: “Most of these were adapter and orchestration bugs. The models were rarely the problem.” And on the human role: “An agent company does not remove you. It moves you from producer to approver.” When he stepped away, the org stopped — publishing was the one human gate and nobody was at it.
The open issues tell the same story. The top feature request since March 7 is native Ollama / local LLM support (#187, 34 comments) — today’s answer is “route through OpenCode or OpenClaw.” Chat with agents (#49) has since shipped as experimental Agent Chat. The newest hot thread (#14093, September 26) reports all Claude Code agents failing with acpx_turn_failed — exactly the adapter-boundary bug class the hardening release targeted.
The comparison that comes up most is Paperclip vs. Multica. Our Multica review framed Paperclip as a governance-heavy single-operator AI-company simulator and Multica as a lighter multi-user board. Six months on, Paperclip’s multi-org support, SSO/RBAC, and responsible-user attribution make it a real team product too — but governance is still the dividing line: AI company vs. AI teammates.
Where Paperclip Falls Short
It’s only as good as the adapter. The contract with the runtime — launch it, pass context, capture state — is clean on paper, but identity injection, liveness detection, session resume, and token accounting all cross that boundary, and that’s where operators spend their time. Claude Code and Codex through the ACP engine are the well-trodden path; anything else, expect to read adapter source.
Liveness is heuristic. A runtime that goes quiet during a long think can be treated as dead. The hardening pass addresses parts of this; long-silence runtimes still deserve a keepalive.
Six thousand open issues. Triage is real, but don’t expect a quick response on an edge-case adapter bug.
No first-class local-LLM story. The most upvoted issue in the project’s life is still open. OpenCode and OpenClaw get you there, with an extra hop.
Full-auto defaults. Since October 1, fresh installs run harnesses with permission prompts bypassed — right for the product’s premise, wrong for “try it on my laptop.”
The bottleneck is you. Humans stay at the gate — hiring approvals, review stages, board decisions. That’s a feature. It also means a company you don’t show up to simply fills a queue.
Who Should Use Paperclip?
Good fit:
- You run more than one agent runtime (Claude Code for code, Hermes or OpenClaw for research) and want one place tracking tasks, cost, and history across them
- You have recurring agent work — support triage, weekly reports, PR review — and want it scheduled with run history instead of a crontab nobody documents
- You need budget hard-stops and an audit trail because agents spend real money unattended
- You’re building a solo-operator business where agents produce and you approve — and you’ll be at the gate every morning
Bad fit:
- One coding agent on one repo — the orchestration overhead costs more than it saves
- You want a drag-and-drop workflow builder — Paperclip is an org, not a DAG
- Your runtime isn’t in the adapter list and you won’t write or vet an adapter package
- You need native local-model support without an OpenCode/OpenClaw hop
- You can’t tolerate daily nightlies and occasional breaking releases
FAQ
What is Paperclip, in one sentence?
Paperclip is an open-source (MIT) Node.js server and React UI that orchestrates a team of AI agents as an organization — org chart, tickets, budgets, approvals, scheduled routines, heartbeat scheduler — while delegating the thinking to whatever runtime each agent is wired to. It is at 96.9K GitHub stars as of October 4, 2026.
Does Paperclip run the models itself?
No. It schedules, routes, records, and governs. Each agent uses an adapter to a runtime — Claude Code, Codex, Gemini CLI, Cursor, OpenCode, Pi, Hermes, Grok Build, Kimi Code, OpenClaw via gateway, or any process/HTTP endpoint. Model choice belongs to the runtime; the Hermes field report recommends letting the runtime own it.
How is Paperclip different from OpenClaw or Claude Code?
Different layers. OpenClaw and Claude Code are agents — they do work. Paperclip uses them and manages the company they work in. The README: “It orchestrates them into a company — with org charts, budgets, goals, governance, and accountability.”
Can I use Paperclip with local models like Ollama?
Not directly — that’s the project’s most-requested feature since March (issue #187). The working path is the OpenCode adapter with provider/model routing, or OpenClaw pointed at a local endpoint and connected through the OpenClaw Gateway adapter.
How much does Paperclip cost?
Self-hosted is free under MIT and needs no Paperclip account. A hosted Paperclip Cloud is on a waitlist at paperclip.ing with no public pricing as of October 2026. Your real cost is whatever the agents’ models burn — which is what the per-agent, per-project, and per-goal budgets cap.
Is it production-ready?
For Claude Code / Codex through the ACP engine, operators run it 24/7, and the October 1 release is explicitly a reliability pass. For other adapters, budget time for the boundary bugs above. Daily nightlies and a 6K-issue tracker mean a very active project that is still changing shape.
Bottom Line
Paperclip is the most complete open-source answer so far to “how do I run a dozen agents without babysitting a dozen terminals.” The primitives — atomic checkout, goal ancestry, budget hard-stops, revisioned governance, routines with run history — are the right ones, and the product is honest about not being an agent framework. Ninety-seven thousand stars and 2,700 merged PRs say the bet is landing.
The caveat comes straight from people who’ve run it in anger: the hard problems live on the adapter boundary, and the further you stray from Claude Code and Codex, the more of that boundary you own. Start with test-drive, one agent, one routine, and the smallest loop that closes. Build the org chart after the loop works, not before.
Repo: github.com/paperclipai/paperclip · 96.9K ⭐ · MIT · +10.7K stars this week · latest v2026.1001.0
Sources
- paperclipai/paperclip on GitHub — README, FAQ, issue tracker (accessed October 4, 2026)
- Paperclip v2026.1001.0 release notes — October 1, 2026
- Paperclip docs: Adapters Overview
- Felipe Fontoura, “I Built a Company of AI Agents in Paperclip. Here Is What Broke.” — October 1, 2026
- Issue #187: Support for Local LLMs via Ollama