TL;DR

CLI-Anything (HKUDS/CLI-Anything) is an Apache-2.0 project from the University of Hong Kong’s data-intelligence lab (the group behind LightRAG and RAG-Anything). Its pitch is a sentence: today’s software serves humans, tomorrow’s users will be agents. Instead of teaching an agent to click through GIMP’s menus from screenshots, you point a coding agent at GIMP’s source and it generates cli-anything-gimp, a Python CLI with --json output, a REPL, a test suite and a SKILL.md, all built on the real application backend.

It is two products in one repo:

  • The generator — a Claude Code plugin (also packaged for Cursor, Pi, OpenClaw, OpenCode, Codex, Hermes, Copilot CLI) that runs a seven-phase pipeline from codebase analysis to PyPI publishing.
  • CLI-Hub — pip install cli-anything-hub gives you a package manager for the 79 harnesses already in the registry (Blender, FreeCAD, LibreOffice, OBS, Kdenlive, Obsidian, n8n, Zotero, Godot, QGIS, Ollama, ComfyUI…) plus 24 public third-party CLIs, and a meta-skill so an agent can install what it needs on its own.

Key facts as of 2026-09-30:

  • Stars: 51,067; repo created 2026-03-08; still on GitHub’s weekly trending page in late September
  • Harnesses: 79 in-repo, 2,464 tests reported passing; cli-anything-hub 0.4.1 on PyPI
  • Tech report: arXiv 2606.03854, “CLI-Anything: Towards Agent-Native Computer Use”
  • Requires: Python 3.10+, a frontier model (the README names Claude Opus 4.6, Sonnet 4.6, GPT-5.4), and the target’s source code
  • Price: free; you pay for the agent tokens that generate the harness

Verdict: the strongest argument yet that CLI, not screenshots, is the right interface for agent computer use, and CLI-Hub is genuinely useful today. But treat “zero compromises, real backend” as a goal rather than a guarantee: three open PRs from September show shipped harnesses where the render step never actually calls Blender, Kdenlive or SoX.

Quick reference

Repogithub.com/HKUDS/CLI-Anything
Stars51,067 (2026-09-30), created 2026-03-08
LicenseApache-2.0
LanguagePython (Click ≥ 8.0, pytest); one harness (Sketch) is Node.js
Install (use CLIs)pip install cli-anything-hub then cli-hub install <name>
Install (build CLIs)Claude Code: /plugin marketplace add HKUDS/CLI-Anything → /plugin install cli-anything
Agent skillnpx skills add HKUDS/CLI-Anything --skill cli-hub-meta-skill -g -y
Registry79 CLI-Anything harnesses + 24 public CLIs at hkuds.github.io/CLI-Anything
PaperarXiv:2606.03854

What CLI-Anything is

The problem statement comes from the tech report. The dominant approach to “computer use” is a GUI agent: screenshot, find the button, click, repeat. The authors argue this “fundamentally misaligns with agent capabilities” — pixel coordinates break when the interface changes, and the agent is forced to emulate human perception instead of using its strengths: structured data and deterministic commands.

CLI-Anything’s answer is a harness: a command-line program in front of the real application that exposes its functionality as structured commands. Five rules are set by HARNESS.md, the project’s methodology document:

  1. Authentic integration. The CLI writes valid project files (ODF for LibreOffice, MLT XML for Kdenlive, SVG for Inkscape) and delegates rendering to the real application. No Pillow stand-in for GIMP, no custom renderer for Blender.
  2. Dual interaction modes. Run the bare command and you get a stateful REPL; pass subcommands and it behaves like a normal CLI for scripts and pipelines.
  3. Consistent UX. Every harness shares repl_skin.py: same banner, prompt, history, progress indicators.
  4. Agent-native output. --json on every command; --help and which are the discovery mechanism.
  5. No graceful degradation. The backend application is a hard dependency; tests fail, not skip, when it is missing.

Rule 5 is the one to remember when we get to limitations.

Three things line up. The “agent-native CLI” idea has become its own genre: a May HN thread on the principles hit 112 points, and derivatives like CLI-Anything-WEB (CLIs for web apps via traffic capture) cite this repo as the origin. The skills ecosystem matured — every harness now ships a canonical SKILL.md under skills/, installable with npx skills add, so Claude Code, Codex and OpenClaw can discover a tool without an MCP server. And the registry crossed the threshold where an agent can plausibly find what it needs: 79 harnesses spanning CAD, DAWs, note apps, GIS, game engines and infra tools.

How the generator works

Install the plugin in Claude Code, then point it at a source tree or a GitHub URL:

# Claude Code
/plugin marketplace add HKUDS/CLI-Anything
/plugin install cli-anything

# Build a full harness (all 7 phases)
/cli-anything ./gimp
/cli-anything https://github.com/blender/blender

# Later: expand coverage, re-test, or validate against HARNESS.md
/cli-anything:refine ./shotcut "picture-in-picture compositing"
/cli-anything:test ./inkscape
/cli-anything:validate ./audacity

The phases, from the plugin README:

PhaseWhat happens
1Codebase analysis — backend engine (MLT, GEGL, bpy), file formats, headless entry points
2Architecture — command groups mirroring the app’s domains
3Implementation — core/ modules, utils/ backends, the Click CLI
4–5Test planning, then unit tests with synthetic data and E2E tests against the real app
6pytest -v results written to TEST.md
6.5SKILL.md generated from the Click decorators, setup.py and README
7PyPI packaging under cli_anything.* and pip install -e .

The output lands in <software>/agent-harness/cli_anything/<software>/ as core/, utils/ (backend adapters), tests/, a packaged skills/SKILL.md and the Click entry point. The real gimp_cli.py is ordinary, readable code: the JSON switch is a module-level flag, errors become {"error": ..., "type": ...} objects in JSON mode, and the group auto-saves the project on exit:

@click.group(invoke_without_command=True)
@click.option("--json", "use_json", is_flag=True, help="Output as JSON")
@click.option("--project", "project_path", type=str, default=None)
@click.option("--dry-run", "dry_run", is_flag=True, default=False)
@click.pass_context
def cli(ctx, use_json, project_path, dry_run):
    global _json_output
    _json_output = use_json

@project.command("new")
@click.option("--width", "-w", type=int, default=1920, help="Canvas width")
@click.option("--height", "-h", type=int, default=1080, help="Canvas height")
@click.option("--mode", type=click.Choice(["RGB", "RGBA", "L", "LA"]), default="RGB")
def project_new(width, height, mode, ...):

That matters because the generated code is yours to maintain. There is no runtime framework; the plugin is Markdown instructions (HARNESS.md plus per-phase guides loaded on demand) and a small Python skill generator.

Using a generated CLI

Whatever platform built it, the result is the same shape:

cd gimp/agent-harness && pip install -e .
cli-anything-gimp project new --width 1920 --height 1080 -o poster.json
cli-anything-gimp --json --project poster.json layer add -n "Background" --type solid --color "#1a1a2e"
cli-anything-gimp            # bare command → REPL

The README’s LibreOffice example shows the stateful pattern end to end: create a .json project, add a heading and a table, then export render output.pdf calls libreoffice --headless and the test suite checks the %PDF- magic bytes rather than trusting the exit code. That output-verification rule is one of the “critical lessons” in HARNESS.md, alongside the rendering gap (GUI apps apply effects at render time, so a naive exporter silently drops them) and 29.97 fps timecode rounding.

CLI-Hub: the part you can use today

If you never run the generator, CLI-Hub is still worth ten minutes:

pip install cli-anything-hub
cli-hub list                 # 79 harnesses + 24 public CLIs
cli-hub search image
cli-hub install obsidian
cli-hub launch obsidian --help

Since 0.2.0 the hub installs from pip, npm, brew or bundled system tools, so it also fronts third-party agent CLIs that CLI-Anything never generated.

For agents, install the meta-skill and let them shop for themselves:

npx skills add HKUDS/CLI-Anything --skill cli-hub-meta-skill -g -y

Then prompt: “Find appropriate CLI software in CLI-Hub and complete the task: transpose this MuseScore file up a fifth and export MP3.” The agent picks a harness from the live catalog, installs it, reads that harness’s SKILL.md, and runs. The skill is also on ClawHub and SkillHub.

OpenClaw users get a native SKILL.md for the generator too: copy openclaw-skill/SKILL.md from the repo into ~/.openclaw/skills/cli-anything/ and invoke @cli-anything build a CLI for ./gimp.

Community reaction

The repo never had a big HN moment — two March submissions topped out at a handful of points — so the 51K stars came from GitHub trending, Chinese developer communities (the README ships in Chinese, Japanese and German) and the skills ecosystem.

The most useful outside data point is a benchmark by the author of safari-mcp, who built the Safari harness that merged as PR #212. Before shipping it he measured MCP against a subprocess-per-call CLI on the same backend:

  • Per-call latency (list_tabs, warm cache): MCP over persistent stdio 119 ms median vs CLI 3,023 ms — 25× faster for MCP, because every CLI call pays the npx spawn.
  • Five-op reactive workflow: MCP 2.7 s vs CLI 15.3 s; a shell pipeline did not help (15.2 s).
  • Token overhead per API call: 84 MCP tool definitions cost 7,986 tokens in the prompt; the CLI path needs only the bash tool definition, 95 tokens — 84× fewer.

His conclusion: with an MCP server available, use MCP; the CLI is for agents that do not speak MCP (Codex CLI, Copilot CLI, older frameworks), for CI and cron where jq-pipeable JSON matters, and for long sessions where tool-definition tokens dominate cost. That is a fair summary of CLI-Anything’s real niche — not an MCP replacement, but a lower-overhead, more portable interface.

The broader HN discussion on agent-native CLIs raised objections worth carrying over. One commenter did not want them to proliferate at all: “I’d rather we design CLIs for human use and programmatic (automation) use first… Too many tools stray so wildly from UNIX principles.” Another noted a well-behaved CLI should simply check isatty() and drop interactivity. And a maintainer of a deliberately interactive-only CLI reported that Claude wrote a Python pseudo-shell around it and drove it anyway — a reminder that agents can already use most CLIs; the harness’s value is reliability and --json, not access.

Honest limitations

The “real backend” guarantee is not uniformly true. On 2026-09-09 a contributor opened three PRs, all still open, documenting shipped harnesses that violate HARNESS.md’s first rule:

  • Blender (#459): cli-anything-blender render execute out.png generates the bpy script, prints command: blender --background --python /tmp/_render_script.py, exits 0, and never runs it. The complete headless implementation exists in utils/blender_backend.py; nothing outside tests/ imports it.
  • Kdenlive (#461): the export group has presets and xml but no render. export presets advertises h264_hq and h265_hq for output no command can produce. melt_backend.render_mlt() is written and unused.
  • Audacity (#460): export render out.flac --preset flac writes RIFF/WAV bytes to a .flac file and reports format: FLAC. Both branches of the format check call write_wav. sox_backend.convert_format() exists and is never imported.

The pattern is the same in all three: the generator produced a correct backend adapter and a correct test for it, then wired the CLI command to a stub. That is the failure mode you would expect from an LLM pipeline optimising for “tests pass” — so “2,464 tests, 100% pass rate” tells you the tests pass, not that every export produces a real file. Run one real end-to-end export before trusting a harness’s exit code.

Frontier models only. The README is explicit: weaker models “may produce incomplete or incorrect CLIs that require significant manual correction.” A full seven-phase run on a large codebase is a long, expensive agent session, and /refine is “often needed” before production quality.

Source code required. Binary-only software degrades “substantially”. Packaging closed-source APIs and web services is an unchecked roadmap item.

Subprocess latency. A fresh process per call costs seconds when the backend is heavy. The stateful REPL and project files help, but a reactive multi-step loop is still slower than a persistent MCP session.

Python-only, and uneven test depth. Every harness is a Click package under cli_anything.*; Node or Go stacks inherit a Python dependency. Several harnesses (Zoom, NotebookLM, 3MF) have zero E2E tests against the real backend.

Security of generated code. The changelog records a GIMP Script-Fu path injection (fixed in March), Sketch token-file traversal (May) and a defusedxml sweep for XML parsing. Harnesses take untrusted paths and project files; review them like any other generated code.

Who should use CLI-Anything

  • Agents in Codex CLI, Copilot CLI or cron where MCP is unavailable and --json is the interface. Install CLI-Hub, pick harnesses.
  • Teams with internal tools agents need to drive: the generator produces a maintainable CLI, tests and skill file from an existing codebase.
  • OpenClaw and Claude Code users who want agents to discover tooling: the meta-skill is a cheap add.
  • Not for: anyone with a maintained MCP server for the app (use it), anyone needing binary-only or SaaS wrapping today, or anyone planning to skip the “check the bytes” step.

Comparison with alternatives

CLI-AnythingGUI agents (OpenAI CUA, Anthropic Computer Use, cua)MCP servers
InterfaceStructured CLI + --jsonScreenshots + clicksPersistent JSON-RPC tools
Binary-only appsPoorlyYesDepends on API
Per-call overheadProcess spawn (seconds if backend heavy)Vision model per step~100 ms over stdio
Prompt costbash tool only (~95 tokens)Screenshot tokens per stepAll tool schemas per call
Discovery--help, SKILL.md, CLI-HubNoneTool list
PortabilityAny agent with a shellNeeds a CUA harnessMCP-capable clients only
DeterminismHighLowHigh
Build effortOne agent run + refinementNoneManual

For a deeper look at the GUI side of that table, see our review of trycua/cua; for the skills format that CLI-Anything now emits, see Superpowers.

FAQ

Is CLI-Anything free? Yes, Apache-2.0. The generator is a plugin for your existing coding agent, so the only cost is the tokens it burns building a harness. CLI-Hub and every registry harness are free pip installs.

Do I need Claude Code to use it? No. Claude Code is the reference platform, but there are packages for Cursor, Pi, OpenClaw, OpenCode, Codex, Hermes, Reasonix, Qodercli and GitHub Copilot CLI, several marked experimental or community-maintained. To use generated CLIs you need only Python 3.10+ and the target application.

How is this different from MCP? MCP keeps a server alive and exposes typed tools; CLI-Anything produces a normal command-line program. MCP wins on latency (about 25× in the safari-mcp benchmark) when the client supports it; the CLI wins on portability, prompt size (no tool schemas in every request), and in CI or cron. Both can coexist for the same app.

Does a harness reimplement the application? It is not supposed to: the methodology requires real project files and the real application for rendering. In practice, check — at least three shipped harnesses (Blender render, Kdenlive render, Audacity non-WAV export) currently stop short of invoking the backend, with fixes pending in open PRs.

Which software is already covered? 79 harnesses, including Blender, GIMP, Inkscape, Krita, Audacity, LibreOffice, OBS Studio, Kdenlive, Shotcut, Draw.io, FreeCAD, QGIS, Godot, Obsidian, Joplin, Zotero, Calibre, n8n, ComfyUI, Ollama, ChromaDB, WireMock, RenderDoc, Zoom and Exa. cli-hub list prints the current set.

Can it wrap a SaaS or a closed-source app? Not well. The pipeline analyses source; binary-only targets degrade substantially, and closed-source API packaging is an unchecked roadmap item. Harnesses for Exa, Zoom, MiniMax and Novita exist because those have documented HTTP APIs.

How do agents find the right CLI on their own? Install cli-hub-meta-skill via npx skills add. It tells the agent to search the live CLI-Hub catalog, install a match, and read that harness’s SKILL.md, which is generated from its Click decorators and so stays in sync with the real commands.

Bottom line

CLI-Anything makes a convincing case, backed by a tech report and 79 worked examples, that agents should drive software through structured commands rather than screenshots. CLI-Hub and the SKILL.md-per-harness convention are the practical wins: an OpenClaw or Claude Code agent can now discover, install and use a Blender or LibreOffice CLI without anyone writing an MCP server.

The caveat is the one the project’s own methodology states: never trust an export because it exited 0. Three open September PRs show the generator can produce a perfect backend adapter, a passing test suite and a CLI command that calls neither. Use the harnesses, run one real export, read the bytes — and if you build your own, budget for /refine.

Sources