TL;DR
e2e (tester-army/e2e) is an open-source end-to-end testing framework for web and mobile apps. You write ordinary TypeScript tests with locators and assertions, and where a selector would be brittle you write a goal in plain English instead: agent.act('upgrade the workspace to the Pro plan'). An LLM agent drives the app to reach the goal. Once a later assertion confirms the result, the framework records the actions, and the next run replays them without calling a model until the app changes.
Facts as of October 11, 2026:
- 8,824 GitHub stars, 427 forks, Apache-2.0 license, TypeScript. The repo was created on July 22, 2026, and it is on GitHub’s weekly trending list as of today.
- Latest release: e2e 0.19.0 (October 9, 2026), shipped together with
@e2e-dev/web0.14.0 and@e2e-dev/mobile0.11.0. The project says plainly that APIs “can still change between minor releases” on the way to 1.0. - 29 open issues (109 closed).
- npm: the
e2epackage was downloaded 218,726 times between October 3 and October 9. - Requirements: Node.js 24.8+ (or 22.22.3+ on Node 22). Mobile needs Xcode with an iOS simulator, or the Android SDK with an emulator. Windows users run it inside WSL.
- Built by TesterArmy, a Y Combinator company (P26 batch) that sells a hosted agentic testing platform. The framework is the open-source core; you bring your own model.
Our verdict: the deterministic half works out of the box. In our lab, scaffolding, installing Chromium and running a locator-only suite took under a minute, with no account and no API key. The agent half is a real idea, not a gimmick: “record once, replay free, re-plan when the UI changes” answers the main objection to AI testing, which is paying tokens on every CI run. We could not verify the agent half ourselves (it needs a model key), and the 0.x label is honest. Good for teams already on Playwright who want AI steps only where selectors keep breaking; too early to bet a large suite on.
What problem e2e solves
End-to-end tests are valuable and miserable to maintain: selectors break when a button is renamed, and AI-generated content is hard to assert with exact strings.
TesterArmy’s co-founder Oskar Kwaśniewski summed up the pitch in the company’s Launch HN thread in June: “static tests are very brittle: you rely on selectors, need wait times, and can’t really test a lot of dynamic content (think AI chats/interactions).”
The obvious counter-argument came from the same thread. One commenter, poisonborz, wrote that E2E tests “are now quick to write due to LLMs, and are then deterministic AND cheap to run”, and asked how an agent running “the whole time for each test” could compete on token cost. Another, Eridrus, said he was “not super excited about using some 3rd party SaaS as a critical part of my testing.”
The open-source e2e framework, which went viral in early October, reads like an answer to both objections: it runs in your own CI with your own model, and its replay cache means an agent step costs tokens once, not on every run.
How a test looks
Here is the example from the README. It mixes three kinds of step in one test:
// tests/checkout.e2e.ts
import { test, expect } from 'e2e';
test('a member upgrades to Pro', async ({ app, agent, screen }) => {
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');
});
agent.act(goal)hands one goal to the agent, which operates the app until the goal is reached.agent.assert(condition)asks the model to judge the current screen. Useful for things like “the summary mentions the refund”, where an exact string would be fragile.expect(locator)is a classic deterministic assertion, with Testing Library-style queries (getByRole,getByLabel,getByText,getByTestId).
Goals can take parameters, which keeps them stable for the cache:
await agent.act('sign up for a free trial as {name} with email {email}', {
params: { name: 'Ada Lovelace', email: '[email protected]' },
});
await agent.assert('the welcome screen greets Ada by name');
A test with no agent steps needs no model at all. That matters: you can adopt e2e as a plain Playwright-based runner and add AI steps one at a time.
The replay cache: the part that matters
According to the caching docs, the cache records the actions of an agent.act() step only after a later check verifies the result: a locator assertion, agent.assert, or agent.waitFor. An act with no verification after it is reported as unconfirmed and not saved. That rule is the important design decision: the framework never caches an agent run it cannot prove worked.
On the next run, the replay:
- checks that the starting screen is the same route,
- finds each recorded control by role, name, test id and surrounding context, and repeats the action,
- checks that the end state changed the way the recording did (controls appeared, went away, or changed state),
- finishes without a model call, or hands the step to the agent, from the current screen, with a note of where the replay stopped.
The run summary reports the split, for example Cache 4 replayed · 1 handed off · 1 missed. Values that change on every run, such as timestamps and fresh email addresses, can be wrapped in unique() so they don’t cause a miss each time. Changing the model does not invalidate the cache; renaming the test, changing the instruction or switching engines does.
agent.assert, agent.waitFor and agent.extract always run live, so a suite that leans on AI assertions keeps paying for them. The docs’ advice is direct: “Use expect for exact checks, and give each act one goal.”
What else is in the box
e2e ships as several packages:
| Package | What it does |
|---|---|
e2e | The SDK, runner and CLI |
@e2e-dev/web | Chromium, Firefox and WebKit through Playwright |
@e2e-dev/mobile | iOS and Android simulators and emulators through agent-device |
@e2e-dev/github | Posts results as a pull request comment |
@e2e-dev/kernel | Kernel hosted browsers |
@e2e-dev/eas | EAS hosted iOS simulators and Android emulators |
@e2e-dev/smol | Each attempt in a branch of a warm Chromium microVM |
@e2e-dev/decision | Decision-model executors for bounded actions |
Other features worth knowing about, all from the docs:
- Model choice. Any AI SDK provider works, including local models through Ollama, plus subscription logins (ChatGPT, GitHub Copilot, OpenCode Console, SuperGrok) via
npx e2e login. e2e explore. Give the agent a goal and no test file; it plans, operates the app and writes an assessment.- Secrets handling. A
Secretis a handle the model never sees in plaintext; the runner types it. After a secret fill, no more screenshots go to the model for that attempt, and known secret values are redacted to<secret:name>in model input, reports and logs. The security page is unusually candid about what still leaks: videos, binary downloads, your app’s own log output, and transformed values like the last four characters of a key. - Coding-agent integration.
npx e2e initinstalls an agent skill and registers ane2e mcpserver for Claude Code and Cursor; the docs ship offline innode_modules/e2e/docs. - What the model sees. By default, the accessibility tree, the redacted URL and earlier steps. A screenshot is added only when the text snapshot is not enough, or when you set
vision. That hybrid approach is cheaper than vision-only agents, which Kwaśniewski called out on HN as the main difference from vision-based competitors.
Hands-on test (October 11, 2026)
Our lab is a CPU-only GitHub-hosted runner: a fresh node:22-bookworm container (x86_64, 2 CPUs, 8 GB RAM, no GPU), no model API key and no TesterArmy account. That means we could test everything e2e does without a model, and confirm what happens when an agent step has none. We ran six steps against commit 449fa93670ee. All six ran, in 55.3 seconds in total.
| Step | Result | Time |
|---|---|---|
npx [email protected] init --yes | Scaffolded config, example test, skill and MCP files | 7.4 s |
npm install, e2e --version, telemetry status | 0.19.0 on Node v22.23.3; telemetry enabled | 9.6 s |
npx e2e-web install chromium --with-deps | Chrome Headless Shell 153.0.8010.12, 114.3 MiB | 24.7 s |
Write a demo page and a locator-only test; e2e list | Both tests listed | 1.8 s |
npx e2e run (no model calls) | 2 passed, run duration 2.24 s | 5.9 s |
| Agent step with no API key | Test failed with MODEL_PROVIDER_FAILED, as expected | 5.9 s |
What we learned from the record:
init --yes writes more than a config. It created package.json, e2e.config.ts and tests/example.e2e.ts, plus .agents/skills/e2e/ (9 files), a symlink from .claude/skills/e2e, .mcp.json, .cursor/mcp.json and 10 new .gitignore entries. That is convenient if you use Claude Code or Cursor and noise if you don’t. The non-interactive default picks the Vercel AI Gateway with gateway('openai/gpt-6-luna-fast').
The deterministic path is fast and needs nothing. Our test served a one-button counter page and clicked “Increment” twice with screen.getByRole('button', 'Increment').click(), then asserted Count: 2 on the status element. The counter test passed in 932 ms and the generated example in 247 ms. The run header still printed the configured model (model gateway/openai/gpt-6-luna-fast) even though no model was called; the docs say tests without agent steps make no model calls.
Without a key, agent steps fail fast and clearly. The agent.act('press the increment button once') step failed after 299 ms with “Unauthenticated request to AI Gateway”, an instruction to set AI_GATEWAY_API_KEY, a pointer to the trace file, and Cache 1 missed. The step itself exited 0 only because we piped the output through tail; the test failed. No hang, no confusing stack trace.
Telemetry is on by default. e2e telemetry status printed Status: enabled and said it sends “the command, the versions, the OS, and summaries of runs and MCP sessions. Never test names, app data, or credentials.” Turn it off with npx e2e telemetry disable or E2E_TELEMETRY_DISABLED=1 (we set the variable for every step after the check).
What this test does not tell you: how reliable agent.act is on a real app, what a typical step costs in tokens, how often a replay hands off after a UI change, or anything about mobile (no simulator in a Linux container). Those questions decide whether e2e is worth adopting.
Community reaction
e2e topped GitHub’s daily trending list on October 5, gaining 1,430 stars that day according to AICoder. Reactions quoted on TesterArmy’s e2e page focus on the replay cache: “the agent uses tokens only once, it remembers what it did for all future runs.”
The harder questions came during the platform’s Launch HN in June (132 points, 69 comments), before the open-source framework existed:
- Cost. pranshuchittora reported that a basic test of booking.com took about 3 minutes, and that Playwright MCP “easily consumes 1M+ tokens for a test with ~20 steps (including image input).” The open-source framework’s text-first snapshots and replay cache are the direct response.
- “Is this a problem?” Several commenters argued LLMs already make classic E2E tests easy to write, a fair point for stable apps.
- Vendor lock-in. Commenters didn’t want a third-party SaaS in their test path. e2e is Apache-2.0, runs in your CI, and works with local models, so that objection mostly goes away for the framework, though not for the hosted platform.
Honest limitations
- Pre-1.0 churn. The project says APIs and config can change between minor releases, and it shipped 0.19.0 less than three months after the repo was created. Pin exact versions.
- The package name.
e2eis so generic that searching for help is hard. Search for “tester-army e2e”. - AI assertions always cost. Only
actsteps are cached. A suite built onagent.assertpays for a model call on every run. - Non-determinism moves, it doesn’t vanish. Replays are deterministic, but every cache miss or hand-off puts an LLM back in the loop, in CI, where you least want surprises.
- Telemetry on by default. Documented and easy to disable, but on.
- Mobile setup is heavier. Xcode and a simulator, or the Android SDK and an emulator, or a hosted service.
- Company incentives. The framework is the open core of a commercial platform. Expect the hosted product to get some features first.
How it compares
| Tool | Approach | Model needed | Best for |
|---|---|---|---|
| e2e (TesterArmy) | Locators + natural-language goals, verified replay cache, web and mobile | Only for agent steps | Teams mixing stable and fragile flows across web and mobile |
| Playwright | Deterministic code, selectors | No | Stable web apps; the engine e2e’s web package is built on |
| Stagehand (Browserbase) | AI actions inside browser automation scripts | Yes, for AI steps | Web automation and scraping more than test suites |
| Midscene.js | Vision-driven UI automation | Yes | Visual UI automation across platforms |
| Maestro / Detox | Deterministic mobile flows | No | Mobile-only teams who want no model in CI |
If your Playwright suite is stable, e2e adds little. If a few volatile flows keep breaking, or you test AI features whose output changes every run, it is the most complete open-source attempt we have seen at making agent steps affordable in CI.
FAQ
Is e2e free?
Yes. The framework is Apache-2.0 and runs in your own CI. You pay only for the model you configure, and only for agent steps; tests that use locators alone make no model calls. TesterArmy’s hosted platform is a separate, paid product.
Can I run e2e with a local model?
Yes. The docs list Ollama (through ollama-ai-provider-v2) and any server with an OpenAI Responses-compatible API. How well a small local model handles multi-step goals is a separate question that the project does not benchmark.
Does e2e replace Playwright?
No. The web engine runs on Playwright, and a migration guide covers running both side by side and moving one file at a time.
How does e2e keep my passwords away from the model?
Declare credentials as secrets. The model sees only the secret’s name and purpose; the runner types the value, turns screenshots off for the rest of the attempt, and redacts the value in model input, reports and logs. Videos and binary downloads are not masked, so check them before sharing.
What happens when my UI changes and a cached step no longer matches?
The replay stops at the first control or end state that doesn’t match, and the agent takes over from the current screen with a record of what already ran. A replay that ends in a repair evicts the old entry, and the next verified pass records a new one.
Verdict
e2e gets the core design right: deterministic assertions by default, agent steps where they earn their cost, and a cache that only records what a later check has proven. The part we could test, scaffolding and a model-free suite, worked in under a minute with clear output and a clear failure when no model was configured. The part that matters most, how well the agent and the replay cache hold up on a real, changing app, still needs testing in your own project. Start with one flaky flow, pin the version, and turn telemetry off if that matters to you.
Sources
- tester-army/e2e on GitHub: README, stars, license, releases, issues
- e2e documentation
- Quickstart
- Caching (replay cache)
- Security
- How agent steps work
- CLI reference
- e2e 0.19.0 release
- e2e on npm
- Launch HN: TesterArmy (YC P26), June 18, 2026
- TesterArmy e2e product page
- AICoder: e2e tops GitHub daily trending (October 5, 2026)