TL;DR

e2e (tester-army/e2e) is an open-source end-to-end testing framework for web and mobile apps. You write ordinary TypeScript tests with locators and assertions, and where a selector would be brittle you write a goal in plain English instead: agent.act('upgrade the workspace to the Pro plan'). An LLM agent drives the app to reach the goal. Once a later assertion confirms the result, the framework records the actions, and the next run replays them without calling a model until the app changes.

Facts as of October 11, 2026:

  • 8,824 GitHub stars, 427 forks, Apache-2.0 license, TypeScript. The repo was created on July 22, 2026, and it is on GitHub’s weekly trending list as of today.
  • Latest release: e2e 0.19.0 (October 9, 2026), shipped together with @e2e-dev/web 0.14.0 and @e2e-dev/mobile 0.11.0. The project says plainly that APIs “can still change between minor releases” on the way to 1.0.
  • 29 open issues (109 closed).
  • npm: the e2e package was downloaded 218,726 times between October 3 and October 9.
  • Requirements: Node.js 24.8+ (or 22.22.3+ on Node 22). Mobile needs Xcode with an iOS simulator, or the Android SDK with an emulator. Windows users run it inside WSL.
  • Built by TesterArmy, a Y Combinator company (P26 batch) that sells a hosted agentic testing platform. The framework is the open-source core; you bring your own model.

Our verdict: the deterministic half works out of the box. In our lab, scaffolding, installing Chromium and running a locator-only suite took under a minute, with no account and no API key. The agent half is a real idea, not a gimmick: “record once, replay free, re-plan when the UI changes” answers the main objection to AI testing, which is paying tokens on every CI run. We could not verify the agent half ourselves (it needs a model key), and the 0.x label is honest. Good for teams already on Playwright who want AI steps only where selectors keep breaking; too early to bet a large suite on.

What problem e2e solves

End-to-end tests are valuable and miserable to maintain: selectors break when a button is renamed, and AI-generated content is hard to assert with exact strings.

TesterArmy’s co-founder Oskar Kwaśniewski summed up the pitch in the company’s Launch HN thread in June: “static tests are very brittle: you rely on selectors, need wait times, and can’t really test a lot of dynamic content (think AI chats/interactions).”

The obvious counter-argument came from the same thread. One commenter, poisonborz, wrote that E2E tests “are now quick to write due to LLMs, and are then deterministic AND cheap to run”, and asked how an agent running “the whole time for each test” could compete on token cost. Another, Eridrus, said he was “not super excited about using some 3rd party SaaS as a critical part of my testing.”

The open-source e2e framework, which went viral in early October, reads like an answer to both objections: it runs in your own CI with your own model, and its replay cache means an agent step costs tokens once, not on every run.

How a test looks

Here is the example from the README. It mixes three kinds of step in one test:

// tests/checkout.e2e.ts
import { test, expect } from 'e2e';

test('a member upgrades to Pro', async ({ app, agent, screen }) => {
  await app.open('/settings/billing');

  await agent.act('upgrade the workspace to the Pro plan');
  await agent.assert('the invoice preview shows a prorated amount');

  await expect(screen.getByRole('status')).toContainText('Pro');
});
  • agent.act(goal) hands one goal to the agent, which operates the app until the goal is reached.
  • agent.assert(condition) asks the model to judge the current screen. Useful for things like “the summary mentions the refund”, where an exact string would be fragile.
  • expect(locator) is a classic deterministic assertion, with Testing Library-style queries (getByRole, getByLabel, getByText, getByTestId).

Goals can take parameters, which keeps them stable for the cache:

await agent.act('sign up for a free trial as {name} with email {email}', {
  params: { name: 'Ada Lovelace', email: '[email protected]' },
});
await agent.assert('the welcome screen greets Ada by name');

A test with no agent steps needs no model at all. That matters: you can adopt e2e as a plain Playwright-based runner and add AI steps one at a time.

The replay cache: the part that matters

According to the caching docs, the cache records the actions of an agent.act() step only after a later check verifies the result: a locator assertion, agent.assert, or agent.waitFor. An act with no verification after it is reported as unconfirmed and not saved. That rule is the important design decision: the framework never caches an agent run it cannot prove worked.

On the next run, the replay:

  1. checks that the starting screen is the same route,
  2. finds each recorded control by role, name, test id and surrounding context, and repeats the action,
  3. checks that the end state changed the way the recording did (controls appeared, went away, or changed state),
  4. finishes without a model call, or hands the step to the agent, from the current screen, with a note of where the replay stopped.

The run summary reports the split, for example Cache 4 replayed · 1 handed off · 1 missed. Values that change on every run, such as timestamps and fresh email addresses, can be wrapped in unique() so they don’t cause a miss each time. Changing the model does not invalidate the cache; renaming the test, changing the instruction or switching engines does.

agent.assert, agent.waitFor and agent.extract always run live, so a suite that leans on AI assertions keeps paying for them. The docs’ advice is direct: “Use expect for exact checks, and give each act one goal.”

What else is in the box

e2e ships as several packages:

PackageWhat it does
e2eThe SDK, runner and CLI
@e2e-dev/webChromium, Firefox and WebKit through Playwright
@e2e-dev/mobileiOS and Android simulators and emulators through agent-device
@e2e-dev/githubPosts results as a pull request comment
@e2e-dev/kernelKernel hosted browsers
@e2e-dev/easEAS hosted iOS simulators and Android emulators
@e2e-dev/smolEach attempt in a branch of a warm Chromium microVM
@e2e-dev/decisionDecision-model executors for bounded actions

Other features worth knowing about, all from the docs:

  • Model choice. Any AI SDK provider works, including local models through Ollama, plus subscription logins (ChatGPT, GitHub Copilot, OpenCode Console, SuperGrok) via npx e2e login.
  • e2e explore. Give the agent a goal and no test file; it plans, operates the app and writes an assessment.
  • Secrets handling. A Secret is a handle the model never sees in plaintext; the runner types it. After a secret fill, no more screenshots go to the model for that attempt, and known secret values are redacted to <secret:name> in model input, reports and logs. The security page is unusually candid about what still leaks: videos, binary downloads, your app’s own log output, and transformed values like the last four characters of a key.
  • Coding-agent integration. npx e2e init installs an agent skill and registers an e2e mcp server for Claude Code and Cursor; the docs ship offline in node_modules/e2e/docs.
  • What the model sees. By default, the accessibility tree, the redacted URL and earlier steps. A screenshot is added only when the text snapshot is not enough, or when you set vision. That hybrid approach is cheaper than vision-only agents, which Kwaśniewski called out on HN as the main difference from vision-based competitors.

Hands-on test (October 11, 2026)

Our lab is a CPU-only GitHub-hosted runner: a fresh node:22-bookworm container (x86_64, 2 CPUs, 8 GB RAM, no GPU), no model API key and no TesterArmy account. That means we could test everything e2e does without a model, and confirm what happens when an agent step has none. We ran six steps against commit 449fa93670ee. All six ran, in 55.3 seconds in total.

StepResultTime
npx [email protected] init --yesScaffolded config, example test, skill and MCP files7.4 s
npm install, e2e --version, telemetry status0.19.0 on Node v22.23.3; telemetry enabled9.6 s
npx e2e-web install chromium --with-depsChrome Headless Shell 153.0.8010.12, 114.3 MiB24.7 s
Write a demo page and a locator-only test; e2e listBoth tests listed1.8 s
npx e2e run (no model calls)2 passed, run duration 2.24 s5.9 s
Agent step with no API keyTest failed with MODEL_PROVIDER_FAILED, as expected5.9 s

What we learned from the record:

init --yes writes more than a config. It created package.json, e2e.config.ts and tests/example.e2e.ts, plus .agents/skills/e2e/ (9 files), a symlink from .claude/skills/e2e, .mcp.json, .cursor/mcp.json and 10 new .gitignore entries. That is convenient if you use Claude Code or Cursor and noise if you don’t. The non-interactive default picks the Vercel AI Gateway with gateway('openai/gpt-6-luna-fast').

The deterministic path is fast and needs nothing. Our test served a one-button counter page and clicked “Increment” twice with screen.getByRole('button', 'Increment').click(), then asserted Count: 2 on the status element. The counter test passed in 932 ms and the generated example in 247 ms. The run header still printed the configured model (model gateway/openai/gpt-6-luna-fast) even though no model was called; the docs say tests without agent steps make no model calls.

Without a key, agent steps fail fast and clearly. The agent.act('press the increment button once') step failed after 299 ms with “Unauthenticated request to AI Gateway”, an instruction to set AI_GATEWAY_API_KEY, a pointer to the trace file, and Cache 1 missed. The step itself exited 0 only because we piped the output through tail; the test failed. No hang, no confusing stack trace.

Telemetry is on by default. e2e telemetry status printed Status: enabled and said it sends “the command, the versions, the OS, and summaries of runs and MCP sessions. Never test names, app data, or credentials.” Turn it off with npx e2e telemetry disable or E2E_TELEMETRY_DISABLED=1 (we set the variable for every step after the check).

What this test does not tell you: how reliable agent.act is on a real app, what a typical step costs in tokens, how often a replay hands off after a UI change, or anything about mobile (no simulator in a Linux container). Those questions decide whether e2e is worth adopting.

Community reaction

e2e topped GitHub’s daily trending list on October 5, gaining 1,430 stars that day according to AICoder. Reactions quoted on TesterArmy’s e2e page focus on the replay cache: “the agent uses tokens only once, it remembers what it did for all future runs.”

The harder questions came during the platform’s Launch HN in June (132 points, 69 comments), before the open-source framework existed:

  • Cost. pranshuchittora reported that a basic test of booking.com took about 3 minutes, and that Playwright MCP “easily consumes 1M+ tokens for a test with ~20 steps (including image input).” The open-source framework’s text-first snapshots and replay cache are the direct response.
  • “Is this a problem?” Several commenters argued LLMs already make classic E2E tests easy to write, a fair point for stable apps.
  • Vendor lock-in. Commenters didn’t want a third-party SaaS in their test path. e2e is Apache-2.0, runs in your CI, and works with local models, so that objection mostly goes away for the framework, though not for the hosted platform.

Honest limitations

  • Pre-1.0 churn. The project says APIs and config can change between minor releases, and it shipped 0.19.0 less than three months after the repo was created. Pin exact versions.
  • The package name. e2e is so generic that searching for help is hard. Search for “tester-army e2e”.
  • AI assertions always cost. Only act steps are cached. A suite built on agent.assert pays for a model call on every run.
  • Non-determinism moves, it doesn’t vanish. Replays are deterministic, but every cache miss or hand-off puts an LLM back in the loop, in CI, where you least want surprises.
  • Telemetry on by default. Documented and easy to disable, but on.
  • Mobile setup is heavier. Xcode and a simulator, or the Android SDK and an emulator, or a hosted service.
  • Company incentives. The framework is the open core of a commercial platform. Expect the hosted product to get some features first.

How it compares

ToolApproachModel neededBest for
e2e (TesterArmy)Locators + natural-language goals, verified replay cache, web and mobileOnly for agent stepsTeams mixing stable and fragile flows across web and mobile
PlaywrightDeterministic code, selectorsNoStable web apps; the engine e2e’s web package is built on
Stagehand (Browserbase)AI actions inside browser automation scriptsYes, for AI stepsWeb automation and scraping more than test suites
Midscene.jsVision-driven UI automationYesVisual UI automation across platforms
Maestro / DetoxDeterministic mobile flowsNoMobile-only teams who want no model in CI

If your Playwright suite is stable, e2e adds little. If a few volatile flows keep breaking, or you test AI features whose output changes every run, it is the most complete open-source attempt we have seen at making agent steps affordable in CI.

FAQ

Is e2e free?

Yes. The framework is Apache-2.0 and runs in your own CI. You pay only for the model you configure, and only for agent steps; tests that use locators alone make no model calls. TesterArmy’s hosted platform is a separate, paid product.

Can I run e2e with a local model?

Yes. The docs list Ollama (through ollama-ai-provider-v2) and any server with an OpenAI Responses-compatible API. How well a small local model handles multi-step goals is a separate question that the project does not benchmark.

Does e2e replace Playwright?

No. The web engine runs on Playwright, and a migration guide covers running both side by side and moving one file at a time.

How does e2e keep my passwords away from the model?

Declare credentials as secrets. The model sees only the secret’s name and purpose; the runner types the value, turns screenshots off for the rest of the attempt, and redacts the value in model input, reports and logs. Videos and binary downloads are not masked, so check them before sharing.

What happens when my UI changes and a cached step no longer matches?

The replay stops at the first control or end state that doesn’t match, and the agent takes over from the current screen with a record of what already ran. A replay that ends in a repair evicts the old entry, and the next verified pass records a new one.

Verdict

e2e gets the core design right: deterministic assertions by default, agent steps where they earn their cost, and a cache that only records what a later check has proven. The part we could test, scaffolding and a model-free suite, worked in under a minute with clear output and a clear failure when no model was configured. The part that matters most, how well the agent and the replay cache hold up on a real, changing app, still needs testing in your own project. Start with one flaky flow, pin the version, and turn telemetry off if that matters to you.

Sources