THE VERIFICATION LAYER FOR YOUR SOFTWARE FACTORY

Black-box tests for your web apps, written and maintained as code by AI agents, and run on infra that lets you ship fast.

Book a demo

Product walkthrough

See Empirical in action · 12:51

Trusted by fast growing teams

  • Wayground
  • Atlan
  • Shopflo
  • DPDzero
  • 100ms
  • Leap
  • Sortment
  • Shipsy
  • SciSpace
  • Mili
Wayground

“We didn't have to write code, we didn't have to maintain the code. Tests were generated from a prompt.”

— Sandeep Bantia, CTO, Wayground

All customer stories →

01 · Black box

Tests that only see what users see

Our tests don't know or care about your source code. They drive a real browser against your app (click, type, and check what's on the screen) so they stay put when the code underneath changes. Migrating to Rust? Your tests don't know, and don't change.

  • No hooks, mocks, or test-only backdoors in your app
  • Refactors and rewrites don't break the suite
  • Any stack, any framework, any language
Your app · rewrite to Rust
api/checkout.ts-412
api/checkout.rs+388
db/orders.ts-156
db/orders.rs+171

148 files changed

Your tests · 0 files changed
await page.getByRole("button", { name: "Pay now" }).click();
await expect(page.getByText("Order confirmed")).toBeVisible();

✓ 212 passed · before and after

02 · Infra

Fast infra for browser tests

Playwright on CI splits tests by file order, so one slow shard holds up the whole run. We schedule every test from its duration history, longest first, so shards finish together.

  • Duration-aware sharding, retries, and parallel workers
  • No CI config or runners to maintain
  • Video, trace, and logs for every test

Race CI against Empirical →

PLAYWRIGHT + CI · FILE ORDER
#1 21m
#2 17m
#3 14m
#4 13m

WALL CLOCK: 21 MIN

EMPIRICAL · LONGEST FIRST
#1 17m
#2 16m
#3 16m
#4 16m

WALL CLOCK: 17 MIN

Same 16 tests · 4 shards · illustrative durations

03 · Agent

A suite that maintains itself

When your app changes, the agent figures out whether a failure is a bug, a flake, or a test that drifted, then fixes the test or tells you about the bug, in Slack, with evidence.

  • Writes new tests from a prompt, a PR, or a ticket
  • Triages every failure and repairs drifted tests
  • Every change lands as a PR, reviewed by an engineer

See it work in Slack →

Activity · after a deploy
  1. 09:12 DEPLOY checkout v2.4 shipped to staging
  2. 09:31 RUN 3 failures in checkout/
  3. 09:33 AGENT Triaged: 2 tests drifted (button renamed), 1 real regression
  4. 09:34 AGENT Posted to #releases: coupon field rejects valid codes
  5. 09:41 AGENT Opened PR #1260: updated 2 tests, re-ran green
  6. 09:58 HUMAN Reviewed and merged

04 · Composable

Primitives you can build on

Everything in the dashboard is a primitive with an API and a CLI. Trigger runs from your pipeline, hand work to the agent from your own coding agent, and pull results wherever you need them.

  • CLI and API for every action in the product
  • Agent skill for Claude Code, Codex, and friends
  • Results in GitHub checks, Slack, and your tracker
~/app — zsh80×6
$ empirical session -x "add a test for coupon codes at checkout"
✓ PR #1262 opened · 1 test added · run passed

$ empirical api api/test-runs/4821/status
{ "status": "failed", "passed": 211, "failed": 1 }

Stories and writing