End-to-end tests · web, Android, CLI tools and Electron

Plain-English end-to-end tests.AI writes each test once. Every run after that uses none.

What you get

So a 20-test run costs 2¢ and gives the same answer every time. Web, Android, CLI tools and Electron apps, in our cloud or on your own machine. For developers and teams without a QA engineer.

Start

100 credits a month · no card

The problem

  1. Flaky end-to-end tests nobody has time to fix.
  2. AI testing tools that charge 12–40¢ every single run.
  3. Tests locked in a vendor's cloud. When Octomind shut down, its users lost their runs.

See a run

Show a screen of the app

Pick a screen. Write and Run are the real run in Fig. 1.

Test my app · describe it

Say what should work

Your words Dictate

Creating a project is the core action of the app, so we want to be sure it is really saved and not just shown. Log in as the seeded user (use the shared login flow, flows/login.test.md). On the dashboard click "Create project"; a dialog titled "New project" should open. Type Q3 roadmap into the "Project name" field and click "Create". A message saying "Project created" should appear and the projects list should show "Q3 roadmap". Then reload the page: after the reload the projects list must still show "Q3 roadmap", otherwise the project was never persisted.

9 AI calls for $0.0015, lint-cleanDraft the steps

tests/create-project.test.mddraftdrafted

  1. ---
  2. name: A new project is saved
  3. tags: [smoke, projects]
  4. start: /login
  5. setup:
  6. - request: POST /__test/seed
  7. ---
  8. 1. Use: flows/login.test.md
  9. 2. Click "Create project"
  10. 3. Expect: a dialog titled "New project" is open
  11. 4. Fill "Project name" with Q3 roadmap
  12. 5. Click "Create"
  13. 6. Expect: a message says "Project created"
  14. 7. Expect: the projects list shows "Q3 roadmap"
  15. 8. Reload the page
  16. 9. Expect: the projects list shows "Q3 roadmap"

Live run · replay-only · Acme Shop · local

Replaying 1 testNothing failed

A new project is saved13 steps

#StepStatusTime
1.1Go to /loginqueuedrunningpassed52 ms
1.2Fill "Email" with {{params.email}}queuedrunningpassed34 ms
1.3Fill "Password" with {{params.password}}queuedrunningpassed50 ms
1.4Click "Log in"queuedrunningpassed74 ms
1.5Expect: the page heading is "Dashboard"queuedrunningpassed5 ms
2Click "Create project"queuedrunningpassed52 ms
3Expect: a dialog titled "New project" is openqueuedrunningpassed4 ms
4Fill "Project name" with Q3 roadmapqueuedrunningpassed31 ms
5Click "Create"queuedrunningpassed512 ms
6Expect: a message says "Project created"queuedrunningpassed5 ms
7Expect: the projects list shows "Q3 roadmap"queuedrunningpassed3 ms
8Reload the pagequeuedrunningpassed39 ms
9Expect: the projects list shows "Q3 roadmap"queuedrunningpassed971 ms

This run

Trigger
cli · replay-only
AI calls
0 · replays use no AI
Cost
$0.00
Engine
0.1.0

Results › example › Add to cart

Add to cart healed

The one line that matters

Step 2 was repaired: the 'Add to cart' button is now 'Add to bag'. Review the fix.

1 fix to review AIwaitingapproved

Never applied silently. Until you approve, the test keeps its original step. A heal can change how a step is done, never what is checked. Heals are free.

Before · your step

2. Click "Add to cart"

After · AI repair

2. Click "Add to bag"

What is checked: unchangedApprove

Proof

After: the cart shows 1 item

feat: discount codes #42

Open 3 commits into main from feature/discounts

optestra bot commented · edited

Optestra: running 2 tests against the preview…

Optestra: Failed · demo-shop · staging

PassedHealedFailedFlakyBlocked
10100

2 tests in 16.1s · 0 AI calls · $0.00

What went wrong

Failed: Expected order total '$90.00', found '$100.00'

product bug · affects 1 test: Discount code takes 10% off

✕Tests 1 failed, 1 passed

✓build Successful

Fig. 1 A real replay on the demo shop: passed in 2.0 s, 13 steps, 0 AI calls, $0.00.Writing it the first time, from the paragraph in Write, took 9 AI calls for $0.0015 at list price (Bench, 2026-10-08).

Write once.Replay for pennies.

Three steps. Only the second one uses AI.

  1. 1Write

    Write it, or say it.

    Type steps in plain English, one per line. Or describe the test in a paragraph, or speak it, and the app drafts the steps for you to read.

  2. 2Record

    AI writes it once.

    The first run works out each step in a locked-down browser and saves exactly what it did, as a file in your repo. That is the only run that uses AI.

  3. 3Replay

    Every run after replays it.

    No AI, no surprises. Each Expect: line is checked by code, and a run of up to 20 tests costs 2 credits (2¢).

New · version comparison

Did v2 breakwhat v1 did?

Run the same plain-English tests against two versions of your CLI tool or app, and get the regressions in one report. Free on your own machine.

optestra compare --base correct --head broken-fewer-results

Per test: the measure in each version
Testcorrectbroken-fewer-results
tests/sort-orders.test.md✓ order rows: 5✗ order rows: 3 (−40%)
1 regression(s) between correct and broken-fewer-results.

Measure regressions:
  tests/sort-orders.test.md  order rows: 5 → 3 (-40%)
Fig. 2 Real output: the engine's compare on its demo shop, a correct build against one that returns fewer orders. Both builds passed every check the test asked for. Comparing the versions caught what nobody wrote a check for: the order list went from 5 rows to 3.

What you can test.In our cloud or on your machine.

Web, Android, CLI tools and Electron apps, in our cloud or on your own machine. Native desktop apps are coming, on your own machine.

  • Websites

    Any public address or preview deploy.

    Chromium, Firefox and WebKit at desktop, tablet and phone sizes. Saved logins, 2FA codes and a hosted test inbox.

    In our cloud
    Yes
    On your machine
    Yes
  • Android apps

    Upload your APK. A fresh emulator for every test.

    Android 13 to 17, phone or tablet. In the cloud, 2 credits a test (at least 10 a run).

    In our cloud
    Yes
    On your machine
    Yes
  • CLI tools

    Test a command-line tool in plain English, and compare two versions.

    Version comparison is free on your own machine.

    In our cloud
    Coming soon (paid plans)
    On your machine
    Yes
  • Electron and Tauri apps

    Electron apps on every OS, Tauri apps on Windows.

    On your own machine.

    In our cloud
    No
    On your machine
    Yes
  • Native desktop apps

    Windows, macOS and Linux apps.

    Coming, on your own machine.

    In our cloud
    No
    On your machine
    Coming

Runs on your own machine can show up in your dashboard, free.

Who it's for.If you ship, and nobody tests.

Developers and teams without a QA engineer.

  • Solo developers

    Cover your critical flows in an afternoon. Run them on every push for a few cents.

    Start free
  • Startup teams with no QA

    Anyone can write a test in English. Every pull request is checked before it merges.

    Start free
  • Agencies

    One readable test suite per client site, checked on every deploy. No seat pricing.

    Start free
  • Coming from Octomind, Cypress, Playwright or mabl

    We'll migrate your first 10 tests free.

    Migrate my first 10 tests

What "checked" means

A row forevery line.

This is the check table of the run in Fig. 1, straight from the app. Each line in your test is a row: what the test said, what was expected, what the page showed. Code evaluates it on every replay, and your AI never sees the verdict.

The same run's "What was checked" table: five lines from the test, each passed, with the expected and the found value, and "AI calls and cost: 0 calls, $0".
Fig. 3 The app's result screen for the run in Fig. 1.

Measured.With the date and the method.

Bench: fixture apps with real bugs, and a score for what our AI got wrong. Real numbers from 2026-10-08.

0

False passes

on step-written tests, in 15 replays of buggy builds.

0

AI calls

on every replay of an unchanged app. AI only writes a test or proposes a heal.

11/11

Step-written tests

authored correctly by the hosted AI.

Single runs on fixture apps · DeepSeek-V4.1-Flash via OpenRouter and DeepInfra · engine commit 5ab6cf2 · every number, including the ones that don't flatter us

Open-source engine (MIT) on GitHub

What we won'tcompromise.

Six promises. Each one is a design decision, not a slogan.

  1. Green means green.

    Every Expect line becomes a typed check that plain code evaluates. A model never decides pass or fail, and a check is tested to be able to fail.

  2. No AI, no surprises, on every replay.

    A replay uses no model: same steps, same speed, same result, at the price of a browser for a few seconds.

  3. Your tests are yours.

    Plain Markdown files in your repository. Every website test is also a Playwright spec that runs without us, and the engine is open source (MIT).

  4. No silent fixes.

    When your app changes, a heal is proposed with its diff and its proof. Nothing is fixed silently, and no heal can change what is checked.

  5. Clear about data.

    Our AI is pinned to one provider with zero data retention and no fallbacks. Secrets are typed by the runner and never seen by a model.

  6. Nothing to set up.

    Nothing to install and no command line in our cloud. Paste a URL, write a sentence, press Run.

Cheap to start.Cheap to keep.

1 credit is 1 cent. Every feature is on every plan, except cloud CLI runs, which are coming soon to paid plans.

  • Free100 credits a month, no card
  • From $8a month for plans with our AI
  • 2¢for a 20-test cloud run
  • $0for runs on your own machine, forever

Founding members: the first 100 subscribers keep their price for life and get 25% more credits every month.

See pricing

Switching tools?

We'll migrate your first 10 tests,free.

Send us your existing tests (Octomind export, Cypress, Playwright, mabl) or a list of your key flows. We rewrite your 10 most important as plain-English tests, run them against your app, and hand them over within 2 working days.

Questions

The honestanswers.

Does every run use AI?

No. The first run works out the steps and records them. Later runs replay the recording with no AI, so a cloud run costs a couple of credits and the same result every time. AI comes back only for a new or changed step, or when your app changed and a step needs healing, and then it proposes a fix for you to review. How a run works.

Can the AI make a failing test pass?

No. Each Expect line becomes a typed check that plain code evaluates, and the verdict is computed from the checks. The AI works out how to do a step, never whether it passed. A heal can change how a step is done, never what is checked, and it is shown to you for review. Checks and verdicts.

What if you disappear?

Nothing breaks. Your tests are Markdown files in your own repository, the engine that runs them is open source (MIT) and runs on your machine or in your CI, and every recorded website test is also a plain Playwright spec that runs without Optestra. The cloud, the app and hosted AI are the paid parts; your tests don't depend on them. If Optestra disappears.

Playwright's test agents are free. Why would I pay?

You might not. Playwright's planner, generator and healer are good, free and open, and if your team is happy owning Playwright code they may be all you need. We are different in what you keep: the test is an English file anyone can read, each Expect is a checked-by-code assertion, heals are proposals you review, and the same file runs on Android. We also host the browsers and emulators, PR checks and inboxes, so there is nothing to set up. And because every website test exports as a Playwright spec, choosing us doesn't close the other door. Playwright test agents vs us.

More questions

Say what should work.We'll test it.

Built so your first passing test takes about five minutes. 100 free credits a month, no card.

Not ready yet? Get one email a month with what shipped.