Testing 102: Fast unit tests, honest end-to-end tests

This is the second post in our testing series. The first one laid out the map. This one is about where to put the effort now that a coding agent can bring the whole system up by itself.
For twenty years, advice on the test mix has rested on one fact: the higher up the stack a test runs, the more it costs. End-to-end tests needed a shared staging environment, someone to keep it alive and someone to read the failure. So teams kept a handful of them and pushed everything else down into unit tests. That's the test pyramid, and the reasoning behind it was sound.[1]
That cost has moved. A coding agent can run docker compose up, apply the migrations, seed the data, drive a browser through a purchase order, read the logs when it fails, fix the code and run it all again. It will do that at 3am without complaining that staging is slow. The layer that used to be the most expensive to run is now one an agent can run as often as you let it.
That changes the argument, but not in the direction people expect. Unit tests don't matter less. They matter differently: they are the agent's inner loop, and they have to be fast. Integration and end-to-end tests become the main automation, the place where you find out whether a change actually works. And between the agent and a full end-to-end run there is almost always one thing in the way: the login screen.
Unit tests are the agent's inner loop
A unit test still answers the most precise question in the suite: is this piece of logic right? When it fails, it points at one function and one assertion. For a person that's convenient. For an agent it's the difference between a two-line fix and an afternoon of guessing, because an agent reasons from the failure message it's given. A red end-to-end test says the order never reached Awaiting approval. A red unit test says line_total(100, Decimal("10.00")) returned 1000.00 when it should have returned 950.00.
Unit tests also carry the rules you can't afford to get wrong: prices, quantities, dates, rounding, unit conversions, approval thresholds. Every one of those rules has boundaries, and an end-to-end journey will only ever walk through one or two of them. The unit layer is where you check all of them, including the one at exactly 100 units.
And they push back on design. Code that's hard to unit test is usually code where the business rule is tangled up with the database call. Left alone, an agent will happily write that code. A suite that expects small, pure, testable rules is one of the few pressures that stops it.
Speed matters more when the agent is the one waiting
A person runs the suite a few times an hour. An agent runs it every time it changes something, which on a real task can be dozens of times. If the suite takes three minutes, twenty iterations is an hour of the agent sitting idle and an hour of you waiting for the pull request. Worse, a slow suite teaches the agent, and the harness around it, to skip the run or to run a narrow slice and hope.
Most Python suites run serially, because parallelism in pytest is opt-in and keeping a suite parallel-safe is work nobody finishes. That's the gap we built rstest to close: a pytest-compatible runner that is parallel by default, with nothing to migrate. On SQLAlchemy's suite it took a 533-second serial run down to 53 seconds; on aiohttp, 197 seconds down to 68. Two of its features line up exactly with how agents work: --changed runs only the tests affected by the diff, and --watch reruns on every save.
# inner loop: only what the change touched, in parallel
rstest --changed
# before the agent hands the change back: the whole suite, still in parallel
rstest -n autoPut that in the agent's loop and the inner cycle is seconds, not minutes. Fast unit tests keep the agent moving. What they can't tell you is whether the product works.
Three problems a parallel suite brings to the surface
Parallel by default is the right default, but it surfaces problems a serial suite was hiding. For a person they're annoying. For an agent they're worse: each one is a failure it can't reproduce or can't attribute, so it guesses, and an agent's guess usually arrives as a confident code change. Three come up again and again, and rstest has a command for each.
It fails in CI and passes locally. In a parallel run, which tests share a worker, and in what order, decides which leaked state reaches which test. CI has a different core count from your laptop, so it builds a different schedule, and the failure never shows up locally. rstest records every parallel run's per-worker assignment and order in a small journal. Upload it when a CI job fails, and rstest replay pins that exact schedule back on your machine, forcing the recorded worker count even on a laptop with fewer cores.
# CI: keep the schedule of a failed run
- run: rstest -n auto
- uses: actions/upload-artifact@v4
if: failure()
with:
name: rstest-replay
path: .rstest_cache/replay/latest.json# locally, or in the agent's sandbox
gh run download <run-id> -n rstest-replay -D ci-replay
rstest replay --journal ci-replay/latest.jsonEach worker's assignment and order are reproduced exactly, which is what state-leak failures depend on; a true timing race between workers is still best-effort. For an agent, that's the difference between "can't reproduce, adding a retry" and a red test it can actually debug.
It only fails in parallel. Most of these come down to isolation: a test that changes a module-level setting and never resets it, a fixture that writes to a fixed file path or port, a session-scoped fixture that one test mutates while another assumes it's untouched, a test that only passes because an earlier test in the same file set something up. rstest audit runs the suite in parallel, repeatedly if you ask, since a race doesn't fire every time, compares the results against serial runs and classifies every parallel-only failure: isolation leak, wall-clock dependency, order dependency within a file, or a test that's flaky even on its own. For the isolation and wall-clock failures it prints a ready-to-paste conftest.py block that marks exactly those tests serial as a stopgap, and names the real fix for each.
rstest audit --audit-repeat 5This matters for agents in particular. An agent that sees a test fail under parallelism will try to fix the code under test. audit tells it the code is fine and the fixture is the problem.
Was it my change, or another test? When a test goes red after a change, there are two very different explanations: the change broke it, or some other test leaves state behind that this one now trips over. rstest bisect separates them. It first runs the failing test alone. If it fails by itself, it's a plain bug: either the change is wrong or the test's expectation is. If it only fails after the rest of the suite has run, bisect narrows the tests that ran before it down to the smallest set that still breaks it, and prints a command that reproduces the failure with just those tests.
rstest bisect tests/test_orders.py::test_reservation_leaves_stock_untouchedFor an agent, the first answer means "fix your change, or fix the test" and the second means "leave this test alone and fix the one polluting it". Getting that wrong is how an agent ends up loosening a perfectly good assertion.
State that outlives the test
Of all the problems that come with adding unit and integration tests at scale, leaked state costs the most time, because it almost never fails the test that caused it. A test starts a background thread and never joins it, opens a file or a socket and never closes it, or builds a connection pool in a fixture with no teardown. That test passes. A later test on the same worker fails, with a stray thread racing it, an exhausted pool or a "too many open files" error, and takes the blame. In a serial suite the victim is at least always the same test. In a parallel suite it moves between runs. A suite that grows at agent speed collects these faster than anyone can review them.
rstest measures this per test. It counts live threads and open file descriptors before setup and again after teardown, so a test that cleans up after itself nets zero and a test that doesn't is named. --doctor reports the leaks; --fail-on-leak turns them into a gate that fails the build.
rstest -n 0 --fail-on-leakRun the gate at -n 0 so the order is fixed. A session-scoped fixture's resources are charged to whichever test uses it first, and in parallel that can be a different test on every run. The fix is almost always the same: move the resource into a fixture that releases it in teardown.
@pytest.fixture
def carrier_client():
client = CarrierClient(base_url=FAKE_CARRIER_URL)
yield client
client.close() # runs even when the test failsIntegration tests have a second kind of leak that no thread or file count will show: data. A test that inserts a supplier and leaves it behind changes what the next test sees, and under parallelism, what a test on another worker sees too. The standard fix is to run each test inside a transaction and roll it back at the end, so every test starts from the same seeded database:
@pytest.fixture
def session(db):
connection = db.connect()
transaction = connection.begin()
yield Session(bind=connection)
transaction.rollback() # undoes everything the test wrote
connection.close()Between the two, a test leaves nothing behind that the next one can trip over, and audit and bisect have far less to find.
Unit tests prove the parts. The bugs agents write live between them.
Look at what actually breaks when an agent makes a change it believes is finished. It renames a field in the API response and updates every test that mentions it, but not the frontend that reads it. It adds a column and forgets the migration. It reads a new setting from an environment variable that exists on its machine and nowhere else. It changes a message on the producer side of a queue, and the consumer, three repositories away, quietly starts dropping events.
Every one of those passes the unit tests. Every one of them is invisible from inside the agent's context window, because the agent sees files, not a running system. The only way to catch them is to put the pieces together and run them. That's why integration and end-to-end tests should carry the main automation: they are the only layers that check the assumptions one part makes about another.
This isn't a new position. "Write tests. Not too many. Mostly integration." is Guillermo Rauch's line, and Kent C. Dodds built a whole testing philosophy on it, on the grounds that integration tests give the most confidence for the effort.[2] What held teams back from going further up the stack was never the value of those tests. It was the cost of running them and the pain of keeping them green. Agents change both.
What changes when the agent can bring the whole stack up
The old end-to-end setup was one long-lived staging environment shared by everyone. Tests collided with each other's data, someone deployed a half-finished branch, and a red run could mean a bug, a noisy neighbour or a certificate that expired on Tuesday. Google reported that about 1.5% of all its test runs came back flaky, and that almost 16% of its tests showed some flakiness.[3] Full-stack tests are where the classic causes, async waits, concurrency and order dependencies, pile up.[4]
An agent doesn't need staging. It needs a stack it can own: started from nothing, used by exactly one run and thrown away afterwards. Google's engineering book makes the same case: a shared staging environment is a major source of unreliability and slow turnaround, and the answer is a hermetic system under test, one the test brings up for itself so the result depends on the code and nothing else.[5] Once the stack is hermetic, the end-to-end loop looks just like the unit loop, only slower:
- Start the stack from the commit under test.
- Wait until every service reports healthy.
- Apply migrations and seed known data.
- Run the journeys.
- On failure, collect logs, traces and screenshots and hand them back to the agent.
- Tear everything down.
For the purchase-to-delivery system from the first post, the whole stack fits in one Compose file. The health checks matter: they're what turns "the container started" into "the service is ready", which is the difference between a real failure and a race.[6]
services:
db:
image: postgres:16
environment:
POSTGRES_PASSWORD: test
healthcheck:
test: ["CMD", "pg_isready", "-U", "postgres"]
interval: 2s
retries: 30
queue:
image: rabbitmq:3.13
healthcheck:
test: ["CMD", "rabbitmq-diagnostics", "-q", "ping"]
interval: 5s
retries: 30
idp: # a real OIDC server with fake users; more on this below
image: ghcr.io/navikt/mock-oauth2-server:3
ports: ["8080:8080"]
api:
build: ./api
environment:
DATABASE_URL: postgresql://postgres:test@db/postgres
OIDC_ISSUER: http://idp:8080/purchasing
depends_on:
db: { condition: service_healthy }
queue: { condition: service_healthy }
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 2s
retries: 30
web:
build: ./web
ports: ["3000:3000"]
depends_on:
api: { condition: service_healthy }And a session fixture that owns its lifecycle, so the agent never has to remember to start or stop anything:
import subprocess
from uuid import uuid4
import pytest
@pytest.fixture(scope="session")
def stack():
compose = ["docker", "compose", "-p", f"e2e-{uuid4().hex[:8]}"]
subprocess.run([*compose, "up", "--build", "--wait"], check=True)
try:
subprocess.run([*compose, "exec", "-T", "api", "python", "-m", "app.seed"], check=True)
yield Stack(web_url="http://localhost:3000")
finally:
with open("test-results/stack.log", "w") as log:
subprocess.run([*compose, "logs", "--no-color"], stdout=log)
subprocess.run([*compose, "down", "-v"], check=True)The last step is the one people skip, and for an agent it matters most. A person looking at a red end-to-end test can open the browser and poke around. An agent can only work with what you hand it. Give it the service logs, a Playwright trace and a screenshot of the page at the moment it failed, and a vague "timed out waiting for Awaiting approval" becomes "the API returned 422 because unit_of_measure is now required". That's the precision of a unit test, recovered at the top of the stack.
pytest tests/e2e --tracing retain-on-failure --screenshot only-on-failure --output test-resultsA stack an agent can actually use has a few properties, and most of them are about removing the steps that used to live in someone's head:
- One command, no manual steps. If the README says "then log into the admin panel and create a supplier", the agent will either stall or improvise.
- Health checks on every service, so "up" means ready, not started.
- Seed data created by the run, not borrowed from a database someone else is also using.
- Nothing shared between runs. A unique project name per run, and no fixed host ports if you want several agents working side by side.
- Artefacts on disk: logs, traces and screenshots written somewhere the agent can read them.
- Teardown that always runs, pass or fail.
- Fakes at the boundary for the systems you don't own, like the carrier API, the payment provider or the customer's ERP, with the real contracts checked separately.
Build that and the agent will get the stack up, open the browser, and stop at the login page.
Authentication is the wrinkle
Authentication is the wrinkle in all of this. It's worth being specific about why, because each reason has a different fix.
The login lives outside your stack. With single sign-on, the browser leaves your app for Entra ID, Okta or Google: a system you don't own, can't start in a container, and that is actively designed to stop automated clients. Bot detection, CAPTCHAs and device checks are doing their job when they block your agent.
MFA is built to need a human. A push notification to a phone is the whole point of the second factor. No amount of prompting gets an agent past it, and that's correct.
Tokens expire mid-run. Access tokens often last minutes. A long journey or a retry loop crosses the boundary, and the failure shows up as a random 401 halfway through a test that has nothing to do with auth.
Redirects and issuers are exact-match. The identity provider only knows the redirect URIs someone registered, and a stack on a random port isn't one of them. Even with an identity provider in the stack, the browser reaches it as localhost:8080 while the API reaches it as idp:8080. The token's issuer says one and the API expects the other, and a perfectly valid token gets rejected.
Credentials are the thing you least want to hand an agent. A real test account on your real identity provider is a real credential. Put it in the agent's environment and it can end up in a prompt, a log or a transcript.
Don't switch authentication off
The tempting fix, and the one an agent will reach for on its own if you let it, is a flag that skips authentication in test mode:
def current_user(request) -> User:
if settings.AUTH_DISABLED: # green tests, untested product
return User(id="test", tenant="test", roles=["admin"])
return verify_token(request.headers["Authorization"])Don't. It removes exactly the code paths that break most often in production: signature and expiry checks, audience validation, role mapping, tenant scoping. Every test now runs as an all-powerful admin in a single tenant, so the suite goes green against a system that doesn't exist anywhere else. The one kind of bug it was best placed to catch, a buyer seeing another company's orders, is now guaranteed to get through. And a flag like that has a way of shipping.
What works instead
Put an identity provider in the stack. Run a real OIDC server with fake users: mock-oauth2-server,[7] Keycloak or Dex. The app goes through the same redirect, receives a signed token, and validates the signature, issuer, audience and expiry exactly as it does in production. Only the users are fake. Give the identity provider one hostname that both the browser and the services resolve, and the issuer problem goes away.
Seed a user per role. A buyer with a €10,000 approval limit, a finance approver, and a buyer from a different tenant. That last one is the user who proves isolation, and it only exists if authentication is real.
Log in once per role, then reuse the session. One test covers the login journey itself through the UI. Every other test starts already signed in, from a saved browser state.[8] That's faster, and it takes the most fragile steps out of every journey that isn't about logging in.
from playwright.sync_api import expect
ROLES = ["buyer", "approver", "other_tenant_buyer"]
@pytest.fixture(scope="session")
def auth_state(stack, browser, tmp_path_factory):
states = {}
for role in ROLES:
context = browser.new_context()
page = context.new_page()
page.goto(stack.web_url)
page.get_by_role("button", name="Sign in").click() # redirects to the idp in the stack
idp_login(page, stack.users[role]) # fills the idp's own form
page.wait_for_url("**/orders")
states[role] = tmp_path_factory.mktemp("auth") / f"{role}.json"
context.storage_state(path=states[role])
context.close()
return states
@pytest.fixture
def buyer_page(browser, auth_state):
context = browser.new_context(storage_state=auth_state["buyer"])
yield context.new_page()
context.close()
def test_buyer_cannot_open_another_tenants_order(buyer_page, stack):
buyer_page.goto(f"{stack.web_url}/orders/{stack.orders['other_tenant']}")
expect(buyer_page.get_by_text("Order not found")).to_be_visible()Test MFA on purpose, in one place. In the hermetic stack, most test users don't have a second factor. One dedicated test enrols a TOTP factor with a secret it generates itself and computes the codes with pyotp, so the MFA path is covered without anyone's phone.
Use client credentials for API journeys. Not every end-to-end test needs a browser. For API-level journeys, a service client in the test identity provider gets a token through the client-credentials grant, scoped to the test tenant. No login form involved.
Treat test credentials as credentials. Short-lived tokens, narrow scopes, the test tenant only, injected into the environment at run time and never committed or pasted into a prompt. The OAuth security best practice is explicit about restricting tokens to the audience and scope they need,[9] and that applies to the agent's tokens as much as anyone's. The simplest rule: nothing the agent holds should work against production.
Keep the real identity provider in a separate, scheduled suite. You still need to know that the real Entra ID or Okta integration works. Cover it with a small suite against a dedicated test tenant, run by CI on a schedule with credentials the agent never sees, and read by a person when it goes red. That's the one part of the stack that stays out of the agent's loop.
Putting the layers together
| Loop | When it runs | What runs | Time | Who waits |
|---|---|---|---|---|
| Inner | Every edit | Affected unit tests, in parallel | Seconds | The agent |
| Middle | Before a commit | Full unit suite and integration tests against real dependencies in containers | A minute or two | The agent |
| Outer | Before handing back a pull request | End-to-end journeys on a hermetic stack with an identity provider inside it | Minutes | The agent |
| Scheduled | Nightly | Real identity provider and partner sandboxes | Longer | CI, then a person |
The pyramid still describes how many tests of each kind you write. Unit tests remain the cheapest to write and the most precise, so you still want a lot of them. What has changed is how often the top of the pyramid runs. When an agent can run the full journeys before every pull request, end-to-end testing stops being a nightly ritual owned by one team and becomes the gate every change passes through.
A practical starting point:
- Make the unit suite fast enough that the agent never skips it: parallel by default, and only what changed in the inner loop.
- Gate on leaked state:
--fail-on-leakfor threads and files, and a rolled-back transaction per test for data. - Reproduce before you fix: replay CI-only failures, audit parallel-only ones, and bisect a red test before deciding whether the change or the test is wrong.
- Put integration tests on every seam, with the real database and queue in containers.
- Make the whole stack start with one command, hermetic, healthy and seeded.
- Put an identity provider in the stack, and never switch authentication off.
- Log in once per role and reuse the session everywhere else.
- Hand the agent logs, traces and screenshots on every failure.
- Keep the real identity provider and partner sandboxes in a small scheduled suite, with credentials the agent never holds.
What's next in this series
Next up is mutation testing. Fast unit tests and honest end-to-end tests only help if they would fail when the code is wrong, and once an agent writes most of them, nobody is reading every assertion. The third post, coming soon, introduces fermut, our open-source mutation tester built to run inside the agent's loop.
Later in the series, two threads from this post get their own posts. The first is the seams you can't bring up in a container, the customer's ERP, the carrier and the bank, and how contract tests stand in for them without drifting from reality. The second is what happens to flaky tests when an agent runs your end-to-end suite a hundred times a day: every flake becomes a false lead the agent will try to fix, so flakiness stops being an annoyance and starts costing real work.
If you've solved the login problem differently, tell us. We'd like to hear how.
The testing series
- Testing 101 for the Age of Coding Agents
- Testing 102: Fast unit tests, honest end-to-end tests (this post)
- Testing 103: Mutation testing with fermut (coming soon)
References
- M. Fowler, "TestPyramid," 2012. martinfowler.com/bliki/TestPyramid.html
- K. C. Dodds, "Write tests. Not too many. Mostly integration.," 2019. kentcdodds.com/blog/write-tests
- J. Micco, "Flaky Tests at Google and How We Mitigate Them," Google Testing Blog, 2016. testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
- Q. Luo, F. Hariri, L. Eloussi, D. Marinov, "An Empirical Analysis of Flaky Tests," FSE 2014. doi.org/10.1145/2635868.2635920
- J. Graves, "Larger Testing," in T. Winters, T. Manshreck, H. Wright (eds.), Software Engineering at Google, O'Reilly, 2020, ch. 14. abseil.io/resources/swe-book/html/ch14.html
- Docker, "Control startup and shutdown order in Compose." docs.docker.com/compose/how-tos/startup-order
- NAV, "mock-oauth2-server." github.com/navikt/mock-oauth2-server
- Playwright, "Authentication." playwright.dev/python/docs/auth
- T. Lodderstedt, J. Bradley, A. Labunets, D. Fett, "Best Current Practice for OAuth 2.0 Security," RFC 9700, 2025. www.rfc-editor.org/rfc/rfc9700



