Testing 101 for the Age of Coding Agents

This is the first post in a series on software testing. Before going deep on any one technique, it helps to have the whole map in front of you.
"Do you have tests?" is the wrong question. Almost every team does. The better question is which questions your tests actually answer. A suite made entirely of one kind of test can be large, green and fast, and still leave a whole category of bug completely unguarded. Dijkstra put the limit plainly in 1970: testing can "show the presence of bugs, but never to show their absence".[1] The practical answer is to point different kinds of tests at different kinds of bugs.
Each kind of test answers one question. Unit tests ask whether a piece of logic is right. Integration tests ask whether the pieces fit together. End-to-end tests ask whether a real user can get the job done. Property-based tests ask whether a rule holds for every input, not just the three you thought of. Mutation tests ask whether your tests would notice if the code were wrong. Behavioral tests ask whether the system does what the business actually asked for.
To keep this concrete, every example below uses the same small system: a purchase-to-delivery service. A buyer raises a purchase order, it gets priced, stock gets reserved, someone approves it, and the goods ship on pallets. It's small enough to hold in your head and realistic enough to have real bugs.
Unit tests: is this piece of logic right?
A unit test exercises one small piece of code, usually a single function or class, in isolation. No database, no network, no clock. It runs in milliseconds, which is why you can have thousands of them and run them on every save. Almost every unit testing framework today, from JUnit to pytest's assertion style, descends from Kent Beck's SUnit for Smalltalk.
The system: the pricing rule. Orders of 100 units or more get 5% off.
from decimal import Decimal
def line_total(qty: int, unit_price: Decimal) -> Decimal:
total = qty * unit_price
if qty >= 100:
total *= Decimal("0.95")
return total.quantize(Decimal("0.01"))
def test_bulk_orders_get_five_percent_off():
assert line_total(200, Decimal("10.00")) == Decimal("1900.00")
def test_small_orders_pay_list_price():
assert line_total(10, Decimal("10.00")) == Decimal("100.00")What it catches: arithmetic mistakes, wrong branches, rounding errors, regressions when someone refactors the pricing code.
What it misses: everything between the units. line_total can be perfect while the order service calls it with the quantity in boxes instead of units. Keep that in mind; we'll come back to this exact function in the mutation testing section, because these two tests have a gap in them.
Integration tests: do the pieces fit together?
An integration test puts two or more real components together and checks the seam between them: your code and a real database, your service and a message queue, your connector and a (sandboxed) ERP API. They're slower than unit tests, but they catch the bugs that live in the assumptions each side makes about the other.
The system: stock reservation against a real Postgres. The rule is that a reservation either fully succeeds or leaves stock untouched. Mocking the database here would test your mock, not your transaction handling, so we start a real one in a container.
import pytest
from testcontainers.postgres import PostgresContainer
@pytest.fixture(scope="session")
def db():
with PostgresContainer("postgres:16") as pg:
engine = create_engine(pg.get_connection_url())
run_migrations(engine)
yield engine
def test_failed_reservation_leaves_stock_untouched(db):
repo = InventoryRepo(db)
repo.add_stock("SKU-42", on_hand=5)
with pytest.raises(InsufficientStock):
repo.reserve("SKU-42", qty=8)
assert repo.available("SKU-42") == 5 # nothing half-writtenWhat it catches: broken SQL, missing migrations, transactions that don't roll back, serialization mismatches, a client library that behaves differently from its documentation.
What it misses: whether the whole flow works for a user. And they cost more to run, so you want fewer of them, placed exactly on the seams that matter.
End-to-end tests: can a real user get the job done?
An end-to-end (E2E) test drives the whole deployed system the way a user would: through the browser or the public API, with every service running. It's the closest thing to a real person clicking through the product.
The system: the buyer journey. Log in, raise a purchase order, submit it for approval, and see the right status.
from playwright.sync_api import expect
def test_buyer_can_raise_a_purchase_order(page, staging):
page.goto(staging.url)
page.get_by_label("Email").fill(staging.buyer.email)
page.get_by_label("Password").fill(staging.buyer.password)
page.get_by_role("button", name="Sign in").click()
page.get_by_role("link", name="New purchase order").click()
page.get_by_label("Supplier").select_option("Acme Fasteners")
page.get_by_label("SKU").fill("SKU-42")
page.get_by_label("Quantity").fill("250")
page.get_by_role("button", name="Submit for approval").click()
expect(page.get_by_text("Awaiting approval")).to_be_visible()What it catches: the things nobody owns. A frontend that sends a field the backend renamed last sprint, a missing environment variable, an auth redirect loop, a queue consumer that was never deployed.
What it misses: precision. When an E2E test fails, it tells you something is broken, not what. They're also the slowest and the flakiest tests you'll own; the classic study of flaky tests found the top three causes were async waits, concurrency and test-order dependencies, exactly the things a full-stack test is full of.[2] Keep a handful that cover the journeys your business can't live without, and push everything else down to cheaper layers.
Property-based tests: does the rule hold for every input?
The tests above are example-based: you pick an input and assert an output. The weakness is that you pick inputs you already thought of, and bugs live in the ones you didn't. A property-based test flips it around. You describe a rule that must always hold, and the framework generates hundreds of inputs trying to break it. The idea was popularised by QuickCheck for Haskell,[3] and its roots go back to fuzzing, where feeding random input to standard UNIX utilities crashed 25-33% of them on every version tested.[4] When it finds a failure, it shrinks it down to the smallest input that still fails.
The system: splitting a shipment into pallets. Whatever the quantity and pallet size, no unit may be lost or invented, and no pallet may be empty or overfilled.
from hypothesis import given, strategies as st
@given(
qty=st.integers(min_value=0, max_value=100_000),
pallet_size=st.integers(min_value=1, max_value=500),
)
def test_splitting_never_loses_or_invents_units(qty, pallet_size):
pallets = split_into_pallets(qty, pallet_size)
assert sum(pallets) == qty
assert all(0 < p <= pallet_size for p in pallets)Tests like this tend to find the edges almost immediately: a quantity of zero that produces one empty pallet, a quantity that's an exact multiple of the pallet size that produces an extra one, an off-by-one in the remainder. Nobody writes those example tests by hand until after the incident.
Good properties are usually one of a few shapes: conservation (nothing is lost, like above), round-trips (parse(render(order)) == order for an EDI message), invariants (stock on hand is never negative), or equivalence (the new fast allocator gives the same answer as the old slow one).
What it misses: anything you can't phrase as a rule. It's also only as good as the input generator; if your strategy never produces a negative quantity, you'll never learn what happens with one.
Mutation tests: would your tests notice if the code were wrong?
Every technique so far tests the code. Mutation testing tests the tests. A tool makes small, deliberate changes to your source code (flip >= to >, change + to -, replace a return value with None), and reruns your suite against each "mutant." If a test fails, the mutant is killed: good, your tests noticed. If everything stays green, the mutant survived: your tests would not have caught that bug. The idea dates to DeMillo, Lipton and Sayward in 1978,[5] and has a large research literature behind it.[6]
The system: back to line_total. Run a mutation tool such as mutmut against it and one mutant survives:
- if qty >= 100:
+ if qty > 100:Both of our unit tests still pass. We tested 200 units and 10 units, and never tested 100, the exact number the business rule is about. Line coverage said 100%. Mutation testing says the boundary was never checked. That gap isn't an anecdote: across 31,000 test suites for five large Java systems, coverage was only low-to-moderately correlated with how many mutants a suite killed once suite size was controlled for,[7] and killing mutants does track finding real faults.[8] The fix is one test:
def test_discount_starts_at_exactly_one_hundred_units():
assert line_total(100, Decimal("10.00")) == Decimal("950.00")What it catches: weak assertions, missing boundary cases, tests that execute code without checking its result. It's the most honest answer to "are my tests any good?"
What it misses: it can't tell you about behavior you never wrote. And it's expensive: every mutant is a test run, so you scope it to the code that matters or to what changed in a pull request.
Behavioral tests: does it do what the business asked for?
Behavioral tests describe what the system should do in the language of the people who asked for it, not in terms of functions and classes. The best-known form is behavior-driven development (BDD), introduced by Dan North:[9] scenarios written as Given / When / Then that a buyer, a finance lead and an engineer can all read and agree on before any code is written.
The system: the approval policy.
Feature: Purchase order approval
Scenario: Orders above the buyer's limit need a second approver
Given a buyer with an approval limit of €10,000
When they submit a purchase order for €12,500
Then the order status is "Awaiting approval"
And the finance lead is notified
Scenario: Orders within the limit are approved immediately
Given a buyer with an approval limit of €10,000
When they submit a purchase order for €4,000
Then the order status is "Approved"Tools like pytest-bdd, behave or Cucumber bind each line to a step function, so the specification is the test. When the policy changes, the conversation happens on this file, not in a ticket that drifts away from the code.
There's a second meaning of behavioral testing that matters more every month: testing systems whose output isn't deterministic, like language models and agents. You can't assert on an exact string, so you assert on behavior instead. The CheckList approach from NLP research is a good starting vocabulary:[10] invariance tests (renaming the supplier in an email must not change the delivery date an agent extracts), directional tests (a later confirmed date must never move the order earlier), and minimum functionality tests (a plain "we can ship on the 14th" must always be read correctly). For anyone building agents, this is where most of the testing effort ends up.
What it misses: behavioral tests are only as good as the scenarios someone thought to write. They sit on top of the other layers; they don't replace them.
The whole map on one page
| Type | Question it answers | Example in our system | Speed | Typical tools |
|---|---|---|---|---|
| Unit | Is this logic right? | Bulk discount pricing | Milliseconds | pytest, JUnit, Jest |
| Integration | Do the pieces fit? | Stock reservation in Postgres | Seconds | pytest + Testcontainers |
| End-to-end | Can a user get the job done? | Buyer raises a PO | Seconds to minutes | Playwright, Cypress |
| Property-based | Does the rule hold for every input? | Pallet splitting conserves units | Seconds | Hypothesis, fast-check |
| Mutation | Would the tests notice a bug? | The >= boundary in pricing | Minutes to hours | mutmut, PIT, Stryker |
| Behavioral | Does it do what the business asked? | Approval limits | Varies | pytest-bdd, behave, Cucumber |
How they fit together
These aren't competing philosophies; they're layers, and each one covers for the blind spot of another. The classic advice is the test pyramid, which Martin Fowler credits to Mike Cohn:[11] many unit tests at the base, fewer integration tests in the middle, a small number of end-to-end tests at the top. That shape still holds for a good reason: cost and precision. The cheaper and more precise a test is, the more of them you can afford.
Property-based, mutation and behavioral testing cut across the pyramid rather than sitting on it. Property-based testing makes any layer search a wider space of inputs. Mutation testing grades the layers you already have. Behavioral tests make sure the whole thing is pointed at the right requirements in the first place.
A practical starting point for most teams:
- Unit-test the rules that carry money, quantities and dates.
- Integration-test every seam where you've been burned before: the database, the queue, the external API.
- Keep a short list of end-to-end journeys that must never break, and treat a flaky one as a bug, not as noise.
- Add a property wherever you catch yourself writing the fourth example test for the same function.
- Run mutation testing on changed code in pull requests, so test quality is checked where it's cheapest to fix.
- Write behavioral scenarios for any policy a non-engineer owns.
And make the suite fast enough that people actually run it. A suite nobody runs locally protects nobody. (That's why we open-sourced rstest, a parallel-by-default pytest runner.)
Why so many kinds of tests when an agent writes the code?
Fair question. If a coding agent can write a feature in minutes, and write the tests for it too, why keep six different kinds? Because agents change who writes the code, not what can go wrong with it. If anything, they make the testing portfolio matter more, for four reasons.
Code got cheap. Verification didn't. An agent produces more changes than a person can read carefully, so review attention becomes the scarce resource, and tests are the part of review that scales. Code that looks right is exactly what language models are best at producing. When researchers prompted GitHub Copilot with security-relevant scenarios, roughly 40% of the 1,689 programs it wrote were vulnerable.[12] Plausible isn't the same as correct.
Tests are the agent's feedback loop. An agent works by writing code, running the tests, reading the failure and trying again. The benchmarks measure agents the same way: HumanEval scores generated code against hidden unit tests,[13] and SWE-bench counts an issue as resolved only when the repository's tests pass.[14] So an agent is only as good as the tests it's pointed at. When EvalPlus extended HumanEval's tests 80-fold, it caught wrong code the original tests had accepted, cut reported pass rates by up to 28.9%, and even changed which models ranked higher.[15] If tests curated by researchers let wrong code through, so do yours.
Green is a target, and targets get gamed. Tell an agent to make the suite pass and it will look for the shortest path: special-case the exact input the test uses, loosen an assertion, skip the awkward test. That isn't malice, it's optimisation. The different kinds of test resist it in different ways. A property can't be satisfied by hard-coding one example. Mutation testing notices when assertions stop checking anything. A behavioral scenario owned by the business can't be quietly rewritten to match the code without someone seeing it in review.
Same author, same blind spot. When one agent writes the implementation and the tests from the same prompt, both inherit the same misunderstanding. Tests generated by reading the code mostly confirm what the code already does. Deciding what the correct output should be is the oracle problem, which testing research has long treated as the hard part.[16] The fix is independent oracles: properties that come from domain rules, scenarios that come from the people who own the policy, and integration and end-to-end tests that hit real systems the agent never saw in its context window.
The good news is that agents make the expensive techniques cheap. Writing input strategies, killing surviving mutants one by one, turning a policy document into scenarios: tedious for people, cheap for an agent. The trick is to trust the output only after it clears a filter. Meta's TestGen-LLM kept a generated test only if it built, passed reliably and increased coverage; in one evaluation, 75% of candidates built, 57% passed reliably and 25% increased coverage.[17] The filter did as much work as the model.
So the portfolio doesn't shrink; it shifts. People spend less time typing unit tests and more time on what an agent can't supply for itself: the rules worth stating as properties, the scenarios that capture what the business actually wants, and a mutation score that says whether the suite is really watching.
What's next in this series
This post is the map. Next up: how to improve your coding agents with this map in hand. We'll cover how to give an agent independent oracles instead of letting it grade its own homework, how to put mutation scores and property tests in the loop so the agent gets feedback it can't talk its way around, and how to structure prompts, context and review so the tests it writes catch real bugs instead of restating the code.
After that, the series goes deep on single techniques, with real code and real failures: how to find properties worth testing when your domain is messy, how to make mutation testing cheap enough to run on every pull request, where integration tests should stop and contract tests should start, and how we test agents whose outputs are never exactly the same twice.
If there's a testing question you'd like us to cover, tell us. The best posts in a series like this usually start as someone else's hard problem.
References
- E. W. Dijkstra, "Notes on Structured Programming," EWD249, 1970. www.cs.utexas.edu/~EWD/ewd02xx/EWD249.PDF
- Q. Luo, F. Hariri, L. Eloussi, D. Marinov, "An Empirical Analysis of Flaky Tests," FSE 2014. doi.org/10.1145/2635868.2635920
- K. Claessen, J. Hughes, "QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs," ICFP 2000. doi.org/10.1145/351240.351266
- B. P. Miller, L. Fredriksen, B. So, "An Empirical Study of the Reliability of UNIX Utilities," Communications of the ACM, 1990. doi.org/10.1145/96267.96279
- R. A. DeMillo, R. J. Lipton, F. G. Sayward, "Hints on Test Data Selection: Help for the Practicing Programmer," IEEE Computer, 1978. doi.org/10.1109/C-M.1978.218136
- Y. Jia, M. Harman, "An Analysis and Survey of the Development of Mutation Testing," IEEE Transactions on Software Engineering, 2011. doi.org/10.1109/TSE.2010.62
- L. Inozemtseva, R. Holmes, "Coverage Is Not Strongly Correlated with Test Suite Effectiveness," ICSE 2014. doi.org/10.1145/2568225.2568271
- R. Just, D. Jalali, L. Inozemtseva, M. D. Ernst, R. Holmes, G. Fraser, "Are Mutants a Valid Substitute for Real Faults in Software Testing?," FSE 2014. doi.org/10.1145/2635868.2635929
- D. North, "Introducing BDD," Better Software, 2006. dannorth.net/blog/introducing-bdd
- M. T. Ribeiro, T. Wu, C. Guestrin, S. Singh, "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList," ACL 2020. aclanthology.org/2020.acl-main.442
- M. Fowler, "TestPyramid," 2012. martinfowler.com/bliki/TestPyramid.html
- H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, R. Karri, "Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions," IEEE S&P 2022. arxiv.org/abs/2108.09293
- M. Chen et al., "Evaluating Large Language Models Trained on Code," 2021. arxiv.org/abs/2107.03374
- C. E. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?," ICLR 2024. arxiv.org/abs/2310.06770
- J. Liu, C. S. Xia, Y. Wang, L. Zhang, "Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation," NeurIPS 2023. arxiv.org/abs/2305.01210
- E. T. Barr, M. Harman, P. McMinn, M. Shahbaz, S. Yoo, "The Oracle Problem in Software Testing: A Survey," IEEE Transactions on Software Engineering, 2015. doi.org/10.1109/TSE.2014.2372785
- N. Alshahwan et al., "Automated Unit Test Improvement using Large Language Models at Meta," FSE 2024 (Industry). arxiv.org/abs/2402.09171



