Testing 103: Mutation testing with fermut

This is the third post in our testing series. The first one laid out the map, and the second put unit and end-to-end tests in the coding agent's loop. This one asks the question underneath both: would those tests notice if the code were wrong?
Code coverage answers a narrow question: which lines did my tests run? It says nothing about whether those tests would notice if the line were wrong. You can sit at 100% coverage with assertions that execute a line and never check its behavior.
Mutation testing answers the question you actually care about. It introduces small, deliberate breaks in your code (flip a <= to <, a + to -, a True to False), re-runs your suite, and watches what happens. A mutant your tests kill is a line your suite genuinely defends. A mutant that survives is a gap: the code changed and your tests shrugged.
It's the most honest answer to "are my tests any good?" And for years it's been too slow to run anywhere near a pull request.
Here is the example from the first post again, because it shows exactly what is going on. The rule: orders of 100 units or more get 5% off. Two unit tests cover it:
from decimal import Decimal
def line_total(qty: int, unit_price: Decimal) -> Decimal:
total = qty * unit_price
if qty >= 100:
total *= Decimal("0.95")
return total.quantize(Decimal("0.01"))
def test_bulk_orders_get_five_percent_off():
assert line_total(200, Decimal("10.00")) == Decimal("1900.00")
def test_small_orders_pay_list_price():
assert line_total(10, Decimal("10.00")) == Decimal("100.00")Both tests pass, and line coverage reads 100%: every line ran. Now run fermut on it:
$ fermut run src/ --tests tests/
SURVIVED src/pricing.py:5 [boundary-shift] `>=` → `>`
SURVIVED src/pricing.py:5 [number-shift] `100` → `101`
SURVIVED src/pricing.py:5 [number-shift] `100` → `99`Here is what happened. fermut made small edits to line_total, ran the tests that cover each edited line, and reported the edits nobody noticed. All three survivors sit on the same line, and they all move the discount threshold. Start the discount at 101 units, or at 99, and both tests stay green. The tests checked 200 units and 10 units, both far from the boundary, and never checked the numbers next to 100, which is the one quantity the rule is about.
The fix from the first post was one test at exactly 100 units. Add it and run again:
def test_discount_starts_at_exactly_one_hundred_units():
assert line_total(100, Decimal("10.00")) == Decimal("950.00")SURVIVED src/pricing.py:5 [number-shift] `100` → `99`Two of the three die. The third shows the other side of the boundary: if the discount started at 99 units, nothing would fail, because no test checks that 99 units pays list price. One more test closes it:
def test_ninety_nine_units_pay_list_price():
assert line_total(99, Decimal("10.00")) == Decimal("990.00")Now every mutant is killed. Coverage never moved: it was 100% before these two tests and 100% after. Only the mutation results told us the boundary was unchecked, and on which side.
Finding that on one function is easy. Finding it across a real codebase, on every change, is where mutation tools have always run out of time.
Today we're open-sourcing fermut: agent-first mutation testing for Python, with a mutator written in native Rust.
Why you need mutation testing, and why it's becoming non-negotiable
Mutation testing was always the right idea and a hard sell. It told you something coverage couldn't, but it cost too much to run, so most teams filed it under "nice in theory" and moved on. The economics that made it optional are changing, and agentic coding is the reason.
Start with the failure mode it catches. Coverage is a quantity metric, and quantity is easy to fake, accidentally or on purpose. A test that calls a function and never asserts on the result still counts as coverage. A test that re-implements the logic it's checking still counts as coverage. You can hit any number you like and learn nothing about whether your suite would catch a regression. Mutation score is a quality metric: it can only go up if the tests would actually fail when the behavior breaks. That's the difference between "the line ran" and "the line is defended."
Why chase a higher score? Because every surviving mutant is a bug you could ship without a single red test. In line_total, the 100 → 101 survivor means a customer ordering exactly 100 units gets charged full price, and the suite says nothing. This isn't only intuition. Across 31,000 test suites for five large Java systems, coverage was only low-to-moderately correlated with how many mutants a suite killed once suite size was controlled for,[1] so a high coverage number tells you little about how good a suite is. Killing mutants, on the other hand, does track finding real faults.[2] A rising mutation score is evidence that the suite catches more of the bugs that actually happen. A rising coverage number is evidence that more lines ran.
Now add agents to the picture. Two things change at once.
The volume of code outruns human review. When an agent writes a module and its tests in one turn, nobody is reading every assertion. You need an automated signal for whether those tests are real, not just a green checkmark and a coverage percentage that looks fine.
Coverage is exactly the wrong target to hand an agent. Point an agent at "raise coverage to 90%" and it will get there, often with tests that exercise lines without verifying them, because that's the cheapest way to satisfy the metric. The optimizer does what you measured, not what you meant. Mutation score is far harder to game: the only reliable way to kill a mutant is to write a test that genuinely detects the broken behavior. It's the kind of objective, hard-to-fake feedback signal an agent loop actually needs to improve against.
So the case for mutation testing is no longer "your tests might be weaker than they look." It's that agents are now generating most of the tests, coverage is the metric they'll happily satisfy without doing the work, and mutation score is the check that tells you, and the agent, whether the safety net has holes. The only thing standing in the way was speed. That's the part we set out to fix.
What makes it fast
Three things, in order:
- The mutator runs in Rust on ruff's parser, with no slow Python AST round-trips to generate mutants.
- Every mutant is then type-checked by
tyand dropped if it's type-invalid before it ever costs a pytest run. On type-hinted code that removes 20–40% of the naive mutants that older tools would have dutifully executed and counted. - Survivors run against only the tests whose coverage actually touched the mutated line, not the whole suite every time.
The test runner itself stays Python. Your suite is unchanged; fermut just stops wasting it on mutants that can't teach you anything.
How fermut is built
Every mutant goes through the same pipeline. Parse once, throw away as much as possible before paying for a test run, run only the tests that can see the change, and sort what comes back. The cache and the agent loop sit around that pipeline, so a second run only pays for what changed.
The coverage filter is the whole game
Without it, every mutant runs the full suite, which is exactly why the older tools don't finish on real code. With it, the numbers change shape. Measured on the same machine, fermut 0.5.0 vs mutmut 3.6.0:
| Repo | mutmut | cosmic-ray | fermut |
|---|---|---|---|
| typer (~2,200 mutants) | 298s / 70.5 | DNF (>15 min) | 164s / 71.5 |
| pyjwt | 45s / 75.4 | DNF | 42s / 82.2 |
| werkzeug (~15k mutants) | n/a | n/a | ~6 min |
(time / mutation score)
On typer, fermut is both faster and scores higher. cosmic-ray doesn't finish a real suite inside a 15-minute cap at all: process-per-mutant over the full suite with no coverage filter simply doesn't scale. Strip out the one-time coverage generation and typer's steady-state mutation work is about 74s versus mutmut's ~298s, roughly 4× faster. Whole-repo werkzeug, at 15,000 mutants, finishes in about six minutes.
A coding agent can now drive the loop
The point isn't just speed. It's that mutation testing stops being a quarterly audit and becomes something a coding agent runs turn by turn. fermut exposes the workflow as composable commands the agent calls directly:
fermut next → rank survivors by value, pick the one worth fixing
fermut explain → hint + source context + a killing-test skeleton
(the agent writes the test; it's already a model
with the repo in context)
fermut run → confirm the mutant dies and the suite stays green
fermut score → the reward delta vs the prior runThat's a closed loop with an objective reward at the end of it: pick a survivor, write a test, confirm the kill, read the score delta, repeat.
And it ships ready for that loop, not as a DIY integration you have to assemble:
- An MCP server (
fermut mcp) exposes those steps as native JSON-RPC tools, with no CLI scraping and no temp-filejqjuggling. Structured results, deterministic survivor IDs, JSON shapes treated as API. - A Claude Code skill in the box: a battle-tested playbook for running, triaging, and killing survivors, so the agent knows the workflow without you teaching it. Install it with
fermut install-skills, or add the Claude Code plugin with/plugin marketplace add KovantAI/fermut. - A structural (AST-hash) cache that survives short edits: typer goes from ~164s cold to about 8s warm, so "agent writes a test → re-measure" lands in single-digit seconds. In our agent-loop benchmark, fermut's re-measure cost on typer was about 56s versus mutmut's ~251s. That's the difference between a tool an agent calls every iteration and one it calls once a day.
For non-agent drivers, such as a CI job or a human at a terminal, fermut autofix runs generate-verify-keep against fermut's own model. Over MCP it's deliberately omitted: the agent's own model, already holding the repo in context, writes a better test for free.
Getting started
$ uv tool install fermut # prebuilt wheel from PyPI, no Rust needed
$ fermut init # detects source / tests / runner / coverage, writes fermut.toml
$ fermut run # first survivor report in under five minutesComing from another tool? fermut migrate {mutmut,cosmic-ray} translates the config and rewrites the ignore pragmas for you.
For pull requests, there is a GitHub Action. KovantAI/fermut@v0.5.0 mutates only the lines the PR changed, gates on a score you set with fail-under, and posts the report as a PR comment.
Honest scope
The data above is preliminary. fermut is at 0.5.0, still pre-1.0 and pre-optimization, so expect these numbers to move.
A few caveats worth stating plainly. On tiny pure-Python libraries with already-fast suites (more-itertools, for example), mutmut's simpler model still wins: the ty + coverage overhead doesn't pay for itself when running the full suite is already cheap. Mutation scores across tools aren't directly comparable either: different operator catalogues mean different denominators. And there's no LSP extension or bisect yet.
The thesis is narrower than "fermut is faster than everything": it's that the coverage filter, type-aware pruning, and a cache built for short edits are what let mutation testing keep up with an agent that's writing tests in real time. That's the workload we built it for.
Try it
It's open source now. Try it on a repo where you've never trusted your coverage number, and see what survives.
https://github.com/KovantAI/fermut
What's next in this series
Mutation testing tells you where a boundary went unchecked. It can't tell you which boundaries exist in the first place. That's the next post: property-based testing, and how to find properties worth stating when the domain is messy, like units that must be conserved when a shipment is split into pallets, EDI messages that must survive a round-trip, and stock that must never go negative. After that, we put mutation scores and properties together and look at how to stop an agent grading its own homework.
If mutation testing turns up something surprising in your suite, tell us. We'd like to hear what survived.
The testing series
- Testing 101 for the Age of Coding Agents
- Testing 102: Fast unit tests, honest end-to-end tests
- Testing 103: Mutation testing with fermut (this post)
References
- L. Inozemtseva, R. Holmes, "Coverage Is Not Strongly Correlated with Test Suite Effectiveness," ICSE 2014. doi.org/10.1145/2568225.2568271
- R. Just, D. Jalali, L. Inozemtseva, M. D. Ernst, R. Holmes, G. Fraser, "Are Mutants a Valid Substitute for Real Faults in Software Testing?," FSE 2014. doi.org/10.1145/2635868.2635929



