Mutation testing a localisation service
A few days ago I asked why one of our test runs took three hours. The answer was mutation testing. I knew the name and had a rough idea of what it was. I had never watched it run on something I own. This post is what I learned, written down while it is still fresh.
The product#
At RODMENA we are building L10n, a localisation service. It holds translation strings and language tags, works out plural rules for every locale, counts words in a message, and controls who can see and change which project. On paper it is a boring piece of software. Nobody gets hurt if a word count is off by one.
That is what I thought too. Then I looked at what sits underneath it. There are tenants who must never see each other’s data. There are permission grants. There is an audit log that has to show tampering if someone edits a row behind our back. Translation text arrives from the outside world, so anything it can do to the server it will do. The word counting is the friendly face of the product. The rest is the kind of code that ends up in a breach report.
What mutation testing does#
The idea is old. It comes from the 1970s, and there are tools for most languages: PIT for Java, Stryker for JavaScript, mutmut for Python, which is what we use.
The tool makes a small change to your code. It might turn < into <=, flip True to False, change + 1 to - 1, or delete a line. That changed copy is called a mutant. Then it runs your tests against the mutant.
If a test fails, good. The tests noticed something was wrong, and we say the mutant was killed.
If every test still passes, the mutant survived. Your code is now broken in a small and specific way and your test suite has nothing to say about it. Someone has to look at each survivor and decide. Either a test is missing, or the change makes no difference that anyone could ever observe.
The tool repeats this hundreds or thousands of times, one mutant after another. That is why it is slow. Our word counting module produced 817 mutants on its last run. Some of the functions it touches are used almost everywhere, so each of those mutants re-runs a large part of the test suite. Three and a half hours on three workers.
It costs machine time and nothing else. The tool is an ordinary program running on the laptop. The expensive part is the human, or in our case the AI agent, who has to go through the survivors afterwards.
Coverage is a different question#
Most teams measure test coverage, and I did for years. Coverage tells you which lines your tests executed. It does not tell you whether the tests would notice if those lines were wrong.
You can have 100% coverage with tests that assert almost nothing. Mutation testing asks the question I actually care about: if this line were wrong, would anyone find out before a customer did?
What it found for us#
I expected it to find nothing interesting. It found plenty, and some of it was embarrassing.
Missing tests in the permissions code. A batch of survivors in the grant and revocation logic were real gaps. Nothing was broken today. Nothing would have warned us when it broke tomorrow.
A test that could never fail. We use deliberate faults to prove that a test goes red when it should. One of those faults replaced a function with another one. A refactor a few days earlier had made those two the same function, so the fault was swapping a thing for itself, and the test had been passing quietly ever since. We now have a tool that applies every deliberate fault in a fresh process and checks that something actually changes.
A score that was too good. One flaky test was killing mutants by accident. It failed for reasons unrelated to the mutant, so mutmut counted the mutant as caught. When we replayed those mutants one at a time, a large share of them survived. The kill rate we had been proud of was partly luck.
A trap in our own equivalence checks. Some mutants only change the case of an SQL keyword. To prove that changes nothing, the team compared PostgreSQL query plans before and after. They ran EXPLAIN with NULL parameters. PostgreSQL then folds column = NULL to a constant false and drops the rest of the WHERE clause, so a mutant that really did change the query could still show an identical plan. The fix is EXPLAIN (GENERIC_PLAN) together with one deliberate control that must show a different plan. I had accepted the weaker evidence three times before somebody noticed.
None of these were customer-facing bugs on the day we found them. Every one of them was a place where our tests, or our evidence, said something that was not true.
Where it is worth it#
After three hours I wanted a rule, so here is the one I settled on.
Mutation testing earns its cost on code where a silent mistake is expensive and there is nothing else watching it. For us that means tenancy, permissions, logins, encryption and the audit log. A gap there is a leak or a hidden break-in, and a normal test suite can look perfectly healthy while it happens.
It earns much less on code that already has an independent judge. Our word counting is compared with the official Unicode test data, with ICU, and with a frozen copy of the previous implementation. The last rewrite was checked against the old code on 150,000 random texts with zero differences. When you have that kind of oracle, a mutation run on top tells you very little you did not already know, and it still costs three hours.
So we keep it for the security core and for new API surfaces, and we run it once per phase on the code that changed since the last phase. We skip it where a reference implementation already does the checking.
What I took away#
I came into this thinking a localisation service did not need bank-grade testing. I was wrong about where the risk was. The translation strings are harmless. The plumbing around them is not.
Mutation testing has not found a dramatic bug for us. What it has done is show me where our tests were lying. Finding that out from a tool on my laptop is a lot cheaper than finding it out from an incident report.
If you have never tried it, pick the one module you would least like to see in the news and run a mutation tool against that module. Look at what survives. It is usually educational.