Writing · Testing · 11 min read
A test that always fails isn't flaky
Failure rate ranks the wrong tests. What flakestat measures instead, why its verdicts wait for evidence, and the times it fooled itself before it could fool anyone else.
A flaky test passes and fails on the same code. Everyone has met one, and everyone has the same ritual for it: rerun the job, watch it go green, merge. The ritual works, in the sense that the build goes green. What it doesn't do is tell you anything. A rerun proves that flakiness exists. It doesn't say how much, or in which tests, or whether the one you just retried is covering for a real bug.
I built flakestat to answer those questions from data a CI system already produces: the JUnit XML that pytest, Jest, go test and almost every other runner can write. It is one static Go binary with no dependencies and no account, and test results never leave your machine. This post is about the decisions inside it, because nearly all of them turned out to be about the same thing: not fooling yourself.
1Failure rate ranks the wrong tests
The obvious way to find flaky tests is to sort by how often they fail. It's also wrong, in a way that wastes exactly the people you're trying to help. A test that fails every single run isn't flaky. It's broken, and it is completely deterministic about it. Sort by failure rate and it goes straight to the top of the list, and whoever is hunting nondeterminism spends their morning on a test that needs a different kind of fix.
Flakiness is inconsistency. So flakestat doesn't count failures; it counts changes. Walk a test's results in order and mark every place where the outcome flips, pass to fail or fail to pass. The share of neighbouring pairs that disagree is the flip rate.
top of the list: test_export_csv
Look at what each ranking does to test_export_csv. By failure rate it's the worst test in the suite. By flip rate it scores zero, because it never changes its mind, and flakestat files it separately as consistently-failing, since hiding a broken test would be worse than the flake. test_ws_reconnect passed for a while and then failed for a while, and it scores low too: a single change of state looks much more like a regression than like chance.
This isn't a hypothetical. Across the two test suites I validated flakestat on with known answers, nine tests failed a hundred times out of a hundred. A detector that ranks by failure rate puts all nine at the top of its flaky list. flakestat scored them 0.00 and filed them as broken.
2Not every disagreement is worth the same
In a burst, where flakestat runs your suite twenty times in a row on one checkout, any disagreement is flakiness by definition, because nothing changed except chance. History is murkier. A test that failed on Tuesday's commit and passed on Wednesday's might be flaky, or it might be a bug that someone fixed on Wednesday.
So every observation carries its commit, and a flip between two different commits counts as weaker evidence: a third of a same-commit flip, by default. Recency matters too, so a test that was flaky last month and has since been fixed fades back towards stable. Put together, the score is a weighted average of transitions:
t[i] = 1 if outcome[i] != outcome[i + 1] else 0 # did it change its mind?
e[i] = 1 if commit[i] == commit[i + 1] else 1 / 3 # across commits it might be a fix
score = sum(r[i] * e[i] * t[i]) / sum(r[i]) # r: recency, aged in commits
That last comment, "aged in commits", hides the first time the tool fooled me. The recency weights originally decayed per observation. Their total could never exceed about 3.3 transitions' worth, so a hundred runs bought almost exactly as much certainty as ten: on a test failing a quarter of the time, the spread of scores across repeated trials was 0.246 at a hundred runs and 0.255 at ten. Ninety extra runs, no extra information. Ageing by commit fixed it, because a burst shares one commit and gets to use all of its evidence.
spread, aged per run –spread, aged per commit –
3No verdict without evidence
The second mistake was subtler. Classifying on the score itself meant that a test failing 2% of the time was called flaky in a quarter of trials at ten runs, and in none at twenty. The verdict got better with less data, which is the opposite of how evidence is supposed to work.
The fix is to classify on how low the score could plausibly be, rather than on the score. Below five runs a test isn't scored at all. Above that, flakestat computes a Wilson lower bound on the score, and the bound, not the estimate, has to clear the threshold. A small sample can't reach high confidence however flaky it looks.
runs 0 score – lower bound – verdict insufficient data
Set it to fail rarely and watch the verdict hold back while the interval is wide. The thresholds aren't arbitrary either. Stable is below 0.05 and flaky is above 0.10, and the 0.10 was measured, not chosen: at a hundred runs, a threshold of 0.15 caught a test that fails 10% of the time in 45% of trials, while 0.10 catches it in 80%, and still flags stable and broken tests in none.
4How long to hunt
flakestat looks for flakes in two ways, and they answer different questions. A hunt runs your suite N times right now, on one checkout, so any disagreement is flakiness by definition. History ingests the JUnit reports CI already writes and scores them over weeks, which is the only way to catch a test that only fails on one runner image, or under one night's load. A hunt is the strong signal for "is this test flaky?" and "did my fix work?"; history is the one that finds what nobody thought to hunt for.
The obvious question about a hunt is how many runs it needs. A test that fails with probability p shows at least one disagreement in N runs unless every run agrees: the chance of a hunt seeing anything is 1 − (1 − p)N − pN. That number is unforgiving for rare flakes. A test that fails 2% of the time slips through a twenty-run hunt about two times in three. And seeing a disagreement is only the start: a verdict needs enough of them to lift the Wilson bound past 0.10.
runs hover the chartfails 2% –fails 10% –
5Only compare what's comparable
One rule has now been broken three times, the same way, at three different levels of the code:
Two outcomes are evidence of nondeterminism only if everything that could legitimately change the result was held constant.
The first time, it was branches. A test passing on main and failing on an unfinished feature branch, interleaved in time, read as constant disagreement. The second time it was platforms, and the figure below is that case: a test that always fails on Windows and always passes everywhere else scored 0.45, and the explanation printed beneath it claimed direct evidence of nondeterminism. The test is perfectly deterministic on every platform.
flips –score –
The third time was the display layer, which kept grouping by branch after scoring had moved on. The verdict was right, but the evidence printed under it was wrong. That is arguably worse than being wrong outright, because the evidence is what a sceptical reader checks. Now there is exactly one definition of an execution context (branch, operating system, architecture, runtime and its version), and everything that shows transition evidence calls it. Duplicating the rule is what let it drift.
6The apparatus is part of the experiment
Two more failures taught me to distrust the measuring instrument as much as the thing it measures.
Running a suite with --parallel puts several copies of it in one working directory, and any suite that uses a fixed port, path or database collides with itself. Against spf13/cobra this produced a confident flaky verdict, a score of 0.60, on a test that passed eight times out of eight when run alone. At the time, the README's headline example used --parallel 4. Now anything found in parallel is re-run sequentially and demoted to suspect if it doesn't reproduce. Demoted, not deleted: five clean runs are weak evidence of innocence, not proof.
Then there's ingesting the same CI report twice. You'd expect duplicates to inflate confidence, with more agreeing observations and a tighter bound. They do the reverse. A copy sorts right next to its original and always agrees with it, so duplication injects fake agreements:
observations 12flips –flip rate –
In the case recorded in flakestat's design notes, a test failing four runs in twelve dropped from a score of 0.64 to 0.30 when its history was ingested twice, and its lower bound from 0.51 to 0.23. Duplication isn't a bookkeeping nuisance. It's a way of losing flaky tests quietly, with more apparent support for the wrong answer. Deduplication now runs when reading as well as writing, keyed on what was executed rather than when it was ingested. When in doubt it keeps the record, because missing a duplicate costs a little accuracy while a false dedup destroys evidence and says nothing.
7JUnit XML is not a standard
The whole tool rests on reading reports that every ecosystem can produce, and that's both its biggest advantage and its most fragile part. JUnit XML is a convention, not a standard: there's no official schema, and every runner speaks its own dialect.
| Ecosystem | How it writes JUnit XML |
|---|---|
| Python | pytest --junitxml=out.xml |
| JavaScript, TypeScript | jest-junit, vitest --reporter=junit |
| Go | gotestsum --junitfile out.xml |
| Rust | cargo nextest run --profile ci |
| Java | Surefire and Gradle, natively |
| Ruby, PHP, .NET | rspec_junit_formatter, phpunit --log-junit, dotnet test --logger junit |
The parser is deliberately tolerant: a root wrapper that's there or isn't, nested suites, a class name that's missing or copied into the name, missing timings, and invalid characters inside captured stack traces. Surefire's own rerun markers, flakyFailure and flakyError, are read as the flakiness signals they already are.
A score also needs a stable name for the thing it scores. A test's id is a hash of its suite, class name and name, with the readable parts kept alongside for display. Parameterised tests get two ids: the exact one, test_upload[s3], and a normalised one with the parameters stripped. That lets flakestat say "this one case is flaky" and "this whole family is flaky", which are different bugs with different fixes.
8From a verdict to a decision
A list of flaky tests is only useful if it changes what happens next, so the last two commands are about action. flakestat check is a CI gate with exit codes that mean something: zero when nothing changed, one when a new test crossed the flaky threshold, two when an existing flake got worse. It also writes GitHub Actions annotations and a markdown summary for the pull request.
flakestat quarantine emits the skip list in the format each runner actually accepts, because a generic list nobody can consume is the most common way this kind of tool fails: pytest node ids for --deselect, a regular expression for go test -skip, Jest's ignore patterns, or a plain YAML file for anything else. Quarantine is a holding pen, not a fix: the tests in it can still be hunted and tracked, and when they recover they come out.
9Don't take my word for it
The validation I trust most isn't a test case I wrote. ConduitIO's contributors reported a set of flaky tests in their own issue tracker, named them, and later fixed them. flakestat was run with an identical protocol, fixed in advance, on either side of their fix. Before it, three of the four tests in scope scored flaky at 0.67 with high confidence. At the fix itself, all four were stable at 0.00. I didn't create the bug, label it or repair it. flakestat was handed observations from both sides of a commit it had no part in, and separated them.
The protocol I'd written down first, one test repetition per process, found nothing at all in a hundred executions. That's because these were process-global-state bugs: the pollution has to happen inside one process before a later test trips over it, which Conduit's own flake-hunting job accounts for and my first protocol didn't. That null result is published next to the one that worked, as is a test run I discarded because its fixture was contaminated. Reporting only the arm that succeeded would have been exactly the kind of adjustment the whole exercise exists to police.
I'm also clear about what it hasn't shown. The default weight for cross-commit evidence is a defensible guess rather than a finding. Below a failure rate of about 5%, sampling stops being the right instrument. And flakestat has not yet found a flaky test that its maintainers didn't already know about. Those are open problems, written down where anyone can check them, which is the point.
10Try it
brew install rowhitswami/tap/flakestat # or npm i -D flakestat, or pip install flakestat
flakestat hunt --runs 20 --junit 'reports/junit-{run}.xml' \
-- pytest --junitxml='{junit}'
Run your suite a few times and see what disagrees. Then ask flakestat explain why a test was flagged. It will show you the evidence, and it will say so when there isn't enough of it.