Rohit Swami
India Resume ↗

Writing · Testing · 9 min read

A test that always fails isn't flaky

Failure rate ranks the wrong tests. What flakestat measures instead, why its verdicts wait for evidence, and the times it fooled itself before it could fool anyone else.

A flaky test passes and fails on the same code. Everyone has met one, and everyone has the same ritual for it: rerun the job, watch it go green, merge. The ritual works, in the sense that the build goes green. What it doesn't do is tell you anything. A rerun proves that flakiness exists. It doesn't say how much, or in which tests, or whether the one you just retried is covering for a real bug.

I built flakestat to answer those questions from data a CI system already produces: the JUnit XML that pytest, Jest, go test and almost every other runner can write. It is one static Go binary with no dependencies and no account, and test results never leave your machine. This post is about the decisions inside it, because nearly all of them turned out to be about the same thing: not fooling yourself.

1Failure rate ranks the wrong tests

The obvious way to find flaky tests is to sort by how often they fail. It's also wrong, in a way that wastes exactly the people you're trying to help. A test that fails every single run isn't flaky. It's broken, and it is completely deterministic about it. Sort by failure rate and it goes straight to the top of the list, and whoever is hunting nondeterminism spends their morning on a test that needs a different kind of fix.

Flakiness is inconsistency. So flakestat doesn't count failures; it counts changes. Walk a test's results in order and mark every place where the outcome flips, pass to fail or fail to pass. The share of neighbouring pairs that disagree is the flip rate.

top of the list: test_export_csv

Fig. 1 Six tests, twenty runs each, all on one commit. Switch the ranking and watch the broken test fall from first to last. Verdicts use a Wilson lower bound on the flip rate, as flakestat's do, without its recency weighting.

Look at what each ranking does to test_export_csv. By failure rate it's the worst test in the suite. By flip rate it scores zero, because it never changes its mind, and flakestat files it separately as consistently-failing, since hiding a broken test would be worse than the flake. test_ws_reconnect passed for a while and then failed for a while, and it scores low too: a single change of state looks much more like a regression than like chance.

This isn't a hypothetical. Across the two test suites I validated flakestat on with known answers, nine tests failed a hundred times out of a hundred. A detector that ranks by failure rate puts all nine at the top of its flaky list. flakestat scored them 0.00 and filed them as broken.

2Not every disagreement is worth the same

In a burst, where flakestat runs your suite twenty times in a row on one checkout, any disagreement is flakiness by definition, because nothing changed except chance. History is murkier. A test that failed on Tuesday's commit and passed on Wednesday's might be flaky, or it might be a bug that someone fixed on Wednesday.

So every observation carries its commit, and a flip between two different commits counts as weaker evidence: a third of a same-commit flip, by default. Recency matters too, so a test that was flaky last month and has since been fixed fades back towards stable. Put together, the score is a weighted average of transitions:

t[i] = 1 if outcome[i] != outcome[i + 1] else 0     # did it change its mind?
e[i] = 1 if commit[i] == commit[i + 1] else 1 / 3   # across commits it might be a fix
score = sum(r[i] * e[i] * t[i]) / sum(r[i])         # r: recency, aged in commits

That last comment, "aged in commits", hides the first time the tool fooled me. The recency weights originally decayed per observation. Their total could never exceed about 3.3 transitions' worth, so a hundred runs bought almost exactly as much certainty as ten: on a test failing a quarter of the time, the spread of scores across repeated trials was 0.246 at a hundred runs and 0.255 at ten. Ninety extra runs, no extra information. Ageing by commit fixed it, because a burst shares one commit and gets to use all of its evidence.

3No verdict without evidence

The second mistake was subtler. Classifying on the score itself meant that a test failing 2% of the time was called flaky in a quarter of trials at ten runs, and in none at twenty. The verdict got better with less data, which is the opposite of how evidence is supposed to work.

The fix is to classify on how low the score could plausibly be, rather than on the score. Below five runs a test isn't scored at all. Above that, flakestat computes a Wilson lower bound on the score, and the bound, not the estimate, has to clear the threshold. A small sample can't reach high confidence however flaky it looks.

fails

runs 0 score – lower bound – verdict insufficient data

Fig. 2 Each run fails at the chance you set. The dot is the flip rate so far, the bar is its 95% Wilson interval, and the verdict reads the bar's left end. The dashed line is where the flip rate settles for that chance, 2p(1 − p).

Set it to fail rarely and watch the verdict hold back while the interval is wide. The thresholds aren't arbitrary either. Stable is below 0.05 and flaky is above 0.10, and the 0.10 was measured, not chosen: at a hundred runs, a threshold of 0.15 caught a test that fails 10% of the time in 45% of trials, while 0.10 catches it in 80%, and still flags stable and broken tests in none.

4Only compare what's comparable

One rule has now been broken three times, the same way, at three different levels of the code:

Two outcomes are evidence of nondeterminism only if everything that could legitimately change the result was held constant.

The first time, it was branches. A test passing on main and failing on an unfinished feature branch, interleaved in time, read as constant disagreement. The second time it was platforms, and the figure below is that case: a test that always fails on Windows and always passes everywhere else scored 0.45, and the explanation printed beneath it claimed direct evidence of nondeterminism. The test is perfectly deterministic on every platform.

flips –score –

Fig. 3 The same twenty-four runs, compared two ways. Arcs mark flips between neighbours. Only one of the two views is measuring the test.

The third time was the display layer, which kept grouping by branch after scoring had moved on. The verdict was right, but the evidence printed under it was wrong. That is arguably worse than being wrong outright, because the evidence is what a sceptical reader checks. Now there is exactly one definition of an execution context (branch, operating system, architecture, runtime and its version), and everything that shows transition evidence calls it. Duplicating the rule is what let it drift.

5The apparatus is part of the experiment

Two more failures taught me to distrust the measuring instrument as much as the thing it measures.

Running a suite with --parallel puts several copies of it in one working directory, and any suite that uses a fixed port, path or database collides with itself. Against spf13/cobra this produced a confident flaky verdict, a score of 0.60, on a test that passed eight times out of eight when run alone. At the time, the README's headline example used --parallel 4. Now anything found in parallel is re-run sequentially and demoted to suspect if it doesn't reproduce. Demoted, not deleted: five clean runs are weak evidence of innocence, not proof.

Then there's ingesting the same CI report twice. You'd expect duplicates to inflate confidence, with more agreeing observations and a tighter bound. They do the reverse. A copy sorts right next to its original and always agrees with it, so duplication injects fake agreements:

observations 12flips –flip rate –

Fig. 4 Twelve runs with four failures. Ingest the same report twice and every copy lands beside its original: the flips stay put while the agreements double.

In the case recorded in flakestat's design notes, a test failing four runs in twelve dropped from a score of 0.64 to 0.30 when its history was ingested twice, and its lower bound from 0.51 to 0.23. Duplication isn't a bookkeeping nuisance. It's a way of losing flaky tests quietly, with more apparent support for the wrong answer. Deduplication now runs when reading as well as writing, keyed on what was executed rather than when it was ingested. When in doubt it keeps the record, because missing a duplicate costs a little accuracy while a false dedup destroys evidence and says nothing.

6Don't take my word for it

The validation I trust most isn't a test case I wrote. ConduitIO's contributors reported a set of flaky tests in their own issue tracker, named them, and later fixed them. flakestat was run with an identical protocol, fixed in advance, on either side of their fix. Before it, three of the four tests in scope scored flaky at 0.67 with high confidence. At the fix itself, all four were stable at 0.00. I didn't create the bug, label it or repair it. flakestat was handed observations from both sides of a commit it had no part in, and separated them.

The protocol I'd written down first, one test repetition per process, found nothing at all in a hundred executions. That's because these were process-global-state bugs: the pollution has to happen inside one process before a later test trips over it, which Conduit's own flake-hunting job accounts for and my first protocol didn't. That null result is published next to the one that worked, as is a test run I discarded because its fixture was contaminated. Reporting only the arm that succeeded would have been exactly the kind of adjustment the whole exercise exists to police.

I'm also clear about what it hasn't shown. The default weight for cross-commit evidence is a defensible guess rather than a finding. Below a failure rate of about 5%, sampling stops being the right instrument. And flakestat has not yet found a flaky test that its maintainers didn't already know about. Those are open problems, written down where anyone can check them, which is the point.

7Try it

brew install rowhitswami/tap/flakestat      # or npm i -D flakestat, or pip install flakestat

flakestat hunt --runs 20 --junit 'reports/junit-{run}.xml' \
  -- pytest --junitxml='{junit}'

Run your suite a few times and see what disagrees. Then ask flakestat explain why a test was flagged. It will show you the evidence, and it will say so when there isn't enough of it.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.