A passing test tells you one of two things. Either the behaviour is correct, or the test cannot tell the difference. Those look identical from the outside, and a green suite gives you no way to distinguish them. Over one week of work on this codebase I found five tests that were passing while asserting nothing useful, and I wrote three of them myself. Here is each one, how it hid, and the one habit that caught every single one.
One result, two causes, and nothing downstream can tell them apart.
The behaviour is correct
The code does the thing, the assertion checked that it did, and the run is evidence.
The test cannot tell the difference
The assertion never ran, or it would hold under the bug too. The run is evidence of nothing.
A check that cannot go red is not a check
We have a coverage gate that keeps per-package floors as a ratchet. Its own file header said it enforced three things, and the first was that every package stays above its floor. It did enforce that. What it did not do was notice a package with no floor at all, which it printed as (no floor) and passed.
So a package could join the test suite, get measured, and slide from 90% to nothing without failing anything. By the time we looked, seven packages had arrived that way.
The gate was not broken. It was doing exactly what it said, and what it said had a hole in it. A rule that reads "nothing may fall below its floor" is silent about what happens to something with no floor, and silence in a gate means allow.
The fix was ten lines. The interesting part is how it stayed hidden: everybody read the header, agreed the gate enforced floors, and nobody asked the gate to prove it by failing.
Guarded by a condition that is sometimes false
A test for shell PATH recovery asserted that the login shell's entries were merged into a bare PATH. The assertion was written like this:
if os.Getenv("SHELL") != "" && got == bare { … }
On Windows, and on any runner image that does not export SHELL, the left side is false, the whole condition is false, and the one assertion that verifies the behaviour evaluates to nothing. The test passes. It has always passed. It was passing on a platform where the code under test returns before doing any of the work being tested.
This is the most common shape of the problem, and it is worth learning to see. Any assertion guarded by an environment check is an assertion that some environments do not make. The test still reports success in those environments, and success is indistinguishable from the real thing on a dashboard.
The fix was not a bigger timeout, which was the obvious reading of the symptom. It was giving the test a shell it controls, so the behaviour is deterministic on every platform, and turning the platform difference into an assertion rather than a hole.
Right answer, wrong reason
This is the subtlest one, and I shipped it before catching it.
We had a list that promises "newest first" and sorts on a timestamp, falling back to comparing identifiers when two timestamps are equal. On a coarse clock, two items created in the same tick get identical timestamps and the fallback decides the order, which is not an ordering in time at all.
I wrote a test with two items, aaa and zzz, and asserted the newer one came first. It passed. It would also have passed with the bug still in, because I had picked the identifiers such that the broken fallback produced the same answer as a working time comparison. The test could not distinguish them.
Sorting newest-first on a coarse clock, where equal timestamps fall back to comparing identifiers.
So when your code has a fallback path, a default, or any second route to the same output, choose inputs where the two routes disagree. Otherwise you are testing that something produced the answer, not that the right something did.
Never compute the expected value with the function under test
We stream model output to a UI, and each chunk carries the byte offset of the text before it, so a view that missed an event can tell a duplicate from a gap. Bytes matter: measuring in UTF-16 units would disagree with the Go side on the first non-ASCII character.
So I wrote a test with a multi-byte string, and computed the expected offset like this:
apply([delta('héllo', 0), delta('!', byteLength('héllo'))])
Spot it? byteLength is the function being tested. Break it, and the expected value breaks the same way, and both sides agree at the new wrong number. The test passes under the exact defect it was written to catch.
The fix is to write the number out, with a comment saying where it comes from:
// 6: "héllo" is five characters and six bytes.
apply([delta('héllo', 0), delta('!', 6)])
A literal cannot move with the implementation. That is the entire point of it.
Make the test fail on purpose
Every one of the problems above was found the same way, and none of them were found by review. The habit is one step, and it takes about thirty seconds:
Before you trust a new test, break the thing it is testing and watch it go red. Not conceptually. Actually edit the code, run the test, read the failure message, and put the code back.
If it does not fail, you have learned something far more valuable than a passing run: the test does not test what you think. That is true whether the code is right or wrong, and you cannot find it out any other way.
Run before you trust any new test.
Now check the four questions below. Failing is necessary, not sufficient.
The test does not test what you think, whether or not the code is right. You cannot learn this any other way.
What to check when it does fail
We build this into guards now. When a test exists specifically to prevent a class of bug, the pull request has to say how it was made to fail. It costs half a minute and it is the only technique I know that distinguishes a test from a decoration.
Three of the five examples in this post are tests I wrote myself and nearly shipped. The habit is not a comment on anyone's care. It is a comment on the fact that a passing test looks exactly the same either way.