HYDRA blog All posts
field notes · testing

A test that passes under
the bug it exists to catch

A green test means the behaviour is correct, or it means the test cannot tell the difference. Those look identical from outside. Here are five that were passing while asserting nothing useful, three of which I wrote myself, and the thirty-second habit that caught every one.

7 min read·five real examples· Sep 18, 2026

A passing test tells you one of two things. Either the behaviour is correct, or the test cannot tell the difference. Those look identical from the outside, and a green suite gives you no way to distinguish them. Over one week of work on this codebase I found five tests that were passing while asserting nothing useful, and I wrote three of them myself. Here is each one, how it hid, and the one habit that caught every single one.

What a green test actually reports

One result, two causes, and nothing downstream can tell them apart.

✓ PASS

The behaviour is correct

The code does the thing, the assertion checked that it did, and the run is evidence.

The test cannot tell the difference

The assertion never ran, or it would hold under the bug too. The run is evidence of nothing.

both render as the same green tick
Every example below is the right-hand branch. The only way to learn which branch you are on is to make the test fail on purpose, which is the habit in section 05.
01 · the gate that documented a rule it did not enforce

A check that cannot go red is not a check

We have a coverage gate that keeps per-package floors as a ratchet. Its own file header said it enforced three things, and the first was that every package stays above its floor. It did enforce that. What it did not do was notice a package with no floor at all, which it printed as (no floor) and passed.

So a package could join the test suite, get measured, and slide from 90% to nothing without failing anything. By the time we looked, seven packages had arrived that way.

The gate was not broken. It was doing exactly what it said, and what it said had a hole in it. A rule that reads "nothing may fall below its floor" is silent about what happens to something with no floor, and silence in a gate means allow.

The fix was ten lines. The interesting part is how it stayed hidden: everybody read the header, agreed the gate enforced floors, and nobody asked the gate to prove it by failing.

02 · the assertion that evaporates

Guarded by a condition that is sometimes false

A test for shell PATH recovery asserted that the login shell's entries were merged into a bare PATH. The assertion was written like this:

if os.Getenv("SHELL") != "" && got == bare { … }

On Windows, and on any runner image that does not export SHELL, the left side is false, the whole condition is false, and the one assertion that verifies the behaviour evaluates to nothing. The test passes. It has always passed. It was passing on a platform where the code under test returns before doing any of the work being tested.

This is the most common shape of the problem, and it is worth learning to see. Any assertion guarded by an environment check is an assertion that some environments do not make. The test still reports success in those environments, and success is indistinguishable from the real thing on a dashboard.

The fix was not a bigger timeout, which was the obvious reading of the symptom. It was giving the test a shell it controls, so the behaviour is deterministic on every platform, and turning the platform difference into an assertion rather than a hole.

03 · a second mechanism reaching the same answer

Right answer, wrong reason

This is the subtlest one, and I shipped it before catching it.

We had a list that promises "newest first" and sorts on a timestamp, falling back to comparing identifiers when two timestamps are equal. On a coarse clock, two items created in the same tick get identical timestamps and the fallback decides the order, which is not an ordering in time at all.

I wrote a test with two items, aaa and zzz, and asserted the newer one came first. It passed. It would also have passed with the bug still in, because I had picked the identifiers such that the broken fallback produced the same answer as a working time comparison. The test could not distinguish them.

The same test, two choices of identifier
newer = zzzfallback also says zzz first · proves nothing
newer = aaafallback says zzz first · only a real comparison passes
Identical test, identical assertion, opposite diagnostic value. The only difference is whether the wrong mechanism and the right one disagree on this input.
The same assertion, two choices of input

Sorting newest-first on a coarse clock, where equal timestamps fall back to comparing identifiers.

newer = zzz
The broken fallback also puts zzz first, so the assertion holds either way. Proves nothing.
newer = aaa
The fallback would put zzz first and fail. Only a real time comparison passes. Now it is a test.
Identical test, identical assertion, opposite diagnostic value. The input is what decides whether the two mechanisms disagree, and only a disagreement is something a test can observe.

So when your code has a fallback path, a default, or any second route to the same output, choose inputs where the two routes disagree. Otherwise you are testing that something produced the answer, not that the right something did.

04 · the test that moves with the code

Never compute the expected value with the function under test

We stream model output to a UI, and each chunk carries the byte offset of the text before it, so a view that missed an event can tell a duplicate from a gap. Bytes matter: measuring in UTF-16 units would disagree with the Go side on the first non-ASCII character.

So I wrote a test with a multi-byte string, and computed the expected offset like this:

apply([delta('héllo', 0), delta('!', byteLength('héllo'))])

Spot it? byteLength is the function being tested. Break it, and the expected value breaks the same way, and both sides agree at the new wrong number. The test passes under the exact defect it was written to catch.

The fix is to write the number out, with a comment saying where it comes from:

// 6: "héllo" is five characters and six bytes.
apply([delta('héllo', 0), delta('!', 6)])

A literal cannot move with the implementation. That is the entire point of it.

05 · the habit

Make the test fail on purpose

Every one of the problems above was found the same way, and none of them were found by review. The habit is one step, and it takes about thirty seconds:

Before you trust a new test, break the thing it is testing and watch it go red. Not conceptually. Actually edit the code, run the test, read the failure message, and put the code back.

If it does not fail, you have learned something far more valuable than a passing run: the test does not test what you think. That is true whether the code is right or wrong, and you cannot find it out any other way.

The thirty-second habit

Run before you trust any new test.

01
Break the thing it tests. Actually edit the code. Not conceptually.
02
Run the test and read what comes back.
03
Put the code back.
→ it went red

Now check the four questions below. Failing is necessary, not sufficient.

→ it stayed green

The test does not test what you think, whether or not the code is right. You cannot learn this any other way.

Three of the five examples in this post were mine, and none of them were caught by review. Every one was caught by this.

What to check when it does fail

Does it fail for the right reason? Read the message. A test that fails because of a nil panic rather than the assertion you care about is still blind to the real defect.
Does it fail on every platform? An assertion behind an environment check does not exist where that check is false.
Would a second mechanism rescue it? Defaults, fallbacks and tiebreaks all quietly produce plausible answers.
Did the expected value come from the code under test? If so, the two will always agree.

We build this into guards now. When a test exists specifically to prevent a class of bug, the pull request has to say how it was made to fail. It costs half a minute and it is the only technique I know that distinguishes a test from a decoration.

Three of the five examples in this post are tests I wrote myself and nearly shipped. The habit is not a comment on anyone's care. It is a comment on the fact that a passing test looks exactly the same either way.