HYDRA blog All posts
field notes · release engineering

Twenty-four green builds
that shipped nothing

Every check passed. Every pull request was green. And for twenty-four consecutive runs the release channel those checks exist to protect produced no artifact at all. Here is the mechanism, the two confident wrong diagnoses I made on the way, and the three changes that stop it recurring.

8 min read·every number measured· Sep 9, 2026

Our edge release channel was dead for twenty-four consecutive builds. Every pull request in that window was green. Every check people actually look at passed, every time, while the thing those checks exist to protect produced nothing at all. Nobody noticed for the better part of a week. This is the story of how that happens, because the mechanism is boring and general and almost certainly sitting in your pipeline too.

01 · the symptom

Green does not mean shipped

The edge channel builds from develop on every merge and publishes a prerelease. It is how you get a build of what just landed without waiting for a release cut. One morning someone asked for an edge build and there wasn't one. Not a stale one. None.

Edge Build, most recent runs at the time
failurec5760e8
failure489bef1
failure…22 more…
Twenty-four in a row. The last successful edge build was old enough that the version it published no longer matched anything anyone was running.
Edge Build · 25 consecutive runs

Each cell is one run of the workflow that publishes the edge prerelease.

last green buildthe morning someone asked
Not one of these red runs appeared on a pull request. The channel was dead for the whole width of this chart, and every merge in that window still reported green.

The reason nobody saw it is one line of configuration, and it is the same line in most repositories: Edge Build is not a required check. A required check blocks a merge and appears on every pull request. A non-required workflow runs, fails, and reports to a page nobody opens unless they are already looking for it. GitHub will happily show you a green tick on the PR while a workflow triggered by that same merge burns down thirty seconds later.

One merge, two kinds of workflow

Both run. Only one of them can reach you.

A required check
Blocks the merge until it finishes, and renders on the pull request whether it passes or not.
→ you find out before you merge
Edge Build · not required
Triggered by the same merge, fails about thirty seconds later, and reports to its own page in the Actions tab.
→ you find out when someone asks for a build
The green tick on the pull request is telling the truth. It is a statement about the required checks, and it says nothing at all about the workflow the merge kicked off.

So the first lesson is not technical at all. Ask which of your workflows can fail without anyone being told. Every one of those is a channel that can quietly stop existing.

02 · two confident wrong answers

I diagnosed it twice, and was wrong twice

I want to keep these in, because the useful part of a debugging story is not the answer. It is the shape of the reasoning that produced the wrong ones.

Wrong answer one: upstream shipped a breaking change

The obvious suspect. Our release tool is pinned by a GitHub Action that resolves a version at run time, so a new upstream release enters the pipeline with nobody choosing it. Clean theory. It took one command to kill: the last successful run had used the same version as the failing ones. Upstream had not moved between the last green build and the first red one.

Wrong answer two: the tag name became invalid

A change in that window altered how the edge version string is computed. Plausible: the release tool validates semver, and an invalid tag name would explain a hard failure. I built the old name and the new name and ran both against the pinned binary. Both failed. The name was never the issue.

Both hypotheses were reasonable, both were about the right area of the system, and both were wrong. What finally worked was not a better theory. It was diffing the actual command line that ran on the last green build against the one that ran on the first red build, and reading the difference instead of reasoning about it.

03 · what actually happened

The fix that broke it was a real fix

The difference was a single flag. A previous change had removed --snapshot from the build, correctly: snapshot mode makes the tool compute its own version and ignore the tag, which is why every prerelease archive we had ever shipped was stamped 0.0.0-SNAPSHOT-none and could not say which commit it came from. Removing it was the right call.

It also removed the thing that had been hiding a second bug. The build runs on a branch, so no tag points at the commit being built. The tool needs one. In snapshot mode it logged ignoring errors because this is a snapshot and carried on. Without snapshot mode it did what it should always have done, which is fail.

Measured against the pinned binary, three runs
--snapshot, no tagsucceeds, version 0.0.0-SNAPSHOT-none
no snapshot, no tagfails
no snapshot, annotated tagsucceeds, version 1.4.3-edge.gc5760e8
The bug was always there. For months it was being swallowed by a flag that was itself a bug. Fixing the visible defect exposed the hidden one, and the pipeline went from silently wrong to loudly broken. That is an improvement, and it did not look like one.

One detail worth keeping: the tag has to be annotated, not lightweight. The tool reads the tag's message, and a lightweight tag's message is the same empty string the error complains about. git tag -a, not git tag.

04 · the fix that was not a fix

Verifying locally verifies the mechanism, not the environment

I wrote the fix, tested it against the real binary on my machine rather than reasoning about it, and watched it produce a correctly versioned build. Then it failed on the runner.

fatal: empty ident name (for <runner@runnervm…cloudapp.net>) not allowed

An annotated tag writes a tag object, and a tag object carries a tagger. A tagger needs a name and an email. My laptop has had a git identity configured since before this project existed. A fresh CI runner has none, and git's fallback derives an empty name from the OS.

The part that stings: I could not reproduce it by unsetting the global config. Git's fallback still derived a usable name from the operating system. You have to force it:

git -c user.name= -c user.email=x@y tag -a t1 -m x

What your shell hands you that a runner will not

Every row is something a locally verified fix can silently depend on.

your machinea fresh runner
git identityconfigured years agonone, and the fallback derives an empty name
$PATHeverything you have ever installedthe image default
$SHELLset by your terminaloften unset entirely
$HOMEyours, with its dotfilesa bare runner home
fetched tagswhatever you have pulledonly what the checkout asked for
The tagger name came from this table. Unsetting the global config does not reproduce it, because git still derives a usable name from the operating system; you have to pass an empty identity explicitly to see what CI sees.

So the fix became inline identity flags on the tag command, which keeps it out of every other step in the job that touches git. And the rule I actually took away: when a CI fix is verified locally, ask what your shell is handing you that a runner will not. A git identity. A populated PATH. A $SHELL. A $HOME. An already-fetched tag. Each one is a thing your test silently depends on and the runner silently lacks.

05 · what we changed

Check the artifact, not the log

The build going green is not evidence that it produced the right thing. We have learned this repeatedly and it keeps being true: a green pipeline has shipped a bundle with the wrong name, a version string of dev · none, a frozen macOS bundle version, a binary compiled by no CI job, and, in one memorable case, a routing registry that shipped in no install path at all so every installed user had blank prices.

So the verification was not "the run is green." It was downloading the published asset and running it:

hydra 1.4.3-edge.g0b60cc5
commit: 0b60cc5
built: 2026-09-08T20:20:10Z

A real version, on a real commit. The first one in twenty-five builds.

Three changes, not one

Fixing the immediate bug is the least interesting part. What stops the next one:

A test that fails at pull request time when a prerelease workflow names a tag it never creates. The class of bug, not the instance.
An exact version pin on the release tool in all three publishing workflows, including the stable one. It was running whatever had shipped most recently: on the day we pinned it, that was a version released the previous day that nobody had chosen or read the notes for. A range like ~> v2 would not have helped, because the behaviour change that broke us was a minor bump inside v2.
A note in the repository's own instructions saying that a non-required workflow can fail invisibly, so the next person does not have to rediscover it.

If you take one thing from this: go and look at your release pipeline right now and answer two questions. Which workflows in it can fail without telling anyone? And when did you last download something it published and run it? Not read the log. Run the file.