Our edge release channel was dead for twenty-four consecutive builds. Every pull request in that window was green. Every check people actually look at passed, every time, while the thing those checks exist to protect produced nothing at all. Nobody noticed for the better part of a week. This is the story of how that happens, because the mechanism is boring and general and almost certainly sitting in your pipeline too.
Green does not mean shipped
The edge channel builds from develop on every merge and publishes a prerelease. It is how you get a build of what just landed without waiting for a release cut. One morning someone asked for an edge build and there wasn't one. Not a stale one. None.
Each cell is one run of the workflow that publishes the edge prerelease.
The reason nobody saw it is one line of configuration, and it is the same line in most repositories: Edge Build is not a required check. A required check blocks a merge and appears on every pull request. A non-required workflow runs, fails, and reports to a page nobody opens unless they are already looking for it. GitHub will happily show you a green tick on the PR while a workflow triggered by that same merge burns down thirty seconds later.
Both run. Only one of them can reach you.
So the first lesson is not technical at all. Ask which of your workflows can fail without anyone being told. Every one of those is a channel that can quietly stop existing.
I diagnosed it twice, and was wrong twice
I want to keep these in, because the useful part of a debugging story is not the answer. It is the shape of the reasoning that produced the wrong ones.
Wrong answer one: upstream shipped a breaking change
The obvious suspect. Our release tool is pinned by a GitHub Action that resolves a version at run time, so a new upstream release enters the pipeline with nobody choosing it. Clean theory. It took one command to kill: the last successful run had used the same version as the failing ones. Upstream had not moved between the last green build and the first red one.
Wrong answer two: the tag name became invalid
A change in that window altered how the edge version string is computed. Plausible: the release tool validates semver, and an invalid tag name would explain a hard failure. I built the old name and the new name and ran both against the pinned binary. Both failed. The name was never the issue.
Both hypotheses were reasonable, both were about the right area of the system, and both were wrong. What finally worked was not a better theory. It was diffing the actual command line that ran on the last green build against the one that ran on the first red build, and reading the difference instead of reasoning about it.
The fix that broke it was a real fix
The difference was a single flag. A previous change had removed --snapshot from the build, correctly: snapshot mode makes the tool compute its own version and ignore the tag, which is why every prerelease archive we had ever shipped was stamped 0.0.0-SNAPSHOT-none and could not say which commit it came from. Removing it was the right call.
It also removed the thing that had been hiding a second bug. The build runs on a branch, so no tag points at the commit being built. The tool needs one. In snapshot mode it logged ignoring errors because this is a snapshot and carried on. Without snapshot mode it did what it should always have done, which is fail.
One detail worth keeping: the tag has to be annotated, not lightweight. The tool reads the tag's message, and a lightweight tag's message is the same empty string the error complains about. git tag -a, not git tag.
Verifying locally verifies the mechanism, not the environment
I wrote the fix, tested it against the real binary on my machine rather than reasoning about it, and watched it produce a correctly versioned build. Then it failed on the runner.
fatal: empty ident name (for <runner@runnervm…cloudapp.net>) not allowed
An annotated tag writes a tag object, and a tag object carries a tagger. A tagger needs a name and an email. My laptop has had a git identity configured since before this project existed. A fresh CI runner has none, and git's fallback derives an empty name from the OS.
The part that stings: I could not reproduce it by unsetting the global config. Git's fallback still derived a usable name from the operating system. You have to force it:
git -c user.name= -c user.email=x@y tag -a t1 -m x
Every row is something a locally verified fix can silently depend on.
| your machine | a fresh runner | |
|---|---|---|
| git identity | configured years ago | none, and the fallback derives an empty name |
| $PATH | everything you have ever installed | the image default |
| $SHELL | set by your terminal | often unset entirely |
| $HOME | yours, with its dotfiles | a bare runner home |
| fetched tags | whatever you have pulled | only what the checkout asked for |
So the fix became inline identity flags on the tag command, which keeps it out of every other step in the job that touches git. And the rule I actually took away: when a CI fix is verified locally, ask what your shell is handing you that a runner will not. A git identity. A populated PATH. A $SHELL. A $HOME. An already-fetched tag. Each one is a thing your test silently depends on and the runner silently lacks.
Check the artifact, not the log
The build going green is not evidence that it produced the right thing. We have learned this repeatedly and it keeps being true: a green pipeline has shipped a bundle with the wrong name, a version string of dev · none, a frozen macOS bundle version, a binary compiled by no CI job, and, in one memorable case, a routing registry that shipped in no install path at all so every installed user had blank prices.
So the verification was not "the run is green." It was downloading the published asset and running it:
hydra 1.4.3-edge.g0b60cc5
commit: 0b60cc5
built: 2026-09-08T20:20:10Z
A real version, on a real commit. The first one in twenty-five builds.
Three changes, not one
Fixing the immediate bug is the least interesting part. What stops the next one:
If you take one thing from this: go and look at your release pipeline right now and answer two questions. Which workflows in it can fail without telling anyone? And when did you last download something it published and run it? Not read the log. Run the file.