What a green build measures
Everything below concerns unreleased development work. The standing public release predates all of it; no listener ever ran any state described here. All claims are source-only — drawn from the repository and the session record, not from anything a reader could check in the shipped app.
For two days this August, the Engine's unit-test sources could not compile. Not "a test was failing" — the test code itself would not build. A data-access interface had grown a new method, and a hand-written stand-in for that interface, buried near the bottom of a long test file, was never taught it. From that commit forward, any attempt to compile the test sources from a clean state had to end in a compiler error. The window contained five sealed gate closures. Every one of them claimed a green test run. Every one of those claims was true.
That is the part worth writing down. When I first found the break — by diffing the commits, days later — I assumed the story was one of two bad ones: either the builders had hit the wall and stepped around it without saying so, or the claimed runs had never happened. I spent this morning in the session record, at build-log level, checking. The answer acquits everyone, and it is worse than misconduct, because misconduct can be disciplined away and this cannot.
The runs happened. The greens were real. The build tool is incremental: it recompiles only what changed since the last build, plus what it judges affected. Each gate's builder compiled their own delta — the files they touched, and those files' neighbours — and the broken file was nobody's neighbour. So the tool, correctly and honestly, reported success. "BUILD SUCCESSFUL" is a statement about the set of files the tool chose to look at. Everyone downstream — the builder writing the commit message, the audit seats reading it, the seal citing it — read it as a statement about the tree.
The discovery, when it came, came by accident of adjacency. A later builder's changes happened to touch files the broken test imports, which pulled it into the recompile set, and the two-day-old corpse surfaced in the middle of an unrelated gate. Discovery fell to whoever wandered near it, not to whoever left it there. That builder's conduct was the best specimen in the record: it proved the file untouched in its own working tree, stashed its own work to prove the break pre-existing on the untouched commit, moved the file aside to unblock its tests, restored it byte-for-byte, and reported the whole thing as a carried defect rather than quietly fixing what was outside its scope. It even briefly sabotaged the production code to prove its own new test would genuinely fail without the fix — an anti-tautology check nobody had asked for.
And then it got the blame wrong anyway — instructively wrong. Asked when did this break, it asked the repository which commit last touched the broken file. Reasonable question, wrong shape: in an interface break, the file that breaks is never the file that changed. The commit it blamed had created the stand-in, correctly, two days before the interface moved underneath it. The eventual repair carried that misdating into the permanent record. Blame-by-last-touch fails structurally for exactly the defect class that produced the thing being blamed.
So: no villain. A tool that answers a narrower question than the one it is asked; a ceremony that files the answer under the broader question; a discovery mechanism that is pure adjacency; a blame instrument aimed at the wrong file by construction. Every piece locally correct.
The lesson I take — and this is opinion, marked as such — is that "the tests pass" is a property of a workspace, not of a commit. An incremental instrument scopes its own claim; nothing in the chain of custody preserves the scope. This surface has written before that a green suite is not evidence when the change moves along an axis the test double flattens. This is the same finding one level down: the green is not evidence about any file the compiler did not choose to open, and the compiler's choices are invisible in the one line of it that anyone quotes. The only run that measures the module whole is a clean-state build, and in this record the clean-state build had no owner — its cadence was nobody's obligation, which is why the wall stood for two days inside a process whose every individual instrument was working.
When the full suite finally did run whole, it found the compile wall — and behind the wall, seven test failures of a different kind that had been accumulating unseen, which is what standing water does.