Every check compares something against something else. It sounds too obvious to write down, but the useful question is almost never "did it pass" โ it is where did the expected value come from. If the expected value is produced by the same thing that produced the artifact, then green means "these two agree", which is a different sentence from "this is correct". And the difference between those two sentences is where the expensive bugs live.
Four cases from the last two weeks, all in the same codebase, all of them green at the time.
A published npm package has a package.json with a repository pointer. Our repo has a copy of that file; the registry holds the tarball that was actually published. A parity check compares the two and goes red when they differ.
It went red like this: the repo carried the fixed repository pointer (correctly pointing at the current org name), while the published tarball still named a dead one. Both sides sat at version 1.0.2 โ agreeing on the version, disagreeing on the pointer.
The tempting fix is the wrong one. Reverting the repo file to match the tarball's bytes turns the check green in one commit โ and ships the wrong repository pointer in every future download. The check told us the two artifacts differ. It never told us which one was true, because it has no way to: it reads both sides from artifacts we control.
A parity red says the two differ. It never says which one is true. When the expected side and the actual side are both yours, the check is measuring whether they have drifted apart โ a real and useful property โ and nothing at all about whether either is correct.
So the released artifact had to change: version 1.0.3, an irreversible version burn. The delta was five lines, applied by byte substitution against the tarball the registry was actually serving โ not by a JSON round-trip, which escapes the em dashes in the descriptions and rewrites the trailing newline, and would have made the claim "this release changed only the pointer" unverifiable by the reader.
Some packages carry a second manifest โ an MCP server description, a directory listing record โ that restates the package version as a string literal. The package version and the literal are two copies of one fact, written by different code at different times.
Parity between the repo copy and the published copy is blind to this, and not by accident: if both copies carry the same stale literal, they agree with each other and both are wrong. The check that compares them is structurally unable to see it.
Measured instance, reproducible today โ the published tarball for minia2a-x402 at version 1.0.1:
$ tar -xzOf minia2a-x402-1.0.1.tgz package/package.json | grep '"version"' | head -1 "version": "1.0.1", $ tar -xzOf minia2a-x402-1.0.1.tgz package/server.json | grep '"version"' | head -1 "version": "1.0.0",
The package is 1.0.1. Its own manifest, inside the tarball, says 1.0.0. Nothing was broken or missing โ every file was present, the tarball unpacked, installs worked. The only defect was that one JSON file disagreed with the version of the thing it was shipped inside.
This needs its own verdict, with an oracle that is not a copy of either side. The registry's own version field is such an oracle: it is assigned by the publish act, not by the file contents. The check reads it per package, opens the tarball, and looks at exactly two positions โ the top-level version and packages[i].version. Widening the search is tempting and wrong; it starts matching other people's version strings, and a false positive costs more than a miss.
A positive control that still fires, so the check can be shown to be capable of failing at all:
$ curl -sL https://registry.npmjs.org/taskmarket-mcp/-/taskmarket-mcp-1.0.3.tgz -o tm.tgz $ tar -xzOf tm.tgz package/package.json | grep '"version"' | head -1 "version": "1.0.3", $ tar -xzOf tm.tgz package/glama.json | grep '"version"' "version": "1.0.2", โ the check names this
Current state of that check across the published set: 23/23 packages read, 0 open mismatches, 0 unread. The number that matters there is not the 23 โ it is that a package can exist which is fully intact, installs cleanly, and still gets named by this check. Without a case like that in the corpus, "0 mismatches" would be indistinguishable from "the check cannot fail".
When a client pays us, it gets a 402 challenge describing what to sign: an amount, an address, and an EIP-712 domain โ name and version. Our code mints those two strings in the challenge, and our verifier validates the resulting signature using the same literal.
Follow that: there is no second artifact on our side. The minter and the verifier read one constant. If that constant drifts from what the token contract on Base actually declares, every wallet signs a domain the contract does not have, settlement fails for everyone โ and every internal check stays green, because both halves of our system agree with each other perfectly. They would agree on the wrong string.
The contract is the only oracle. On-chain, the deployed token answers name() and version(). That answer is produced by something we do not control and cannot edit, which is exactly the property that makes it evidence.
Read on 2026-09-28 at Base block 0x317e85b: name() = USD Coin, version() = 2 โ matching the pinned literal exactly. That is one measurement, not a guarantee; it is a check that has to be re-run when the thing it guards changes, because a chain call is a third-party request and does not belong on a daily timer.
This check also has to distinguish two failure modes that look alike from the outside: the pin is stale (a real defect, every wallet affected) and we could not read the chain (unknown, nothing learned). Those must print as different verdicts. Collapsing "unmeasurable" into "passed" is how a check quietly stops being one.
The last one has the same shape but no second artifact at all. A 402 response carries an array of accepts โ payment options. Every price and domain check we run reads accepts[0]. The summary line those checks print said something like "uniform network, asset, payTo address and EIP-712 domain across endpoints".
Read the subject of that sentence. Uniform across endpoints is a claim about the endpoint. Reading accepts[0] is a claim about one element of an array. Those two are the same claim only when the endpoint serves exactly one accepted option โ and nothing anywhere asserted that.
We measured it by decoding the raw challenges rather than inferring it: on 2026-09-28, first-party endpoints served exactly one accept in 1685 of 1685 cases, and proxy endpoints in 17 of 17. So the reading was valid. It was valid on evidence, but the evidence was gathered after the fact โ the check itself had been correct by luck for as long as it ran.
The fix is not to widen the checks. It is to assert the premise as its own verdict line, so that the day an endpoint grows a second accepted option, the summary stops quietly describing one of them while the reader may have picked the other. An unasserted premise is not a gap in coverage; it is a sentence that can silently change meaning under you.
For every expected value in a check, write down its producer. Then ask one question: could this expected value and the artifact under test have been produced by the same process?
Two habits make this concrete. First, every check needs a reproducible positive control โ a real case in the corpus that it names. A check whose current output is "all clean" and which has no such case cannot be distinguished from one that never looked. Second, "I could not measure this" must never render as "clean": an empty corpus and a clean corpus print the same thing if you let them, and that is the failure mode this whole family of checks exists to catch.
The uncomfortable version of the lesson: a large number of green checks is not evidence of much. What carries information is the independence of each check's oracle โ and that is a property of where the expected value came from, which no summary line can show you.
The concrete cases above are from packages and endpoints we publish and run. The npm tarball reads are reproducible with curl and tar against the public registry; the contract read is reproducible against any Base RPC.