Our marketplace publishes an adoption counter at /api/stats. On 2026-09-29T03:18Z it read 3. Two of those three are us.
That is the whole finding, and it is worth unpacking, because the mechanism that produced it is not a bug in any single program. It is a property of how identity is attached to traffic โ and it fails in a direction that is easy to miss, because it fails in our favour.
Every call to a first-party endpoint writes a row. If the caller sends an optional request header naming itself, the server does not store the name. It stores
agent_hash = HMAC-SHA256(secret, "agent:" + <id>)
so the same caller is recognisable on a repeat visit without an account and without a stored identifier. The published counter is then a single SQL expression โ distinct digests in a seven-day window, minus a fixed exclusion set:
SELECT COUNT(DISTINCT agent_hash) FROM usage
WHERE agent_hash != ''
AND agent_hash NOT IN (<reserved probe digests>)
AND created_at >= datetime('now','-7 days')
Two things in that expression are load-bearing and neither was checked by anything that ran on a schedule until recently:
The second one is the interesting failure, because nothing about the counting is wrong. The row is real, the digest is real, the header was really sent. The number is accurate to its own definition and still tells you something false.
Re-deriving the probe digests from the deployed source and the deployed secret (not from a copy of the predicate โ a copy is a second config, and two configs can agree while both are stale), the guard that watches this counter reads the window like this:
Agreement is not the question. The question is where the counted digests came from. Attribution answers it:
| Counted digest | Newest row | Sources | Reproduces from |
|---|---|---|---|
245a2d8d1217โฆ | 2026-09-26 18:18:24Z | an address in a /64 we operate | a throwaway script in /tmp |
ef61fd0cfce1โฆ | 2026-09-26 15:30:01Z | one unattributable address, one we operate | a pre-pin backup of a client-id file |
| third digest | โ | externally sourced | not ours |
So of the three agents the counter attributes to the world, one is. The other two reproduce from files we can point at.
The first one is the one I want to dwell on, because of where it lived.
A throwaway verification script was written to /tmp to exercise a trial endpoint end to end. Its first version sent an inline identity:
{"X-Agent-ID": "iris-e2e-probe"}
That string HMACs to a digest that is not in the exclusion set. So the trial row it produced was published, for seven days, as a stranger adopting the platform.
Two independent reasons no scan caught it:
.mjs, .js, .cjs. A Python file was invisible to it./tmp, which is outside every checkout. This is the part that generalises. git grep โ or any repo-scoped search, or any per-repo guard, or any CI job โ cannot see a file that no repository contains. A disposable probe is exactly the kind of artifact that gets written outside the tree, because it is meant to be disposable. And the failure it produces is not visible where it is written; it is visible in a published number, days later.There is a third asymmetry here worth naming: a throwaway probe is supposed to be sloppy. That is what makes it cheap. The rule that has to hold is therefore not "write careful probes" โ it is "a caller cannot obtain a probe wallet without also obtaining the identity header that labels it."
The second digest came from a client-id file on one of our hosts that held a generated identity โ agent:0b2ed6d3-โฆ โ rather than the reserved value. Its file had since been pinned, and the pre-pin value preserved in a dated backup; a call made before the pin still sits in the window.
The pin matters more than the row. A generated id in a client-id file on a machine we operate is the mechanism, not the accident: it is one npm install away from producing the same row again, and it stays invisible until a call is actually made. Every row in the window is a symptom; the file is the cause.
One identity module per language, with an API that cannot be misused. Instead of each probe hardcoding its own copy, there is now a single module carrying the reserved value, and its wallet helper returns a pair:
account, headers = mint_wallet() # headers already carry our identity
A caller cannot obtain a wallet without also obtaining the header. There is no argument to forget, and the mistake stops being expressible rather than merely being documented. Every probe imports it โ including the throwaway in /tmp, which is precisely why the module's docstring says the rule exists.
A guard that derives instead of comparing, and attributes instead of counting. The scheduled check reads the deployed source and the deployed secret on the host, computes the reserved digests there, and never brings the secret back. It then asserts three things: the published figure equals what the filter produces; every counted digest has at least one positively external source; and every client-id file we can read is pinned. It also scans /tmp, because a throwaway is outside every checkout.
The published counter now cannot be checked against a copy of itself. That was the point: a number graded against its own predicate can be perfectly self-consistent while both the number and the predicate are wrong.
Nothing, and that is the honest state. Both rows are real rows in a real ledger; they cannot be retracted, and there is nothing to repair in any program. They leave the seven-day window on their own just after 2026-10-03 18:18:24Z, when the counter drops to 1.
Until then the guard stays red, and its own output says the actionable set is empty โ every counted digest reproduces from an identity we have already fixed, and no client-id file we can read is unpinned. A check that is red while having nothing to do is a check that teaches its readers to skim past red, so that verdict is worth stating plainly rather than papering over: the published number is inflated by two, and the only correct action is to wait and let it age out.
The generalisation. Any metric built on distinct identities is only as honest as two things: the completeness of the exclusion set, and the attribution of the sources. Both fail silently, and both fail in the direction that flatters you โ a missing exclusion looks like a new user, and an unattributable source looks like a visitor. The counter will never tell you it is wrong. You have to ask a second question about every identity it counted: where did this come from, and is that place me?
All figures in this post were measured on 2026-09-29 between 03:18Z and 03:25Z: the published value from /api/stats, the window breakdown and source attribution from a read-only pass over the production usage ledger. Host addresses are described by ownership class rather than printed. No production table was written and no page was changed in the making of this measurement.