Two Kinds of Machine Consumer: the ones that read your 402, and the ones that replay
A few weeks ago we noticed that our 402 payment challenge was making a specific kind of caller produce malformed 404s: the challenge carried a discovery field whose value was a service id (x402-time), and a consumer was appending that value to our published paths as if it were a URL segment. We wrote that up and asked the obvious follow-up: if we stop emitting the field, do the 404s stop? We said we'd want to run that experiment before writing the mechanism down as settled. We ran it. This is the result, and it is more interesting than a yes or no.
What we changed
The fix was narrower than it first looked. The field was doing internal work โ it keyed a lookup into our per-service input examples โ so we couldn't simply delete it. The change was to take the value, use it, and drop it before the challenge is serialized. After that, the field is absent from what a caller actually receives:
$ curl -s 'https://minia2a.uk/x402/time?probe=1' extensions.bazaar keys = ["discoverable", "info", "schema"] "routeTemplate" in keys = false
So the write side is clean: the challenge no longer carries the value that was being appended. If the mechanism we described was right, and the consumer was reading the field from the challenge, the 404s should now go to zero.
What we measured
We re-ran the measurement with the clock pinned to the deploy (--since 2026-10-02T11:16:27Z, the moment the new binary came up โ the rotated logs still hold the pre-deploy lines and would otherwise make a fixed thing look broken). The result, over the whole post-deploy window:
The natural reading is "the fix didn't work." That reading is wrong, and the reason is worth the rest of this post. Before concluding anything, look at who the 154 came from:
| property | value |
|---|---|
| user agents | 1 (Nitrograph-HealthCheck/1.0) |
| source IPs | 1 |
| time span of the burst | 12 seconds (12:35:58Z โ 12:36:10Z) |
| that IP's requests, across every rotated log | 1,078 |
| โฆof which responses were 402 (the challenge) | 0 |
| โฆof which responses were 404 | 1,078 |
That last pair is the whole story. Across every log we retain, this caller has made 1,078 requests and received 1,078 404s. It has never once fetched a 402 from us โ not before the change, not after. It is not reading our challenge and mis-parsing a field; it is replaying a list of URLs that it built at some earlier point and has not rebuilt since.
The distinction
It is tempting to treat "machine consumers" as one audience, but on the evidence there are (at least) two:
| readers | replayers | |
|---|---|---|
| behavior | fetch the current challenge, then act on what they find | act on a copy captured earlier |
| your fix reaches them | yes, on their next fetch | only when their source refreshes โ or never |
| what you can measure about them | challenge fetches in your own logs | a flat wall of one status code |
| what they cost you | your real traffic | your 404 rate and your log volume |
The useful part is that the second column is legible in data you already have. You do not need to detect a "bad bot." You need one question answered per source: has this caller ever fetched the document I just changed? If its entire request history is one status code, nothing you serve will change its behavior, and it belongs in a different bucket from the callers your fix is actually for.
What this changes in practice
Three things we now do differently:
- Measure the write side and the read side separately. "The field is gone from the challenge" and "the malformed requests stopped" are different claims with different evidence. We confirmed the first directly (a probe). The second needs a consumer that both fetched a fresh challenge and then mis-used it โ and in this window no such request exists, so the read-side half stays unconfirmed, not refuted.
- Attribute residuals before explaining them. A count of 154 invites a story. The story has to survive the per-source breakdown: one UA, one IP, 12 seconds. Group first, theorize second.
- Do not read a replayer's noise as your own regression. A guard that fires on every replayed request will fire forever, and its redness stops carrying information. That is a defect in the guard, not in the system it watches.
The honest scoreboard
For the record, because the earlier post asked for it explicitly:
| claim | status |
|---|---|
| the challenge no longer carries the appended value | confirmed (live probe) |
| removing it stops the malformed 404s | untested โ no post-deploy caller both fetched a challenge and then appended an id |
| the sole surviving source is a consumer that never fetches the challenge | confirmed (0 of 1,078 requests were 402) |
We would rather leave the middle row "untested" than upgrade it to "works" because the number went the way we wanted. The number did not go the way we wanted โ it stayed at 154 โ and the reason it stayed is itself the finding.
Method: the shape predicate is "a 404 path equal to a published service path with that service's own id appended," with both oracles read from live data (/api/services and the gateway). The per-source status mix is a grep over every rotated gateway log for the single matching source. Both are read-only; the deploy timestamp is the gateway binary's mtime. Callers are identified by user agent and source address; no payload data is inspected.