The 429 Your Agent Client Can't Read
A rate-limit response has two readers: a human debugging it in a terminal, and a retry loop parsing it in a program. Most APIs write it for the first reader and forget the second โ and the second is the one that decides whether the next second brings one request or ten thousand.
The one field a retry loop actually reads
When a service answers 429 Too Many Requests, the mechanism the caller is supposed to use to back off is the Retry-After header. It is in the HTTP spec for exactly this (RFC 9110 ยง10.2.3), and the retry machinery in common clients already honors it without the application doing anything: Python's urllib3 and requests Retry classes respect it by default, and every serious HTTP gateway and service mesh has a knob for it.
What none of that machinery does is parse your JSON body. A generic client does not know that your error document carries a field named retryAfter, or retry_after, or retryAfterSeconds โ you invented the name, and only clients that read your docs will ever look inside. So a retry hint that exists only in the body is addressed to a reader that, in practice, is not there.
The asymmetry in one line: a header is read by default; a body field is read only if the caller already knows its name. Putting your retry hint in the body is not wrong โ it is optional to the one reader that acts on it.
What we found in our own API
We went looking at this because of a monitoring bug, not a client one. While reviewing why one of our own internal checks had gone quiet on a large sweep, we read our rate-limit response at the source. The anonymous gate on our catalog routes answers:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
{"error":"rate limit exceeded","retryAfter":60}
That is retryAfter โ a body field, in seconds. There is no Retry-After header on the response. A caller that reads the JSON and knows our field name backs off exactly right; a caller using stock retry middleware sees a bare 429 with no timing and does whatever its default is, which is often "retry immediately, then again."
We also noticed a second, quieter inconsistency: we do send a Retry-After header on our database-unavailable path, and the two responses are otherwise similar enough that an integrator would reasonably assume the same shape applies to both. One carries the standard signal, the other carries a private one.
We found this while reading our own source, which is the cheap way. The expensive way to find it is to watch an agent client quietly turn a soft limit into a hard one.
Why it costs you, not just your caller
The bill for an unreadable retry hint comes back to the server. A client that cannot tell "wait 60 seconds" from "wait forever" usually picks one of two bad defaults: retry tightly until something changes, or give up and never come back. The first is a self-inflicted flood โ the exact traffic your limiter exists to stop โ and it inflates your own 429 counters until they stop distinguishing a real abuser from an ordinary client that simply could not read you. The second loses you the client silently.
There is a second-order cost, too. When your monitoring and your product share an origin, a rate limit is also a fact your own checks have to interpret correctly. A sweep that hammers its own edge until it is throttled learns nothing about the endpoints and something misleading about the response codes โ a 429 read as a defect is a phantom failure, and a 429 read as "no data" is a false all-clear. We have hit both shapes. The fix on the client side is to treat a throttle as a distinct outcome from a defect; the fix on the server side is to make the throttle impossible to misread.
A short checklist for agent-facing APIs
- Send
Retry-After. It is one line, it is in the spec, and every retry library honors it. Body fields are for humans and for callers who read your docs; the header is for the caller who did not. - If you also put the hint in the body, match the header. Two sources of the same number that can disagree is worse than one. If they can only ever be equal, say so in the docs and in a test.
- Do not let 429 mean several different things. "Slow down," "you already registered," and "daily quota spent" are three different actions โ back off, stop, come back tomorrow. If they share a status code, distinguish them in a field and in your docs, and keep the machine-readable one stable.
- Say what the limit is scoped to. Per IP, per wallet, per key, per route group โ the caller cannot infer it, and guessing wrong is how a correctly-written client still gets throttled by a limit it never read.
- Do not page on 429. It is a request for patience, not an outage. A monitor that treats it as a failure will lie to you the first time your own traffic is healthy but fast.
None of this is exotic. It is the difference between a limit that shapes traffic and a limit that a client simply cannot see.
We operate an x402 pay-per-call API marketplace and hit our own edge with the same clients agents do. The 429 shape quoted above is read from the deployed gateway source, not from memory. No numbers in this post are invented; where we describe our own behavior, it is the behavior of the running service.