Idempotency keys: the missing piece in most agent tool-call stacks
Agents retry. Networks time out. Without idempotency keys on side-effecting tools, that combination charges cards twice and sends duplicate emails.
An agent calls charge_card. The request reaches your payment API, the charge succeeds, and then
the response times out on the way back — maybe a proxy hiccup, maybe the model’s own runtime
restarting mid-turn. From the agent’s point of view, the tool call failed. It does what agents do:
it retries. The card gets charged twice.
Nothing in that sequence is a bug in the traditional sense. Every component behaved correctly given what it could see. The failure is architectural: nobody made “call this tool twice with the same intent” a safe operation. That’s what idempotency keys fix, and most agent stacks we see in the wild don’t have them anywhere near the tool layer.
Why agents retry more than you think
Retries aren’t just an error-handling feature you bolted on — they’re structural to how agentic loops work:
- The orchestration layer retries on timeouts, rate limits, and 5xxs, often with no visibility into whether the underlying side effect already landed.
- The model retries inside its own reasoning. If a tool result comes back malformed, empty, or ambiguous, a capable model will often just try again rather than surface the ambiguity to a human.
- Humans retry. A user who doesn’t see confirmation within a few seconds clicks the button again, or re-sends the same instruction in a new message that the agent interprets as a fresh request.
- Multi-agent handoffs retry. A supervisor that re-delegates a sub-task after a worker agent stalls has no idea whether the worker’s tool call already completed.
Each of these is reasonable in isolation. Stacked on top of a non-idempotent tool, they turn a transient blip into a duplicated charge, a duplicate shipment, or a second copy of an email that was supposed to go out once.
The fix: idempotency keys, generated once, checked server-side
The pattern is old — Stripe and most payment APIs have shipped it for years — but it rarely gets carried into the tool layer of agent systems. The rule:
- Generate the key at the point of intent, not at the point of the HTTP call. When the agent decides “charge this customer $49,” that decision gets a stable key immediately — a UUID derived from the conversation turn, the tool name, and the arguments. If the call is retried, the key doesn’t change.
- Pass the key as a tool argument, not a header the runtime injects invisibly. The model (or your tool-calling wrapper) should include it explicitly so it survives serialization, logging, and replay.
- The receiving service is the source of truth. It stores
(key → result)and, on a repeat key, returns the original result without re-executing the side effect. The tool layer should never assume the caller only tries once. - Keys expire, but not too fast. 24 hours is a reasonable default for most business operations — long enough to cover retries across timeouts, deploys, and human re-sends; short enough that the dedup table doesn’t grow forever.
{
"tool": "charge_card",
"arguments": {
"customer_id": "cus_8g2k",
"amount_cents": 4900,
"idempotency_key": "turn_9f3a-charge_card-4b7e1c"
}
}
Which tools actually need this
Not every tool call needs a key — reads are naturally idempotent, and some writes are cheap to duplicate and easy to reconcile. Prioritize:
- Anything that moves money: charges, refunds, payouts, invoice creation.
- Anything that’s externally visible and hard to unsend: emails, SMS, Slack posts, support ticket creation.
- Anything that mutates shared state non-trivially: inventory decrements, seat allocation, infra provisioning (spinning up a second VM because the first API call “failed” is an expensive way to find out it didn’t).
Idempotency keys are cheap to add and free to not need — a duplicate key lookup costs a few milliseconds. Skipping them costs you a support queue full of double-charged customers and an on-call engineer manually diffing logs to figure out which of the two calls actually landed.
Where this lives in the stack
sequenceDiagram
participant Agent
participant Runtime as Tool Runtime
participant API as Payment API
Agent->>Runtime: charge_card(key=turn_9f3a...)
Runtime->>API: POST /charges (Idempotency-Key: turn_9f3a...)
API-->>Runtime: 200 charge_id=ch_1
Note over Runtime,API: timeout on the way back
Agent->>Runtime: retry charge_card(key=turn_9f3a...)
Runtime->>API: POST /charges (Idempotency-Key: turn_9f3a...)
API-->>Runtime: 200 charge_id=ch_1 (same result, no new charge)
The key point: the dedup check belongs in the service that owns the side effect, not in the agent runtime. The runtime’s job is to generate a stable key and pass it through every retry path — the service’s job is to make repeat keys a no-op. Get that division right once, per tool, and an entire class of “the agent did something twice” incidents stops happening.
Want something like this built for your team?
Get a quote →