How an Unsorted Redis Set Fragmented Our Claude Prompt Cache Into 40 Variants Over One Weekend
02:12 UTC, Saturday. A Datadog monitor named anthropic_api_cost_per_hour fires
for the first time since it was created. Hourly Claude spend has gone from a steady $7-$9 to
$61. Nobody's phone had gone off for anything else that night. No outage, no error rate spike,
no customer ticket. Just a bill climbing in the background while the on-call engineer's dog
walk got interrupted.
the scramble
First theory: a traffic spike. A support-triage feature had launched to a new enterprise tenant that week, and more tenants usually means more tokens. Pulling request volume from the API gateway rules that out in under five minutes, the request rate is flat, within 4% of the same hour the previous Saturday.
Second theory: someone shipped a bigger prompt, or swapped to a more expensive model, over the
Friday afternoon deploy. The diff for that deploy touches a feature-flag service, nothing in
the Claude client, the model string is unchanged (claude-opus-4-8), and a quick
call to Anthropic's token-counting endpoint confirms the assembled system prompt is the same
length, token for token, as it was the week before. Not a bigger prompt. Not a pricier model.
That rules out the two obvious explanations in under fifteen minutes, and the cost graph keeps climbing anyway. Whatever this is, it isn't about how much content is in the prompt. It's about what's happening to that content on the way to the API.
the hunt
The system prompt for the triage assistant is large: style guide, output schema, and a
per-tenant block listing which integrations (CRM lookup, refund processor, shipment tracker,
a dozen others) are enabled for that account. At roughly 2,400 tokens it clears Anthropic's
1,024-token minimum for prompt caching by a wide margin, so the whole block sits behind a
single cache_control breakpoint:
response = client.messages.create(
model="claude-opus-4-8",
system=[
{
"type": "text",
"text": system_prompt,
"cache_control": {"type": "ephemeral"},
}
],
messages=messages,
)
Every response from that call carries a usage block that splits input tokens into three
buckets: input_tokens, cache_creation_input_tokens, and
cache_read_input_tokens. A healthy cache looks like a handful of
cache_creation_input_tokens entries (the first call for a given prompt, billed at
roughly 1.25x the base input rate) followed by a long run of cheap
cache_read_input_tokens hits, billed at roughly a tenth of the base rate.
Pulling an hour of usage logs for the triage service shows almost nothing in
cache_read_input_tokens. Nearly every single request is paying the
cache_creation_input_tokens price, as if each one were the first request Anthropic
had ever seen for that tenant. For a service handling thousands of requests an hour against a
small, stable set of tenant configs, that's the actual anomaly: the cache isn't slow, isn't
evicting early, it's barely landing a hit at all.
To see why, the team adds a one-line debug log: a SHA-256 fingerprint of the exact bytes of the assembled system prompt, tagged with the worker's process ID, on every request. Within twenty minutes of redeploying with that log line, one tenant alone has produced 23 distinct fingerprints for what should be one unchanging block of text. Diffing two of them side by side finds the same content, word for word, in a different order:
worker pid 4821:
- crm_lookup
- refund_processor
- shipment_tracker
worker pid 5114:
- shipment_tracker
- crm_lookup
- refund_processor
Anthropic's prompt cache keys on an exact match of the prefix bytes up to the cache breakpoint. It has no idea these two blocks mean the same thing. To the cache, they're two unrelated prompts that happen to share most of their words, and only one of them, at best, can ever be a hit.
the find
The feature-flag migration that shipped four days earlier replaced a Postgres query
(SELECT tool_name FROM tenant_tools WHERE tenant_id = %s ORDER BY tool_name) with a
Redis-backed toggle store, so enabling or disabling an integration for a tenant could take
effect without a database write. The new lookup:
enabled_tools = redis_client.smembers(f"tenant:{tenant_id}:tools")
tools_block = "\n".join(f"- {t}" for t in enabled_tools)
system_prompt = TEMPLATE.format(tools=tools_block)
SMEMBERS returns a Python set, and a set of strings
iterates in whatever order the interpreter's hash table happens to place them in. Since Python
3.3, string hashing is randomized per process by default (PYTHONHASHSEED,
specifically to make hash-flooding attacks against dict and set keys harder), so the same three
tool names hash to different slots, and therefore iterate in a different order, in every
separate Python process. Within one worker, the order is stable, because the hash seed doesn't
change for the life of that process. Across the fleet, every worker has its own seed, so every
worker builds its own byte-distinct "version" of what is logically the same prompt.
The old Postgres query carried an implicit ORDER BY. The new one carried no
ordering guarantee at all, because nothing about SMEMBERS or Python's
set type promises one. The bug had existed since the Wednesday deploy, quietly
halving the cache hit rate across six steady-state worker processes, a real but easy-to-miss
cost bump that hadn't crossed any alert threshold. Friday evening's autoscaling event, timed
for an expected weekend traffic bump, took the fleet from 6 workers to 14. More workers meant
more distinct hash seeds, which meant more distinct prompt variants competing for the same
tenant's traffic, which meant the hit rate didn't just degrade, it collapsed.
the fix
The immediate fix is one line: stop trusting a Python set's iteration order for
anything that becomes part of a cached, byte-sensitive prompt. Sort it.
enabled_tools = redis_client.smembers(f"tenant:{tenant_id}:tools")
tools_block = "\n".join(f"- {t}" for t in sorted(enabled_tools))
system_prompt = TEMPLATE.format(tools=tools_block)
Disabling hash randomization fleet-wide with PYTHONHASHSEED=0 was raised and
rejected in the same incident channel: it would fix this one symptom while reopening the
hash-flooding protection that flag exists for, on a service that also builds dict and set keys
from request-derived data elsewhere in the codebase. The fix belongs at the data layer instead.
Every place a collection feeds a cached prompt now gets sorted, or otherwise serialized
deterministically, before it's used.
Two guardrails went in alongside the fix. First, the fingerprint log from the investigation
stayed in production as a real check: an alert fires if a tenant produces more than one distinct
system-prompt fingerprint within a five-minute window, which should never happen for an
unchanged config. Second, a Datadog metric now tracks
cache_read_input_tokens / (cache_read_input_tokens + cache_creation_input_tokens)
per tenant, with an alert below 0.8, catching cache fragmentation directly instead of waiting
for it to show up as a dollar figure.
the aftermath
-
Prompt caching keys on exact bytes, not meaning. Anything assembled from an unordered
collection, a
set, a dict built from a hash map response, a parallel fetch that resolves out of order, is a cache-fragmentation risk the moment it touches a cached prefix. -
Python's hash randomization is a real, intentional security feature, and a real, easy trap
for anyone who assumes a
set's iteration order is stable across processes. It is stable within one process and nowhere else. -
A migration that swaps a data source rarely gets scrutinized for ordering guarantees it
never explicitly depended on until something downstream breaks. The Postgres
ORDER BYwasn't in the code for correctness, it was just how Postgres happened to return rows, and nobody wrote it down as a requirement because nobody knew it was one. - Watch the cache hit ratio directly, not just the bill. The dollar amount is a lagging indicator that depends on traffic volume and worker count; the hit ratio is the actual health signal and would have caught this four days earlier, before an autoscaling event turned a slow leak into a spike.
Nothing about the tenant's tools had changed in four days. Not one integration was added, removed, or reconfigured. The only thing that changed was the order three strings came back in, a detail so far below the API that nobody thought to ask about it, and Claude's cache cared about that order more than anyone building the prompt ever had.