How an Unsorted Redis Set Fragmented Our Claude Prompt Cache Into 40 Variants Over One Weekend
← Back
October 9, 2026Claude9 min read

How an Unsorted Redis Set Fragmented Our Claude Prompt Cache Into 40 Variants Over One Weekend

Published October 9, 20269 min read

02:12 UTC, Saturday. A Datadog monitor named anthropic_api_cost_per_hour fires for the first time since it was created. Hourly Claude spend has gone from a steady $7-$9 to $61. Nobody's phone had gone off for anything else that night. No outage, no error rate spike, no customer ticket. Just a bill climbing in the background while the on-call engineer's dog walk got interrupted.


the scramble

First theory: a traffic spike. A support-triage feature had launched to a new enterprise tenant that week, and more tenants usually means more tokens. Pulling request volume from the API gateway rules that out in under five minutes, the request rate is flat, within 4% of the same hour the previous Saturday.

Second theory: someone shipped a bigger prompt, or swapped to a more expensive model, over the Friday afternoon deploy. The diff for that deploy touches a feature-flag service, nothing in the Claude client, the model string is unchanged (claude-opus-4-8), and a quick call to Anthropic's token-counting endpoint confirms the assembled system prompt is the same length, token for token, as it was the week before. Not a bigger prompt. Not a pricier model.

That rules out the two obvious explanations in under fifteen minutes, and the cost graph keeps climbing anyway. Whatever this is, it isn't about how much content is in the prompt. It's about what's happening to that content on the way to the API.


the hunt

The system prompt for the triage assistant is large: style guide, output schema, and a per-tenant block listing which integrations (CRM lookup, refund processor, shipment tracker, a dozen others) are enabled for that account. At roughly 2,400 tokens it clears Anthropic's 1,024-token minimum for prompt caching by a wide margin, so the whole block sits behind a single cache_control breakpoint:

anthropic client call, system block
response = client.messages.create(
    model="claude-opus-4-8",
    system=[
        {
            "type": "text",
            "text": system_prompt,
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=messages,
)

Every response from that call carries a usage block that splits input tokens into three buckets: input_tokens, cache_creation_input_tokens, and cache_read_input_tokens. A healthy cache looks like a handful of cache_creation_input_tokens entries (the first call for a given prompt, billed at roughly 1.25x the base input rate) followed by a long run of cheap cache_read_input_tokens hits, billed at roughly a tenth of the base rate.

Pulling an hour of usage logs for the triage service shows almost nothing in cache_read_input_tokens. Nearly every single request is paying the cache_creation_input_tokens price, as if each one were the first request Anthropic had ever seen for that tenant. For a service handling thousands of requests an hour against a small, stable set of tenant configs, that's the actual anomaly: the cache isn't slow, isn't evicting early, it's barely landing a hit at all.

To see why, the team adds a one-line debug log: a SHA-256 fingerprint of the exact bytes of the assembled system prompt, tagged with the worker's process ID, on every request. Within twenty minutes of redeploying with that log line, one tenant alone has produced 23 distinct fingerprints for what should be one unchanging block of text. Diffing two of them side by side finds the same content, word for word, in a different order:

same tenant, two worker processes, diffed
worker pid 4821:
  - crm_lookup
  - refund_processor
  - shipment_tracker

worker pid 5114:
  - shipment_tracker
  - crm_lookup
  - refund_processor

Anthropic's prompt cache keys on an exact match of the prefix bytes up to the cache breakpoint. It has no idea these two blocks mean the same thing. To the cache, they're two unrelated prompts that happen to share most of their words, and only one of them, at best, can ever be a hit.


the find

The feature-flag migration that shipped four days earlier replaced a Postgres query (SELECT tool_name FROM tenant_tools WHERE tenant_id = %s ORDER BY tool_name) with a Redis-backed toggle store, so enabling or disabling an integration for a tenant could take effect without a database write. The new lookup:

prompt builder, as shipped
enabled_tools = redis_client.smembers(f"tenant:{tenant_id}:tools")
tools_block = "\n".join(f"- {t}" for t in enabled_tools)
system_prompt = TEMPLATE.format(tools=tools_block)

SMEMBERS returns a Python set, and a set of strings iterates in whatever order the interpreter's hash table happens to place them in. Since Python 3.3, string hashing is randomized per process by default (PYTHONHASHSEED, specifically to make hash-flooding attacks against dict and set keys harder), so the same three tool names hash to different slots, and therefore iterate in a different order, in every separate Python process. Within one worker, the order is stable, because the hash seed doesn't change for the life of that process. Across the fleet, every worker has its own seed, so every worker builds its own byte-distinct "version" of what is logically the same prompt.

The old Postgres query carried an implicit ORDER BY. The new one carried no ordering guarantee at all, because nothing about SMEMBERS or Python's set type promises one. The bug had existed since the Wednesday deploy, quietly halving the cache hit rate across six steady-state worker processes, a real but easy-to-miss cost bump that hadn't crossed any alert threshold. Friday evening's autoscaling event, timed for an expected weekend traffic bump, took the fleet from 6 workers to 14. More workers meant more distinct hash seeds, which meant more distinct prompt variants competing for the same tenant's traffic, which meant the hit rate didn't just degrade, it collapsed.


the fix

The immediate fix is one line: stop trusting a Python set's iteration order for anything that becomes part of a cached, byte-sensitive prompt. Sort it.

prompt builder, after
enabled_tools = redis_client.smembers(f"tenant:{tenant_id}:tools")
tools_block = "\n".join(f"- {t}" for t in sorted(enabled_tools))
system_prompt = TEMPLATE.format(tools=tools_block)

Disabling hash randomization fleet-wide with PYTHONHASHSEED=0 was raised and rejected in the same incident channel: it would fix this one symptom while reopening the hash-flooding protection that flag exists for, on a service that also builds dict and set keys from request-derived data elsewhere in the codebase. The fix belongs at the data layer instead. Every place a collection feeds a cached prompt now gets sorted, or otherwise serialized deterministically, before it's used.

Two guardrails went in alongside the fix. First, the fingerprint log from the investigation stayed in production as a real check: an alert fires if a tenant produces more than one distinct system-prompt fingerprint within a five-minute window, which should never happen for an unchanged config. Second, a Datadog metric now tracks cache_read_input_tokens / (cache_read_input_tokens + cache_creation_input_tokens) per tenant, with an alert below 0.8, catching cache fragmentation directly instead of waiting for it to show up as a dollar figure.


the aftermath

23 Distinct byte-identical-content prompt variants found for one tenant
$3,140 Excess Claude API spend from Wednesday's deploy to Saturday's fix
6 → 14 Worker count change that turned a quiet regression into an alert
4 days The bug ran before any monitor noticed it
  • Prompt caching keys on exact bytes, not meaning. Anything assembled from an unordered collection, a set, a dict built from a hash map response, a parallel fetch that resolves out of order, is a cache-fragmentation risk the moment it touches a cached prefix.
  • Python's hash randomization is a real, intentional security feature, and a real, easy trap for anyone who assumes a set's iteration order is stable across processes. It is stable within one process and nowhere else.
  • A migration that swaps a data source rarely gets scrutinized for ordering guarantees it never explicitly depended on until something downstream breaks. The Postgres ORDER BY wasn't in the code for correctness, it was just how Postgres happened to return rows, and nobody wrote it down as a requirement because nobody knew it was one.
  • Watch the cache hit ratio directly, not just the bill. The dollar amount is a lagging indicator that depends on traffic volume and worker count; the hit ratio is the actual health signal and would have caught this four days earlier, before an autoscaling event turned a slow leak into a spike.

Nothing about the tenant's tools had changed in four days. Not one integration was added, removed, or reconfigured. The only thing that changed was the order three strings came back in, a detail so far below the API that nobody thought to ask about it, and Claude's cache cared about that order more than anyone building the prompt ever had.

Share this
← All Posts9 min read