How an Unpinned Server Actions Key Turned a Routine Deploy Into 47 Duplicate Orders
14:32 UTC. Slack: #support-urgent lights up with four tickets in six minutes, all
some version of "charged me twice for one order." Stripe's dashboard shows a cluster of
duplicate charges starting at 14:26, same customer, same cart total, eleven seconds apart. A
deploy to main went out at 14:25. Nobody on call believes those two facts are
unrelated, and nobody yet knows why.
the setup
Checkout runs as a Next.js Server Action, placeOrder, called directly from the
cart's submit button. It inserts the order row, then calls Stripe to capture the charge:
'use server';
export async function placeOrder(cartId: string, formData: FormData) {
const cart = await getCart(cartId);
const order = await db.insert(orders).values({
cartId,
total: cart.total,
status: 'pending',
}).returning();
const charge = await stripe.paymentIntents.create({
amount: cart.total,
currency: 'usd',
customer: cart.customerId,
});
await db.update(orders).set({ status: 'charged', paymentIntentId: charge.id })
.where(eq(orders.id, order[0].id));
return { orderId: order[0].id };
}
No idempotency key anywhere in that path, because in eighteen months of production traffic it
had never once needed one. The submit button disables itself during the pending transition via
useFormStatus, which covers the obvious double-click case. The gap nobody had
reason to think about was a deploy landing in the middle of an open checkout tab.
the scramble
First theory, because it's always the first theory: a double-click slipping past the disabled state. Session replay for the first affected customer rules that out inside two minutes, one click, one spinner, one error toast, then a second click eleven seconds later. Not a double click, a retry.
Second theory: a race in the INSERT, two concurrent requests both reading
pending before either writes charged. Also wrong, pg_stat_activity
shows the two inserts for each duplicated order are seconds apart, not concurrent, and each one
completes fully before the next starts. Two separate, successful, sequential requests. That's
not a race condition, that's two legitimate-looking submissions.
The detail that breaks the case open: every affected session has a Sentry breadcrumb timestamped between 14:25:40 and 14:31:50, the exact window the 14:25 deploy was rolling out. All forty-seven of them. Nothing before, nothing after.
the hunt
Sentry's breadcrumb for the first failed attempt is the same error on every single affected session:
Error: Failed to find Server Action "7f3a9c1e2b4d...".
This request might be from an older or newer deployment.
That's a known Next.js error, not a bug in application code. A Server Action's id is a hash that's signed with an encryption key generated at build time. If that key isn't pinned across deployments, every build mints a new one, and a client who loaded the page under the old build is holding an action reference the new build's servers can't verify. On Vercel, the swap from old deployment to new isn't instantaneous, a tab that's been open for even a few minutes can submit against a build that no longer exists by the time the request lands.
grep -r NEXT_SERVER_ACTIONS_ENCRYPTION_KEY across the repo and the Vercel
environment variables returns nothing. Every deploy has been generating its own random key this
entire time, and it had simply never mattered before, because nobody had hit submit on a
years-old session in the exact ninety-second window a deploy was draining old instances. The
real question still isn't answered yet: that error happens before the action body ever runs, so
it shouldn't create an order at all. Where did the duplicate charge actually come from?
The answer is in the client code, not the server:
export async function submitWithRetry(fn: () => Promise, attempts = 2): Promise {
try {
return await fn();
} catch (err) {
if (attempts <= 1) throw err;
await new Promise((r) => setTimeout(r, 800));
return submitWithRetry(fn, attempts - 1);
}
}
Checkout wraps placeOrder in this. It retries on literally any thrown error,
including "failed to find this action," with no check for whether the first attempt might have
been a dispatch failure versus a real one, and no idempotency key distinguishing attempt one
from attempt two. 800ms after the stale-action error, the wrapper calls placeOrder
again. By then the page has silently picked up the new deployment's action id on the next
interaction, so attempt two resolves clean, inserts a fresh order, and charges the card. Attempt
one never ran placeOrder's body at all, so that part is harmless on its own. What
turns it into a duplicate is the customer: seeing a generic "something went wrong, please try
again" toast after attempt one's visible failure, several of them clicked the button again
themselves, not realizing the wrapper's own retry was already in flight. Two independent,
fully successful submissions, each with its own order row and its own Stripe charge, for one
cart.
the find
Two separate decisions combined to cause this, and either one alone would have been fine. Leaving the Server Actions encryption key unpinned across deployments is a known footgun for any multi-instance or rolling deployment, Next.js's own docs call out setting it explicitly for exactly this reason, but on its own it just produces an error toast, not money moving twice. Retrying blindly on any thrown error with no idempotency key is also, on its own, a latent risk that had simply never been exercised. Together, a deploy-timing error that looks exactly like a transient network blip triggered a retry path that had no way to tell "this attempt never touched the database" apart from "this attempt may have landed," and a user who also, reasonably, tried again.
the fix
First, pin the encryption key so a deploy stops invalidating in-flight action references at all:
vercel env add NEXT_SERVER_ACTIONS_ENCRYPTION_KEY production
# 32-byte base64 key, same value across every environment and every future build
Second, and the part that actually closes the gap regardless of what else goes wrong mid-request, give the mutation an idempotency key and make the database the source of truth for "has this cart already been charged," not the client:
'use server';
export async function placeOrder(cartId: string, idempotencyKey: string, formData: FormData) {
const cart = await getCart(cartId);
const existing = await db.query.orders.findFirst({
where: eq(orders.idempotencyKey, idempotencyKey),
});
if (existing) return { orderId: existing.id };
const [order] = await db.insert(orders).values({
cartId,
total: cart.total,
status: 'pending',
idempotencyKey,
})
.onConflictDoNothing({ target: orders.idempotencyKey })
.returning();
if (!order) {
const row = await db.query.orders.findFirst({ where: eq(orders.idempotencyKey, idempotencyKey) });
return { orderId: row!.id };
}
const charge = await stripe.paymentIntents.create({
amount: cart.total,
currency: 'usd',
customer: cart.customerId,
idempotencyKey,
});
await db.update(orders).set({ status: 'charged', paymentIntentId: charge.id })
.where(eq(orders.id, order.id));
return { orderId: order.id };
}
The key is generated once per cart on the client, stored alongside the cart in
sessionStorage, and reused across every retry for that cart whether the retry comes
from the wrapper or from the customer clicking again. A unique constraint on
orders.idempotency_key backs it in Postgres, so even a retry that somehow raced past
the application-level check gets caught at the database. Passing the same key to Stripe's
paymentIntents.create gets the same protection on the charge itself, independent of
the order row.
Third, submitWithRetry stopped retrying blind. It now only retries on errors that
are unambiguously pre-dispatch, a TypeError from fetch itself failing
to reach the server, not on any error the server actually returned:
export async function submitWithRetry(fn: () => Promise, attempts = 2): Promise {
try {
return await fn();
} catch (err) {
const isNetworkFailure = err instanceof TypeError;
if (!isNetworkFailure || attempts <= 1) throw err;
await new Promise((r) => setTimeout(r, 800));
return submitWithRetry(fn, attempts - 1);
}
}
the aftermath
- An unpinned Server Actions encryption key is silent until a deploy lands inside an open session's window, then it surfaces as "Failed to find Server Action," which reads exactly like a flaky network error and gets treated like one.
- That error happens at dispatch, before the action body runs, so it is not itself a duplicate write. The duplicate only exists because the retry path had no way to distinguish a request that never touched the database from one that might have.
- Any mutation reachable by both an automatic retry wrapper and a user's own second click needs an idempotency key that survives across both, generated once per logical attempt and enforced with a database constraint, not just checked in application code.
- A generic catch-and-retry on any thrown error is a reasonable default for read paths. On a write path it needs to know the difference between "the server never saw this" and "the server might have already acted on this," and those are not the same category of failure.
Pinning the key made the error stop happening. The idempotency key is what actually made it safe for it to happen anyway, since the next edge case that produces the same symptom, a dropped connection mid-response, a cold start timeout, anything, won't be a new incident. It'll just be a retry that returns the same order id it returned the first time.