How Rotating S3 Presigned URLs Silently Defeated Next.js Image Optimization's Cache
09:41 UTC, Tuesday. A billing alert from Vercel lands in the platform team's Slack channel: "Image Optimization usage is approaching your plan limit." The plan limit is 50,000 source images a month. The alert says the account used 2.3 million in the last 24 hours. Nobody had touched the image pipeline. Nobody had touched the plan. The only thing that had shipped recently was a new product gallery page, live since Monday morning.
the scramble
First theory: a scraper. The gallery page is public, product photos are easy to hot-link, and
a bot crawling every product ID in sequence would explain a sudden jump in image requests.
Pulling the Vercel function logs for /_next/image rules that out inside ten
minutes, the referer header on almost every request is the app's own domain, and the request
rate tracks normal daytime traffic, not a crawler's steady drumbeat.
Second theory: the gallery page itself was the problem, maybe it rendered more image variants
per view than intended. The on-call engineer checks the diff for Monday's deploy. The
<Image> component's sizes prop is unchanged, still three
breakpoints, the same as every other image on the site. Forty products per gallery page,
three sizes each, that's 120 image requests per page view at worst, nowhere near enough to
explain 2.3 million unique source images in a day with the gallery's actual traffic.
Both obvious explanations are gone within fifteen minutes, and the number on the Vercel dashboard keeps climbing. Whatever is driving this, it isn't more images and it isn't a bot. It's something about how the existing images are being counted.
the hunt
Vercel's Image Optimization API doesn't bill per request. It bills per unique source image it has to fetch and transform. A healthy app re-fetches the same source rarely, the optimizer caches the transformed output at the edge and serves repeat requests for the same source URL, width, and quality straight from that cache. The volume alone doesn't explain the bill. What needs explaining is why almost none of these requests are hitting that cache.
Pulling a sample of the url= query parameters the optimizer is being asked to
fetch for a single avatar, the same user's avatar, across five page loads inside two minutes,
shows five different strings:
https://assets-private.ourapp.com.s3.amazonaws.com/avatars/8841.jpg
?X-Amz-Algorithm=AWS4-HMAC-SHA256
&X-Amz-Credential=AKIA.../20261006/us-east-1/s3/aws4_request
&X-Amz-Date=20261006T094012Z
&X-Amz-Expires=900
&X-Amz-SignedHeaders=host
&X-Amz-Signature=7f2a9c... (load 1)
...same path, X-Amz-Date=20261006T094047Z, X-Amz-Signature=b81e4d... (load 2)
...same path, X-Amz-Date=20261006T094119Z, X-Amz-Signature=029cf7... (load 3)
Same bucket, same key, same bytes on disk. Different query string every time. Dead end first: the team suspects an overeager cache-purge webhook, something firing on a schedule and evicting the edge cache before it can ever get used. Checking the optimizer's response headers for a repeat request to the exact same URL rules that out, when the query string is identical, the cache hits fine and returns in under 20ms. The problem isn't eviction. The problem is that the URL is never identical twice.
That one line of logs is the whole bug. Next.js's image optimizer, and Vercel's hosted version
of it, treats the full request URL, including every query parameter, as part of the cache key.
A presigned S3 URL carries a fresh X-Amz-Date and X-Amz-Signature on
every call, by design, that's what makes it a presigned URL instead of a permanent link. Which
means every single page view of the same unchanged avatar was being treated as a request for a
brand-new image the optimizer had never seen before.
the find
Two weeks earlier, a SOC 2 audit finding flagged that user avatars and product photos sat in a
publicly readable S3 bucket, no auth required if you had the URL. The fix, shipped quietly and
correctly from a security standpoint, moved everything to a private bucket and started
generating a presigned URL server-side for every image, with a 15-minute expiry, passed
straight into the src prop:
// server component, product gallery
const presignedUrl = await s3.getSignedUrl('getObject', {
Bucket: 'assets-private-ourapp',
Key: `products/${product.imageId}.jpg`,
Expires: 900,
});
return ;
At the traffic the site was doing before Monday, this was a real inefficiency, every page view was already paying for a fresh optimization instead of a cached one, but the absolute numbers stayed small enough to hide inside normal usage. Nobody had a monitor on cache hit ratio for the image optimizer specifically, only a monthly usage total that nobody checks mid-cycle. The gallery launch didn't introduce the bug. It multiplied an existing one: forty product photos rendered per page view, each one a guaranteed cache miss, across a traffic volume high enough to turn a slow leak into something that blew through a 50,000-image monthly plan limit in under a day.
the fix
The access control and the cache key needed to be separated. A presigned URL is the right tool for proving the request is authorized. It is the wrong tool for identifying the image to the optimizer, because its identity changes every time it's generated. The fix routes image requests through a stable, internal path, and does the S3 fetch and the authorization check server-side, where the signature never has to leave the backend:
export async function GET(
req: NextRequest,
{ params }: { params: { assetId: string } }
) {
const authorized = await checkAccess(req, params.assetId);
if (!authorized) {
return new Response('Forbidden', { status: 403 });
}
const object = await s3.getObject({
Bucket: 'assets-private-ourapp',
Key: `products/${params.assetId}.jpg`,
});
return new Response(object.Body as ReadableStream, {
headers: {
'Content-Type': 'image/jpeg',
'Cache-Control': 'public, max-age=31536000, immutable',
},
});
}
The <Image> call changes from a signed, expiring URL to a plain, stable path:
return (
);
product.imageId never changes for a given photo, so the optimizer's cache key is
stable now, no matter who's viewing it or when, until the underlying asset actually gets
replaced. The route handler still checks authorization on every call, nothing about access
control got weaker, it just stopped leaking into the one place that needed to stay constant.
The immutable cache directive is safe here specifically because asset IDs are
never reused for a changed image, a new upload gets a new ID.
the aftermath
- A presigned URL and a cache key are different problems wearing the same string. Anything that intentionally changes on every call, a signature, a timestamp, a nonce, will defeat any cache that treats the full URL as its identity, no matter how good that cache otherwise is.
- A security fix and a performance regression can ship in the same commit without anyone noticing, because nobody reviews a SOC 2 remediation for cache behavior. The review caught the access control problem it was meant to catch and missed the one it accidentally created.
- Usage caps that only alert at a monthly total, not a daily rate of change, hide a slow leak until something unrelated, in this case a product launch, amplifies it past the threshold. A day-over-day anomaly alert would have caught this two weeks earlier, while the fix was still a one-line git blame away.
- Keep authorization server-side and identity stable client-side. Any time a URL that gets cached, by a CDN, an image optimizer, a browser, also needs to prove the request is allowed, put the proof behind a stable endpoint instead of inside the URL itself.
The avatars never moved. The product photos never moved. For two weeks, the only thing wrong was a query string that looked identical in a browser's network tab and was, byte for byte, never the same request twice. The optimizer did exactly what it was built to do: cache what it could prove was the same image. It just never got the chance.