How a Cookie Domain Attribute Added for SSO Let a Staging Session Authenticate as Admin on Production for 47 Minutes
09:14 UTC. The security on-call phone goes off with a single audit-log line:
admin_user=priya.k, action=bulk_refund, count=38, source_ip=203.0.113.44.
Priya hasn't logged in in three weeks. She's on parental leave. The source IP doesn't match
any known VPN range, office network, or device she's ever used. Her account has no MFA challenge
recorded in the last thirty days. Something just acted as her, and it didn't need to ask twice.
the scramble
First theory: credential compromise. The on-call engineer force-resets Priya's password, revokes her active sessions, and checks the SSO provider's login history for anything unusual. Nothing. No failed logins, no password reset requests, no new device registration. Whatever is holding this session didn't authenticate through the login form at all.
Second theory: a leaked session token from a phishing kit or a browser extension harvesting cookies. The team pulls the session ID from the audit log and checks its issued-at timestamp against the session store. The session is eleven days old, far outside the 24-hour token lifetime every production login is supposed to get. Either the session TTL is broken everywhere, which would mean far more than 38 refunds by now, or this one session was issued somewhere that doesn't enforce the same lifetime.
That's the detail that redirects the whole investigation: not where the session came from, but why it's still alive.
the hunt
Pulling the full session record instead of just the ID:
SELECT session_id, user_email, issued_at, issuer_host, scope
FROM sessions
WHERE session_id = 'sess_8f2a1c...';
session_id | user_email | issued_at | issuer_host | scope
----------------+-------------------+---------------------+-----------------------+--------
sess_8f2a1c... | priya.k@ourapp.com| 2026-09-26 14:02:11 | staging.ourapp.com | admin
issuer_host is staging.ourapp.com. This session was never issued by
production at all. It was issued by staging, eleven days ago, and it's been sitting valid this
whole time because staging's session TTL was set to 30 days for QA convenience. Nobody
expected a staging-issued session to ever be presented to a production endpoint, so nobody thought
to cap it tighter.
The next question is mechanical: browsers don't send a staging.ourapp.com cookie to
app.ourapp.com unless something explicitly told them to. Checking the
Set-Cookie header staging actually issues:
Set-Cookie: session=sess_8f2a1c...; Domain=.ourapp.com; Path=/; Secure; HttpOnly; SameSite=Lax
Domain=.ourapp.com. Not staging.ourapp.com, the parent domain with a
leading dot, which per RFC 6265 scopes the cookie to every subdomain of
ourapp.com, including app.ourapp.com and admin.ourapp.com
in production. A browser holding this cookie sends it to any of them.
Git blame on the auth middleware points to a PR from three weeks earlier, titled
“SSO: share session across app and admin subdomains.” Before that change, each
production subdomain issued and checked its own host-scoped cookie, and a login on
app.ourapp.com didn't carry over to admin.ourapp.com, a real
usability complaint from the support team. The fix widened the cookie's Domain
attribute to the shared parent so one login would work on both. It shipped to production and to
staging in the same deploy, because staging runs the same auth middleware image as production by
design, to keep the two environments behaviorally identical.
What nobody accounted for: staging and production aren't just behaviorally identical, they share
the literal parent domain, ourapp.com, under one wildcard TLS certificate
(*.ourapp.com) issued years ago for convenience. A cookie scoped to
.ourapp.com doesn't distinguish staging from production at all. To the browser,
they're just two subdomains of the same registrable domain, exactly as eligible for the cookie as
each other.
The last piece is why this specific session got used for something damaging instead of sitting
inert. Staging's seed data includes test admin accounts created from a script that generates
emails as {'{'}firstname{'}'}.{'{'}lastinitial{'}'}@ourapp.com for realism — and
one of those seeded accounts, created before the seeding script added a +staging
suffix convention, happened to collide exactly with Priya's real production email. A QA contractor
logged into staging using that seeded fixture on September 26th to test the new SSO flow. The
resulting cookie, scoped to the parent domain, was valid wherever
priya.k@ourapp.com had admin rights, which included production.
the find
Three independent decisions, each reasonable alone, combined into one failure: the cookie
Domain attribute was widened to support cross-subdomain SSO; staging and
production were deliberately kept on the same registrable domain and the same wildcard cert for
operational simplicity; and a seed script generated a test account whose email happened to collide
with a real admin's. None of those three decisions touches authentication logic. The bug isn't in
how sessions are validated, it's in which hosts a valid session is allowed to reach,
a property that lives entirely in one cookie attribute nobody treated as a security boundary.
the fix
Immediate containment: revoke every session with issuer_host = 'staging.ourapp.com'
and force a session-secret rotation, invalidating all tokens site-wide:
DELETE FROM sessions WHERE issuer_host LIKE '%staging%';
-- 212 stale staging sessions found, 1 with production-scoped admin rights
UPDATE app_config SET session_secret = gen_random_uuid() WHERE key = 'session_secret';
-- invalidates every signed session cookie currently outstanding, prod and staging
The real fix is scoping the cookie to exactly the hosts that need it, instead of the entire parent domain, and keeping staging off that shared scope entirely:
// before: scopes to every subdomain of ourapp.com, including staging
res.cookie('session', token, {
domain: '.ourapp.com',
secure: true,
httpOnly: true,
sameSite: 'lax',
});
// after: explicit host list, no leading-dot wildcard
const PROD_SSO_HOSTS = ['app.ourapp.com', 'admin.ourapp.com'];
res.cookie('session', token, {
domain: PROD_SSO_HOSTS.includes(req.hostname) ? 'ourapp.com' : req.hostname,
secure: true,
httpOnly: true,
sameSite: 'strict',
});
Setting domain to the bare registrable domain (ourapp.com, no leading
dot) still shares the cookie across app and admin as required, but only
when the issuing host is actually one of the two production hosts on the allow list — staging
falls through to its own host-scoped cookie and can never produce a token that production accepts.
Staging was also moved off the shared wildcard cert onto its own registrable domain,
ourapp-staging.internal, so the two environments can no longer share a cookie jar even
if a future change widens the domain again. The seed script was patched to always append
+staging to generated emails, and a CI check now fails any deploy where a staging
config sets a cookie Domain matching production's registrable domain.
the aftermath
-
A cookie's
Domainattribute is an access-control boundary, not a convenience setting. Widening it for a legitimate SSO use case has the same blast radius as widening any other authorization scope, and deserves the same review. -
Sharing a registrable domain or wildcard certificate between staging and production to cut
ops overhead quietly erases the isolation between them for anything that scopes
itself to the parent domain: cookies, some CORS configurations, and any storage that
keys off
document.domain. -
Seed data that generates realistic-looking emails can collide with real users by accident.
A namespacing convention (
+staging, a dedicated test domain) isn't cosmetic, it's the second-to-last line of defense once environment isolation has already been weakened. -
The session's own
issuer_hostwas logged the whole time and nobody looked at it until the TTL anomaly forced the question. Any session store that spans more than one environment should alert the moment a session issued by staging is presented anywhere near production, rather than waiting for a human to notice the age doesn't add up.
The login form never stopped checking passwords. The boundary that failed was supposed to sit between two environments, and it turned out to run through a single leading dot in a header neither team was reviewing as a security control.