How a Rolling Deploy's Readiness Window Pinned Every gRPC Client to One Backend Pod for 6 Hours
13:52 UTC. p99 latency on pricing-svc alerts at 1.8 seconds, four times its usual
ceiling. The dashboard shows six healthy pods, six pods passing their readiness probe, six pods
with room on their CPU requests. One of them is pegged at 97%. The other five are idling under 12%.
the setup
checkout-api calls pricing-svc over gRPC for every cart total, roughly
4,000 RPCs a minute at midday. Both run in the same namespace, and pricing-svc is
exposed through a standard ClusterIP Service, resolved once by DNS to a single virtual IP with
kube-proxy handling the fan-out to whichever pods are in the endpoint list.
const client = new PricingServiceClient(
'pricing-svc.default.svc.cluster.local:50051',
grpc.credentials.createInsecure()
);
// one channel per pod, reused for every call it makes
That last line is the part nobody thought hard about. gRPC runs on HTTP/2, and HTTP/2 multiplexes
many concurrent RPCs over a single TCP connection. A gRPC client opens one connection to a target
and keeps it open, reusing it for every subsequent call instead of opening a new one per request.
For a REST client that reconnects constantly, ClusterIP round robin evens out over thousands of
short-lived connections. For a gRPC client, the load balancing decision happens exactly once, at
connection time, and then holds for as long as that connection stays up. checkout-api
runs eight pods. Each one makes that decision independently, once, on startup.
the scramble
First theory: bad node. The hot pod, pricing-svc-7d4f-x2k9, is cordoned and drained,
forcing it onto a different node entirely. Eleven minutes later the alert fires again, same
symptom, different pod, pricing-svc-7d4f-m8p1, now pegged while the rest idle.
Different node, different hardware, same shape of failure. That rules out anything node-specific
and burns fifteen minutes doing it.
Second theory: the HPA is fighting itself, scaling on a metric that doesn't reflect real load.
kubectl describe hpa pricing-svc shows six of six replicas, target CPU utilization
at 60%, current average across the deployment sitting at 22%. The average looks fine because five
pods are nearly idle. Averages hide a skew this sharp, and the HPA has no reason to add capacity
when its own number says the fleet is underloaded.
the hunt
Instead of watching pods cycle, the next step is watching the connections themselves. Shelling
into a checkout-api pod and checking its open sockets to the pricing service:
ESTAB 0 0 10.2.4.19:44822 10.2.6.31:50051
One line. One connection. Every one of the eight checkout-api pods shows the exact
same thing: a single established socket to a single backend IP, held open since the pod started.
Cross-referencing those destination IPs against the pod list settles it: all eight client pods
are connected to the same two pricing-svc pods, and nothing is connected to the other four at all.
A Prometheus query against grpc_server_handled_total, split by pod, shows just how
lopsided it got: those two pods have handled 83% of all RPCs since the timestamp on a
checkout-api deploy from six hours earlier, a routine dependency bump with nothing
in the diff touching networking. The deploy itself didn't cause an error. It just decided, once,
where six hours of subsequent traffic would go.
the find
The rollout replaced all eight checkout-api pods inside a forty-second window,
standard behavior for a deployment with no maxUnavailable tuning. At the moment
those eight new pods came up and each opened its one gRPC connection, pricing-svc
was mid-rollout too, from an unrelated change earlier that morning. Only two of its six pods had
passed the readiness probe and made it into the Service's endpoint list. kube-proxy's iptables
rules can only route to what's registered, so every new connection from every
checkout-api pod landed on one of those two.
By the time the other four pricing-svc pods turned ready, thirty seconds later,
it didn't matter. The gRPC client's default load balancing policy is pick_first:
resolve the target once, connect to the first address, and stay connected until that connection
breaks. It doesn't re-resolve, doesn't rebalance, doesn't notice four more pods just joined the
pool. The connection from that forty-second window was healthy, so it just kept being used,
for six hours, until midday traffic pushed those two pods past the point where 1.8-second p99
was survivable.
the fix
Immediate: restart the two hot pricing-svc pods. Their connections drop, the eight
checkout-api pods reconnect, and this time all six backends are already ready, so
the new connections spread out. Latency back to baseline in under two minutes. That's a mitigation,
not a fix, since the exact same readiness race can happen on the next deploy.
The real fix is making the client aware there's more than one backend to pick from. That means a headless Service, so DNS returns every pod IP instead of one virtual IP, paired with a client-side load balancing policy that actually uses all of them:
apiVersion: v1
kind: Service
metadata:
name: pricing-svc-headless
spec:
clusterIP: None
selector:
app: pricing-svc
ports:
- port: 50051
const client = new PricingServiceClient(
'dns:///pricing-svc-headless.default.svc.cluster.local:50051',
grpc.credentials.createInsecure(),
{ 'grpc.lb_policy': 'round_robin' }
);
With round_robin, the client resolves every backend IP and opens a connection to
each one, spreading calls across all of them instead of pinning to whichever pod answered first.
As a second layer, pricing-svc now caps how long any single connection can live,
so even a client stuck on an old load balancing policy is forced to reconnect and re-resolve
periodically instead of holding one connection indefinitely:
const server = new grpc.Server({
'grpc.max_connection_age_ms': 5 * 60 * 1000,
'grpc.max_connection_age_grace_ms': 30 * 1000,
});
the aftermath
- A ClusterIP Service load balances TCP connections, not requests. That distinction only matters for protocols that hold a connection open and reuse it, which is exactly what HTTP/2 and gRPC are built to do.
-
pick_firstis a reasonable default for a client with one backend. Against a Service backed by multiple pods, it turns whatever pod answered first during a narrow readiness window into the only pod that matters, indefinitely. - Average CPU across a deployment can look completely healthy while two pods carry the entire fleet. Per-pod metrics on anything that does client-side connection reuse are not optional, they're the only view that shows this failure mode at all.
- A connection age limit on the server is a cheap insurance policy even after fixing the client. It bounds how long any future imbalance, caused by a bug nobody's found yet, can last before the system corrects itself on its own.
Nothing about that rollout failed. Both deployments finished, both health checks passed, and kube-proxy did exactly what a ClusterIP Service is supposed to do. The problem was a forty-second window where "correct" and "evenly distributed" briefly meant two different things, and gRPC's connection reuse turned that window into six hours.