How a Rolling Deploy's Readiness Window Pinned Every gRPC Client to One Backend Pod for 6 Hours
← Back
September 29, 2026Kubernetes8 min read

How a Rolling Deploy's Readiness Window Pinned Every gRPC Client to One Backend Pod for 6 Hours

Published September 29, 20268 min read

13:52 UTC. p99 latency on pricing-svc alerts at 1.8 seconds, four times its usual ceiling. The dashboard shows six healthy pods, six pods passing their readiness probe, six pods with room on their CPU requests. One of them is pegged at 97%. The other five are idling under 12%.


the setup

checkout-api calls pricing-svc over gRPC for every cart total, roughly 4,000 RPCs a minute at midday. Both run in the same namespace, and pricing-svc is exposed through a standard ClusterIP Service, resolved once by DNS to a single virtual IP with kube-proxy handling the fan-out to whichever pods are in the endpoint list.

checkout-api/pricing-client.ts
const client = new PricingServiceClient(
  'pricing-svc.default.svc.cluster.local:50051',
  grpc.credentials.createInsecure()
);
// one channel per pod, reused for every call it makes

That last line is the part nobody thought hard about. gRPC runs on HTTP/2, and HTTP/2 multiplexes many concurrent RPCs over a single TCP connection. A gRPC client opens one connection to a target and keeps it open, reusing it for every subsequent call instead of opening a new one per request. For a REST client that reconnects constantly, ClusterIP round robin evens out over thousands of short-lived connections. For a gRPC client, the load balancing decision happens exactly once, at connection time, and then holds for as long as that connection stays up. checkout-api runs eight pods. Each one makes that decision independently, once, on startup.


the scramble

First theory: bad node. The hot pod, pricing-svc-7d4f-x2k9, is cordoned and drained, forcing it onto a different node entirely. Eleven minutes later the alert fires again, same symptom, different pod, pricing-svc-7d4f-m8p1, now pegged while the rest idle. Different node, different hardware, same shape of failure. That rules out anything node-specific and burns fifteen minutes doing it.

Second theory: the HPA is fighting itself, scaling on a metric that doesn't reflect real load. kubectl describe hpa pricing-svc shows six of six replicas, target CPU utilization at 60%, current average across the deployment sitting at 22%. The average looks fine because five pods are nearly idle. Averages hide a skew this sharp, and the HPA has no reason to add capacity when its own number says the fleet is underloaded.


the hunt

Instead of watching pods cycle, the next step is watching the connections themselves. Shelling into a checkout-api pod and checking its open sockets to the pricing service:

kubectl exec checkout-api-9f3a-vv2q -- ss -tnp | grep :50051
ESTAB  0  0  10.2.4.19:44822  10.2.6.31:50051

One line. One connection. Every one of the eight checkout-api pods shows the exact same thing: a single established socket to a single backend IP, held open since the pod started. Cross-referencing those destination IPs against the pod list settles it: all eight client pods are connected to the same two pricing-svc pods, and nothing is connected to the other four at all.

A Prometheus query against grpc_server_handled_total, split by pod, shows just how lopsided it got: those two pods have handled 83% of all RPCs since the timestamp on a checkout-api deploy from six hours earlier, a routine dependency bump with nothing in the diff touching networking. The deploy itself didn't cause an error. It just decided, once, where six hours of subsequent traffic would go.


the find

The rollout replaced all eight checkout-api pods inside a forty-second window, standard behavior for a deployment with no maxUnavailable tuning. At the moment those eight new pods came up and each opened its one gRPC connection, pricing-svc was mid-rollout too, from an unrelated change earlier that morning. Only two of its six pods had passed the readiness probe and made it into the Service's endpoint list. kube-proxy's iptables rules can only route to what's registered, so every new connection from every checkout-api pod landed on one of those two.

By the time the other four pricing-svc pods turned ready, thirty seconds later, it didn't matter. The gRPC client's default load balancing policy is pick_first: resolve the target once, connect to the first address, and stay connected until that connection breaks. It doesn't re-resolve, doesn't rebalance, doesn't notice four more pods just joined the pool. The connection from that forty-second window was healthy, so it just kept being used, for six hours, until midday traffic pushed those two pods past the point where 1.8-second p99 was survivable.


the fix

Immediate: restart the two hot pricing-svc pods. Their connections drop, the eight checkout-api pods reconnect, and this time all six backends are already ready, so the new connections spread out. Latency back to baseline in under two minutes. That's a mitigation, not a fix, since the exact same readiness race can happen on the next deploy.

The real fix is making the client aware there's more than one backend to pick from. That means a headless Service, so DNS returns every pod IP instead of one virtual IP, paired with a client-side load balancing policy that actually uses all of them:

k8s/pricing-svc-headless.yaml
apiVersion: v1
kind: Service
metadata:
  name: pricing-svc-headless
spec:
  clusterIP: None
  selector:
    app: pricing-svc
  ports:
    - port: 50051
checkout-api/pricing-client.ts, after
const client = new PricingServiceClient(
  'dns:///pricing-svc-headless.default.svc.cluster.local:50051',
  grpc.credentials.createInsecure(),
  { 'grpc.lb_policy': 'round_robin' }
);

With round_robin, the client resolves every backend IP and opens a connection to each one, spreading calls across all of them instead of pinning to whichever pod answered first. As a second layer, pricing-svc now caps how long any single connection can live, so even a client stuck on an old load balancing policy is forced to reconnect and re-resolve periodically instead of holding one connection indefinitely:

pricing-svc/server.ts
const server = new grpc.Server({
  'grpc.max_connection_age_ms': 5 * 60 * 1000,
  'grpc.max_connection_age_grace_ms': 30 * 1000,
});

the aftermath

83% Of all pricing-svc RPCs handled by 2 of 6 pods for six hours
1.8s p99 latency at the point the imbalance became visible
5 min max_connection_age_ms added as a floor on any future skew
3 Other internal gRPC clients found using ClusterIP + pick_first
  • A ClusterIP Service load balances TCP connections, not requests. That distinction only matters for protocols that hold a connection open and reuse it, which is exactly what HTTP/2 and gRPC are built to do.
  • pick_first is a reasonable default for a client with one backend. Against a Service backed by multiple pods, it turns whatever pod answered first during a narrow readiness window into the only pod that matters, indefinitely.
  • Average CPU across a deployment can look completely healthy while two pods carry the entire fleet. Per-pod metrics on anything that does client-side connection reuse are not optional, they're the only view that shows this failure mode at all.
  • A connection age limit on the server is a cheap insurance policy even after fixing the client. It bounds how long any future imbalance, caused by a bug nobody's found yet, can last before the system corrects itself on its own.

Nothing about that rollout failed. Both deployments finished, both health checks passed, and kube-proxy did exactly what a ClusterIP Service is supposed to do. The problem was a forty-second window where "correct" and "evenly distributed" briefly meant two different things, and gRPC's connection reuse turned that window into six hours.

Share this
← All Posts8 min read