← all discussions

Are Kubernetes readiness probes being asked to do too much?

ararunv4 hours ago7 replies

I've been reading the probe documentation again.

A lot of configurations seem to treat readiness as a generic "is the application healthy?" signal. But readiness really answers something narrower: should this pod receive traffic?

I'm wondering whether combining dependency health into readiness actually creates more failure modes than it prevents.

Discussion

mhmhoffman3 hours ago

Database connectivity is the common example.

If every pod marks itself unready because the database has a temporary problem, you've turned a database outage into an application outage too.

Reply
sasamir_k2 hours ago

But if requests require the database anyway, what do you gain by continuing to send traffic?

Reply
mhmhoffman2 hours ago

Depends on the application. Not every endpoint necessarily requires the database.

More importantly, readiness affects rollout behavior. That's where it gets interesting.

Reply
kckchen48 minutes ago

The docs actually warn about part of this. Readiness controls whether the endpoint stays in matching Services. It isn't intended as a general dependency health mechanism.

Reply
nunullroute20 minutes ago

Fair. I'd still check a local cache warmup in readiness, just not remote deps.

Reply
opops_marguerite35 minutes ago

I've been burned by this exact pattern twice, so let me write down what actually happened, because I think the failure mode is more subtle than "don't check dependencies".

The first time was a payments service. Readiness hit /healthz, and /healthz did three things: pinged Postgres, pinged Redis, and called an internal fraud-scoring API. All reasonable on paper. Then the fraud-scoring team deployed a bad build that added ~4s latency. Our probe timeout was 1s.

Within about 40 seconds every single pod in the deployment went unready. The Service had zero endpoints. The ingress started returning 503 for everything — including endpoints that never touched the fraud API, like the receipts page and the webhook receiver.

So a latency regression in a non-critical dependency of one code path became a full outage of a service with eleven other code paths.

The second time was worse because it was during a rollout:

readinessProbe:
  httpGet:
    path: /healthz
    port: 8080
  periodSeconds: 5
  timeoutSeconds: 1
  failureThreshold: 3
strategy:
  rollingUpdate:
    maxUnavailable: 0
    maxSurge: 25%

New pods couldn't become ready because Redis was briefly overloaded. With maxUnavailable: 0 the rollout just sat there. That part is fine — that's the rollout protecting you. But then someone "fixed" it by bumping maxUnavailable to 50% to push the release through, and now half the old pods were gone too.

What we do now:

1. Readiness checks only things the pod itself owns: has it loaded config, is the cache warm, is the listener up, is it draining. 2. Dependency health is exported as metrics and alerts. It never decides routing. 3. Each endpoint degrades on its own. If fraud scoring is down, checkout returns a specific error or queues; receipts keep working. 4. Startup probes cover slow init so we don't have to inflate readiness thresholds.

The rule we wrote on the wiki is roughly: "readiness answers whether *this pod* can serve, not whether *the world* is healthy". If every replica would fail the check at the same moment for the same reason, it's not a readiness signal — it's a global outage detector, and Kubernetes is the wrong tool to act on it.

There's one exception I'd still defend: a hard local dependency like a sidecar database proxy that lives in the same pod. If the proxy is dead, this pod genuinely can't serve. That's local, so it belongs in readiness.

Reply
tltlin12 minutes ago

Small addition: readiness also affects PodDisruptionBudgets during node drains. Unready pods already count as disrupted, so a flapping readiness check can block cluster upgrades in ways that are very confusing at 2am.

Took us an afternoon to figure out why kubectl drain was hanging forever.

Reply