I've been burned by this exact pattern twice, so let me write down what actually happened, because I think the failure mode is more subtle than "don't check dependencies".
The first time was a payments service. Readiness hit /healthz, and /healthz did three things: pinged Postgres, pinged Redis, and called an internal fraud-scoring API. All reasonable on paper. Then the fraud-scoring team deployed a bad build that added ~4s latency. Our probe timeout was 1s.
Within about 40 seconds every single pod in the deployment went unready. The Service had zero endpoints. The ingress started returning 503 for everything — including endpoints that never touched the fraud API, like the receipts page and the webhook receiver.
So a latency regression in a non-critical dependency of one code path became a full outage of a service with eleven other code paths.
The second time was worse because it was during a rollout:
readinessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3
strategy:
rollingUpdate:
maxUnavailable: 0
maxSurge: 25%
New pods couldn't become ready because Redis was briefly overloaded. With maxUnavailable: 0 the rollout just sat there. That part is fine — that's the rollout protecting you. But then someone "fixed" it by bumping maxUnavailable to 50% to push the release through, and now half the old pods were gone too.
What we do now:
1. Readiness checks only things the pod itself owns: has it loaded config, is the cache warm, is the listener up, is it draining.
2. Dependency health is exported as metrics and alerts. It never decides routing.
3. Each endpoint degrades on its own. If fraud scoring is down, checkout returns a specific error or queues; receipts keep working.
4. Startup probes cover slow init so we don't have to inflate readiness thresholds.
The rule we wrote on the wiki is roughly: "readiness answers whether *this pod* can serve, not whether *the world* is healthy". If every replica would fail the check at the same moment for the same reason, it's not a readiness signal — it's a global outage detector, and Kubernetes is the wrong tool to act on it.
There's one exception I'd still defend: a hard local dependency like a sidecar database proxy that lives in the same pod. If the proxy is dead, this pod genuinely can't serve. That's local, so it belongs in readiness.