Are Kubernetes readiness probes being asked to do too much?
I've been reading the probe documentation again. A lot of configurations seem to treat readiness as a generic "is the application healthy?" signal. But readiness really answers something narrower: should this pod receive traffic?
Most discussed
dockerWhat exactly belongs inside a production container?
The interesting part of native sidecars isn't startup
When does a smaller image actually improve security?
Are Kubernetes resource limits still misunderstood?
Why do so many Dockerfiles invalidate the build cache unnecessarily?
The classic mistake is COPY . . before installing dependencies. Any source change re-runs the whole install step.
Do AI coding agents require a different isolation model?
Coding agents run arbitrary commands, install packages and hit the network. A normal container shares the host kernel. Is that enough boundary for code nobody reviewed?
When should you avoid multi-stage builds?
Multi-stage builds are the default advice now. Are there cases where a single stage is actually better?
Should applications know they're running on Kubernetes?
Some apps call the Kubernetes API directly for config or leader election. Others stay completely unaware. Which ages better?
How much should platform teams hide from developers?
Golden paths hide YAML and cluster details. Great until something breaks and developers can't debug what they can't see.
What does "distroless" actually buy you?
Distroless images have no shell and no package manager. Beyond a smaller CVE list, what do they really give you?
Is YAML complexity really Kubernetes' problem?
People blame YAML, but the complexity is in the API surface. Swap YAML for anything else and you still have the same hundreds of fields.
PodDisruptionBudgets protect you from voluntary disruptions only — and that's the trap
Re-read the disruptions page this morning. The doc is very clear that PDBs only apply to *voluntary* disruptions — drains, evictions via the API. Yet in almost every postmortem I've seen, someone says "but we had a PDB" about a node that died. How do you communicate this to teams that treat PDBs as an availability guarantee?
Do you alert on probe failures or on user-facing SLOs?
Our on-call gets paged when liveness probes fail more than N times. Half of those pages have zero user impact. Considering moving everything to SLO burn-rate alerts. Anyone done this migration?
cgroup v2 memory.high vs memory.max — is Kubernetes using the right one?
The cgroups page now assumes v2 everywhere. Under v2 there's memory.high (throttle/reclaim) and memory.max (OOM kill). Pod memory limits map to memory.max. MemoryQoS was supposed to use memory.high to give a soft ceiling before the kill. Where does that stand, and does anyone run with it enabled?
User namespaces for pods: anyone running hostUsers: false in production?
Setting hostUsers: false maps root inside the container to an unprivileged UID on the host. On paper it removes a whole class of container escape impact. What breaks?
Is it time to stop writing new Ingress objects?
The Ingress page now points people to Gateway API, and the Ingress API is frozen. For greenfield clusters, is there any reason left to choose Ingress?
Bridge vs host networking for latency-sensitive containers
Benchmarked a UDP service under Docker's default bridge vs --network host. Host mode cut p99 by ~40µs. Is that the NAT/iptables path, or veth overhead?
Docker build attestations: is anyone actually verifying provenance at deploy time?
BuildKit can attach SBOM and SLSA provenance attestations with --sbom and --provenance. Generating them is easy. I've yet to meet a team that rejects deploys based on them. Who's closing the loop?
Base image freshness policies — how old is too old?
Docker Scout has an "up-to-date base images" policy. What threshold do you use before forcing a rebuild?
kube-state-metrics, metrics-server, cAdvisor — which one answers which question?
New team members get confused by three overlapping metric sources. Here's how I explain it: metrics-server is for autoscaling only, cAdvisor is container resource usage, kube-state-metrics is object state. Is that accurate enough?
Rootless Docker in CI: worth the friction?
Moved our self-hosted runners to rootless mode. Builds work, but port binding below 1024 and some overlay storage quirks bit us. Curious if others stuck with it.
Compose Watch replaced our bind-mount dev setup
We used bind mounts for local dev and fought file permission and performance issues on macOS for years. Compose's develop.watch with sync and rebuild actions fixed most of it. Anyone hit limitations?
In-place pod resize changes how we think about vertical scaling
Resizing CPU/memory without restarting the pod used to be impossible. Now it's documented as a regular task. Does this make VPA in auto mode viable for stateful workloads?
Running local LLMs with Docker Model Runner — where does it fit?
Docker Model Runner lets you pull and run models as OCI artifacts. For local dev it's convenient. Is anyone using the same artifacts in a cluster, or is this strictly a laptop tool?
MCP Gateway: one place to put policy for agent tool calls?
Docker's MCP Gateway runs MCP servers in containers behind a single endpoint. The appeal for us is a single choke point for logging and secret injection. Does this become the 'API gateway' for agents?
