docker

Why do so many Dockerfiles invalidate the build cache unnecessarily?

The classic mistake is COPY . . before installing dependencies. Any source change re-runs the whole install step.

coldboot2 days ago4 replies▲ 36
sandboxing

Do AI coding agents require a different isolation model?

Coding agents run arbitrary commands, install packages and hit the network. A normal container shares the host kernel. Is that enough boundary for code nobody reviewed?

rohitm3 days ago6 replies▲ 58
docker

When should you avoid multi-stage builds?

Multi-stage builds are the default advice now. Are there cases where a single stage is actually better?

emilyk4 days ago2 replies▲ 22
platform

Should applications know they're running on Kubernetes?

Some apps call the Kubernetes API directly for config or leader election. Others stay completely unaware. Which ages better?

danielc5 days ago3 replies▲ 29
platform

How much should platform teams hide from developers?

Golden paths hide YAML and cluster details. Great until something breaks and developers can't debug what they can't see.

stefanr1 week ago4 replies▲ 47
supply-chain

What does "distroless" actually buy you?

Distroless images have no shell and no package manager. Beyond a smaller CVE list, what do they really give you?

packetdrop1 week ago3 replies▲ 33
kubernetes

Is YAML complexity really Kubernetes' problem?

People blame YAML, but the complexity is in the API surface. Swap YAML for anything else and you still have the same hundreds of fields.

tomh2 weeks ago4 replies▲ 41
sre

PodDisruptionBudgets protect you from voluntary disruptions only — and that's the trap

Re-read the disruptions page this morning. The doc is very clear that PDBs only apply to *voluntary* disruptions — drains, evictions via the API. Yet in almost every postmortem I've seen, someone says "but we had a PDB" about a node that died. How do you communicate this to teams that treat PDBs as an availability guarantee?

sre_lena2 hours ago3 replies▲ 58
sre

Do you alert on probe failures or on user-facing SLOs?

Our on-call gets paged when liveness probes fail more than N times. Half of those pages have zero user impact. Considering moving everything to SLO burn-rate alerts. Anyone done this migration?

pager_dev9 hours ago1 replies▲ 33
linux

cgroup v2 memory.high vs memory.max — is Kubernetes using the right one?

The cgroups page now assumes v2 everywhere. Under v2 there's memory.high (throttle/reclaim) and memory.max (OOM kill). Pod memory limits map to memory.max. MemoryQoS was supposed to use memory.high to give a soft ceiling before the kill. Where does that stand, and does anyone run with it enabled?

coldboot5 hours ago4 replies▲ 72
linux

User namespaces for pods: anyone running hostUsers: false in production?

Setting hostUsers: false maps root inside the container to an unprivileged UID on the host. On paper it removes a whole class of container escape impact. What breaks?

nullroute1 day ago2 replies▲ 44
networking

Is it time to stop writing new Ingress objects?

The Ingress page now points people to Gateway API, and the Ingress API is frozen. For greenfield clusters, is there any reason left to choose Ingress?

netops_priya6 hours ago3 replies▲ 66
networking

Bridge vs host networking for latency-sensitive containers

Benchmarked a UDP service under Docker's default bridge vs --network host. Host mode cut p99 by ~40µs. Is that the NAT/iptables path, or veth overhead?

packet_tom12 hours ago1 replies▲ 29
supply-chain

Docker build attestations: is anyone actually verifying provenance at deploy time?

BuildKit can attach SBOM and SLSA provenance attestations with --sbom and --provenance. Generating them is easy. I've yet to meet a team that rejects deploys based on them. Who's closing the loop?

secops_jules3 hours ago3 replies▲ 61
supply-chain

Base image freshness policies — how old is too old?

Docker Scout has an "up-to-date base images" policy. What threshold do you use before forcing a rebuild?

arunv2 days ago1 replies▲ 24
observability

kube-state-metrics, metrics-server, cAdvisor — which one answers which question?

New team members get confused by three overlapping metric sources. Here's how I explain it: metrics-server is for autoscaling only, cAdvisor is container resource usage, kube-state-metrics is object state. Is that accurate enough?

obs_maya7 hours ago2 replies▲ 41
containers

Rootless Docker in CI: worth the friction?

Moved our self-hosted runners to rootless mode. Builds work, but port binding below 1024 and some overlay storage quirks bit us. Curious if others stuck with it.

devtools_amy10 hours ago2 replies▲ 37
devops

Compose Watch replaced our bind-mount dev setup

We used bind mounts for local dev and fought file permission and performance issues on macOS for years. Compose's develop.watch with sync and rebuild actions fixed most of it. Anyone hit limitations?

devtools_amy1 day ago2 replies▲ 48
cloud-native

In-place pod resize changes how we think about vertical scaling

Resizing CPU/memory without restarting the pod used to be impossible. Now it's documented as a regular task. Does this make VPA in auto mode viable for stateful workloads?

kchen14 hours ago1 replies▲ 55
ai-infra

Running local LLMs with Docker Model Runner — where does it fit?

Docker Model Runner lets you pull and run models as OCI artifacts. For local dev it's convenient. Is anyone using the same artifacts in a cluster, or is this strictly a laptop tool?

ml_ops_dev4 hours ago2 replies▲ 52
agents

MCP Gateway: one place to put policy for agent tool calls?

Docker's MCP Gateway runs MCP servers in containers behind a single endpoint. The appeal for us is a single choke point for logging and secret injection. Does this become the 'API gateway' for agents?

agentsmith_r8 hours ago2 replies▲ 46