Day-1 Kubernetes feels like progress. Day-2 is where clusters age: versions lag, node images drift, certs expire, and “temporary” privilege escalations become permanent.
Schedule upgrades as a product. Keep add-ons version-compatible, test in a non-prod twin, and document rollback. Watch disk pressure, DNS, and etcd/API latency as leading indicators.
Enforce resource requests/limits and pod disruption budgets so one noisy workload cannot take down a node pool. Policy-as-code stops privileged containers from creeping back in.
A production Kubernetes base includes the ops story—not only the YAML that first passed a demo.
Separate platform add-ons from product workloads so an ingress or DNS failure does not require app owners to debug everything.
Watch for noisy neighbours: unlimited bursts without quotas create mysterious latency incidents.
Practice node replacement and PodDisruptionBudgets before a real zone event forces the lesson.
Key takeaways
- Upgrades, certificates, etcd/backup (or control-plane DR), and capacity are the usual day-2 killers.
- Standardize add-on versions and ownership or clusters drift into pets.
- Cost and reliability both need resource requests/limits that reflect reality.
FAQ
How often should we upgrade clusters?
Follow your provider’s support window with a planned cadence. Skipping versions increases risk more than frequent, rehearsed upgrades.
What should be in a cluster runbook?
Upgrade steps, certificate renewal, node failure response, admission webhook outages, and how to drain or restore critical namespaces.