July 23, 2026
Machine Preservation on Failure in Gardener
When a node fails in a Kubernetes cluster, the normal response is immediate replacement: Machine Controller Manager (MCM) terminates the failed machine and creates a new one. This is the right default for self-healing clusters, but it leaves operators with a narrow window — or none at all — to investigate what went wrong.