Перейти к содержимому

Recovering from CAPI Disasters in Day 2 Operations

Cloud Native Lima

0:00 / 0:00

Recovering from CAPI Disasters in Day 2 Operations

32 просмотра · 2 недели назад
Cloud Native Lima
308 подписчиков
32 просмотра · 2 недели назад
It's Friday at 4:47 PM. Your workload clusters are healthy, pods are running smoothly, and users are happy. Just as you're about to call it a day, an alert pops up: your Cluster API management cluster isn't responding, and it must be back up by Monday. What do you do now? This talk explores what actually happens when a Cluster API management cluster fails in production. We'll begin by examining how CAPI manages the cluster lifecycle through reconciliation, then escalate through real-world failure modes — certificate expiry, lost kubeconfigs, and even a complete management cluster loss. The second half walks through a real-world disaster recovery scenario using Velero: backing up CAPI objects, restoring with the correct order and pause semantics, and finally how controllers re-adopt existing workload clusters without disruption. You'll leave with a practical recovery playbook and a clearer understanding of what production disaster recovery on Kubernetes truly demands.