A proposed four-day workshop on recovery mechanisms
We adapt the architecture and failure exercises to one defined workload and failure model. Each experiment states its conditions so results are not treated as universal timing guarantees.
Day 1
Workload health and recovery definitions
Explain probes, controller ownership, and the difference between replacement and application availability. Define what the sample service must do before recovery is considered complete.
Hands-on exercises
- Establish workload and client observations before the failure.
- Compare a container restart with a missing-node scenario.
Day 2
Detection, eviction, and reconciliation
Connect controller configuration, node conditions, tolerations, and workload replacement. Inspect the sequence instead of attributing every delay to the same timer.
Hands-on exercises
- Build a timeline for a controlled node interruption.
- Identify the conditions governing the next replacement action.
Day 3
Placement and the returning traffic path
Examine capacity, affinity, discovery, endpoint updates, and routing during recovery. Determine why a new Pod can exist without restoring the original service path.
Hands-on exercises
- Diagnose a replacement blocked by placement constraints.
- Compare Pod readiness with client-visible recovery.
Day 4
State, permissions, and repeatable experiments
Review storage access, credential dependencies, and the permissions needed for diagnosis. Change one bounded assumption and compare the complete recovery sequence.
Hands-on exercises
- Investigate a storage dependency that delays a replacement workload.
- Repeat the experiment and document its conditions and limitations.
Your workloads, failure assumptions, and recovery responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.