A proposed four-day workshop on Kubernetes investigation
We establish any missing foundations, then ask engineers to justify progressively more of their investigation. The exercises show understanding and remaining support needs, not an automatic on-call qualification.
Day 1
Workload symptoms and release failures
Explain Pod states, container restarts, probes, and rollout behavior. Investigate similar symptoms with different causes before selecting a corrective action.
Hands-on exercises
- Compare an image-startup failure with an application readiness failure.
- Record evidence before a controlled release recovery.
Day 2
Configuration and controller decisions
Connect Helm or Kustomize output to live resources, controllers, and node behavior. Identify which component is responsible when the observed state does not match the request.
Hands-on exercises
- Find a configuration difference behind an unexpected workload change.
- Build a timeline for a missing replica and its replacement.
Day 3
Requests, dependencies, and scheduling
Trace DNS, Services, ingress, policy, and placement through the failed service. Choose observations that separate an unavailable backend from an inaccessible one.
Hands-on exercises
- Locate the failed stage of a client request.
- Explain why a replacement Pod cannot reach a suitable running state.
Day 4
Recovery, state, and incident records
Examine storage, credentials, resource pressure, autoscaling, and permissions during recovery. State what must be observed before declaring the service restored.
Hands-on exercises
- Verify a correction against the original service symptom and its dependencies.
- Have another engineer continue an investigation from the recorded evidence.
Your services, incident questions, and on-call responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.