A recommended four-day workshop for escalation teams
This proposed agenda uses representative incidents and explicit investigation boundaries. Each exercise produces evidence that another engineer can follow.
Day 1
Workload state and failure evidence
We examine containers, controllers, probes, and rollout behavior. You will distinguish process failure, missing readiness, and an incomplete deployment.
Hands-on exercises
- Build an incident timeline from a failing workload.
- Collect the evidence needed to distinguish two plausible causes.
Day 2
Architecture, configuration, and reproduction
You will compare Helm and Kustomize configuration with the live resources. We trace controller, control-plane, and node responsibilities through the incident.
Hands-on exercises
- Identify a configuration difference that changes the failure.
- Reproduce a bounded fault and record what the result establishes.
Day 3
Network and placement investigations
We inspect DNS, Services, ingress, network policies, service mesh interactions, and scheduling constraints. The focus is a hypothesis-driven sequence rather than an indiscriminate diagnostic dump.
Hands-on exercises
- Distinguish a failed lookup from a blocked application connection.
- Investigate a Pending workload and identify the relevant capacity or policy owner.
Day 4
Storage, access, metrics, and escalation
You will connect storage and credential lifecycles to application behavior. We review metrics, autoscaling, authentication, and RBAC as evidence for a mitigation or escalation.
Hands-on exercises
- Trace a storage-related failure from claim status to workload symptoms.
- Prepare a report with reproduction conditions, observations, and the next responsible team.
Your customer incident classes, access constraints, and engineering boundaries can shape the agenda. Get in touch to tailor the workshop to your escalation work.