Kubernetes training for customer production escalation engineers

We offer private, instructor-led training for escalation engineers investigating production-blocking Kubernetes incidents in customer environments.

Your team knows the product and support process. An urgent symptom can cross several technical owners, while access and evidence remain incomplete.

LearnKube uses controlled workload and infrastructure failures to connect observations to hypotheses, permitted investigations, mitigation decisions, and actionable engineering escalations.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Lead a bounded investigation by defining the failing behavior and collecting relevant status, so the customer understands what the team is trying to establish.
  • Preserve recovery options by distinguishing evidence collection from disruptive changes, so an attempted mitigation does not erase the information needed for diagnosis.
  • Isolate a likely cause by correlating workload, network, storage, and permission evidence, so the team can identify the responsible layer.
  • Prepare a useful escalation by recording observations, hypotheses, and reproduction conditions, so product or platform engineers can continue the investigation.

Kubernetes Pod status and events describe different parts of execution. A storage claim adds another lifecycle involving a provisioner and underlying storage. Customer incidents can cross these boundaries. Engineers need to correlate the observed states and changes before assigning the problem to the application, infrastructure, or product implementation.

I ask what the evidence rules out before accepting a root-cause statement. A restart that removes the symptom can also remove the best clue, so the investigation needs a record before the next change.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Customer task or requirementKubernetes decisionPractice
Define the incident boundaryWorkload and dependency stateBuild a symptom timeline
Gather useful evidenceEvents, logs, and permissionsSelect the next observation
Compare possible causesApplication and infrastructure relationshipsChallenge an initial hypothesis
Escalate the investigationReproduction and ownershipPrepare an engineering report

Portworx escalation engineers work directly with enterprise customers on critical container-storage incidents. Their responsibilities include root-cause investigation, escalation to engineering, and coaching front-line technical services engineers.

For teams handling that work, we recommend a four-day workshop on Kubernetes evidence and failure relationships. Product-source debugging remains a separate specialist skill.

This proposed agenda uses representative incidents and explicit investigation boundaries. Each exercise produces evidence that another engineer can follow.

Day 1

We examine containers, controllers, probes, and rollout behavior. You will distinguish process failure, missing readiness, and an incomplete deployment.

  • Build an incident timeline from a failing workload.
  • Collect the evidence needed to distinguish two plausible causes.

Day 2

You will compare Helm and Kustomize configuration with the live resources. We trace controller, control-plane, and node responsibilities through the incident.

  • Identify a configuration difference that changes the failure.
  • Reproduce a bounded fault and record what the result establishes.

Day 3

We inspect DNS, Services, ingress, network policies, service mesh interactions, and scheduling constraints. The focus is a hypothesis-driven sequence rather than an indiscriminate diagnostic dump.

  • Distinguish a failed lookup from a blocked application connection.
  • Investigate a Pending workload and identify the relevant capacity or policy owner.

Day 4

You will connect storage and credential lifecycles to application behavior. We review metrics, autoscaling, authentication, and RBAC as evidence for a mitigation or escalation.

  • Trace a storage-related failure from claim status to workload symptoms.
  • Prepare a report with reproduction conditions, observations, and the next responsible team.

Your customer incident classes, access constraints, and engineering boundaries can shape the agenda. Get in touch to tailor the workshop to your escalation work.

When an engineer attributes a restart to storage without evidence, the instructor can compare events, termination details, and volume state. The group can refine the hypothesis before recommending a customer change.

Course was great, but the labs are top notch. They make you think about the problem and how dot ifx.

— Jason, F5.

Tell us which customer failures reach your engineers, what evidence they can access, and where they escalate. We will recommend investigations and technical explanations for those boundaries.