Kubernetes on-call troubleshooting training for application engineers

Private, instructor-led training for engineers who deploy Kubernetes applications and need to investigate failures during their operating or on-call responsibilities.

Familiar commands and uneven self-study can leave gaps between a symptom and its cause. Restarting a workload does not explain why service failed or whether recovery is complete.

LearnKube connects workload mechanisms to a sequence of incident questions. Engineers collect evidence, compare explanations, and verify a bounded recovery in a representative environment.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain a service symptom through workload state and dependency behavior, so the team can predict which evidence matters next.
  • Diagnose a likely cause using Pod events, logs, endpoints, and request results, so engineers can distinguish competing explanations.
  • Validate recovery against the original failure condition, so the team can identify incomplete restoration or a reason to stop the change.
  • Make triage repeatable through a concise incident record and investigation sequence, so another engineer can continue with the evidence already collected.

The team already observes failing requests or unhealthy Pods, but similar symptoms can come from different layers. Pod diagnosis and Service diagnosis inspect different relationships. Engineers need to connect the user's symptom to workload state, endpoints, and dependencies before selecting a recovery action and deciding which observations would demonstrate that it worked.

I want the engineer to preserve the clue before changing the system. A restart can restore service, but it can also remove the evidence needed to explain recurrence or decide whether the same action is safe.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
What stopped working?Service and dependency behaviorDefine the observed symptom
Which layer explains it?Events and resource relationshipsCompare competing causes
Has service recovered?Application success conditionsRepeat a representative request
Can someone continue?Evidence and action historyWrite an investigation handoff

Rebellion Defense wanted engineers to progress beyond deploying and interacting with a cluster. Its buyer specifically asked about preparation for on-call triage, with participants bringing different levels of self-study.

LearnKube delivered private Kubernetes training. Nigel described its value as a foundation:

It provided a very good grounding in the topic

— Nigel, Rebellion Defense.

We establish any missing foundations, then ask engineers to justify progressively more of their investigation. The exercises show understanding and remaining support needs, not an automatic on-call qualification.

Day 1

Explain Pod states, container restarts, probes, and rollout behavior. Investigate similar symptoms with different causes before selecting a corrective action.

  • Compare an image-startup failure with an application readiness failure.
  • Record evidence before a controlled release recovery.

Day 2

Connect Helm or Kustomize output to live resources, controllers, and node behavior. Identify which component is responsible when the observed state does not match the request.

  • Find a configuration difference behind an unexpected workload change.
  • Build a timeline for a missing replica and its replacement.

Day 3

Trace DNS, Services, ingress, policy, and placement through the failed service. Choose observations that separate an unavailable backend from an inaccessible one.

  • Locate the failed stage of a client request.
  • Explain why a replacement Pod cannot reach a suitable running state.

Day 4

Examine storage, credentials, resource pressure, autoscaling, and permissions during recovery. State what must be observed before declaring the service restored.

  • Verify a correction against the original service symptom and its dependencies.
  • Have another engineer continue an investigation from the recorded evidence.

Your services, incident questions, and on-call responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer identifies a cause from one status field, the instructor can introduce a second scenario with the same symptom. Comparing the evidence makes the missing diagnostic question visible.

Tell us which workloads your team supports, what it finds difficult to diagnose, and where its on-call responsibility ends. We will recommend investigations and recovery exercises at the appropriate depth.