Kubernetes onboarding for SREs entering supported on-call work

We offer private, instructor-led training for SREs preparing to contribute to Kubernetes investigations and on-call work with support from their team.

Learners can already recognize workload resources and perform basic deployments. A production symptom requires them to choose evidence, explain uncertainty, and identify when another engineer needs to act.

LearnKube connects familiar commands to a bounded service investigation. Guided examples progress toward an explanation of the likely cause, next check, and escalation boundary.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Orient to the operating task by connecting known workloads to their signals and dependencies, so the learner understands which part of the investigation they can undertake.
  • Practice a bounded investigation with guidance on observations and competing explanations, so the learner can explain why the next check is useful.
  • Demonstrate a reasoned handoff through an evidence summary and an escalation decision, so the team can identify what the learner understands and where support remains necessary.
  • Retain the investigation method through annotated practice and relevant references, so a later shift or colleague can follow the reasoning behind the steps.

An SRE can know workload commands without knowing which signal explains an unavailable service. Pod state and events describe different problems from Service selection and endpoints. Before taking a supported action, the learner needs to connect those observations. The course uses a bounded failure to practice an evidence summary, a next check, and a clear escalation boundary.

I want a new on-call engineer to explain what the evidence rules out before they propose a fix. The next check and escalation boundary tell me more about their understanding than a memorized restart command.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Role milestoneRequired knowledgePractice/check
Locate the affected workloadResources and dependency pathsDefine the investigation scope
Choose useful evidenceState, events, and endpointsExplain the next check
Make a supported recommendationCompeting causes and permissionsPresent a reasoned handoff
Prepare for later repetitionRunbook conditions and referencesRecord what remains uncertain

CrowdStrike's Temporal Platform team operates stateful infrastructure on Kubernetes. Its learning path combines existing Kubernetes fundamentals with upgrades, observability, runbooks, and on-call participation under the guidance of experienced engineers.

Responsibility develops toward component ownership and helping new colleagues. For engineers preparing for the Kubernetes part of this work, we recommend a four-day workshop built around supported tasks and explanations of the evidence.

This proposed agenda begins with the learner's current operating knowledge. It develops a bounded investigation and handoff, with further production responsibilities determined by the team's own support and readiness process.

Day 1

We review images, workloads, Services, and probes through a guided application change. Learners compare the resource state with an actual request before explaining which observation needs attention.

  • Inspect a supplied release and explain the difference between running, ready, and reachable.
  • Investigate an unready version with guidance and summarize the evidence for a possible next action.

Day 2

We use Helm and compare Kustomize to locate the configuration behind the workload. Architecture examples connect controllers, API resources, and nodes to the observations a new on-call engineer can collect.

  • Trace an unexpected workload setting back to its configuration source.
  • Explain a replacement Pod and identify which component's evidence would help continue the investigation.

Day 3

We trace DNS, Services, ingress, network policy, and placement through the sample. Service mesh examples identify when an extra component changes the evidence or the responsible team.

  • Choose observations for a supplied connectivity fault and explain why they distinguish possible causes.
  • Present an evidence summary and identify a question that requires a platform or application specialist.

Day 4

We add state, secrets, autoscaling signals, and access to the scenario. Learners practice a limited role task, explain uncertainty, and record which concepts or permissions require further support.

  • Investigate a data, resource, or access symptom without assuming permission to change every component.
  • Write a short handoff covering the observed state, proposed next check, and remaining support needs.

Your learners' Kubernetes foundations, on-call responsibilities, and supported practice arrangements can shape the agenda. Get in touch to tailor the learning path to your team.

When a learner proposes a restart after seeing an unavailable service, the instructor can ask which cause the current evidence supports. A targeted selector or readiness example helps the learner explain a check before proposing an action.

David described what helped him check his understanding during training:

The combination of presentation with exercises to practice, so we can verify we understood the material correctly.

— David, Chief Software Engineer at L3Harris.

Tell us what the learners already understand, which investigations they will join, and how the team will support further practice. We will recommend a course focus and bounded checks for that next responsibility.