Kubernetes upgrade planning training for platform engineers

Private, instructor-led training for engineers who maintain existing Kubernetes clusters and need to plan and validate changes to control planes, add-ons, and nodes.

A supported target version does not establish workload compatibility or maintenance capacity. Upgrade order, removed APIs, controller dependencies, and disruption behavior need separate evidence.

LearnKube connects those mechanisms to a staged maintenance exercise. Engineers identify prerequisites, observe workload effects, and define when to continue, stop, or use an available recovery path.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain upgrade constraints through component compatibility and workload dependencies, so the team can identify a supported sequence.
  • Diagnose a blocked maintenance step using API, controller, eviction, and capacity evidence, so engineers distinguish incompatible resources from unavailable placement.
  • Validate a staged change against workload health and stop criteria, so the team can decide whether continuing is justified.
  • Make maintenance repeatable through a version-specific checklist and recovery plan, so another engineer can review the assumptions before execution.

The team already upgrades components, but a supported version combination does not guarantee an uneventful workload change. Version skew rules govern component compatibility, while node draining interacts with eviction rules and available capacity. Engineers must inspect both before planning the sequence, then define workload observations and a recovery path supported by their actual platform.

I want the recovery path named before the version changes. Reverting an application's Deployment is not the same as downgrading a control plane, and those operations must not be treated as interchangeable steps in a runbook.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
Is the sequence supported?Version and API compatibilityReview upgrade prerequisites
Why did draining stop?Evictions and disruption budgetsInspect blocked maintenance
Can replacements run?Capacity and placement constraintsObserve replacement readiness
What justifies continuing?Health and recovery criteriaReview staged change evidence

BMW already operated an EKS platform with networking, identity, delivery, and monitoring integrations. Its engineers wanted deeper understanding of upgrade approaches and other operating decisions without repeating familiar material.

LearnKube delivered private Kubernetes training. Pawel described how he expected the understanding to support daily work:

You will get a superb understanding of k8s mechanism/principles. It will definitely help you during your daily work with k8s...

— Pawel, BMW.

We use the team's current platform and proposed version change to choose representative examples. The workshop develops an operating explanation rather than assuming every platform offers the same rollback mechanism.

Day 1

Explain probes, replicas, and release behavior before defining maintenance checks. Distinguish workload availability from a successful infrastructure operation.

  • Establish application observations before a controlled change.
  • Identify an upgrade-sensitive API or workload dependency.

Day 2

Connect rendered configuration, control-plane components, add-ons, and node versions. Review a supported sequence and separate an in-place change from a replacement approach.

  • Build a compatibility inventory for the sample environment.
  • Identify stop conditions and a platform-supported recovery option.

Day 3

Examine eviction, disruption budgets, networking continuity, and placement capacity. Investigate why maintenance can block even when remaining nodes look healthy.

  • Rehearse a bounded node drain and inspect blocking evidence.
  • Verify application traffic while replacement Pods obtain capacity.

Day 4

Review storage, credentials, autoscaling, and permissions affected by the change. Compare post-change behavior with the baseline before accepting the maintenance step.

  • Validate a stateful or identity-dependent application after replacement.
  • Produce a staged maintenance record with continue, stop, and recovery criteria.

Your versions, upgrade questions, and component responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer writes “roll back” without naming the operation, the instructor can separate application, node, and control-plane recovery. The team then checks which option the selected platform actually supports.

Tell us which components your team upgrades, which failure paths it cannot confidently explain, and which maintenance duties it owns. We will recommend a focused workshop and validation exercises.