EKS operations training for platform maintenance

We offer private, instructor-led training for engineers who maintain an inherited or established EKS platform and need to explain how its next change will affect it.

Your team knows how to deploy workloads. The next maintenance task needs a clear model of its dependencies. Node replacement, add-on configuration, traffic, storage, and permissions can affect each other during one operation.

LearnKube brings these mechanisms together in one bounded maintenance rehearsal. Engineers compare a recorded baseline with the changed environment, then note when to stop and which dependency to investigate or recover.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain the change's dependencies through workload, node, and controller relationships, so the team can predict which behavior needs observation.
  • Diagnose a blocked operation using update status, workload events, and integration evidence, so engineers can distinguish capacity, permission, and configuration constraints.
  • Validate a bounded change against an application baseline and explicit stop criteria, so the group can judge whether the intended behavior remains intact.
  • Make the procedure repeatable by recording assumptions, observations, and recovery dependencies, so another engineer can explain and rehearse the same maintenance task.

An EKS platform can use managed node updates and add-ons with separate configuration ownership. These features do not make all maintenance actions the same. Engineers need to connect each operation to Pod placement, traffic, state, and supported configuration paths. A rehearsal helps the team observe these dependencies before accepting a procedure for its environment.

I ask which behavior a maintenance operation must preserve, not only whether its API call succeeds. An application baseline and a stopping condition make the rehearsal useful to the next engineer who performs it.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
What changes during maintenance?Nodes and workload placementMap the affected dependencies
Who maintains this configuration?Managed and self-managed componentsInspect the configuration path
Is application behavior preserved?Baselines and replacement effectsCompare before and after
Can another engineer repeat it?Stop and recovery criteriaExplain the operating procedure

BMW's engineers already operated an advanced, multi-cluster EKS environment provisioned through Terraform. Their stack included a custom Helm chart for microservice deployments, Jenkins, ArgoCD, and Datadog.

They wanted more depth in cluster upgrades, releases, networking, and diagnosis. LearnKube delivered Kubernetes training that covered architecture, networking, EKS, scheduling, and hands-on work.

For a team maintaining an inherited or established EKS environment, we recommend the following workshop on dependencies and controlled change. The proposed agenda follows one maintenance operation through its required observations and stop criteria.

This proposed agenda explores one maintenance question through the course modules. We shorten foundation topics the team already knows and select examples at the required operating depth.

Day 1

We use Deployments, probes, and release strategies to define the behavior maintenance must preserve. Engineers compare a healthy baseline with the result of a successful command.

  • Record application and workload observations before a controlled replacement.
  • Explain a change that completes while its application acceptance check fails.

Day 2

We connect Helm resources, controllers, node groups, and selected add-ons to their configuration sources. The team identifies the controlling component and the API used for the change.

  • Map a maintenance task across workload configuration and platform dependencies.
  • Compare a supported configuration change with a manual edit that another controller manages.

Day 3

We connect the application's network path and scheduling requirements to replacement capacity. Participants explain how different update and scaling operations affect workloads.

  • Investigate a workload that cannot obtain suitable replacement capacity.
  • Trace the traffic observations needed while a component or node changes.

Day 4

We review persistent data, credentials, scaling signals, and permissions for the maintenance scenario. Engineers connect these dependencies to a stop condition and a clear recovery question.

  • Rehearse a bounded change and compare its result with the recorded baseline.
  • Explain the procedure to another engineer, including unresolved assumptions and recovery dependencies.

Your EKS compute model, configuration sources, and maintenance responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's operating work.

When an engineer assumes all capacity changes use the same interruption controls, the instructor can compare the documented operations and their effects on workloads. The group explains why its procedure needs specific observations and stop criteria.

Brian offered this advice after an F5 course that included EKS:

The exercises are worth it, very challenging. They make you think.

— Brian, Product Manager at F5.

Tell us what your team runs, which maintenance task is unclear, and who owns the affected components. We will recommend a course focus and exercises for controlled change.

You can buy LearnKube training through AWS Marketplace. Contact us if you want a private offer for the agreed scope.