Kubernetes Karpenter capacity training for platform engineers

Private, instructor-led training for engineers who operate Kubernetes node capacity with Karpenter and need to understand its provisioning and consolidation decisions.

An instance price or utilization chart does not explain the full operating trade-off. Workload requests, placement constraints, disruption controls, and replacement capacity shape which changes are possible.

LearnKube follows a representative capacity decision from workload demand to node behavior. Engineers inspect constraints and compare application effects before accepting a proposed efficiency change.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain capacity decisions through workload constraints and NodePool configuration, so engineers can predict which node options remain eligible.
  • Diagnose blocked provisioning or consolidation using controller events and placement evidence, so the team can identify the actual limiting condition.
  • Validate a capacity change against workload availability and disruption criteria, so a cheaper configuration is not accepted without examining its operating effects.
  • Make capacity reviews repeatable through a recorded decision and comparable workload observations, so another engineer can assess the same trade-offs.

The team already manages node supply, but Karpenter combines workload requirements with NodePool constraints before selecting capacity. Its disruption mechanisms also consider replacement and workload constraints during some changes. Engineers must identify the specific mechanism involved before interpreting an idle node, blocked consolidation, or lower instance price as evidence of a better operating configuration.

I want to know what must move before calling a node unnecessary. A low utilization figure does not explain whether the remaining workloads fit elsewhere or whether their disruption constraints allow the proposed consolidation.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
Which capacity is eligible?NodePool and Pod requirementsCompare intersecting constraints
Why was no node created?Provisioning and provider conditionsInspect controller events
Why is consolidation blocked?Placement and disruption controlsTrace a candidate decision
Is the change acceptable?Cost and workload effectsEvaluate comparable operating evidence

Warner Bros. Discovery's cluster engineering work includes Karpenter tuning, node selection, bin-packing, and staged infrastructure changes. Its operating responsibilities connect efficiency with workload effects and change validation.

For that work, we recommend a four-day workshop on capacity decisions and controlled experiments, rather than a promise of measured savings from attendance.

Participants need Kubernetes scheduling and resource foundations. We adapt capacity examples to the selected Karpenter version and provider configuration, with one workload group as the reference.

Day 1

Explain requests, readiness, replicas, and termination behavior before reviewing capacity. Establish the service conditions that a node change needs to preserve.

  • Record workload constraints and application observations before the experiment.
  • Compare nominal resource use with the capacity required for placement.

Day 2

Connect templates, NodePools, provider resources, and controller status. Explain why the desired node options can differ from what infrastructure can currently supply.

  • Trace a bounded provisioning decision from pending workload to node state.
  • Diagnose an incompatible requirement or provider dependency.

Day 3

Examine topology, affinities, taints, eviction, and replacement capacity. Distinguish consolidation, drift, expiration, and interruptions rather than assuming one disruption policy covers all of them.

  • Inspect why a proposed consolidation can or cannot move its workloads.
  • Observe a controlled replacement and compare application behavior.

Day 4

Connect storage, startup, autoscaling, and permissions to the capacity change. Record service effects alongside the efficiency hypothesis and its recovery conditions.

  • Repeat a bounded capacity adjustment under comparable workload conditions.
  • Document the constraints, observations, and reasons to accept or reject the proposal.

Your NodePools, capacity questions, and operating responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer identifies an apparently idle node, the instructor can examine what prevents its workloads from moving. In feedback on a separate course, Ross identified the element he valued:

Hands-on labs

— Ross, Staff SRE at Matillion.

Tell us which Karpenter configuration you operate, which provisioning or consolidation behavior is unclear, and which workloads your team supports. We will recommend a focused comparison and its operating checks.