Kubernetes autoscaling training for application and platform engineers

Private, instructor-led training for engineers whose existing Kubernetes workloads must respond predictably to changing demand.

A replica limit does not explain when usable capacity arrives. Metric selection, controller timing, startup behavior, and downstream saturation can make more replicas ineffective or harmful.

LearnKube follows a bounded load change through the scaling loop. Engineers inspect the signal, controller decision, workload response, and dependency behavior before proposing a correction.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain replica changes through metrics, controller behavior, and readiness, so engineers can predict why capacity responds late or unevenly.
  • Diagnose ineffective scaling using workload and dependency observations, so the team can distinguish missing signals from downstream saturation.
  • Validate a tuning change under a repeatable load pattern, so engineers can judge service behavior and define when to stop the experiment.
  • Make scaling reviews repeatable through a recorded signal-to-response analysis, so another engineer can evaluate the same assumptions.

The team already configures autoscaling, but HPA decisions depend on the supplied metrics and workload state. KEDA activation adds another decision when workloads scale from zero. Startup delays and downstream limits still affect service throughput. Engineers must connect the selected signal to usable capacity before interpreting oscillation or increasing the maximum replica count.

I ask whether the dependency can accept the extra work before increasing replicas. More consumers can amplify pressure on a saturated service, so a growing backlog does not by itself prove that additional workers will help.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
Is the signal useful?Metrics and demand relationshipCompare load with measurements
When does capacity arrive?Activation, scheduling, and readinessTrace the scaling timeline
Why is throughput unchanged?Dependency capacity and backpressureObserve a saturated backend
Is tuning an improvement?Repeatable service observationsCompare controlled load runs

Vapi's Kubernetes operating responsibilities include custom-metric scaling, backpressure, load testing, and production failure diagnosis. Those tasks connect replica decisions to the service and its dependencies.

For that work, we recommend a four-day workshop on scaling mechanisms and controlled evidence. Its purpose is to explain and evaluate behavior rather than promise a particular latency or capacity result.

Participants need basic workload and metrics knowledge. We adapt the core autoscaling module to the actual HPA or KEDA configuration and a representative load pattern.

Day 1

Connect startup, probes, resource settings, and request handling. Identify when a new replica becomes useful to the service rather than merely existing in the API.

  • Establish a baseline linking client demand to workload behavior.
  • Observe a replica that starts but is not yet ready to accept work.

Day 2

Explain desired replicas, metric sources, controller intervals, and any activation mechanism. Inspect the configuration that actually controls the workload.

  • Trace one scale decision from the supplied metric to desired replicas.
  • Diagnose a missing or unsuitable metric without changing several controls at once.

Day 3

Examine node capacity, traffic distribution, dependency connections, and backpressure. Distinguish an unscheduled replica from one that adds load without adding throughput.

  • Introduce a capacity constraint and inspect the scaling timeline.
  • Observe a downstream bottleneck under increased worker demand.

Day 4

Review stabilization, scale-down effects, persistent work, and permissions. Compare a bounded tuning change with the original load and recovery conditions.

  • Evaluate one change using the same service and dependency observations.
  • Record success, stop, and recovery criteria for a future tuning exercise.

Your scaling signals, demand patterns, and workload responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer raises the replica limit, the instructor can inspect the limiting dependency and expected service response. Mark described the value of examining assumptions in a separate Octopus course:

The questions being asked to confirm our knowledge on matters, especially with gotchas and corner cases.

— Mark, Platform Engineer at Octopus Energy.

Tell us which workloads and autoscalers you run, which demand patterns produce uncertainty, and which dependencies your team owns. We will recommend focused mechanisms and load experiments.