From managed to on-prem Kubernetes: training for platform teams

We offer private, instructor-led training for platform engineers who use a cloud-managed cluster and need to operate Kubernetes on-prem.

Your team knows workload resources and the existing service's controls. The provider's operational work now needs an owner, including control-plane recovery, infrastructure integrations, and node maintenance.

LearnKube connects those responsibilities through a representative workload and a controlled cluster environment. Your engineers will trace failures across the application, control plane, network, and storage.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Operate the destination cluster by identifying control-plane components and infrastructure dependencies, so the team can explain what keeps the Kubernetes API available.
  • Preserve recoverable state by separating cluster-state backups from application-data recovery, so a restoration plan covers both resources and their contents.
  • Diagnose unavailable services by tracing API health, node status, load balancing, and storage, so engineers can identify the failed component instead of only restarting workloads.
  • Assign operational duties by documenting certificates, upgrades, capacity, and recovery procedures, so work previously handled by the provider has a named owner.

On a managed service, the provider performs some cluster operations and supplies integrations. On-prem, the team must identify how the API, nodes, networking, and storage remain available. A LoadBalancer Service still needs an implementation. Recovery also requires etcd backups for Kubernetes state and a separate plan for application data.

A backup file is not proof that a team can recover a cluster. I want the recovery exercise to distinguish the state restored from etcd from the application data that needs its own restoration path.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Current task or requirementNew concept or decisionPractice
Maintain API availabilityControl-plane and etcd healthInvestigate a failed component
Expose application trafficOn-prem load-balancer integrationTrace an external request
Recover cluster stateSnapshot and restore proceduresInspect a recovery sequence
Recover application dataStorage and backup ownershipVerify a restored dataset

We propose a workshop that identifies the operations your current provider performs and assigns them in the on-prem design. The comparison covers the selected distribution and actual infrastructure integrations.

The exercises use a controlled cluster and a representative application. They connect workload symptoms to infrastructure behavior, without assuming a particular cloud provider, hardware platform, or reason for the move.

This proposed agenda keeps familiar workload basics brief. It gives more attention to cluster components, infrastructure dependencies, and recovery responsibilities.

Day 1

We review Pods, Deployments, Services, and release probes against the destination design. You will identify which application behaviors depend on infrastructure supplied by the managed service today.

  • Deploy a representative workload and list its required platform integrations.
  • Introduce a failed release and distinguish application rollback from cluster recovery.

Day 2

You will use Helm and compare Kustomize for environment configuration. We examine etcd, API servers, controllers, kubelets, and high availability, with an on-prem recovery exercise as proposed customization.

  • Prepare repeatable configuration for the workload and its supporting resources.
  • Rehearse a cluster-state recovery sequence in an isolated lab and identify its remaining data dependencies.

Day 3

We connect Pod networking, Services, DNS, and ingress to the selected on-prem load balancer. You will examine network policies, service mesh use cases, affinity, taints, and capacity constraints.

  • Diagnose a Service that lacks a working external traffic path.
  • Inspect why a replacement workload cannot fit the available nodes.

Day 4

You will examine dynamic storage provisioning, secrets, and data recovery. We distinguish HPA decisions from infrastructure capacity and cover authentication, RBAC, and the operational duties that stay with your team.

  • Restore a sample dataset and verify that a replacement application can use it.
  • Apply workload pressure and identify whether additional replicas or infrastructure capacity are required.

Your current provider, on-prem integrations, and recovery responsibilities can shape the agenda. Get in touch to tailor the workshop to your on-prem Kubernetes transition.

When an engineer sees running Pods but no external endpoint, the instructor can trace the Service to the missing load-balancer integration. The exercise exposes work that a managed service previously performed outside the workload manifest.

Aine's advice to someone joining a separate LearnKube course focused on its practical format:

Very good. Hands-on. Practical knowledge of Kubernetes.

— Aine, SRE at Matillion.

Tell us which managed service you use, how the on-prem destination will provide infrastructure services, and who will operate them. We will recommend exercises for those new responsibilities.