From air-gapped Kubernetes pilot to production: training for systems teams

We offer private, instructor-led training for HPC, systems, storage, and network engineers who need to support Kubernetes beyond a small bare-metal pilot.

Your team understands the infrastructure, but Kubernetes experience varies. A working pilot does not explain replacement and recovery, especially when image sources, network access, and storage options are constrained.

LearnKube connects familiar infrastructure duties to a representative workload and its failure paths. Your engineers will examine what the platform needs when an application instance or node becomes unavailable.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Operate representative workloads by understanding images, Pods, controllers, and release behavior, so the team can explain how applications start and recover.
  • Preserve required dependencies by examining approved image sources and persistent storage, so replacement workloads can obtain the software and data they need.
  • Diagnose failures across infrastructure layers by correlating events, logs, network paths, and volume status, so operators can identify the component that prevents recovery.
  • Assign production-support responsibilities by documenting cluster, node, image, network, and data ownership, so the intended operating model has clear actions and escalation paths.

A small bare-metal pilot can depend on images and data already available on its original nodes. Wider operation requires an explicit plan for image availability and persistent storage after replacement. Air-gapped access makes these dependencies important: a controller can request another Pod without guaranteeing that its destination can obtain the required artifacts and data.

The production question is not only whether the workload starts today. I want teams to explain how it starts on a replacement node when registry access, image availability, and storage are constrained.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Current task or requirementNew concept or decisionPractice
Start a pilot workloadControllers and declared stateObserve replica replacement
Supply approved softwareRegistry access and image policiesDiagnose a missing image
Preserve application dataVolume lifecycle and recoveryRestore access after replacement
Support a failed serviceCross-layer operational evidenceLocate the failed dependency

Saudi Aramco's HPC and systems team had a small bare-metal Kubernetes pilot. It wanted to move toward production once the team became comfortable supporting and using the platform.

The intended environment involved air-gapped operation, on-prem storage, and day-to-day diagnosis. LearnKube delivered private Kubernetes courses for HPC and other operations teams.

Asked what he liked most about the course, Ahmed answered:

Materials well prepared, instructor is excellent

— Ahmed, HPC Sys. Admin at Saudi Aramco.

This proposed agenda establishes a common foundation across infrastructure roles. It emphasizes the dependencies and failure behavior that the team must understand before taking on production support.

Day 1

We explain container packaging, Pods, Deployments, Services, and probes. You will connect image policies and workload configuration to the behavior of a service deployed from an approved artifact source.

  • Deploy a representative service with explicit image and configuration requirements.
  • Introduce an unready release and inspect its recovery and rollback options.

Day 2

You will use Helm and compare Kustomize for repeatable resources. We examine API servers, etcd, controllers, kubelets, and high availability through the responsibilities behind application replacement.

  • Prepare deployment resources whose image and configuration dependencies are explicit.
  • Simulate a node failure and trace the requirements of the replacement workload.

Day 3

We connect Pod networking, DNS, Services, and ingress to the selected on-prem environment. You will review network policies, service mesh use cases, affinity, and constraints on available nodes.

  • Diagnose a workload that cannot reach its approved registry or data endpoint.
  • Inspect a placement constraint that prevents a replacement Pod from starting.

Day 4

You will examine persistent data, credentials, and recovery requirements alongside autoscaling metrics. We distinguish workload demand from infrastructure capacity and review authentication, RBAC, and support ownership.

  • Replace a stateful workload and verify access to its required data.
  • Investigate a failed recovery and separate application, storage, and permission evidence.

Your pilot architecture, air-gap constraints, and infrastructure responsibilities can shape the agenda. Get in touch to tailor the workshop to your pilot-to-production Kubernetes preparation.

When an engineer assumes that a cached image will be available after node replacement, the instructor can inspect the image policy and artifact source. Your team can identify what the destination actually needs instead of relying on the original node's state.

Tell us what runs in the pilot, which image and storage constraints apply, and who will support production workloads. We will recommend exercises that connect those responsibilities.