From legacy schedulers to orchestration across Kubernetes and Slurm

We offer private, instructor-led Kubernetes training for data infrastructure engineers who need to support model-training workloads across Kubernetes and Slurm.

Your team already knows the current execution process. The new platform needs decisions about workload placement, data access, retries, and which system handles each failure.

LearnKube connects those decisions to the Kubernetes side of your platform. Your team will get hands-on experience with execution, placement, and recovery through a sample batch workload.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Run Kubernetes batch workloads by defining Jobs and resource requests, so tasks have explicit execution and capacity requirements.
  • Preserve recoverable outputs by separating persistent data from Pod storage, so a replacement task can access its inputs and saved progress.
  • Diagnose delayed execution by inspecting scheduling events, resource availability, and storage access, so engineers distinguish placement problems from application failures.
  • Define cross-platform ownership by documenting submission, data, and recovery boundaries, so operators know which system and team must act.

Legacy schedulers encode assumptions about submission, placement, and completion. A platform spanning Kubernetes and Slurm needs explicit boundaries for those responsibilities. A Kubernetes Job controls Pods within its cluster. Choosing between clusters or execution systems requires another orchestration decision, along with a way to track the resulting work.

With two execution systems, I want a clear owner for the decision to resubmit failed work. If the orchestrator and the workload both retry independently, the platform can spend capacity on duplicate training attempts.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Current task or requirementNew concept or decisionPractice
Submit a batch taskKubernetes Job lifecycleTrack task completion
Select suitable capacityResource requests and affinityDiagnose a Pending Pod
Recover interrupted workCheckpoints and persistent storageResume from saved progress
Coordinate execution systemsSubmission and ownership boundariesTrace an execution handoff

Mistral's data infrastructure initiative includes a strategic transition from legacy scheduling to modern orchestration. Its platform spans Kubernetes and Slurm, with responsibilities for placement across clusters, storage, and model-training job operations.

That environment requires engineers to understand where Kubernetes controls execution and where another system makes the decision. We recommend a four-day Kubernetes workshop focused on job lifecycle, placement, and data recovery within those boundaries.

This proposed agenda focuses on Kubernetes in a mixed execution platform. We spend less time on familiar container and service mesh topics to make room for batch execution and recovery exercises.

Day 1

We show how containers and Pods relate to Jobs, retry limits, and completion status. You will distinguish long-lived platform services from model-training tasks that finish.

  • Run a sample batch task and inspect its completion status.
  • Introduce a failed attempt and observe retry behavior and application output.

Day 2

You will use Helm or Kustomize to configure workloads consistently. We explain control-plane components and reconciliation, including the limits of recovery within one cluster.

  • Template task inputs and resource requirements.
  • Simulate a node failure and inspect how the Job replaces its Pod.

Day 3

We explain resource requests, node affinity, taints, and topology constraints. You will connect Services, DNS, ingress, and network policies to platform interfaces and data endpoints.

  • Diagnose a Pending Pod whose requests do not match available nodes.
  • Trace and repair a blocked connection to a data endpoint.

Day 4

We review storage, secrets, and RBAC, along with how applications handle checkpoints. You will distinguish autoscaling for supporting services from batch concurrency and infrastructure capacity.

  • Restart a checkpoint-aware sample task and recover its saved progress.
  • Compare concurrent task demand with the resources available to the cluster.

Your batch workloads, storage integrations, and cross-platform responsibilities can shape the agenda. Get in touch to tailor the workshop to your orchestration transition across Kubernetes and Slurm.

If an engineer expects Pod replacement to resume model training automatically, the instructor can inspect what the application actually saves. A checkpoint-aware sample shows that Kubernetes recovery and application recovery use different mechanisms.

The technical depth of demos and deep-dive explanations - so much better than any of the many online on-demand classes I've taken.

— Dana, Staff Systems & Infrastructure Engineer at Walmart.

Tell us how you submit work today, which workloads Kubernetes will run alongside Slurm, and who owns placement and recovery. We will recommend a Kubernetes course focus for those responsibilities.