Kubernetes foundations for data and ML engineers

We provide private, instructor-led Kubernetes training for data and ML engineers who want to better understand the platform that supports their workloads.

Your team understands its code, data, and processing needs. A copied workload definition does not explain retries, persistent output, or unavailable capacity. Engineers need a foundation connected to those situations.

At LearnKube, we use a small batch workload and its dependencies to help you connect your data work to Kubernetes. Through guided practice, you will learn how to explain changes in a run and make a plan for ongoing learning.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Get familiar with the workload lifecycle using images, Jobs, and data dependencies, so engineers see what Kubernetes manages for their tasks.
  • Practice running a batch job with clear configuration and storage, so learners can explain its inputs, how it runs, and what output is kept.
  • Demonstrate a reasoned investigation through a retry or placement problem, so engineers can identify the next observation and any platform support needed.
  • Keep a repeatable example with notes and references, so future data work has a solid foundation instead of just copied manifests.

A Kubernetes Job manages attempts to complete a task, subject to its retry limits. Persistent storage has a lifecycle separate from an individual Pod. Data engineers need to connect both mechanisms to their application, especially when repeated execution can create duplicate output. A small batch exercise makes execution, retained data, and application responsibilities clear.

I ask what happens to the output when the task runs again. A replacement Pod can restore execution, but the learner still needs to explain whether the application repeats or preserves its data changes.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Role milestoneRequired knowledgePractice/check
Prepare a batch workloadImages, Jobs, and configurationExplain the task inputs
Complete supported executionCompletion and retry behaviorRun a sample batch
Account for retained outputPod and storage lifecyclesInspect a replacement attempt
Continue platform learningDependencies and responsibility boundariesAnnotate the workload example

Alexander Thamm usually created its own internal training but lacked Kubernetes teaching capacity. Its developers, data scientists, and ML engineers needed a relevant foundation.

LearnKube provided Kubernetes training for the team. Carsten shared his experience and advice for future learners:

The guys really know their stuff and they put a lot of effort into creating the material. Loved it!

— Carsten, Data Engineer at Alexander Thamm.

This proposed agenda uses a small batch task and a supporting service. We shorten familiar application material to make room for Job behavior and data dependencies in Kubernetes.

Day 1

We link images, Pods, Jobs, Deployments, Services, and probes to their different uses. Learners will complete a supported run and learn the difference between finishing a batch and having a long-running service ready.

  • Run a sample batch task with explicit input configuration and explain its completion status.
  • Compare a failed batch attempt with an unready service and identify the controller responsible for each.

Day 2

We use Helm and compare it to Kustomize to set up the example. We also discuss how the API, controllers, and nodes work together to create a replacement workload.

  • Repeat the task with a different input value and explain the resulting resources.
  • Observe a failed attempt and describe what Kubernetes retries and what the application must account for.

Day 3

We look at DNS, Services, ingress, network policies, and scheduling. We talk about service mesh to find important service dependencies, and use examples to show how resource needs match up with available nodes.

  • Diagnose a failed connection to the sample data dependency and explain the evidence.
  • Interpret a Pending workload and identify the capacity or placement question for the platform team.

Day 4

We connect volumes, secrets, resource measurements, authentication, and RBAC to the batch task. Autoscaling discussion distinguishes service scaling from the Job's parallelism and completion requirements.

  • Inspect retained output after another attempt and explain the application's responsibility for duplicate writes.
  • Demonstrate the task with its data and access requirements, then record questions for continued practice.

Your engineers' experience, data workloads, and goals for further practice can shape the agenda. Get in touch to tailor the workshop to your data and ML engineers.

If a learner thinks a retry always gives the right result, the instructor can repeat a small write operation. This shows the difference between finishing a workload and how the application handles repeated work.

Let us know what your engineers already know about Kubernetes, what workloads they need to run, and what platform questions they have. We will suggest a foundation, practical exercises, and a plan for further learning.