A recommended four-day workshop for data infrastructure teams
This proposed agenda focuses on Kubernetes in a mixed execution platform. We spend less time on familiar container and service mesh topics to make room for batch execution and recovery exercises.
Day 1
Define execution and completion
We show how containers and Pods relate to Jobs, retry limits, and completion status. You will distinguish long-lived platform services from model-training tasks that finish.
Hands-on exercises
- Run a sample batch task and inspect its completion status.
- Introduce a failed attempt and observe retry behavior and application output.
Day 2
Repeatable submissions and cluster recovery
You will use Helm or Kustomize to configure workloads consistently. We explain control-plane components and reconciliation, including the limits of recovery within one cluster.
Hands-on exercises
- Template task inputs and resource requirements.
- Simulate a node failure and inspect how the Job replaces its Pod.
Day 3
Placement and data connectivity
We explain resource requests, node affinity, taints, and topology constraints. You will connect Services, DNS, ingress, and network policies to platform interfaces and data endpoints.
Hands-on exercises
- Diagnose a Pending Pod whose requests do not match available nodes.
- Trace and repair a blocked connection to a data endpoint.
Day 4
Persistent progress, demand, and permissions
We review storage, secrets, and RBAC, along with how applications handle checkpoints. You will distinguish autoscaling for supporting services from batch concurrency and infrastructure capacity.
Hands-on exercises
- Restart a checkpoint-aware sample task and recover its saved progress.
- Compare concurrent task demand with the resources available to the cluster.
Your batch workloads, storage integrations, and cross-platform responsibilities can shape the agenda. Get in touch to tailor the workshop to your orchestration transition across Kubernetes and Slurm.