Kubernetes controller and operator training for platform developers

Private, instructor-led training for platform developers who extend Kubernetes and need to reason about a controller's behavior beyond a successful first reconciliation.

Creating a custom resource does not implement its operating contract. Ownership, observed status, repeated events, interruption, and deletion need explicit behavior in the component that acts on it.

LearnKube uses a small, bounded controller example to connect API fundamentals with failure verification. Engineers inspect and modify the control loop, then explain what the component can safely repeat.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Design controller behavior around desired state, ownership, and status, so the component's required actions and completion conditions are explicit.
  • Implement a bounded control loop using supported API and custom-resource mechanisms, so the component can reconcile the selected resource contract.
  • Verify lifecycle failures through repeated events, interruption, and deletion, so the team can detect incorrect repetition or incomplete cleanup.
  • Operate the component with scoped permissions, observable status, and release checks, so engineers can maintain its behavior after the classroom example.

The team already uses workload resources, but an extension needs its own control contract. Custom resources and controllers connect stored intent with implementation, while finalizers can coordinate deletion with required cleanup. Engineers must define repeated-operation, status, ownership, and failure behavior before treating a small successful demonstration as a maintainable platform component.

I ask what the controller does when it sees the same desired state again after losing its previous response. Reconciliation must tolerate repetition, and status must describe observed progress rather than merely the operation it attempted.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Component requirementAPI/control mechanismVerification
Describe the intended serviceCustom-resource contractInspect valid and invalid inputs
Converge after repetitionReconciliation and observed stateRepeat the same intent
Clean up dependent workOwnership and finalizersInterrupt a deletion path
Explain incomplete progressStatus and conditionsCompare actions with observations

Box's platform engineering work includes operators, controllers, internal APIs, and production diagnosis across networking and storage. Those components need both implementation skills and an operating explanation.

For that kind of work, we recommend a scoped four-day workshop using the course's API and CRD customization. The proposed example has a bounded contract rather than the full remit of a production operator project.

Participants need Kubernetes fundamentals and the ability to read and modify an API client. We shorten familiar deployment material and use one small example to make lifecycle behavior visible.

Day 1

Explain desired state, resource identity, status, and the workload behavior the component will manage. Define a small contract before modifying the supplied controller example.

  • Specify the example's valid inputs and observable completion conditions.
  • Compare its desired state with the resources that implement it.

Day 2

Connect watches or polling, repeated actions, and observed state to a control loop. Use templating and architecture knowledge to understand how the controller itself runs.

  • Implement or modify one bounded reconciliation decision.
  • Interrupt the example and inspect its behavior when processing resumes.

Day 3

Examine API access, dependency communication, resource ownership, and placement. Determine which failures should produce another attempt and which need a visible condition.

  • Introduce an unavailable dependency and inspect the resulting status.
  • Verify the permissions required for the component's intended scope.

Day 4

Review persistent effects, finalizers, resource pressure, and release compatibility. Define operating evidence for the lifecycle paths that the example supports.

  • Exercise repeated reconciliation and interrupted cleanup without duplicating the intended effect.
  • Record the contract, failure checks, and known scope of the component.

Your API contract, controller questions, and engineering responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer shows a successful first run, the instructor can repeat it after an interrupted response and inspect ownership and status. T.C. described explanations in an earlier HPE course:

Great explanations during the presentations and tutorials.

— T.C., Software Developer at HPE.

Tell us what your controller needs to manage, which lifecycle behavior is uncertain, and what API programming experience the team has. We will recommend a bounded engineering focus and exercises.