Kubernetes Cilium troubleshooting training for platform engineers

Private, instructor-led training for platform engineers who already operate Kubernetes networking and need to interpret CNI and Cilium behavior beneath familiar resources.

A Service manifest does not describe every runtime forwarding decision. Policy, endpoint identity, node routing, and the scope of network observations can explain different connectivity failures.

LearnKube connects a representative packet path to Cilium's operational evidence. Engineers compare observations, locate a cause, and verify a correction without treating every missing response as the same network problem.

Hands-on learning and the skills engineers take back to work.

Preview: course-wide figures are not yet available.

  • Hands-on learning
    Of instruction time spent on labs and challenges.
  • Troubleshooting confidence
    Of respondents report greater confidence diagnosing Kubernetes problems.
  • Relevant to your work
    Of respondents say the course addressed their engineering responsibilities.
  • Skills put into practice
    Of respondents applied their new skills at work within 90 days.
  • Explain connectivity behavior through CNI, endpoint, and forwarding mechanisms, so the team can identify which implementation serves a request.
  • Diagnose a failed path using agent status, flow observations, and policy evidence, so engineers distinguish unavailable backends from dropped or misrouted traffic.
  • Validate a correction across the affected node and policy boundaries, so a successful local request does not conceal a remaining cross-node failure.
  • Make evidence collection repeatable through a scoped investigation record, so another engineer can identify what was observed and what remains untested.

The team already works with Services and policies, but their behavior depends on the networking implementation. Cilium diagnosis connects Kubernetes resources to agent, endpoint, and forwarding state. Hubble observations also have a collection scope. Engineers must establish where traffic was observed before interpreting an absent event or selecting a policy correction for the affected path.

I ask which agent or relay supplied the evidence before accepting a missing-flow explanation. A request outside the observer's scope can look absent without being dropped, so visibility must be established before policy is blamed.

— Daniele Polencic, LearnKube founder and Kubernetes instructor

Production questionMechanism to understandPractice
Is the endpoint managed?CNI and endpoint stateInspect the responsible agent
Where is traffic observed?Agent and relay scopeCompare observation boundaries
Why was traffic rejected?Identity and effective policyInspect a reported verdict
Does recovery cross nodes?Routing and forwarding pathsRepeat representative connections

Sony Interactive Entertainment's platform work includes multiple Kubernetes networking implementations, eBPF and iptables diagnosis, and Cilium telemetry. That work requires runtime evidence beyond the resource manifests.

For engineers with that responsibility, we recommend a four-day workshop on packet paths, observation scope, and controlled corrections. This recommendation is separate from Sony's earlier LearnKube training.

Participants need basic Kubernetes and IP-networking knowledge. We adapt the networking material to the selected Cilium version and deployment mode, with runtime diagnosis as the main task.

Day 1

Connect probes, ports, labels, and release changes to the request under investigation. Distinguish an unavailable application from a path that fails to reach it.

  • Compare backend health with the result of a representative client request.
  • Identify which endpoints and labels changed during a release.

Day 2

Relate rendered resources to CNI configuration, endpoint state, and node agents. Establish what the available diagnostic tools can observe.

  • Inspect whether the selected workload is managed by the expected networking implementation.
  • Compare node-local observations with the available relay view.

Day 3

Trace same-node and cross-node traffic through the configured data path. Interpret policy verdicts alongside discovery, endpoint identity, and routing evidence.

  • Diagnose a controlled policy denial using the affected endpoints' evidence.
  • Investigate a path that works locally but fails across nodes.

Day 4

Examine how workload replacement, resource pressure, and diagnostic permissions affect the investigation. Define the smallest correction and the paths needed to validate it.

  • Repeat a connectivity check after a bounded configuration or lifecycle change.
  • Produce an evidence record that states collection scope and remaining uncertainty.

Your CNI deployment, connectivity questions, and node responsibilities can shape the agenda. Get in touch to tailor the workshop to your team's work.

When an engineer treats absent telemetry as proof of a drop, the instructor can compare collection scopes before altering access. Kyle described practical investigations in a separate Octopus course:

The labs were very interesting, and weren't the typical "just run get pods". I liked the ones where you had to find the answer inside a pod, or manipulate a pod somehow to get the answer, and put it into the website.

— Kyle, Platform Engineer at Octopus Energy.

Tell us which CNI configuration you operate, which connectivity or visibility problem concerns your team, and which nodes and policies it owns. We will recommend focused investigations.