October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

How to Debug Kubernetes by Tracing Requests and Testing Failure Paths

A practical Kubernetes debugging guide: trace requests through Services and EndpointSlices, interpret Pod and probe evidence, and verify recovery without mistaking a learning project for production.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Kubernetes object can look healthy while the application request still fails. To find the break, trace the request from client to Service, EndpointSlice, target Pod, application, and dependency—and ask at each step: What evidence proves where the failure actually is?

A Flask-and-PostgreSQL project on a local multi-node kind cluster offers a practical learning example. Its reported setup uses two Flask replicas, one PostgreSQL replica, and a PostgreSQL PVC. It is a demonstration, not a production deployment. The useful lesson is its evidence-driven debugging method, not its topology as a template.

As an Amazon Associate I earn from qualifying purchases.

Trace the request path before changing anything

For the example application, the intended path is client → flask-app-svc → Flask → postgres-svc → PostgreSQL. The PVC represents persistence for the database separately from that request path. A failure anywhere along the chain can produce a similar user-facing symptom, so start by locating the earliest point where observed behavior diverges from the intended state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Kubernetes’ debugging guidance recommends first triaging whether the problem appears to involve a Pod, a controller, or a Service. Then inspect the objects and their events rather than treating one status field as a verdict. Kubernetes’ application debugging guide describes this object-oriented approach.

  1. Record the symptom. Note what the client sees, which request fails, and whether the issue is intermittent or consistent.
  2. Find the first unhealthy or unexpected object. Check the relevant workload, Pod, Service, and recent events.
  3. Form competing hypotheses. For example, a Service may have no matching Pods, a port may be wrong, the application may be failing, or its database dependency may be unavailable.
  4. Gather evidence that separates them. Check selectors, EndpointSlices, Pod conditions, logs or application responses, and node or scheduler events as relevant.
  5. Apply the smallest correct fix, then verify recovery at the application level. A changed object is not proof that users can complete the request.

What an EndpointSlice tells you—and what it does not

EndpointSlices record IP addresses for backends associated with a Service. For selector-based Services, the control plane creates slices containing references to matching Pods. They are an important source of backend information for kube-proxy. Kubernetes documents the purpose this way: “The EndpointSlice API is the mechanism that Kubernetes uses to let your Service scale to handle large numbers of backends, and allows the cluster to update its list of healthy backends efficiently.” See the Kubernetes EndpointSlices documentation. The API has been stable since Kubernetes v1.21; the documentation selector showed v1.37 when accessed on 2026-10-07, so check the documentation for your cluster’s version when version-specific behavior matters.

A Service can have a ClusterIP and still have no usable backends. Conversely, populated EndpointSlices show that backend endpoints are associated with the Service; they do not prove that forwarding succeeds, that the application responds correctly, or that its dependencies work. Treat the slice as one piece of evidence in the path, not an end-to-end health check.

  1. Inspect the Service and its selector: kubectl describe service <service-name>.
  2. Inspect associated EndpointSlices: kubectl get endpointslices -l kubernetes.io/service-name=<service-name>.
  3. Compare the Service selector with the intended Pod labels, then inspect the referenced Pods and their events.
  4. Where appropriate, test connectivity to the Service and application behavior from within the cluster. Use the result to narrow the fault; follow through to the application and dependency rather than assuming endpoint presence proves success.

For example, a selector mismatch can leave a Service without usable backends. A targetPort mismatch can leave endpoints populated while traffic still fails. Those observations point to different next checks; neither should be inferred from a ClusterIP alone.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read Pod state and events in context

Pod state narrows the search but does not answer every question. The Kubernetes debugging guide recommends inspecting a Pod and its recent events with kubectl describe pods <name>. In particular, Pending means the Pod cannot be scheduled yet—not that a resource shortage is necessarily the cause. Scheduler messages provide evidence about why scheduling failed.

Likewise, Running does not guarantee that the application is healthy or serving the request. Check readiness and the application’s actual response. For Kubernetes’ triage guidance and examples, see Debugging Applications.

Use probes for distinct questions

Startup, readiness, and liveness probes have different jobs. Kubernetes states: “Readiness probes determine when a container is ready to accept traffic.” When readiness fails, the EndpointSlice controller removes that Pod IP from matching Service EndpointSlices. A liveness failure can trigger a container restart. A configured startup probe delays liveness and readiness checks until startup succeeds. The official Pod Lifecycle documentation explains these behaviors.

  • Startup: Has this application finished initializing?
  • Readiness: Should this container receive Service traffic now?
  • Liveness: Should Kubernetes restart this container?

In the project, the author reports using startup and readiness checks at /health and a liveness check at /. The design goal was to let PostgreSQL trouble make Flask unready for traffic without automatically treating a database outage as a reason to restart Flask. That is one application-specific design choice, not a universal probe recipe. A readiness endpoint that includes dependency health can remove a Pod from traffic during an outage; whether that is appropriate depends on what the application can safely do without the dependency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Match symptoms to evidence, not assumptions

The project reports exercises involving configuration errors, Service routing, PostgreSQL availability and authentication, probes, resource limits, scheduling, storage, RBAC, and node or DNS problems. These are the author’s reported scenarios, not independently verified test results. The distinctions below help avoid jumping from a symptom to an unsupported cause.

Observation What it helps distinguish What it does not prove Useful next evidence
Service exists, but its EndpointSlices lack usable backends Whether matching backend endpoints are being associated with the Service Why they are absent, or whether the application itself would work Service selector, Pod labels and conditions, and related events
EndpointSlices contain backend addresses, but requests fail Whether the Service has associated endpoints Correct port forwarding, application health, or dependency health Service port and targetPort, connectivity, Pod behavior, and application response
Pod is Pending That it has not been scheduled That insufficient CPU or memory is the cause Pod description and scheduler events; check taints, constraints, and quota evidence as indicated
Pod is Running That the container has reached a running state That it is ready or the application request succeeds Probe conditions, events, and an application-level request
Flask cannot complete a database-backed request That the request path or dependency may be failing Whether the cause is network reachability, database availability, credentials, or application behavior Check the Flask response and logs, the database Pod and Service, connectivity, and authentication evidence
PVC remains Pending That storage has not bound as expected That data is backed up or recoverable PVC events and StorageClass configuration; separately verify any backup and restore process

The project’s reported failure classes also include incorrect ConfigMap keys or values, memory enforcement and OOMKilled, CPU throttling, ResourceQuota rejection, taint-related scheduling problems, hostPath limitations, RBAC and identity exercises, and node failures. Each requires evidence from the relevant object or layer. For instance, a quota rejection should be supported by the admission error or events; a suspected memory kill by container state and events—not inferred from a slow response alone.

Verify the repair at the same level as the failure

A corrected selector is not enough if the client still cannot reach the application. After making a fix, check that the relevant Pod or controller has converged, the Service has the intended EndpointSlices, and a representative request succeeds through the intended route. If the original symptom involved PostgreSQL, verify a database-backed operation rather than only a Flask health endpoint that may not exercise the dependency.

The author reports repeating the documented procedure in an isolated namespace. The reported result included two Ready kind nodes, a bound postgres-pvc, successful PostgreSQL and Flask rollouts, populated EndpointSlices, and the application response {"database":"connected","status":"healthy"}, recorded as PASS. This is the author’s reported reproduction, not an independent rerun.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the learning project’s limits in view

The project is explicitly described as a demonstration and learning project, not production-deployed. Its reported architecture has one PostgreSQL replica and no database failover; backup and restore were not tested. Resource values are described as unmeasured local baselines, and the project lacks centralized logs, distributed tracing, and automated alerting. A bound PVC provides persistence semantics, but does not by itself provide a backup or a tested restore path.

The author identifies managed or highly available PostgreSQL, tested backup and restore, measured resource tuning, stronger secret and supply-chain controls, production networking, and observability as future work—not completed capabilities. The cluster is useful for practicing how failures become observable and how to reason through recovery; it should not be mistaken for evidence of production readiness.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.