Recommended Free Tools
A Kubernetes object can look healthy while the application request still fails. To find the break, trace the request from client to Service, EndpointSlice, target Pod, application, and dependency—and ask at each step: What evidence proves where the failure actually is?
A Flask-and-PostgreSQL project on a local multi-node kind cluster offers a practical learning example. Its reported setup uses two Flask replicas, one PostgreSQL replica, and a PostgreSQL PVC. It is a demonstration, not a production deployment. The useful lesson is its evidence-driven debugging method, not its topology as a template.
As an Amazon Associate I earn from qualifying purchases.
Trace the request path before changing anything
For the example application, the intended path is client → flask-app-svc → Flask → postgres-svc → PostgreSQL. The PVC represents persistence for the database separately from that request path. A failure anywhere along the chain can produce a similar user-facing symptom, so start by locating the earliest point where observed behavior diverges from the intended state.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Kubernetes’ debugging guidance recommends first triaging whether the problem appears to involve a Pod, a controller, or a Service. Then inspect the objects and their events rather than treating one status field as a verdict. Kubernetes’ application debugging guide describes this object-oriented approach.
#1 Best Overall
- Record the symptom. Note what the client sees, which request fails, and whether the issue is intermittent or consistent.
- Find the first unhealthy or unexpected object. Check the relevant workload, Pod, Service, and recent events.
- Form competing hypotheses. For example, a Service may have no matching Pods, a port may be wrong, the application may be failing, or its database dependency may be unavailable.
- Gather evidence that separates them. Check selectors, EndpointSlices, Pod conditions, logs or application responses, and node or scheduler events as relevant.
- Apply the smallest correct fix, then verify recovery at the application level. A changed object is not proof that users can complete the request.
What an EndpointSlice tells you—and what it does not
EndpointSlices record IP addresses for backends associated with a Service. For selector-based Services, the control plane creates slices containing references to matching Pods. They are an important source of backend information for kube-proxy. Kubernetes documents the purpose this way: “The EndpointSlice API is the mechanism that Kubernetes uses to let your Service scale to handle large numbers of backends, and allows the cluster to update its list of healthy backends efficiently.” See the Kubernetes EndpointSlices documentation. The API has been stable since Kubernetes v1.21; the documentation selector showed v1.37 when accessed on 2026-10-07, so check the documentation for your cluster’s version when version-specific behavior matters.
A Service can have a ClusterIP and still have no usable backends. Conversely, populated EndpointSlices show that backend endpoints are associated with the Service; they do not prove that forwarding succeeds, that the application responds correctly, or that its dependencies work. Treat the slice as one piece of evidence in the path, not an end-to-end health check.
- Inspect the Service and its selector:
kubectl describe service <service-name>. - Inspect associated EndpointSlices:
kubectl get endpointslices -l kubernetes.io/service-name=<service-name>. - Compare the Service selector with the intended Pod labels, then inspect the referenced Pods and their events.
- Where appropriate, test connectivity to the Service and application behavior from within the cluster. Use the result to narrow the fault; follow through to the application and dependency rather than assuming endpoint presence proves success.
For example, a selector mismatch can leave a Service without usable backends. A targetPort mismatch can leave endpoints populated while traffic still fails. Those observations point to different next checks; neither should be inferred from a ClusterIP alone.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Read Pod state and events in context
Pod state narrows the search but does not answer every question. The Kubernetes debugging guide recommends inspecting a Pod and its recent events with kubectl describe pods <name>. In particular, Pending means the Pod cannot be scheduled yet—not that a resource shortage is necessarily the cause. Scheduler messages provide evidence about why scheduling failed.
Rank #3
Likewise, Running does not guarantee that the application is healthy or serving the request. Check readiness and the application’s actual response. For Kubernetes’ triage guidance and examples, see Debugging Applications.
Use probes for distinct questions
Startup, readiness, and liveness probes have different jobs. Kubernetes states: “Readiness probes determine when a container is ready to accept traffic.” When readiness fails, the EndpointSlice controller removes that Pod IP from matching Service EndpointSlices. A liveness failure can trigger a container restart. A configured startup probe delays liveness and readiness checks until startup succeeds. The official Pod Lifecycle documentation explains these behaviors.
Rank #4
- Startup: Has this application finished initializing?
- Readiness: Should this container receive Service traffic now?
- Liveness: Should Kubernetes restart this container?
In the project, the author reports using startup and readiness checks at /health and a liveness check at /. The design goal was to let PostgreSQL trouble make Flask unready for traffic without automatically treating a database outage as a reason to restart Flask. That is one application-specific design choice, not a universal probe recipe. A readiness endpoint that includes dependency health can remove a Pod from traffic during an outage; whether that is appropriate depends on what the application can safely do without the dependency.
Match symptoms to evidence, not assumptions
The project reports exercises involving configuration errors, Service routing, PostgreSQL availability and authentication, probes, resource limits, scheduling, storage, RBAC, and node or DNS problems. These are the author’s reported scenarios, not independently verified test results. The distinctions below help avoid jumping from a symptom to an unsupported cause.
| Observation | What it helps distinguish | What it does not prove | Useful next evidence |
|---|---|---|---|
| Service exists, but its EndpointSlices lack usable backends | Whether matching backend endpoints are being associated with the Service | Why they are absent, or whether the application itself would work | Service selector, Pod labels and conditions, and related events |
| EndpointSlices contain backend addresses, but requests fail | Whether the Service has associated endpoints | Correct port forwarding, application health, or dependency health | Service port and targetPort, connectivity, Pod behavior, and application response |
Pod is Pending |
That it has not been scheduled | That insufficient CPU or memory is the cause | Pod description and scheduler events; check taints, constraints, and quota evidence as indicated |
Pod is Running |
That the container has reached a running state | That it is ready or the application request succeeds | Probe conditions, events, and an application-level request |
| Flask cannot complete a database-backed request | That the request path or dependency may be failing | Whether the cause is network reachability, database availability, credentials, or application behavior | Check the Flask response and logs, the database Pod and Service, connectivity, and authentication evidence |
PVC remains Pending |
That storage has not bound as expected | That data is backed up or recoverable | PVC events and StorageClass configuration; separately verify any backup and restore process |
The project’s reported failure classes also include incorrect ConfigMap keys or values, memory enforcement and OOMKilled, CPU throttling, ResourceQuota rejection, taint-related scheduling problems, hostPath limitations, RBAC and identity exercises, and node failures. Each requires evidence from the relevant object or layer. For instance, a quota rejection should be supported by the admission error or events; a suspected memory kill by container state and events—not inferred from a slow response alone.
Verify the repair at the same level as the failure
A corrected selector is not enough if the client still cannot reach the application. After making a fix, check that the relevant Pod or controller has converged, the Service has the intended EndpointSlices, and a representative request succeeds through the intended route. If the original symptom involved PostgreSQL, verify a database-backed operation rather than only a Flask health endpoint that may not exercise the dependency.
The author reports repeating the documented procedure in an isolated namespace. The reported result included two Ready kind nodes, a bound postgres-pvc, successful PostgreSQL and Flask rollouts, populated EndpointSlices, and the application response {"database":"connected","status":"healthy"}, recorded as PASS. This is the author’s reported reproduction, not an independent rerun.
Keep the learning project’s limits in view
The project is explicitly described as a demonstration and learning project, not production-deployed. Its reported architecture has one PostgreSQL replica and no database failover; backup and restore were not tested. Resource values are described as unmeasured local baselines, and the project lacks centralized logs, distributed tracing, and automated alerting. A bound PVC provides persistence semantics, but does not by itself provide a backup or a tested restore path.
The author identifies managed or highly available PostgreSQL, tested backup and restore, measured resource tuning, stronger secret and supply-chain controls, production networking, and observability as future work—not completed capabilities. The cluster is useful for practicing how failures become observable and how to reason through recovery; it should not be mistaken for evidence of production readiness.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




