Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Start by deciding whether the process is actually executing JavaScript continuously. A non-yielding synchronous loop can monopolize Node.js’s single JavaScript thread, delaying every other callback and request. Sustained CPU with stalled event-loop work supports that hypothesis; low CPU while requests wait points more often to an asynchronous dependency. Capture evidence safely, profile the hot path, inspect the termination logic, then mitigate before deploying a verified fix.
What an “infinite loop” looks like in production
Node.js runs JavaScript callbacks on one event-loop thread. As Clinic.js puts it, “The event loop is single-threaded: only one operation is processed at a time.” A synchronous loop that never yields prevents timers, promise continuations, socket callbacks and HTTP handlers from running until the function returns.
Do not treat every stall as proof of non-termination. The same symptoms can come from a loop that is merely extremely long, recursion that revisits the same state, repeated retries without effective backoff, traversal of unexpectedly large input, or expensive synchronous parsing performed for every request. A CPU profile shows where samples accumulated; source inspection and a reproduction establish whether execution can terminate.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFirst response: establish scope without destroying evidence
- Record the incident context. Note the affected process or instance, start time, routes or jobs involved, recent deployments and configuration changes, and whether one replica or the whole service is affected. Preserve request IDs, deploy identifiers and existing telemetry according to your incident procedure.
- Compare CPU and waiting signals. Check process CPU, event-loop delay, request latency, throughput, active handles and dependency timing. High CPU alongside delayed callbacks is consistent with synchronous JavaScript. Low CPU with threads waiting on a database, filesystem or remote service suggests an I/O or dependency problem instead.
- Choose one representative process. Avoid collecting repeatedly from every replica. Confirm that your capture method is supported by the deployed Node.js version, operating system, container and process supervisor.
- Protect sensitive data. Diagnostic artifacts can contain URLs, headers, environment details, stack paths and heap information. Store them under your normal access controls and retention policy.
Capture a Node.js diagnostic report
Node.js diagnostic reports are designed for development, test and production problem determination. A report can include JavaScript and native stacks, heap information, libuv handles, platform details and resource data. Some report triggers and options vary by Node.js release, so read the documentation for the exact runtime before enabling them.
#1 Best Overall
Programmatic capture
If your service already exposes an authenticated, restricted diagnostics path, you can generate a report from that path using Node’s diagnostic-report API. Keep the endpoint disabled by default or protected by your incident controls. A minimal example is:
const report = require('node:report');
function writeIncidentReport() {
const filename = `/var/log/my-service/report-${Date.now()}.json`;
report.writeReport(filename);
return filename;
}
// Call only from an authenticated, rate-limited incident pathway.
Do not add a public route that lets arbitrary callers write files or inspect runtime state. If your deployment uses a signal or startup configuration instead, follow the runtime’s documented mechanism and your platform’s signal policy rather than copying a command that may not apply to your process manager.
What to inspect
- JavaScript stacks that repeat the same application frames.
- Native stacks and libuv handles that indicate whether the process is executing or waiting.
- Resource, heap and platform fields that distinguish a CPU problem from memory pressure or an external wait.
- Timestamps and process identity, so the report can be matched to the affected instance.
Profile CPU time to find the hot function
CPU sampling is the fastest way to narrow a large codebase. A flamegraph aggregates sampled call stacks over a time window. Wide application frames are candidates for investigation; repeated frames can point to a loop or repeated computation. Sampling is evidence of where the process spent time, not proof that a function never terminates.
Rank #2
Live process versus reproduction
| Approach | Best use | Trade-off |
|---|---|---|
| Live-process capture | Intermittent incidents that cannot be reproduced | Operational risk and collection overhead must be approved; compatibility depends on runtime and OS. |
| Representative reproduction | Repeatable input or a safe staging copy | Lower production risk, but the input and configuration must match the incident class. |
| Collection on server, visualization off-box | Restricted production environments | Move only the collected profile through an approved secure channel. |
Clinic.js Doctor helps distinguish CPU-bound behavior from waiting patterns, while Clinic.js Flame collects CPU profiles and produces flamegraphs. Its collection-only workflows allow data capture in one environment and visualization elsewhere. The documentation was published several years ago; verify current maintenance and compatibility with your exact Node.js version before an incident.
Visual Studio Code can open JavaScript .cpuprofile files and provide CPU flame views. This is useful when your profiler exports that format, but it does not remove the need to validate runtime and tool compatibility.
Read the profile correctly
- Find the widest frames under the request, job or timer entry point.
- Expand repeated application frames and note the input or route associated with the sample window.
- Separate your code from parser, regex, serialization and dependency frames; an expensive library call may still be triggered by a faulty loop condition.
- Compare more than one capture if the incident is intermittent.
Avoid high-volume synchronous logging while the event loop is already under pressure. Logging can add more blocking work and obscure the original timing.
Rank #3
Trace the hot path back to the bug
Check loop progress
for (let i = 0; i < items.length; i += 1) {
processItem(items[i]);
}
Verify that the control variable changes on every path, that the bound can change in the expected direction, and that an empty or malformed item cannot reset progress. For while loops, identify the state mutation that eventually makes the condition false. For iterators, confirm that next() advances rather than returning the same value forever.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Inspect retries and recursion
- Ensure retry counters are incremented even when an exception is thrown.
- Apply a finite attempt limit and an explicit deadline; exponential backoff without a cap can still create an unbounded job.
- For recursion, prove that every branch moves toward a base case and that input depth is bounded.
- Check whether a callback or promise schedules the same job again after every failure.
Look for input-driven explosions
Large arrays, deeply nested JSON, pathological regular expressions and graph traversals can look infinite in production. Log bounded metadata such as item counts, depth and an identifier—not entire payloads. Reproduce with the same class of input and a time limit.
Check synchronous work inside request handlers
Even a terminating loop can block all requests if it performs CPU-heavy parsing, compression or transformation on the event loop. If the work is legitimately CPU-bound, bound its size, split it into smaller cooperative chunks where correctness permits, or move it to a worker-thread or separate job process. Validate the architecture and observability impact before changing execution models.
Rank #4
Mitigate first, then deploy the code fix
- Contain the workload. Follow your service playbook to shed traffic, disable the affected job, isolate a tenant or route, or remove an unhealthy process from rotation.
- Rollback a suspect change. If timing matches a deployment or configuration change, restore the last known-good version using your normal release controls.
- Bound execution. Add maximum items, recursion depth, attempts and elapsed time. Return a classified error rather than looping indefinitely.
- Fix the state transition. Correct the condition or mutation identified by source inspection. Add a regression test using the incident-shaped input.
- Reproduce before rollout. Run the test under a representative Node.js version and workload, then canary the change while watching CPU, event-loop delay, errors and latency.
Common failure modes and recovery
| Symptom | Likely cause | Next action |
|---|---|---|
| CPU near saturation; timers and requests delayed | Synchronous loop or expensive synchronous function | Capture a report, obtain a CPU profile, inspect the widest application frame. |
| CPU low; requests remain pending | Database, network, filesystem or other asynchronous wait | Trace dependency timing and active handles instead of assuming a loop. |
| Profile points to a parser or regex | Large or pathological input | Bound input, reproduce safely and inspect the caller’s loop. |
| Capture fails or produces unusable data | Unsupported runtime/tool combination or restricted container | Check exact versions and platform permissions; capture a representative reproduction if safe. |
| Incident disappears after restart | State-dependent or input-dependent path | Preserve telemetry and reports from the next occurrence; do not treat restart as a fix. |
Or skip the browser setup
If you need screenshots of dashboards, traces or an incident page while documenting the investigation, ScreenshotNeo provides a single HTTP call rather than a locally managed browser. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. The same endpoint supports full-page or selector captures, device and retina settings, dark mode, custom CSS and JavaScript, waits, request blocking, headers, cookies, authorization, geolocation, PDFs, caching, signed links, asynchronous jobs and bulk capture.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can an asynchronous function still cause an infinite-loop incident?
Yes. A promise or timer may repeatedly schedule work, but each callback can yield between turns. Diagnose both the scheduling cycle and any synchronous segment that consumes the CPU.
Should I restart before collecting a profile?
Only when your incident procedure requires immediate recovery or the process is unsafe. A restart destroys the state that could identify the path, so capture first when operationally safe.
Does a flamegraph prove the loop never ends?
No. It records sampled activity during a time window. Termination must be established by inspecting control flow and testing the relevant input.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

