Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The right Node.js table-extraction method depends on when the table is created. If the <table> is already in the server response, fetch the HTML and parse it with Cheerio. If JavaScript, a click, or a login creates the table in a browser, use Puppeteer (or comparable browser automation) and extract after rendering. Cheerio parses markup; it does not run page JavaScript.
This guide shows both paths, produces structured rows, explains headers and spanning cells, and includes diagnostics for the failures that make scrapers return empty tables.
Choose the capture path first
| Page condition | Recommended approach | What it does |
|---|---|---|
| The table is present in the initial HTTP response | Node.js fetch + Cheerio |
Downloads markup and traverses rows and cells without a browser. |
| The table is inserted by client-side JavaScript | Puppeteer (or another browser automation library) | Loads the page, runs scripts, performs interactions, then reads the rendered DOM. |
| You already have an HTML string or file | Cheerio load |
Parses the supplied markup directly; no network request is needed. |
Do not assume that a page that looks complete in Chrome sends the same HTML to fetch. Open the page’s response body or use an HTTP client first. If the response has no target rows, move to browser automation.
Prepare a Node.js project
Use an ES module project so the examples can run as written:
#1 Best Overall
- CRISP CLARITY: This 23.8″ Philips V line monitor delivers crisp Full HD 1920x1080 visuals. Enjoy movies, shows and videos with remarkable detail
- INCREDIBLE CONTRAST: The VA panel produces brighter whites and deeper blacks. You get true-to-life images and more gradients with 16.7 million colors
- THE PERFECT VIEW: The 178/178 degree extra wide viewing angle prevents the shifting of colors when viewed from an offset angle, so you always get consistent colors
- WORK SEAMLESSLY: This sleek monitor is virtually bezel-free on three sides, so the screen looks even bigger for the viewer. This minimalistic design also allows for seamless multi-monitor setups that enhance your workflow and boost productivity
- A BETTER READING EXPERIENCE: For busy office workers, EasyRead mode provides a more paper-like experience for when viewing lengthy documents
- Create a directory and initialize it:
mkdir table-capture && cd table-capture && npm init -y. - Set
"type": "module"inpackage.json, or save the files with an approach that supports ESM. - Install Cheerio for static parsing:
npm install cheerio. - For rendered pages, install Puppeteer with
npm install puppeteer. Its normal installation downloads a compatible Chrome browser.
Global fetch was added in Node.js v17.5.0/v16.15.0 and became stable in Node.js v21.0.0. Check the version actually used by your deployment with node --version; on older runtimes, use a supported fetch implementation rather than assuming the global is available.
Capture a table that is in the HTTP response
This complete ESM script checks the HTTP status, selects one specific table, and returns each row as an array of cell text. Replace the URL and selector with values from the page you own or are authorized to access.
import * as cheerio from 'cheerio';
const url = 'https://example.com/data';
const response = await fetch(url);
if (!response.ok) {
throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const selector = 'table#results';
const table = $(selector);
if (table.length === 0) {
throw new Error(`No table matched ${selector}`);
}
const rows = table.find('tr').map((_, row) =>
$(row).find('th, td').map((_, cell) => $(cell).text().trim()).get()
).get();
console.log(rows);
The result is an array such as [["Name","Status"],["Job 1","Done"]]. The mapping deliberately includes both th and td, so header rows are retained. It does not guess a schema, expand rowspan or colspan, or preserve links and attributes.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallUse a stable selector
Prefer an ID, meaningful class, data attribute, or a selector scoped beneath a known container. A bare table selector can silently capture a navigation, pricing, or layout table when a page contains several. Check the count with console.log($(selector).length) while developing.
Rank #2
- CRISP CLARITY: This 22 inch class (21.5″ viewable) Philips V line monitor delivers crisp Full HD 1920x1080 visuals. Enjoy movies, shows and videos with remarkable detail
- 100HZ FAST REFRESH RATE: 100Hz brings your favorite movies and video games to life. Stream, binge, and play effortlessly
- SMOOTH ACTION WITH ADAPTIVE-SYNC: Adaptive-Sync technology ensures fluid action sequences and rapid response time. Every frame will be rendered smoothly with crystal clarity and without stutter
- INCREDIBLE CONTRAST: The VA panel produces brighter whites and deeper blacks. You get true-to-life images and more gradients with 16.7 million colors
- THE PERFECT VIEW: The 178/178 degree extra wide viewing angle prevents the shifting of colors when viewed from an offset angle, so you always get consistent colors
Turn rows into objects when the first row is a header
If the first row is a single header row and every data row has the same number of cells, normalize it explicitly:
const matrix = table.find('tr').map((_, row) =>
$(row).find('th, td').map((_, cell) => $(cell).text().replace(/s+/g, ' ').trim()).get()
).get();
const [headers, ...dataRows] = matrix;
const records = dataRows.map((values, index) => {
if (values.length !== headers.length) {
throw new Error(`Row ${index + 2} has ${values.length} cells; expected ${headers.length}`);
}
return Object.fromEntries(headers.map((header, column) => [header, values[column]]));
});
console.log(records);
Do not use this shortcut for complex tables with multiple header rows or spanning cells. In those cases, define the desired output schema and write a normalization pass that accounts for the table’s actual layout.
Preserve links or other cell attributes
Text extraction discards markup. To retain a link, read it while traversing the cell:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →const detailedRows = table.find('tr').map((_, row) => {
return [ ...$(row).find('th, td') ].map(cell => {
const element = $(cell);
return {
text: element.text().replace(/s+/g, ' ').trim(),
href: element.find('a').attr('href') ?? null
};
});
}).get();
Decide whether relative URLs should remain relative or be resolved against the page URL. Resolve them only when your downstream format requires absolute links, and handle malformed values rather than letting one bad cell terminate the entire job.
Rank #3
- Clear visuals. Fluid motion: A 144Hz refresh rate and 1ms MPRT deliver smooth, tear‑free motion across work, gaming, and streaming for clearer, more fluid viewing.
- Eye comfort: TÜV Rheinland 3‑star* certification reduces harmful blue light while preserving stunning color quality without compromise. *TÜV Rheinland 3-star eye comfort certification.
- Wide viewing angle: Get consistent views across a wide 178° /178° viewing angle.
- In-Plane Switching (IPS): See excellent color accuracy and consistency across wide viewing angles with In-plane Switching (IPS) technology.
- Ultra-thin bezels: Maximize your viewing experience with thin bezels.
Handle tables rendered by JavaScript with Puppeteer
Cheerio’s own description is direct: “Cheerio is not a web browser.” It cannot execute the scripts that fetch data and insert rows. Puppeteer controls a browser, so wait for the table (or a more specific row selector) before reading it.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/dashboard', {
waitUntil: 'networkidle2',
timeout: 60_000
});
await page.waitForSelector('table#results tbody tr', { timeout: 30_000 });
const rows = await page.$$eval('table#results tr', tableRows =>
tableRows.map(row =>
[...row.querySelectorAll('th, td')].map(cell => cell.textContent.trim())
)
);
console.log(rows);
} finally {
await browser.close();
}
waitUntil: 'networkidle2' is useful when the page settles, but a selector that represents the actual data is the important synchronization point. Some applications keep analytics connections open forever or render rows in several batches; in those cases, wait for a known row, a loading indicator to disappear, or an application-specific condition instead of relying only on network idle.
Extract rendered HTML, then parse it with Cheerio
Browser evaluation is concise for text. If you need Cheerio’s selectors or the complete document, obtain the rendered markup with Puppeteer’s page.content() and parse it:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsconst renderedHtml = await page.content();
const $ = cheerio.load(renderedHtml);
const rows = $('table#results tr').map((_, row) =>
$(row).find('th, td').map((_, cell) => $(cell).text().trim()).get()
).get();
page.content() returns the full HTML contents, including the document type. This two-stage approach is useful when the same parsing and normalization code must handle both downloaded and browser-rendered input.
Rank #4
- CURVED FOR ENHANCED ENGAGEMENT: An immersive viewing experience with a curved monitor that wraps more closely around your field of vision; It creates a wider view, enhancing depth perception and minimizing peripheral distraction
- SMOOTH PERFORMANCE FOR SEAMLESS CONTENT: Stay in the action when playing games, watching videos, or working on creative projects; The 100Hz refresh rate reduces lag and motion blur so you don't miss a thing in fast-paced moments¹
- MORE GAMING POWER: Gain the edge with optimizable game settings; Color and image contrast can be adjusted to see scenes more vividly and spot enemies hiding in the dark; Game Mode adjusts any game to fill the screen so you can view every detail²
- KEEP IT EASY ON THE EYES: Care for your eyes and stay comfortable, even during long sessions; Advanced eye comfort technology certified by TÜV reduces eye strain by minimizing blue light and reducing irritating screen flicker²
- INCREASED VERSATILITY: Connect to more; Plug devices straight into your monitor for increased flexibility, making your computing environment even more convenient
Synchronize clicks and navigation safely
If a control causes navigation, start the navigation wait and click together so the click cannot outrun the wait:
await Promise.all([
page.waitForNavigation({ waitUntil: 'networkidle2', timeout: 60_000 }),
page.click('button[data-load-results]')
]);
await page.waitForSelector('table#results tbody tr');
For controls that update the current document without navigation, wait for the resulting table, row count, or loading-state change instead.
Local HTML and file input
When the markup is already in memory, skip fetch:
import * as cheerio from 'cheerio';
const html = `<table id="results"><tr><th>Name</th><th>Score</th></tr><tr><td>Ada</td><td>98</td></tr></table>`;
const $ = cheerio.load(html);
const rows = $('#results tr').map((_, row) =>
$(row).find('th, td').map((_, cell) => $(cell).text().trim()).get()
).get();
console.log(rows);
For a file, read it with Node’s filesystem APIs and pass the resulting string to cheerio.load. If you need byte-level encoding control, use a buffer-aware loading method rather than assuming every file is UTF-8.
Free tools Windows power users keep installed
One-click scans. No signup required.
Make extraction reliable in production
Validate the response and content
- Check
response.okbefore callingtext(); a 404 or access-denied page can otherwise look like an empty result. - Log the final URL, status, content type, selector, and matched table count.
- Reject an empty result when rows are expected. An empty array can mean “no records,” “wrong selector,” “JavaScript required,” or “blocked page.” Treat those cases differently.
- Set request and browser timeouts. A scraper without a deadline can consume workers indefinitely.
Normalize deliberately
- Collapse repeated whitespace when visual formatting is irrelevant.
- Keep raw text if exact formatting matters, and store normalized text in a separate field.
- Convert numbers and dates only after considering locale, thousands separators, missing values, and footnote characters.
- Handle
th, multiple header rows,rowspan, andcolspanaccording to the schema your application needs; no generic row mapper can infer every semantic layout.
Choose a parser mode consciously
Cheerio uses parse5 by default and follows HTML parsing rules. It also documents an htmlparser2 option for cases where its parsing behavior or performance characteristics better fit the input. Test either choice against malformed markup and the exact selectors used by your job; parser differences can affect how broken table structure is repaired.
Best Value
- 【INTEGRATED SPEAKERS】Whether you're at work or in the midst of an intense gaming session, our built-in speakers provide rich and seamless audio, all while keeping your desk clutter-free.
- 【EASY ON THE EYES】 Protect your eyes and enhance your comfort with Blue-Light Shift technology. This feature reduces harmful blue light emissions from your screen, helping to alleviate eye strain during long hours of use and promoting healthier viewing habits.
- 【WIDEN YOUR PERSPECTIVE】Our sleek minimal bezel design ensures undivided attention. The nearly bezel-free display seamlessly connects in a dual monitor arrangement, delivering an unobstructed view that lets you focus on more at once, completely distraction-free.
Control browser provisioning
Puppeteer normally downloads a compatible Chrome during installation. Package managers that block dependency-install scripts can skip that download, causing launch failures. puppeteer-core does not download Chrome; use it only when your team manages a local, containerized, or remote browser separately. Pin compatible versions in deployment and verify the browser exists before processing work.
Limit cost and latency
Use Cheerio whenever the response already contains the data: it avoids browser startup, uses less memory, and is usually faster. Reuse a Puppeteer browser for multiple pages when isolation requirements permit, but create a fresh page per job and close pages in a finally block. Wait for the narrowest useful selector instead of an unnecessarily long fixed delay. Block unneeded resources only when doing so cannot prevent the table’s data request.
Troubleshooting empty or incorrect results
| Symptom | Likely cause | Fix |
|---|---|---|
No table matched |
Selector is wrong, markup differs, or the table is client-rendered. | Print a short portion of the response, inspect the selector in browser developer tools, and switch to Puppeteer if rows are absent from the response. |
| HTTP status is 403, 429, or 503 | The server denied, throttled, or temporarily unavailable for the request. | Honor the site’s access rules, slow requests, use permitted authentication, and implement bounded retries for transient failures. Do not attempt to bypass access controls. |
| Browser launches locally but not in deployment | Chrome was not downloaded, the executable is unavailable, or the sandbox/container configuration differs. | Verify installation scripts ran, confirm the executable path and required system dependencies, or use a separately managed browser with puppeteer-core. |
| Rows are present but values are blank | Text is rendered in descendants, replaced by CSS, or the selector matched placeholders. | Use textContent after the data selector is ready, inspect the rendered DOM, and wait for the loading state to finish. |
| Only the first page of a grid is captured | Pagination or virtual scrolling hides the remaining rows. | Click each permitted page and extract it, or scroll and wait for additional rows before reading the DOM. Store page boundaries in your output. |
| Object conversion throws on some rows | A colspan, missing cell, or secondary header changes the column count. | Log the offending row, model spanning cells explicitly, and validate each row before constructing objects. |
| Click-triggered extraction races navigation | The click occurs before the navigation wait is registered. | Use the documented Promise.all pattern, then wait for the target table. |
Or skip the browser setup
If you need a screenshot or rendered capture rather than parsed cell data, ScreenshotNeo accepts one GET request and can render JavaScript pages without you provisioning Puppeteer. Its cleanup step accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
See the complete options and response details in the ScreenshotNeo documentation. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Features include full-page and element capture, device and retina settings, PDF output, custom CSS and JavaScript, waiting rules, request blocking, cookies and headers, caching, signed links, asynchronous webhooks, bulk capture, and a usage API. Sign up for the free 1,000-screenshot plan.
Cheerio or Puppeteer: a practical decision checklist
- Inspect the initial response first.
- Choose Cheerio when the target rows are already there and you need low-overhead parsing.
- Choose Puppeteer when scripts, interaction, authentication, pagination, or virtual scrolling creates the rows.
- Use a shared normalization layer so both paths produce the same schema.
- Validate selectors and column counts, record failures, and close network and browser resources deterministically.
Frequently Asked Questions
Can Cheerio scrape a page that requires JavaScript?
Not by itself. Cheerio parses the markup supplied to it and does not execute browser JavaScript. Render the page with Puppeteer first, then pass the rendered HTML to Cheerio if that fits your workflow.
Which Node.js version should I use for the examples?
Use a runtime with stable global fetch, documented as Node.js v21.0.0 and later. Verify your deployed version; older runtimes need a compatible fetch implementation.
Why do two rows have different numbers of cells?
The table may contain colspan or rowspan cells, multiple header rows, pagination markers, or missing data. Inspect the row structure and normalize those cases explicitly instead of assuming a rectangular matrix.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Is a screenshot the same as extracting table data?
No. A screenshot is a visual capture. Cheerio or Puppeteer extraction returns text and attributes that your program can validate, transform, and store.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

