Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape an HTML table with Cheerio, obtain the page markup, load it with cheerio.load() (or an appropriate loader), select the specific table, traverse its rows and cells, and map the resulting values to real header labels. The short example below handles a regular table; production code must also account for multiple header rows, row and column spans, client-rendered tables, response failures, and untrusted markup.

Install Cheerio and choose a supported Node.js runtime

Install the package in your project:

npm install cheerio

The Cheerio documentation viewed for this guide lists Node.js 22.19 or later as the current requirement. Check the package documentation when you publish or deploy because runtime requirements can change. Cheerio can be imported as an ES module:

import * as cheerio from 'cheerio';

In a CommonJS project, use:

const cheerio = require('cheerio');

Get the HTML before parsing

Fetch a normal HTML response yourself

Fetching separately gives you control over status checks, headers, retries and logging. This complete example fetches a page, verifies the response, selects table#results, and returns the source rows as arrays:

import * as cheerio from 'cheerio';

const response = await fetch('https://example.com/data');
if (!response.ok) {
  throw new Error(`Request failed: ${response.status}`);
}

const contentType = response.headers.get('content-type') || '';
if (!contentType.includes('text/html') && !contentType.includes('application/xhtml+xml')) {
  throw new Error(`Expected HTML, received ${contentType || 'unknown content type'}`);
}

const html = await response.text();
const $ = cheerio.load(html);
const table = $('table#results');
if (!table.length) {
  throw new Error('Results table was not found');
}

const rows = table.find('tr').toArray().map((row) =>
  $(row).find('th, td').toArray().map((cell) =>
    $(cell).text().trim().replace(/s+/g, ' ')
  )
);

console.log(rows);

Always scope row and cell queries to the chosen table. Selecting $('table').first() is fragile on pages containing navigation, comparison, pricing or nested tables.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Load markup you already have

When another request, file or queue has already produced a string, pass it directly to cheerio.load(html). For raw bytes, the loading API also provides loadBuffer. Streaming inputs are handled by decodeStream and stringStream when a stream is more appropriate than buffering the complete document.

Cheerio’s fromURL helper can fetch a URL directly. It follows up to five redirects, rejects non-2xx responses and non-markup content types, chooses XML mode from the response content type, and uses the final URL as the document’s base URI. Those behaviors mean a PDF, JSON endpoint or error page is rejected rather than silently treated as an ordinary HTML table.

Select the intended table reliably

Prefer a stable identifier, class, caption or containing region supplied by the page’s markup:

const table = $('main article table[data-testid="sales"]');
const caption = table.find('caption').first().text().trim();

Cheerio supports CSS-style selectors and relationship selectors. Narrow from a known container, then use traversal methods such as find, children, closest, next and prev. If a selector matches more than one table, decide whether you need all matches or must identify one by its caption or surrounding heading:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const candidates = $('table').filter((_, el) =>
  $(el).find('caption').text().trim() === 'Quarterly revenue'
);
if (candidates.length !== 1) {
  throw new Error(`Expected one revenue table, found ${candidates.length}`);
}

Do not interpolate untrusted text into a selector. Compare untrusted values as data, as in the caption filter above, rather than constructing a selector from them.

Extract rows and cells

Regular tables

For a table with one header row and one cell per column, turn the first row into keys and zip the remaining rows:

function rowsToObjects(table, $) {
  const rows = table.find('tr').toArray().map((row) =>
    $(row).find('th, td').toArray().map((cell) =>
      $(cell).text().trim().replace(/s+/g, ' ')
    )
  );

  if (rows.length < 2) return [];
  const headers = rows[0];
  return rows.slice(1).map((values, rowIndex) => {
    const record = {};
    headers.forEach((header, columnIndex) => {
      record[header || `column_${columnIndex + 1}`] = values[columnIndex] ?? null;
    });
    if (values.length !== headers.length) {
      record._warning = `Row ${rowIndex + 2} has ${values.length} cells; expected ${headers.length}`;
    }
    return record;
  });
}

console.log(rowsToObjects(table, $));

This deliberately assumes a regular grid. A first row is not guaranteed to be the column-heading row: a table can have a title row, multiple header rows, row headers, a footer, or header relationships expressed with scope, id and headers.

Respect header semantics

Inspect the actual markup before choosing headers. Use th elements as headers, and determine whether each is a column or row header from its scope attribute and position. For complex accessible tables, follow the associations represented by headers and matching header id values instead of guessing from visual order. Keep an explicit schema when a column can be empty or unnamed; silently shifting values left produces plausible but incorrect records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand rowspan and colspan

A simple find('th, td') loop returns source cells, not a normalized rectangular grid. A cell with colspan="2" occupies two logical columns, and rowspan="2" continues into the next logical row. If downstream code requires one value per column, expand those spans into a grid before mapping headers. The algorithm is:

  1. Maintain a map of occupied column positions for rows created by earlier rowspan cells.
  2. For each source row, advance to the next free column.
  3. Read the cell’s numeric rowspan and colspan (defaulting each to 1).
  4. Place the cell value in every covered grid position and mark future rows occupied when rowspan is greater than 1.
  5. After all rows are expanded, apply the resulting header grid to the data rows.

Do not claim a rectangular result unless you implement this expansion and test it against the target table. For many jobs, retaining the source cells plus their span attributes is safer than inventing repeated values.

Use declarative extraction when the shape is stable

Cheerio also has an extract method for declarative nested results, including repeated records and attribute values. Row-by-row traversal is usually clearer for irregular tables because you can validate spans, missing cells and header structure at each step. Choose extract when the page’s markup and output shape are stable and the selector map remains easy to review.

When Cheerio cannot see the table

Cheerio parses supplied markup; it does not render a page or execute client-side JavaScript. If the initial HTML contains only an empty table shell and a script later inserts rows, Cheerio will return no data. Check the response body for the expected row text before changing selectors.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When content is rendered in the browser, look for a public data endpoint used by the page and request that endpoint directly when its terms and access controls permit. Otherwise obtain rendered HTML with browser automation such as Puppeteer or Playwright, then pass that HTML to Cheerio. A screenshot is visual evidence, not a structured table; use it to inspect what a user sees, not as a substitute for extracting the underlying data.

Validate results and handle failure cases

  • HTTP failure: check response.ok (or the status returned by fromURL) and record the status before parsing.
  • Wrong content: reject non-markup content types; an API’s JSON or a PDF is not an HTML table.
  • No table: fail loudly when the selector matches zero elements, rather than returning an empty dataset that looks valid.
  • Unexpected table count: assert the number of matches when exactly one table is required.
  • Missing rows or cells: check the minimum row count, compare each row’s width with the logical header width, and preserve nulls for genuinely empty cells.
  • Nested tables: scope each query to the selected table so an inner table’s cells are not mixed into an outer row.
  • Footers and pagination: identify tfoot, “total” rows and “next page” controls explicitly. Cheerio will not click pagination or request subsequent pages.
  • Encoding: use loadBuffer when you need Cheerio to determine encoding from raw bytes rather than decoding prematurely.

Log the final URL, status, content type, selected-table count, row count and a small sample of normalized values. Do not log credentials or entire pages by default.

Security: parsing is not sanitizing

Cheerio’s security guidance notes that scripts and event-handler attributes can remain in parsed and serialized markup. Treat downloaded HTML and extracted attributes as untrusted input. Prefer .text() for data, validate URLs and numbers, and sanitize HTML with a dedicated sanitizer before rendering any serialized markup. Never assume that parsing has made hostile content safe.

Performance, reliability and operating cost

  • Fetch once, parse once, and retain only the selected table when pages are large.
  • Use a timeout, bounded retries with backoff, and a concurrency limit so a slow origin cannot exhaust sockets or memory.
  • Cache responses when the source permits it, but include the source URL and retrieval time in your record so stale data is visible.
  • Prefer a site’s structured endpoint over browser rendering when it supplies the same rows; it is generally simpler to validate and paginate.
  • Keep selector and schema tests with your scraper. A redesign can leave the HTTP request successful while changing every value’s meaning.
  • Measure parse time and memory on your real documents rather than applying an unverified throughput claim.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your immediate need is a clean visual capture of a rendered page—for example, to inspect a table that appears only after scripts run—ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF; it does not turn pixels into structured rows, so use the page’s data endpoint or rendered HTML for extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example cURL (the ScreenshotNeo API docs have the complete option list):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/data -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/data"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/data' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Common problems and fixes

“Results table was not found”

Print the response’s first characters and inspect the saved HTML. You may have received a login page, an anti-bot response, a different locale, or a JavaScript shell. Confirm the selector against the returned markup, not the browser’s post-rendered DOM.

Rows are present but values are shifted

Check for a row header, colspan, hidden cells or a multi-row header. Replace first-row zipping with span-aware grid expansion and header association.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cheerio throws on a URL

Inspect status and content type. fromURL intentionally rejects non-2xx responses and non-markup responses; request the correct HTML URL or use the site’s data endpoint.

Text contains strange spacing

Normalize whitespace only after deciding whether line breaks are meaningful. For numeric data, parse after removing the site’s thousands separators and documenting locale assumptions.

The browser shows rows but the script returns none

The rows are likely inserted by JavaScript. Obtain the endpoint or rendered HTML with a browser tool, then parse that resulting markup with Cheerio.

Frequently Asked Questions

Can Cheerio scrape a table from a PDF?

No. Cheerio parses HTML or XML markup. Extract the PDF with a format-specific tool, or locate the HTML/data endpoint that produced the table.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Cheerio automatically follow table pagination?

No. Pagination is application behavior. Discover the next-page URL or API request and fetch each page under your rate and access limits.

Should I use .html() or .text() for cell values?

Use .text() for ordinary data. Treat .html() as untrusted markup and sanitize it before any rendering.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.