Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a table rendered by JavaScript across multiple pages, use a real browser (such as Playwright), wait for the rows that prove the current page is ready, extract and save those rows, then activate the site’s own next-page control. Repeat until that control is unavailable or disabled. A parser such as pandas.read_html can turn a rendered HTML table into a DataFrame, but it cannot execute JavaScript, click pagination, or preserve browser state by itself.

Choose the least complex access path first

Inspect the target before writing selectors. Determine which of these cases applies:

  • Static HTML: the response already contains the rows. A direct HTTP request and an HTML parser may be enough.
  • JavaScript-rendered rows: scripts fetch data or build the table after navigation. Use browser automation and wait for a row-level condition.
  • Interaction-dependent pagination: clicking Next, changing a page-size control, scrolling, or preserving cookies is required. Keep the browser session open and reproduce that interaction.

If the publisher offers a documented API or export for your intended use, evaluate it before automating the interface. It is usually simpler and less fragile, but the appropriate choice depends on the site’s implementation and terms.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why ordinary requests and read_html fail

Playwright’s navigation guide notes that page.goto() waits for the load event by default. That event only marks a navigation milestone; asynchronous requests can still be fetching rows. A request made with requests or a parser pointed at the initial response sees no data when the table is created later in the page.

pandas.read_html is useful after the browser has produced a genuine <table>. It parses table markup into DataFrames, but it does not wait for JavaScript, click Next, or manage a login/session. Custom grids made from <div> elements require direct DOM extraction instead.

Install Playwright and pandas

The example below uses Python. Create an isolated environment, install the packages, and install a browser binary:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
pip install playwright pandas
playwright install chromium

Pin and record the Playwright and browser versions used in production. The APIs and selectors can change, and this article has not been tested against a particular target site or browser matrix.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a repeatable multi-page scraper

The selectors in this program are deliberately explicit placeholders. Replace them with selectors from the site you are allowed to access. The important control flow is: wait, extract, persist, test the stopping condition, then move to the next page.

from pathlib import Path
import json
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

START_URL = "https://example.com/table"
ROW_SELECTOR = "table#results tbody tr"
NEXT_SELECTOR = "button[aria-label='Next page']"
OUTPUT = Path("rows.jsonl")


def extract_rows(page):
    # Return simple serializable values from the rendered page context.
    return page.locator(ROW_SELECTOR).evaluate_all("""
        rows => rows.map(row => Array.from(row.cells, cell => cell.innerText.trim()))
    """)


def next_is_unavailable(page):
    next_button = page.locator(NEXT_SELECTOR)
    if next_button.count() == 0:
        return True
    return next_button.is_disabled() or next_button.get_attribute("aria-disabled") == "true"


def main():
    all_rows = []
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(START_URL, wait_until="domcontentloaded")

        with OUTPUT.open("w", encoding="utf-8") as f:
            page_number = 1
            while True:
                try:
                    # Prefer a state-based wait over an arbitrary sleep.
                    page.locator(ROW_SELECTOR).first.wait_for(state="visible", timeout=30_000)
                except PlaywrightTimeoutError:
                    raise RuntimeError(f"No visible rows on page {page_number}: {page.url}")

                page_rows = extract_rows(page)
                if not page_rows:
                    raise RuntimeError(f"Empty extraction on page {page_number}: {page.url}")

                for cells in page_rows:
                    record = {"page": page_number, "url": page.url, "cells": cells}
                    all_rows.append(record)
                    f.write(json.dumps(record, ensure_ascii=False) + "n")

                if next_is_unavailable(page):
                    break

                old_first_row = page.locator(ROW_SELECTOR).first.inner_text()
                page.locator(NEXT_SELECTOR).click()
                # Wait for the old page state to be replaced. Adjust this for the site.
                page.wait_for_function(
                    "([selector, oldText]) => { const el = document.querySelector(selector); return el && el.innerText !== oldText; }",
                    [ROW_SELECTOR, old_first_row],
                    timeout=30_000,
                )
                page_number += 1

        browser.close()

    print(f"Saved {len(all_rows)} rows to {OUTPUT}")


if __name__ == "__main__":
    main()

The Page API documents page-context evaluation. Returning strings, numbers, booleans, arrays, and plain objects keeps the result serializable. The JSON Lines output records the source URL and page number, which makes a failed transition diagnosable.

Adapt extraction to the table type

  • Semantic table: use table thead th for headers and tbody tr/td for cells. You can pass the resulting HTML to pandas.read_html when you need DataFrame operations.
  • Custom grid: identify the row container and its cell or column elements, then extract text and attributes with locator.evaluate_all.
  • Virtualized rows: the DOM may contain only visible rows. Scroll in controlled increments, extract each rendered batch, and deduplicate by a stable key; a simple page loop will not recover rows that are never mounted.
  • Links and metadata: extract attributes such as href, data-id, or image URLs alongside visible text.

Handle pagination safely

Next changes the URL

Some sites use a normal link or query parameter. Let the browser follow the link, then wait for a URL or row condition:

with page.expect_navigation(wait_until="domcontentloaded"):
    page.locator("a[rel='next']").click()
page.locator("table#results tbody tr").first.wait_for(state="visible")

Next updates the DOM in place

Do not wait for navigation. Capture a value from the current first row, click, and wait until that value changes (or until a site-specific loading indicator disappears). Controls can appear before their event handlers are ready during hydration, so a visible button alone is not proof that it is usable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no Next button

Infinite scroll requires a different stopping rule. Scroll, wait for the row count to increase, extract only newly seen keys, and stop after the site signals completion or repeated scrolls add no rows. Do not assume a fixed page count.

Capture before transition

Append the current page’s rows before clicking or navigating away. Otherwise an in-place update can replace the only copy of those rows. Save incrementally so a crash does not discard earlier pages.

Parse and normalize the collected data

For a real HTML table, one option is to capture its rendered markup and parse it:

import pandas as pd

html = page.locator("table#results").evaluate("el => el.outerHTML")
frame = pd.read_html(html)[0]
frame = frame.drop_duplicates()
frame.columns = [str(c).strip() for c in frame.columns]

For custom grids, construct records directly and normalize at the boundary: trim whitespace, convert explicit missing markers to None, parse dates with an explicit timezone policy, and retain the original text when conversion is uncertain. Keep a stable primary key if one exists; names alone are often not unique.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate every page and the final dataset

  • Log page number, URL, timestamp, and extracted row count.
  • Check for zero-row pages, repeated header rows, and unexpectedly changing column counts.
  • Detect duplicate primary keys across pages and investigate whether they represent legitimate updates or pagination bugs.
  • Measure required fields for empty values and preserve the raw value for review.
  • Verify that the last page is complete: a disabled/absent Next control, an explicit end marker, or the site’s documented completion signal.
  • Compare a small sample manually with the rendered browser view before scaling up.

Sessions, filters, and request timing

Create a persistent browser context when filters or authentication must survive pagination. Load the filter, wait for the table to reflect it, then begin extraction. If the site uses a search box, wait for the result condition after pressing Enter rather than assuming the request is finished.

Use modest concurrency and delays appropriate to the site. Browser automation is resource-intensive; reuse one context for related pages, close pages you no longer need, and persist batches rather than retaining an unbounded list in memory. Do not bypass authentication, CAPTCHAs, or technical restrictions.

Responsible and permitted collection

RFC 9309, the Robots Exclusion Protocol, explains that robots rules are not access authorization. Check robots instructions as one planning input, but also follow the site’s terms, access controls, applicable law, and any stated license. Use an intended export or documented endpoint where available, identify your crawler when appropriate, and keep request rates conservative.

Troubleshooting common failures

Symptom Likely cause Fix
Rows are missing after goto() Asynchronous rendering has not completed. Wait for a specific row, label, or request-dependent state; do not rely on load or a fixed sleep alone.
Timeout waiting for rows Wrong selector, blocked request, consent wall, or an error state. Inspect a screenshot and page HTML, confirm the selector in DevTools, handle consent, and log the URL and visible error.
Every page contains the same rows Click did not trigger, or the script clicked before hydration. Wait for the control to be actionable, click once, and wait for a row or URL change.
Rows vanish between pages Extraction occurred after an in-place DOM replacement. Extract and persist before the transition, then wait for the new state.
read_html finds no tables The interface is a custom grid or markup was not rendered in the captured HTML. Extract grid cells directly, or evaluate outerHTML after the table is visible.
Duplicate or skipped records Unstable sorting, virtualized rows, or an incorrect end condition. Use a stable sort/key, deduplicate deliberately, log counts, and verify the site’s pagination state.
Click is intercepted Overlay, cookie banner, or sticky element covers the control. Handle the banner legitimately, scroll the control into view, and avoid force-clicking unless you understand the consequence.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

For a screenshot of a rendered page rather than a structured data export, ScreenshotNeo provides a one-call website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the ScreenshotNeo API documentation for all options, including viewport and device presets, full-page lazy-image loading, CSS-selector element capture, dark mode, retina scale, PDF settings, custom CSS/JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and OpenAPI compatibility.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.

Frequently Asked Questions

Can a scraper use only JavaScript disabled in the browser?

Only when the table’s rows are present in the initial HTML. If scripts create or fetch the rows, disabling JavaScript removes the content you need.

Should I scrape the table’s network API instead of the DOM?

Use a documented endpoint or export when the publisher provides one. An internal endpoint may change or require permissions, so treat it as an implementation detail unless its use is authorized.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I know pagination is complete?

Use the site’s own signal—disabled or absent Next control, an explicit end marker, or a documented total—and validate row counts and keys rather than assuming a page count.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.