Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a table already present in a page’s HTML, fetch the page, parse it into a DOM, select its rows with XPath, and turn each row’s cells into PHP arrays. Use DOMDocument with DOMXPath for compatibility with older PHP versions; on PHP 8.4 and later, use DomHTMLDocument when HTML5 parsing fidelity matters. If the table is added by JavaScript after the initial response, a normal HTTP request will not be enough: look for the page’s data endpoint or use browser automation.

What PHP can—and cannot—scrape from a table

PHP can extract a table that arrives in the fetched HTML without running a browser. The basic sequence is: request the page, check the response, parse the HTML, select a table and its rows, then normalize each cell’s text. This is a good fit for server-rendered tables and pages whose HTML includes the data even if styling or other behavior is supplied by JavaScript.

A key distinction is when the table appears. If JavaScript makes a later request and builds the table in the browser, a PHP HTTP client sees only the initial response. First inspect the response HTML and, where appropriate, the page’s network requests for an official or documented data endpoint. Use a browser-capable tool such as Symfony Panther only when the table genuinely depends on browser execution.

Before collecting data, respect the site’s terms, robots policy, authentication boundaries and rate limits. Prefer an official API when it supplies the same information. Scraping is not a way to bypass access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fetch the page and extract rows with DOMDocument

This complete example uses PHP’s cURL extension, checks the HTTP status, sets a timeout and a descriptive User-Agent, and extracts the first table as rows of cell text. Each output row is an indexed array, so the result remains usable even if headings are missing or duplicated.

<?php
declare(strict_types=1);

$url = 'https://example.com/table-page';

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 30,
    CURLOPT_USERAGENT => 'ExampleTableReader/1.0 (contact: [email protected])',
]);
$html = curl_exec($ch);
$status = (int) curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$error = curl_error($ch);
curl_close($ch);

if ($html === false) {
    throw new RuntimeException('Request failed: ' . $error);
}
if ($status < 200 || $status >= 300) {
    throw new RuntimeException("Unexpected HTTP status: $status");
}

// DOMDocument uses an HTML 4 parser. Suppress parser warnings for this
// extraction task, but do not treat this parser as an HTML sanitizer.
$previous = libxml_use_internal_errors(true);
$doc = new DOMDocument();
$loaded = $doc->loadHTML($html, LIBXML_NOERROR | LIBXML_NOWARNING);
$parseErrors = libxml_get_errors();
libxml_clear_errors();
libxml_use_internal_errors($previous);

if (!$loaded) {
    throw new RuntimeException('Could not parse the response as HTML.');
}

$xpath = new DOMXPath($doc);
$table = $xpath->query('(//table)[1]')->item(0);
if (!$table) {
    throw new RuntimeException('No table found in the returned HTML.');
}

$rows = [];
foreach ($xpath->query('.//tr', $table) as $row) {
    $cells = $xpath->query('./th | ./td', $row);
    $values = [];
    foreach ($cells as $cell) {
        $text = preg_replace('/\s+/u', ' ', $cell->textContent ?? '');
        $values[] = trim($text ?? '');
    }
    if ($values !== []) {
        $rows[] = $values;
    }
}

if ($rows === []) {
    throw new RuntimeException('The table was found, but it contained no cells.');
}

// Record provenance and inspect parser warnings in production logs.
$result = [
    'source_url' => $url,
    'retrieved_at' => gmdate(DATE_ATOM),
    'rows' => $rows,
];
if ($parseErrors !== []) {
    error_log('HTML parser reported ' . count($parseErrors) . ' warning(s) for ' . $url);
}

print_r($result);

Replace the example URL and contact string with values appropriate to your application. The example follows redirects and accepts any successful 2xx response; if the target site has different redirect or status expectations, adjust those checks deliberately. Keep the timeout finite so a slow origin does not hold a worker indefinitely.

Why query rows relative to the selected table

The XPath (//table)[1] selects the first table in the document, and .//tr then finds its rows. The relative expression avoids accidentally mixing rows from other tables into the result. To select another table, inspect the page’s structure and use a narrower XPath, for example a table with a known class or identifier. Avoid relying on a broad selector if the page has unrelated layout tables.

Within each row, ./th | ./td reads direct header and data cells in document order. Text is whitespace-collapsed and trimmed. That produces plain text, not the original HTML, links, attributes or nested structure. If those matter, extract the relevant elements or attributes explicitly rather than assuming textContent preserves them.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Turn the extracted rows into records

The raw $rows array is usually the safest first representation: it preserves what was found without assuming the page has a clean schema. If the table has a single header row and every data row has the same number of cells, map values to header labels only after validating that shape.

<?php
$headerIndex = null;
foreach ($rows as $index => $row) {
    // Treat the first row containing a <th> as the header row in a
    // production extractor; this compact example assumes row zero is headers.
    if ($index === 0) {
        $headerIndex = $index;
        continue;
    }

    $headers = $rows[$headerIndex];
    if (count($row) !== count($headers)) {
        // Log and handle the irregular row instead of silently mis-mapping it.
        continue;
    }
    $records[] = array_combine($headers, $row);
}

For real pages, track which cells were th elements rather than assuming the first row is a header: headers can appear in a thead, in multiple rows, or not at all. Normalize header names before using them as keys, and handle duplicates explicitly—PHP associative keys cannot represent two separate columns with the same key. Also decide how empty cells should be represented; an empty string, null and an omitted field have different meanings downstream.

Account for colspan, rowspan and irregular markup

A simple cell-to-array conversion treats each written cell as one value. It does not expand colspan or rowspan into the visual grid a reader sees. For a rectangular table, implement a grid-expansion pass: place each cell at the next unoccupied column, read its span attributes, fill the corresponding grid positions, and carry row spans into subsequent rows. Validate the resulting width and define whether repeated span values should be copied or represented as references.

Do not silently force irregular data into header/value records. A row with a different cell count may be a subheading, a grouped section, a malformed row, or evidence that the site changed its markup. Preserve or log it and make the extraction rule explicit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing a PHP parser or browser tool

Approach Best fit Trade-off
DOMDocument and DOMXPath Server-rendered tables, especially when avoiding Composer dependencies matters Its HTML 4 parsing rules can produce a DOM structure different from an HTML5 browser.
DomHTMLDocument PHP 8.4+ when standards-oriented HTML5 parsing fidelity matters Requires a current PHP runtime; check availability in the deployed environment.
Symfony DomCrawler Convenient CSS or XPath traversal after fetching HTML Adds a dependency; fetching and JavaScript execution remain separate concerns.
Simple HTML DOM Developers who prefer approachable CSS-like selectors Retrieval behavior depends on the setup; use cURL if hosting disables allow_url_fopen.
Symfony Panther or another browser automation tool Tables that exist only after client-side JavaScript runs Browser automation adds operational setup and resource cost compared with parsing a direct HTTP response.

PHP’s documentation recommends DomHTMLDocument for modern HTML instead of DOMDocument when parsing and processing current HTML. DOMDocument::loadHTML uses HTML 4 parsing rules, so malformed or modern markup can be represented differently from what a browser displays. Neither parser should be used as an HTML sanitizer.

<?php
// PHP 8.4+: use the HTML5-oriented parser when its parsing behavior
// is important for your page and runtime.
$document = DomHTMLDocument::createFromString($html);
$xpath = new DOMXPath($document);
$rows = $xpath->query('(//table)[1]//tr');

Confirm the PHP version and class availability on the machine that runs the scraper, not just on a development laptop. The PHP 8.4 HTML5 parser is the relevant option when HTML5-conforming parsing is required; it does not execute JavaScript.

When the table is rendered by JavaScript

  1. Inspect the response. Save or inspect the fetched HTML and search for a distinctive value expected in the table. If the table and values are absent, DOM parsing cannot recover them from that response.
  2. Look for a documented data source. If the page obtains table data from an API or other endpoint, prefer that documented interface, subject to its terms, authentication and rate limits.
  3. Use browser automation when necessary. If the page requires JavaScript execution and offers no suitable data endpoint, a browser-capable tool such as Symfony Panther can render the page before you query the table.
  4. Wait for the real condition. In browser automation, wait for a specific table or data selector instead of relying only on a fixed sleep. Check that the rendered table has the expected rows before extracting.

Browser rendering is slower and more operationally involved than a direct HTTP fetch. Reuse a documented endpoint where possible; use a browser only for the behavior that requires one. A screenshot proves what was visually rendered, but it is not a substitute for structured table extraction.

Validate results and keep extraction reliable

  • Check the HTTP response. Distinguish network errors, redirects and non-success status codes from a valid page with no table.
  • Check for the expected table. A missing table may indicate JavaScript rendering, a changed selector, a block page or a changed site layout.
  • Validate shape and content. Check expected headers, row counts or required fields where appropriate, while allowing legitimate empty cells.
  • Record provenance. Store the source URL and retrieval time alongside extracted data so downstream users can identify when and where it was collected.
  • Log warnings and changes. Parser warnings and schema changes should be observable; do not suppress them and then silently trust malformed output.
  • Be a considerate client. Set timeouts, avoid unnecessary repeated requests, and honor the site’s applicable rules and rate limits.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Practical fix
cURL returns false Connection, TLS, DNS or timeout failure Log curl_error(), verify the URL and server connectivity, and choose a finite timeout appropriate to the job.
HTTP status is not successful The server returned an error, redirect outcome or access response Log the status and inspect the response according to the site’s allowed access methods; do not treat an error page as table data.
No table is found The selector is wrong, the layout changed, or JavaScript adds the table after load Inspect the actual response HTML, narrow or correct the XPath, and use a documented endpoint or browser automation if needed.
Rows are missing or malformed The markup is irregular, the parser’s HTML 4 interpretation differs, or spans distort the visual grid Inspect the parsed structure, consider DomHTMLDocument on PHP 8.4+, and implement explicit span handling where required.
Columns map to the wrong headers Rows have different cell counts, repeated headers, duplicate labels or merged cells Validate row widths and header structure before mapping; preserve anomalous rows for review instead of silently combining them.
Cells contain unexpected whitespace Nested elements or line breaks contribute text nodes Normalize whitespace as in the example, and extract specific child elements if the desired value is not the cell’s full text.

Performance and cost considerations

For a single page, the main work is generally the network request and parsing its response. Avoid refetching the same page without a reason, keep timeouts bounded, and process large collections with a deliberate rate limit and failure policy. Browser automation has additional setup and runtime overhead because it must render pages; reserve it for pages whose data is unavailable in the response or a suitable endpoint. The appropriate infrastructure cost depends on page volume, browser requirements and the target’s allowed request rate; there is no single meaningful price or speed figure for every scraper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extraction quality is also a reliability cost. A selector can continue returning values after the site changes while assigning them to the wrong columns. Validate stable identifying headers and expected structure, log anomalies, and review changes instead of assuming that a successful parse means correct data.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a PHP table parser. Use it when the task is to capture a rendered page as an image or PDF rather than extract structured cell values. One GET request returns a screenshot or PDF; options include full-page capture, waiting for a selector, and executing custom JavaScript. Before capture, it can accept consent banners and remove known consent platforms, newsletter popups and chat widgets; those steps can be turned off. Bot checks, blank pages, timeouts and failed loads are not billed, and cache hits cost nothing. Its MCP server provides take_screenshot, get_page_info and capture_pdf for AI agents.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/table-page -o shot.webp

See the ScreenshotNeo API documentation for request options and response details. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month, with no card.

Frequently Asked Questions

Does DOMDocument run the page’s JavaScript?

No. It parses the HTML supplied to it; it does not run scripts or wait for client-side rendering.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use this approach to sanitize HTML before displaying it?

No. PHP documents that DOMDocument’s HTML parsing is not suitable as an HTML sanitizer; use a purpose-built sanitization approach for untrusted markup.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.