Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a website with PHP, fetch its HTML with cURL, verify both the transport result and HTTP status, parse the response with a DOM or HTML5 parser, and select fields with XPath or CSS selectors. The complete examples below show that workflow, then add pagination, retries, validation, caching, and responsible crawling practices. A normal HTTP request receives the server response; it does not execute JavaScript that later changes the page.

What you need before writing a scraper

  • PHP with the cURL extension enabled. PHP uses libcurl for HTTP and HTTPS requests.
  • A target URL that you are permitted to request, plus the site’s terms and published policies.
  • A parser. PHP’s built-in DOM APIs are enough for many pages; Symfony DomCrawler is convenient in Composer projects.
  • A saved sample response or fixture so selectors can be tested without repeatedly contacting the live site.

Do not treat scraping examples as a way to bypass authentication, CAPTCHAs, paywalls, rate limits, or other access controls. Request only the data you need and identify your client honestly.

1. Fetch a page with PHP cURL

curl_init() creates a cURL handle and curl_exec() performs the request. Set CURLOPT_RETURNTRANSFER when you want the response body as a string. An HTTP 404 or 500 is not, by itself, a cURL execution failure, so check the status code separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
declare(strict_types=1);

$url = 'https://example.com/';
$ch = curl_init($url);

if ($ch === false) {
    throw new RuntimeException('Could not initialize cURL');
}

curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_FOLLOWLOCATION => true,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 30,
    CURLOPT_USERAGENT => 'ExampleResearchBot/1.0 (contact: [email protected])',
]);

$html = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
$error = curl_error($ch);
curl_close($ch);

if ($html === false) {
    throw new RuntimeException("Transport failed: {$error}");
}
if ($status < 200 || $status >= 300) {
    throw new RuntimeException("Unexpected HTTP status: {$status}");
}

echo strlen($html), " bytes receivedn";

Use a strict $html === false check: an empty response is different from a cURL failure. In PHP 8, a successful curl_init() returns a CurlHandle object rather than the resource type used by older releases. Never disable TLS verification to conceal a certificate problem.

2. Parse HTML with DOMDocument and XPath

For many server-rendered pages, DOMDocument plus DOMXPath provides a small dependency-free solution. Real-world HTML can be malformed, so suppress parser warnings deliberately and inspect the resulting tree when a selector behaves unexpectedly.

$dom = new DOMDocument();
libxml_use_internal_errors(true);
$loaded = $dom->loadHTML($html);
$parseErrors = libxml_get_errors();
libxml_clear_errors();
libxml_use_internal_errors(false);

if ($loaded === false) {
    throw new RuntimeException('The response could not be parsed as HTML');
}

$xpath = new DOMXPath($dom);
$headings = $xpath->query('//article//h2');

if ($headings === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($headings as $heading) {
    $text = trim(preg_replace('/s+/', ' ', $heading->textContent));
    if ($text !== '') {
        echo $text, PHP_EOL;
    }
}

DOMDocument::loadHTML() does not implement the HTML5 parsing algorithm exactly. The PHP manual recommends DomHTMLDocument::createFromString() or DomHTMLDocument::createFromFile() for HTML5-conforming parsing; these APIs were added in PHP 8.4, so check the runtime before using them.

Extract links and attributes safely

$links = $xpath->query('//main//a[@href]');
foreach ($links ?: [] as $link) {
    $label = trim(preg_replace('/s+/', ' ', $link->textContent));
    $href = trim($link->getAttribute('href'));
    if ($href !== '') {
        printf("%s => %sn", $label, $href);
    }
}

Selectors should describe stable structure rather than presentation-only classes. Validate that required nodes exist and normalize whitespace before storing data. If a field is absent, record that fact instead of silently shifting values between records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Use Symfony DomCrawler in Composer projects

DomCrawler provides navigation and querying over HTML and XML. Install it with Composer:

composer require symfony/dom-crawler symfony/css-selector

When used outside a full Symfony application, include Composer’s autoloader. CSS selectors require the CssSelector component; XPath works without CSS syntax.

<?php
require __DIR__ . '/vendor/autoload.php';

use SymfonyComponentDomCrawlerCrawler;

$crawler = new Crawler($html, 'https://example.com/');

$titles = $crawler->filter('article h2')->each(
    static fn (Crawler $node): string => trim($node->text())
);

foreach ($titles as $title) {
    if ($title !== '') {
        echo $title, PHP_EOL;
    }
}

$prices = $crawler->filterXPath('//article[@data-product]//span[@data-price]')
    ->each(static fn (Crawler $node): string => trim($node->text()));

DomCrawler is a traversal layer, not a general-purpose DOM editing or re-dumping tool. Its parser may correct malformed markup, so compare unexpected selections with a saved response and the parser behavior you require.

4. Static HTML versus JavaScript-rendered pages

A cURL request receives the initial HTTP response. It does not run browser JavaScript, click consent dialogs, or wait for API calls made after page load. First inspect the downloaded HTML: if the desired text is present there, ordinary parsing is appropriate. If the response contains only an application shell and the data arrives through scripts, identify the underlying documented endpoint when you are authorized to use it, or use a legitimate browser-automation or rendering service. Do not attempt to evade bot checks or authentication.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Pagination, retries, and crawl pacing

Follow pagination deliberately

$next = 'https://example.com/articles';
$seen = [];
$pageLimit = 20;

for ($page = 1; $page <= $pageLimit && $next !== null; $page++) {
    if (isset($seen[$next])) {
        throw new RuntimeException('Pagination loop detected');
    }
    $seen[$next] = true;

    $html = fetch($next); // wrap the cURL example in this function
    $dom = new DOMDocument();
    libxml_use_internal_errors(true);
    $dom->loadHTML($html);
    libxml_clear_errors();
    $xpath = new DOMXPath($dom);

    foreach ($xpath->query('//article') ?: [] as $article) {
        // Extract and validate one record here.
    }

    $node = $xpath->query('//a[@rel="next"]/@href')->item(0);
    $next = $node ? trim($node->nodeValue) : null;
}

Resolve relative links against the site origin, cap the number of pages, and stop when the next link is missing or repeats. A parser should not be allowed to create an unbounded crawl.

Retry only transient failures

Use finite connect and total timeouts. A short exponential backoff can help with temporary network failures or a server response such as 429, but stop when the site denies access or continues throttling. Cache responses when the task permits it, and avoid sending the same request repeatedly while debugging selectors.

6. Robots rules, permissions, and privacy

RFC 9309 defines the Robots Exclusion Protocol. It states: “These rules are not a form of access authorization.” A robots.txt file communicates crawler requests; it does not grant permission and is not a replacement for the site’s terms, contracts, or applicable law. Review those sources separately, use conservative rates, and do not collect personal or sensitive information without a valid basis.

7. Troubleshooting common failures

Symptom Likely cause Fix
curl_exec() returns false DNS, TLS, connection, or timeout failure Log curl_error(), verify the URL and certificate chain, and adjust timeouts; do not turn off TLS verification.
HTML is returned but status is 404 or 500 HTTP error responses are separate from cURL transport errors Read CURLINFO_RESPONSE_CODE and handle non-2xx responses explicitly.
Selectors return zero nodes Wrong structure, changed markup, or JavaScript-generated content Save the response, inspect it, test a simpler XPath, and determine whether the data exists before browser execution.
DOM output differs from a browser loadHTML() uses non-HTML5 parsing rules Use PHP 8.4’s DomHTMLDocument APIs when available or test Symfony’s parser behavior against fixtures.
Requests are blocked or slowed Rate limits, policy restrictions, or bot defenses Stop, review permission and terms, reduce request volume, and never advise bypassing the control.
Memory rises during a large crawl Keeping every response or DOM in memory Process one page at a time, release DOM objects, stream results to storage, and enforce a page limit.

8. Reliability, performance, and cost decisions

  • Transport choice: native cURL offers low-level control and no framework dependency. A Symfony HTTP client can fit better when the rest of the application already uses Symfony components.
  • Parsing choice: native DOM avoids another traversal abstraction; DomCrawler makes XPath and CSS-style navigation easier in Composer projects.
  • Compatibility: choose an HTML5 parser when browser-like tree construction matters, and pin your PHP runtime in deployment so parser behavior does not change unexpectedly.
  • Throughput: concurrency is not automatically better. Respect the target’s limits, reuse cached responses, and measure your own workload rather than relying on undocumented performance rankings.
  • Data quality: store the source URL, retrieval time, HTTP status, and validation errors with each record so an extraction change can be diagnosed later.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When the page needs rendering, consent handling, or reliable screenshots rather than raw HTML extraction, ScreenshotNeo provides a single website screenshot API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options, including full-page capture, selectors, device presets, custom CSS and JavaScript, waits, headers, cookies, geolocation, PDF output, caching, async jobs, bulk capture, and signed links.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Can PHP scrape a page behind a login?

Only when you are authorized and have a permitted, documented authentication flow. Do not use these examples to defeat access controls.

Should I store the entire HTML response?

Keeping a fixture is valuable for selector tests and audits, but apply your retention and privacy requirements. For large crawls, process and persist validated fields incrementally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is DomCrawler a replacement for a browser?

No. It navigates the response tree supplied to it. It does not execute page JavaScript or reproduce all browser behavior.

Frequently Asked Questions

Can PHP scrape a page behind a login?

Only when you are authorized and have a permitted, documented authentication flow. Do not use these examples to defeat access controls.

Should I store the entire HTML response?

Keeping a fixture is valuable for selector tests and audits, but apply your retention and privacy requirements. For large crawls, process and persist validated fields incrementally.

Is DomCrawler a replacement for a browser?

No. It navigates the response tree supplied to it. It does not execute page JavaScript or reproduce all browser behavior.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.