Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchYes—PHP can scrape HTML. A basic scraper requests a permitted page, checks the response, parses its HTML, selects the fields you need, normalizes the values, and stores or outputs them. For a first static page, PHP’s built-in HTTP tools plus DOMDocument and DOMXPath are enough; use Guzzle for a more convenient HTTP client, and Symfony DomCrawler when CSS selectors and higher-level extraction helpers would make the parser easier to maintain.
How PHP web scraping works
Scraping is a data pipeline, not a single command. First, request a page you are allowed to access. Next, verify that the server returned a usable response, parse the HTML, locate the elements containing the data, normalize values such as whitespace and URLs, and then emit or save the records.
- Request: fetch the page with an explicit user agent and a bounded timeout.
- Validate: inspect the HTTP status and content type rather than assuming every response is a page.
- Parse: turn the HTML string into a document tree.
- Select: use XPath or CSS selectors to find the fields.
- Normalize: trim text and resolve relative links against the page URL.
- Output: print, encode, or persist the structured records.
This guide uses a static-page example. The target URL is illustrative: replace it with a page whose access rules permit your use. Do not treat scraping as automatically lawful or permitted. Terms, privacy, copyright, contracts, and applicable jurisdiction can affect what is allowed; robots.txt alone does not grant permission. Use permitted sources and conservative request rates.
Fetch a static page with PHP
PHP’s HTTP stream wrapper can make a request without an additional package. A stream context lets you set a user agent, timeout, and redirect behavior. The example checks both transport errors and the HTTP status before passing the body to a parser.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
<?php
$url = 'https://example.com/catalog';
$context = stream_context_create([
'http' => [
'method' => 'GET',
'timeout' => 15,
'ignore_errors' => true,
'follow_location' => 1,
'max_redirects' => 5,
'header' => "User-Agent: ExampleResearchBot/1.0 (contact: [email protected])rn",
],
]);
$html = file_get_contents($url, false, $context);
if ($html === false) {
throw new RuntimeException('Request failed: ' . error_get_last()['message']);
}
$statusLine = $http_response_header[0] ?? '';
if (!preg_match('~s(d{3})s~', $statusLine, $match)) {
throw new RuntimeException('Could not read HTTP status: ' . $statusLine);
}
$status = (int) $match[1];
if ($status < 200 || $status >= 300) {
throw new RuntimeException("Unexpected HTTP status $status");
}
$contentType = '';
foreach ($http_response_header as $header) {
if (stripos($header, 'Content-Type:') === 0) {
$contentType = trim(substr($header, strlen('Content-Type:')));
break;
}
}
if ($contentType !== '' && stripos($contentType, 'text/html') === false) {
throw new RuntimeException('Expected HTML, received: ' . $contentType);
}
// Continue with the HTML parsing example below.
?>
ignore_errors allows PHP to expose an error response body and status for inspection; it does not make an unsuccessful status successful. The stream wrapper can be configured in php.ini or per-request with a stream context. Check that the PHP installation allows URL streams before relying on this approach.
When to use cURL instead
PHP’s cURL extension gives you more explicit control over request options and is useful when you need concurrent requests. For one simple page, the stream wrapper avoids an extra dependency. Whichever client you choose, set timeouts, inspect the status, and handle redirects and failures deliberately. A successful connection only means a server responded; it does not mean the response is the expected HTML.
Parse HTML with DOMDocument and XPath
DOMDocument and DOMXPath are built into PHP’s DOM extension and expose the document tree directly. XPath is precise and does not require a CSS-selector package.
<?php
libxml_use_internal_errors(true);
$document = new DOMDocument();
$loaded = $document->loadHTML('<?xml encoding="UTF-8">' . $html);
$parseErrors = libxml_get_errors();
libxml_clear_errors();
libxml_use_internal_errors(false);
if (!$loaded) {
throw new RuntimeException('Could not parse the response as HTML');
}
$xpath = new DOMXPath($document);
$records = [];
foreach ($xpath->query('//article') as $article) {
$titleNodes = $xpath->query('.//h2', $article);
$linkNodes = $xpath->query('.//a[@href]', $article);
$title = $titleNodes->item(0)?->textContent ?? '';
$href = $linkNodes->item(0)?->getAttribute('href') ?? '';
$records[] = [
'title' => trim(preg_replace('/s+/u', ' ', $title)),
'url' => $href === '' ? null : resolveUrl($url, $href),
];
}
function resolveUrl(string $base, string $href): string
{
if (parse_url($href, PHP_URL_SCHEME) !== null || str_starts_with($href, '//')) {
return str_starts_with($href, '//')
? ((parse_url($base, PHP_URL_SCHEME) ?? 'https') . ':' . $href)
: $href;
}
$parts = parse_url($base);
$origin = ($parts['scheme'] ?? 'https') . '://' . ($parts['host'] ?? '');
if (str_starts_with($href, '/')) {
return $origin . $href;
}
$path = $parts['path'] ?? '/';
return $origin . rtrim(substr($path, 0, strrpos($path, '/') + 1), '/') . '/' . $href;
}
?>
The helper handles common absolute, protocol-relative, root-relative, and path-relative links. For production use, account for query-only references, fragments, and dot segments as well; a standards-compliant URI-resolution library is preferable when inputs are varied. Do not assume an extracted href is already a complete URL.
libxml_use_internal_errors(true) prevents malformed but recoverable markup from flooding output with parser warnings. HTML on the open web is often imperfect, so inspect parse errors when results look wrong rather than silently assuming the tree matches the source.
Use Symfony DomCrawler for CSS selectors
Symfony describes DomCrawler as easing navigation for HTML and XML documents. It wraps the parsed document with convenient methods for CSS selection, XPath filtering, text and attribute extraction, and iteration. CSS selectors are often easier to read when the target’s markup is familiar.
Rank #2
- Used Book in Good Condition
- Install the components from the directory containing your PHP project’s
composer.json.
composer require symfony/dom-crawler symfony/css-selector
Then give DomCrawler the HTML and extract fields from matching elements:
<?php
require __DIR__ . '/vendor/autoload.php';
use SymfonyComponentDomCrawlerCrawler;
$crawler = new Crawler($html);
$rows = $crawler->filter('article')->each(
static function (Crawler $node): array {
$link = $node->filter('a[href]');
return [
'title' => trim($node->filter('h2')->text('')),
'url' => $link->count() ? $link->attr('href') : null,
];
}
);
print json_encode($rows, JSON_PRETTY_PRINT | JSON_UNESCAPED_SLASHES);
?>
The example checks whether a link exists before asking for its attribute, and supplies an empty default for a missing heading. This matters because real pages do not always use identical markup for every record. Resolve any relative URL before treating it as a destination.
CSS selectors or XPath?
- Choose CSS for concise selectors such as
article h2ora[href], especially when working with designers’ or browser-inspector terminology. - Choose XPath for precise relationships, text-based conditions, and traversal up or across the tree; it is available directly through
DOMXPathor via DomCrawler’sfilterXPath(). - Choose DomCrawler when the convenience of
filter(),filterXPath(),attr(),text(),extract(), andeach()outweighs adding dependencies.
DomCrawler is for navigating and extracting from a document, not for generally re-dumping an arbitrary DOM as formatted HTML. Keep the original response if you need to preserve the source representation.
Choose an HTTP client: built-in PHP, Guzzle, or cURL
Guzzle is an HTTP client installed through Composer. Its handler system can use PHP’s stream wrapper when cURL is unavailable; cURL remains relevant when concurrent requests are needed. Install it with:
composer require guzzlehttp/guzzle
A minimal request with a status check looks like this:
<?php
require __DIR__ . '/vendor/autoload.php';
$client = new GuzzleHttpClient([
'timeout' => 15,
'connect_timeout' => 5,
'allow_redirects' => ['max' => 5],
'headers' => ['User-Agent' => 'ExampleResearchBot/1.0 (contact: [email protected])'],
]);
$response = $client->get('https://example.com/catalog', ['http_errors' => false]);
$status = $response->getStatusCode();
if ($status < 200 || $status >= 300) {
throw new RuntimeException("Unexpected HTTP status $status");
}
$html = (string) $response->getBody();
?>
| Approach | Setup | Selection and navigation | Best fit |
|---|---|---|---|
| PHP streams | Built in when URL streams are enabled | Pair with DOMDocument/DOMXPath; no browser-like navigation helpers | A small, sequential fetch without another HTTP package |
| Guzzle | Composer package | Pair with DOM tools or DomCrawler; client handles HTTP requests | A reusable client with configurable handlers and timeouts |
| cURL | Requires PHP cURL extension | Pair with DOM tools or DomCrawler; request handling is explicit | More control or concurrent requests |
Handle links, forms, and multi-page navigation
Some data requires following ordinary links or submitting a form. Symfony BrowserKit provides a browser-like request model for requests, link clicks, form submissions, JSON requests, and XMLHttpRequest-style requests. It can help when a workflow depends on successive server responses rather than a single page.
BrowserKit simulates browser request behavior; it does not execute arbitrary JavaScript or render a client-side application. If a form depends on JavaScript to create a request or the next page is assembled in the browser, BrowserKit alone may not reproduce what a person sees. Use a site’s permitted API or an authorized rendering method for that case.
For pagination, identify the site’s actual next-page link or documented page parameter, stop when no next page exists, and track records already seen. Apply a request delay appropriate to the target and avoid repeatedly fetching the same pages. Do not guess that a numbered URL pattern is valid without checking the site’s links or documentation.
Why JavaScript-rendered data may be missing
A normal PHP HTTP request receives the server’s response body. If the page fills its content after load using JavaScript, the initial HTML may contain only a shell, and selectors for the visible content will find nothing. Bot-protection systems can also return a challenge or a different response from the expected page. A plain fetch can therefore differ from what a browser displays.
First inspect the returned status, content type, and a small portion of the HTML. If the expected records are absent, check whether the site publishes an official API or data feed that permits your use. Otherwise, use an authorized rendering solution and comply with the target’s access rules. Do not try to evade bot checks or access controls.
Recommended Free Tools
Or skip the browser setup
If the page requires a rendered capture rather than a PHP DOM parse, ScreenshotNeo provides a website screenshot API and MCP server for developers. Its capture pipeline accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Only clean shots are billed: bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome indicated by X-Page-Verdict and X-Billed headers. The MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and other MCP clients.
One GET request can return an image or PDF. The following cURL example saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for authentication, supported formats, and request options. The service also offers full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, HTML/CSS capture, custom CSS and JavaScript, pre-capture clicks, selector hiding, wait conditions, request and resource blocking, custom headers, cookies, user agents and authorization, timezone and geolocation, transparent backgrounds, image resizing, configurable cache TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.
Rank #4
Free includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is included on every plan. Sign up free for 1,000 screenshots a month with no card.
Free tools Windows power users keep installed
One-click scans. No signup required.
Troubleshoot common PHP scraping failures
The request returns false or times out
Check outbound network access, DNS and TLS errors, whether URL streams are enabled, and whether the target is reachable from the PHP runtime. Set a finite timeout and report the transport error. A timeout is a failed request, not a signal to retry indefinitely; use bounded retries only where permitted and appropriate.
You receive an error page or unexpected status
Log the status line and inspect the response content type and a limited body excerpt. A server may return an error page, redirect, login page, or challenge instead of the requested document. Follow redirects only within a reasonable limit and do not assume a 200 response contains the target data.
DOMDocument reports warnings or text looks garbled
HTML can be malformed, and character encoding declarations may be missing or inaccurate. Use libxml’s internal-error mode to capture parser diagnostics, and verify the response’s encoding before parsing. Avoid blindly converting strings: an incorrect conversion can corrupt text. Ensure normalization functions use the intended encoding, such as UTF-8.
A selector returns no nodes or fails after a redesign
Inspect the fetched HTML rather than the browser’s rendered inspector view. Confirm that the content is in the initial response and that the selector matches the actual structure. Prefer stable attributes or semantic relationships over brittle positional selectors. Add checks for missing nodes so a markup change produces a clear warning instead of silently storing empty values.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Links are relative or records are duplicated
Resolve relative references against the source page URL and preserve query strings when they identify records. For pagination, define a stable record key—such as an allowed source ID or normalized URL—and deduplicate before writing. Make the output operation idempotent so a repeated run does not create duplicate rows.
Best Value
Some pages work but JavaScript content is absent
Compare the response HTML with the visible page and check whether an official permitted API exposes the same data. If the page requires browser execution, use an authorized rendering method rather than expecting DOMDocument, Guzzle, or BrowserKit to execute client-side code.
Reliability, performance, and responsible collection
- Bound each request: set connection and overall timeouts, cap redirects, and handle non-success statuses explicitly.
- Keep concurrency conservative: parallelism can reduce elapsed time but increases load on the target. Use cURL or a suitable client for concurrency only when permitted, and keep limits low.
- Make runs resumable: store progress and deduplicate records so a transient failure does not force a full restart.
- Validate the output: record counts, required-field checks, and a sample of normalized records can reveal parser drift early.
- Minimize collection: retrieve only needed fields, use permitted sources, respect access rules, and avoid collecting sensitive personal data without a valid basis.
There is no universal safe request rate or legal rule for every site. Follow the site’s terms and applicable requirements, and contact the site owner when permission or an access method is unclear.
Frequently Asked Questions
Can PHP scrape HTML without installing a package?
Yes. PHP’s HTTP stream wrapper can fetch a page, while the DOM extension provides DOMDocument and DOMXPath for parsing and selection.
Does Symfony BrowserKit run JavaScript?
No. It models requests, clicks, and form submissions; it does not execute arbitrary client-side JavaScript.
Is scraping allowed if a site permits it in robots.txt?
Not necessarily. Robots rules do not by themselves grant permission; terms, privacy, copyright, contracts, and jurisdiction may also matter.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

