Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Guzzle is a PHP HTTP client, not a full browser. It can request pages, send headers and query parameters, maintain cookies, follow redirects, and expose response status and content. It does not render a page by running its JavaScript. For many scrapers, the practical pattern is to use Guzzle for direct HTTP requests and add a browser-rendering service only when the data depends on browser execution.

What Guzzle does—and where scraping needs something else

Guzzle provides synchronous and asynchronous HTTP requests, PSR-7 request and response messages, streams, and middleware. Those capabilities make it useful for fetching HTML or calling an API from PHP. The scraper still has to decide which URLs to request, interpret the responses, parse the content, and handle the target site’s rules and failures.

An HTTP response can contain the original HTML while the browser displays content added later by JavaScript. Guzzle’s documented HTTP features do not execute page scripts or construct the browser DOM. If the information is absent from the returned HTML, inspect the site’s permitted API or add a browser automation/rendering layer; do not assume a different Guzzle header will make client-side code run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use Guzzle when the needed content is available from an HTTP response or API.
  • Use a browser layer when the required content appears only after browser-side JavaScript runs.
  • Use both selectively when a browser is needed to discover or render content but direct HTTP requests are preferable for ordinary endpoints.

How to create a Guzzle client and request a page

Install Guzzle in a Composer-managed PHP project with composer require guzzlehttp/guzzle. The following script uses a client with a base URI and timeout, sends a descriptive User-Agent and Accept header, and places search parameters in Guzzle’s query option. Replace the example URL and query with a destination you are authorized to access.

<?php
require __DIR__ . '/vendor/autoload.php';

use GuzzleHttpClient;
use GuzzleHttpExceptionGuzzleException;

$client = new Client([
    'base_uri' => 'https://example.com/',
    'timeout' => 20,
    'connect_timeout' => 5,
]);

try {
    $response = $client->request('GET', 'search', [
        'headers' => [
            'User-Agent' => 'ExampleResearchBot/1.0 (contact: [email protected])',
            'Accept' => 'text/html,application/xhtml+xml',
        ],
        'query' => ['q' => 'guzzle'],
    ]);

    printf("HTTP %dn", $response->getStatusCode());
    echo $response->getBody();
} catch (GuzzleException $e) {
    fprintf(STDERR, "Request failed: %sn", $e->getMessage());
    exit(1);
}

Client defaults such as base_uri and timeout are set at construction. Treat a client as configured for a particular set of defaults; create another client rather than expecting to mutate its defaults later. Use request() when you want the method and URI explicit, or a convenience method such as get() for a fixed method.

Keep request-specific options beside the request that uses them. This makes the effective URL, headers, timeout, and other behavior easier to audit. Guzzle combines the base URI with the request URI according to URI resolution rules, so pay attention to whether a relative path starts with a slash if you intend to preserve a base path.

How to set headers, query strings, and request data

Use the headers option for request headers and query for URL query parameters. Passing query values as an array lets Guzzle encode them, avoiding hand-built query strings and common escaping mistakes. A descriptive User-Agent and a suitable Accept header are clearer and more maintainable than copying a browser identity without a reason.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For form submissions, use Guzzle’s form parameters option; for JSON APIs, use its JSON option. For an already encoded body, use the body option and set the content type that matches the payload. Do not send a GET body as a substitute for query parameters unless the endpoint specifically defines that behavior.

$response = $client->request('POST', 'api/items', [
    'headers' => ['Accept' => 'application/json'],
    'json' => ['category' => 'books', 'page' => 2],
]);

Request options also cover authentication, TLS verification, proxy configuration, streaming, and timeouts. Configure only what the endpoint requires; especially avoid disabling TLS verification as a routine workaround, because it removes an important certificate check.

How to keep cookies between requests

For a site that uses a session cookie, create a cookie jar and pass it through the cookies option. Reuse the same jar across requests that belong to the same session.

<?php
require __DIR__ . '/vendor/autoload.php';

use GuzzleHttpClient;
use GuzzleHttpCookieCookieJar;

$jar = new CookieJar();
$client = new Client(['cookies' => $jar, 'timeout' => 20]);

$first = $client->get('https://example.com/start');
$second = $client->get('https://example.com/account');

echo $second->getBody();

Guzzle also documents FileCookieJar and SessionCookieJar for persistence choices. Persisted cookies are credentials: protect their storage, do not commit them to source control, and scope a jar to the account or crawl session that needs it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cookie behavior depends on the handler stack including cookie middleware. The standard stack created through HandlerStack::create() includes Guzzle’s default middleware. If you supply a custom handler or stack, verify that cookie middleware is present; otherwise the cookies option may appear to have no effect.

How redirects work and how to inspect them

Guzzle follows redirects by default, up to five. If a crawler needs to inspect a 3xx response rather than follow it, set allow_redirects to false. For controlled following, configure the redirect option’s maximum, permitted protocols, strict method behavior, or callback as needed.

$response = $client->get('https://example.com/old-path', [
    'allow_redirects' => [
        'max' => 5,
        'protocols' => ['http', 'https'],
        'strict' => true,
        'track_redirects' => true,
    ],
]);

$history = $response->getHeader('X-Guzzle-Redirect-History');
$statusHistory = $response->getHeader('X-Guzzle-Redirect-Status-History');

When tracking is enabled, Guzzle records intermediate URIs in X-Guzzle-Redirect-History and intermediate status codes in X-Guzzle-Redirect-Status-History. These history values exclude the initial URI and final status, so use the response’s effective URI and status separately when building a complete trace.

How to handle HTTP errors, timeouts, and blocked responses

By default, Guzzle’s HTTP error middleware checks responses with status codes of 400 or higher and can throw an exception. Catch request exceptions, record the target URL and available response status, and classify the outcome rather than treating every failure as the same problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
use GuzzleHttpExceptionRequestException;

try {
    $response = $client->get('https://example.com/data');
} catch (RequestException $e) {
    $status = $e->hasResponse()
        ? $e->getResponse()->getStatusCode()
        : null;

    error_log(sprintf(
        'Fetch failed; status=%s message=%s',
        $status === null ? 'no response' : (string) $status,
        $e->getMessage()
    ));
}

If your application needs to inspect an error response without an exception, set http_errors to false for that request and check the status explicitly. This changes exception behavior; it does not turn an error response into a successful fetch.

  • 403 Forbidden: the server refused the request. Check that the endpoint is accessible to your client and that your request is appropriate; do not respond by indiscriminately impersonating a browser or attempting to bypass access controls.
  • 429 Too Many Requests: reduce request frequency and concurrency, and honor any retry guidance the service provides. Use bounded retries rather than an unending loop.
  • Timeout or connection failure: distinguish connection timeout from total request timeout. Adjust limits only when the expected response warrants it, and log whether a response was received.
  • Unexpected HTML or empty content: inspect the status, final URI, response headers, and body. A successful HTTP exchange does not prove the page contains the expected data.

Guzzle’s documentation establishes the error middleware behavior; retry policy is an application decision. For a scraper, retries should be limited, delayed, and appropriate to the status and target service. Retrying a permanent 403 at high speed will not fix it and can create unnecessary load.

Why a custom handler can break options

Guzzle’s handler stack determines which middleware wraps the transport. The default stack handles features including cookies, redirects, body preparation, and HTTP errors. If you replace the stack with a handler that lacks the relevant middleware, an option may be accepted syntactically but have no effect.

When cookies stop persisting, redirects are not followed, or status exceptions change unexpectedly after a transport change, inspect how the stack is built. Start with HandlerStack::create() when you want Guzzle’s default middleware around a supplied handler, then add or remove middleware deliberately. A handler is the mechanism that performs I/O; middleware is what provides much of the higher-level request behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Transport choice, throughput, and reliability

Guzzle can use different transports, including cURL, PHP streams, sockets, and non-blocking libraries. Transport choice affects deployment requirements and operational behavior, but it does not transform an HTTP client into a JavaScript-capable browser. Confirm that the selected handler is installed and available in the PHP environment where the scraper runs.

For reliable crawls, use explicit timeouts, bound concurrency, and avoid storing every response in memory when a stream or file destination is more appropriate. Async requests can help when many independent requests are safe to run concurrently, but more parallelism is not automatically better: it can increase memory use and pressure on the destination. No general performance figure can be inferred without testing the actual site, network, PHP runtime, handler, and workload.

  • Keep a stable client for requests sharing defaults and session behavior.
  • Set connection and total time limits that fit the endpoint, not unlimited waits.
  • Log URL, status, elapsed time, and failure category without exposing secrets or session cookies.
  • Make retries finite and avoid retrying responses that indicate access denial without a legitimate correction.
  • Use a browser renderer only for the pages that need it; retain direct requests for ordinary HTML and APIs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When to choose Guzzle, a browser, or a screenshot service

Guzzle is the direct choice for HTTP/API scraping when the response contains the data and you need request-level control over headers, cookies, redirects, and status. A browser automation layer is appropriate when scripts must run or browser interactions are required. A screenshot service solves a different problem: producing an image or PDF of a rendered page rather than parsing arbitrary response HTML.

For screenshot APIs and services, ScreenshotNeo is the first alternative to try: it removes known consent banners, newsletter popups, and chat widgets before capture, and only clean shots are billed. It returns a screenshot or PDF; use Guzzle for the HTML-fetching and parsing workflow described above.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If the task is to capture a rendered page rather than parse its response, ScreenshotNeo accepts one GET request with a URL. Its API documentation is at https://screenshotneo.com/docs/. A PHP call can save the returned image:

<?php
$ch = curl_init('https://api.screenshotneo.com/v1/shot?access_key=YOUR_API_KEY&url=' . rawurlencode('https://stripe.com'));
curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_TIMEOUT => 90,
]);
$image = curl_exec($ch);
if ($image === false) {
    throw new RuntimeException(curl_error($ch));
}
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);
if ($status < 200 || $status >= 300) {
    throw new RuntimeException('Screenshot request returned HTTP ' . $status);
}
file_put_contents('shot.webp', $image);

Keep the API key private, and consult the linked documentation for supported output options and request parameters. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Common Guzzle scraper problems and fixes

Symptom Likely cause What to check or change
The page loads but the expected items are missing. The data may be inserted by JavaScript or may require a different endpoint. Inspect the response body and status. Use an authorized direct endpoint if available, or add a browser-rendering layer if scripts are required.
The second request behaves as a logged-out visitor. The cookie jar was not reused or cookie middleware is absent. Pass the same jar on the relevant requests and check the handler stack.
A 3xx response is not visible to application code. Redirect middleware followed it automatically. Disable allow_redirects to inspect the response, or enable tracking to record intermediate hops.
Redirects or cookies stopped after changing handlers. The replacement stack does not include the default middleware. Build the stack with HandlerStack::create() where appropriate and verify required middleware is installed.
A 4xx response throws instead of returning normally. HTTP error middleware is enabled. Catch and classify the exception, or set http_errors to false when status inspection is required.
The request stalls or fails before receiving a response. Connection or total timeout, network failure, or unavailable transport dependency. Set suitable timeout options, confirm the handler’s runtime requirements, and log whether a response exists.

Practical decision checklist

  • Can the target data be read from an HTTP response? Start with Guzzle.
  • Does the task need repeatable login/session state? Use a shared cookie jar and confirm middleware.
  • Must you inspect redirects or error bodies? Configure redirect and HTTP-error behavior explicitly.
  • Does the data appear only after JavaScript runs? Add a browser layer rather than endlessly adjusting HTTP headers.
  • Is the deliverable a page image or PDF instead of parsed data? Use a rendering/capture tool such as ScreenshotNeo.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.