Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

You can use Scrapy to request a Google Search results page, parse result blocks, and export fields such as title, URL, snippet, and their order in the response. But Scrapy does not make Google’s HTML a stable API or guarantee that Google will serve the page to an automated request. Treat direct scraping as a low-volume learning exercise, not a dependable production ranking service.

This tutorial builds that experiment, adds basic checks and cautious pagination, and explains when to use a structured API instead. The examples collect only ordinary organic-result blocks that the parser recognizes; they do not reliably capture every ad, local result, featured snippet, or other search feature.

Choose your collection method first

“Scraping Google” can mean three different things:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Requesting Google’s HTML: useful for learning Scrapy requests and parsing, but markup and access are variable.
  • Google Custom Search JSON API: returns structured data for a configured Programmable Search Engine; it is not a guaranteed copy of the ordinary live Google results page.
  • A managed SERP API: a vendor returns structured search-result data and manages much of the retrieval infrastructure, usually for a fee and under provider-specific limits and terms.

Important current status: Google says its Custom Search JSON API is closed to new customers. Existing customers have until January 1, 2027, to transition. Its documented former quota and pricing are not a signup option for new users. Check Google’s current documentation before building around an existing account.

Google’s Terms of Service address automated access that violates machine-readable instructions, as well as other restrictions. Whether a particular collection and use is permitted can depend on the method, applicable terms, jurisdiction, and data involved. Public visibility alone does not settle that question. Review the relevant terms and law; do not use this example to evade technical controls.

What the example will collect

The spider will try to extract a query, an ordinal rank, title, destination URL, snippet, timestamp, and source label. Here, rank means the position among organic-result blocks that this parser successfully extracts from one particular response. It is not a universal or definitive Google ranking: results may vary with location, language, device, time, personalization, and which SERP features appear.

Google Search may show organic links alongside ads, local packs, news, images, videos, featured snippets, People Also Ask, and other features. A basic HTML parser should not be assumed to capture those consistently. Even Google’s site: operator is not an exhaustive index or a reliable way to establish ranking; see Google’s search-operator guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. Create a Scrapy project

Use a supported Python installation; Python 3.11 or newer is a reasonable starting point. In a terminal, create and activate an environment, install Scrapy, and create the project:

mkdir google-serp-scraper
cd google-serp-scraper
python -m venv .venv

# macOS or Linux
source .venv/bin/activate

# Windows PowerShell: use this instead
# .venvScriptsActivate.ps1

python -m pip install --upgrade pip
python -m pip install scrapy
scrapy startproject google_serp .

The generated project includes a google_serp package and a spiders directory. Add a structured item in google_serp/items.py:

import scrapy


class SearchResult(scrapy.Item):
    query = scrapy.Field()
    rank = scrapy.Field()
    title = scrapy.Field()
    url = scrapy.Field()
    displayed_url = scrapy.Field()
    snippet = scrapy.Field()
    fetched_at = scrapy.Field()
    source = scrapy.Field()

displayed_url is optional because the visible breadcrumb or URL varies by markup. Keep the raw destination URL if later processing will normalize it.

2. Add a cautious direct-request spider

Create google_serp/spiders/google.py. This is an educational example: the CSS classes are response-specific and may stop matching. The user-agent identifies a client; it does not guarantee access or make a request acceptable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from datetime import datetime, timezone
from urllib.parse import urlencode

import scrapy

from google_serp.items import SearchResult


class GoogleSpider(scrapy.Spider):
    name = "google"
    allowed_domains = ["www.google.com"]

    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "DOWNLOAD_DELAY": 3,
        "RANDOMIZE_DOWNLOAD_DELAY": True,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 1,
        "AUTOTHROTTLE_ENABLED": True,
        "AUTOTHROTTLE_START_DELAY": 3,
        "AUTOTHROTTLE_MAX_DELAY": 30,
        "AUTOTHROTTLE_TARGET_CONCURRENCY": 0.5,
        "RETRY_ENABLED": True,
        "RETRY_TIMES": 2,
        "FEED_EXPORT_ENCODING": "utf-8",
    }

    def start_requests(self):
        queries = ["python web scraping", "scrapy tutorial"]

        for query in queries:
            params = {
                "q": query,
                "hl": "en",
                "gl": "us",
                "num": 10,
            }
            url = "https://www.google.com/search?" + urlencode(params)

            yield scrapy.Request(
                url=url,
                callback=self.parse,
                meta={"query": query},
                headers={
                    "User-Agent": (
                        "Mozilla/5.0 (compatible; ResearchBot/1.0; "
                        "+https://example.com/bot-info)"
                    )
                },
            )

    def parse(self, response):
        query = response.meta["query"]
        page_text = response.text.lower()

        # A 200 response can still contain consent or verification content.
        indicators = (
            "unusual traffic",
            "captcha",
            "not a robot",
            "before you continue to google",
        )
        if any(marker in page_text for marker in indicators):
            self.logger.warning(
                "Verification or consent response for query %r; stopping this page",
                query,
            )
            return

        # Illustrative, fragile selector: inspect and test your actual response.
        result_blocks = response.css("div.MjjYud")
        rank = 0

        for block in result_blocks:
            title = " ".join(
                part.strip()
                for part in block.css("h3::text").getall()
                if part.strip()
            )
            href = block.css("a[href]::attr(href)").get()
            snippet_parts = [
                part.strip()
                for part in block.css("div.VwiC3b ::text").getall()
                if part.strip()
            ]
            snippet = " ".join(snippet_parts)

            if not title or not href:
                continue

            rank += 1
            yield SearchResult(
                query=query,
                rank=rank,
                title=title,
                url=response.urljoin(href),
                displayed_url=None,
                snippet=snippet or None,
                fetched_at=datetime.now(timezone.utc).isoformat(),
                source="direct_html",
            )

        if rank == 0:
            self.logger.warning(
                "No result blocks extracted for %r. Inspect the saved response; "
                "the layout may differ or the page may not be search results.",
                query,
            )

The locale parameters hl=en and gl=us are hints, not a guarantee that the response matches what every English-language user in the United States sees. Network location, cookies, device context, and other factors can still affect results.

3. Run it and export results

From the project directory, run the spider and choose an output format:

# JSON Lines: one result object per line
scrapy crawl google -O results.jsonl

# CSV
scrapy crawl google -O results.csv

# JSON array
scrapy crawl google -O results.json

Scrapy’s -O option overwrites the target file. Feed exports are convenient for a prototype; recurring rank monitoring generally needs a database or another durable store. See the official feed exports documentation and item pipeline documentation.

4. Validate responses instead of trusting the status code

A successful HTTP status does not prove that the response contains search results. Consent screens, unusual-traffic pages, or changed layouts can produce empty output. During development:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Save a response fixture when parsing fails, and inspect it as HTML or text.
  2. Log the number of extracted results per query and flag unexpected zero-result pages.
  3. Test the parser against saved fixtures so markup changes are visible rather than silently corrupting output.
  4. Check whether the response is a verification page before attempting extraction.
  5. Keep retries bounded. Retrying a CAPTCHA or access-denied page is not a remedy.

Scrapy has built-in retry and downloader middleware behavior; consult its downloader middleware documentation. AutoThrottle can adjust crawl delays based on response latency, but it cannot grant access or solve a block; see AutoThrottle.

Symptom Possible explanation Responsible next step
HTTP 429 Request rate is too high or access is restricted Stop or back off; reduce volume and review the permitted collection method.
HTTP 403 or verification page Access denied or automated-access controls triggered Do not brute-force retries, rotate identities to evade controls, or solve CAPTCHAs automatically. Stop direct requests and use an appropriate authorized route.
Consent page Region, consent, or cookie state affects the response Do not treat it as a results page; record collection context and review the applicable requirements.
Zero results with a 200 Selector drift, alternate markup, or non-results content Save and inspect the response, then update and test the parser.
Different results than a browser Location, language, device, time, personalization, or other context differs Record the conditions and treat each observation as context-specific.

5. Add pagination only when the response supports it

Google’s start parameter is often used to request an offset, but do not assume it produces a stable or complete second page. Keep the requested offset separate from the extracted rank, cap the number of pages, stop when a page has no new results, and stop if verification content appears.

For an experiment, the request parameters might include:

params = {
    "q": query,
    "hl": "en",
    "gl": "us",
    "start": 10,
}

A pagination callback should carry the offset in request metadata and increase it only up to a configured maximum. Deduplicate across pages, but do not interpret ten parsed items as proof that ten definitive organic results were returned. Some results may be absent, repeated, or represented by other SERP features.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

6. Preserve collection context and normalize URLs carefully

For reproducible comparisons, store at least the query, UTC timestamp, source method, requested Google host and parameters, language and region hints, parser version, and any relevant device or network-location context you can legitimately record. A useful description of a result is “observed rank for query X under conditions Y at time Z,” not simply “Google rank.”

URLs can contain fragments, tracking parameters, redirect wrappers, or meaningful query parameters. Do not strip all query parameters indiscriminately. Preserve the original value, and use a separate normalized field for deduplication. A conservative normalization can remove fragments and standardize scheme/host casing:

from urllib.parse import urldefrag, urlsplit, urlunsplit


def normalize_url(url):
    url, _fragment = urldefrag(url)
    parts = urlsplit(url)
    return urlunsplit((
        parts.scheme.lower(),
        parts.netloc.lower(),
        parts.path or "/",
        parts.query,
        "",
    ))

Keep both raw_url and normalized_url if you later add URL normalization. Deduplicate by your chosen normalized key while retaining the query and first observed position. Canonicalization is a policy decision, not a license to discard destination parameters that matter.

7. Decide whether direct HTML is the right method

Method Best suited to Main limitations
Direct Google HTML Learning request handling and parsing with a small, controlled experiment Fragile markup, variable responses, access controls, and terms review; high maintenance for reliable monitoring.
Custom Search JSON API Existing eligible users searching through a configured Programmable Search Engine Closed to new customers; results are not necessarily the ordinary live SERP; transition deadline for existing customers is January 1, 2027.
Managed SERP API Structured, recurring SERP collection where location and features matter Cost, provider-specific fields and limits, vendor dependence, and the need to review provider and target-service terms.

If you have an eligible Custom Search API account, Google’s Search API reference documents its request and response fields. It requires a Programmable Search Engine configuration and API key. Do not assume it reproduces ordinary Google.com rankings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a managed SERP provider, Scrapy can still schedule queries, validate and normalize responses, and export or store results. The retrieval endpoint, authentication, schema, quotas, and costs differ by vendor; follow that provider’s current official documentation rather than copying a generic endpoint. Compare cost per successful query, location fidelity, required SERP features, historical retention, failure handling, and exit options. A provider does not by itself establish that a particular use complies with Google’s terms or applicable law.

Production checklist

  • Define whether you need organic links, specific SERP features, or site-restricted search.
  • Review Google’s applicable terms and machine-readable instructions, the provider’s terms, and relevant law.
  • Estimate query volume and acceptable failure rate before selecting direct requests or a paid service.
  • Record timestamp, query, locale, source, and parser or provider version.
  • Test multiple queries and collection contexts; monitor extraction counts and unexpected response types.
  • Use bounded retries and backoff. Stop rather than escalating requests after a block.
  • Keep test fixtures, logs, and an alert for empty or sharply reduced output.
  • Set data-retention rules and protect query data where it could reveal sensitive interests.
  • For recurring jobs, plan for database storage, provider outages, schema changes, and cost review.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.