Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Gemini is not a complete web crawler. In Python, treat scraping as two separate jobs: obtain the page (with your HTTP client or Gemini URL Context), then ask Gemini to extract the fields you need. This separation gives you control over URLs, permissions, retries and storage while leaving interpretation of messy text to the model.

The two operations you must keep separate

1. Fetching content

Your program first obtains HTML or another supported representation. A normal HTTP request gives you control over headers, timeouts, rate limits, caching and error handling. You decide which URLs to request and how to respect a site’s access controls.

2. Extracting information

After fetching, send only the relevant page content to Gemini with a precise schema. Ask for fields, types and missing-value behavior instead of a vague “scrape this page” prompt. Extraction is interpretation; it does not make the initial request legal, reliable or complete.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Option A: fetch a page in Python, then ask Gemini to extract it

The following is a package-level pattern, not a guarantee of current client-library syntax. Confirm the current Gemini SDK and HTTP/HTML-parser documentation before deploying it. The example deliberately separates retrieval, status checking, content selection and model output.

  1. Choose a URL that you are allowed to access and set a clear timeout.
  2. Fetch it and stop on an unsuccessful HTTP status.
  3. Extract the useful text (or selected elements) and cap its size.
  4. Send that text to Gemini with a strict JSON shape.
  5. Validate the returned object before writing it to your database.
import json
import os
import requests
from bs4 import BeautifulSoup
from google import genai

url = "https://example.com/products/widget"

# 1) Fetch
response = requests.get(
    url,
    headers={"User-Agent": "MyResearchBot/1.0 ([email protected])"},
    timeout=30,
)
response.raise_for_status()

# 2) Select readable content; do not send an entire site blindly
soup = BeautifulSoup(response.text, "html.parser")
for node in soup(["script", "style", "noscript"]):
    node.decompose()
text = soup.get_text(" ", strip=True)
text = text[:120_000]                 # choose a limit for your model and task

# 3) Extract a defined structure
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
prompt = f"""
Extract product data from the page text below.
Return JSON only, matching this shape:
{{"name": string|null, "price": number|null, "currency": string|null,
  "availability": string|null, "source_url": "{url}"}}
Use null when a value is not present. Do not infer values.

PAGE TEXT:
{text}
"""
result = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=prompt,
)
record = json.loads(result.text)
print(json.dumps(record, ensure_ascii=False))

The model name and SDK import can change. Pin and review the version you deploy, handle non-JSON responses, and validate types with your preferred schema library. For large pages, select the article or product container rather than truncating in the middle of important content. Keep the original URL, retrieval time, response status and a content hash alongside the extracted record so you can audit changes.

Option B: let Gemini URL Context retrieve supplied URLs

Google describes URL Context as a way to provide additional context to models in the form of URLs. You give Gemini the full, known URL and ask it to analyze the retrieved content. This is useful when your application already knows its targets and does not need custom crawling logic.

  • One request can process up to 20 URLs.
  • Retrieved content is limited to 34 MB per URL.
  • URLs must be publicly accessible; paywalled pages and some content types are unsupported.
  • Google says retrieval may use indexed content first and fall back to a live fetch.
  • Responses can include URL citation annotations and retrieval metadata.
  • It does not follow nested links from a supplied page, so it is not an unrestricted crawler.

A URL Context request should still state the output schema, missing-value rule and comparison logic. Supplying ten product URLs for a comparison is different from asking the model to discover every link on a site; the latter requires your own link-discovery and queueing code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

URL Context versus Python fetching

Question Python fetch, then Gemini Gemini URL Context
Who retrieves? Your application and HTTP stack Gemini’s URL Context tool
URL control Headers, cookies, retries and pacing are yours You provide specific public URLs
Processing limit Set by your infrastructure and model request Up to 20 URLs and 34 MB per URL
Nested links Implement discovery yourself Not followed automatically
Best fit Repeatable pipelines needing audit and custom access logic Known public pages that Gemini can retrieve directly

Do not use Google Search grounding as a crawl-target finder

The Gemini API Additional Terms effective March 23, 2026 prohibit programmatic or automated collection of Grounded Results, Search Suggestions or Links for another purpose, including using links to identify destination pages for crawling or scraping. Fetching a URL your application already knows is a different workflow from collecting search-grounding links to build a crawl queue. Review the terms that apply to your service and geography before shipping.

Gemini CLI is another interface, not a Python crawler

The Gemini CLI web_fetch tool accepts URLs in a prompt and uses Gemini API URL Context. It can be convenient for interactive work, but it is not a Python library and should not be presented as a drop-in replacement for a custom crawler. For scheduled jobs, keep a clear boundary between the CLI workflow and your Python service.

Designing reliable extraction prompts

Specify a contract

Name every field, its type, allowed enum values and the treatment of absent or conflicting text. Tell the model not to infer prices, dates or availability. Include the source URL in the requested object.

Reduce and label input

Remove navigation, scripts and repeated boilerplate when possible. Mark the beginning and end of page text, and identify tables or headings so the model can distinguish labels from values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate and retry safely

Parse JSON, reject unknown or missing required fields, and record the raw response for diagnosis. A retry should use bounded exponential backoff and an idempotent job identifier. Do not silently turn a model failure into an empty record.

Handle dynamic pages

A basic HTTP request may receive an application shell rather than rendered data. If the required content is produced by JavaScript, use a permitted browser-rendering step or a service that renders the page, then pass the resulting content to Gemini. URL Context’s public-access and content-type limits still apply.

Permissions and responsible operation

  • Check the site’s robots.txt; Google documents it as a mechanism for allowing or disallowing crawler access.
  • Read the site’s terms, authentication requirements and rate limits.
  • Obtain permission where required and avoid collecting personal or sensitive data unnecessarily.
  • Throttle requests, cache unchanged pages and identify your client honestly.
  • Robots.txt alone does not resolve contractual, copyright or jurisdiction-specific questions.

Common failures and fixes

Symptom Likely cause Fix
403 or 429 from your fetch Access control or excessive rate Stop, review permission, slow down and use documented authentication; do not evade a block.
HTML contains no products Content rendered after JavaScript Use an authorized renderer or a page’s documented data endpoint, then extract the rendered text.
Gemini invents a value Prompt permits inference or the field is ambiguous Require null for absence, provide the exact schema and validate against page evidence.
Context or request too large Page exceeds your selected limit or URL Context’s 34 MB ceiling Select the relevant element, split work into bounded chunks or reduce the URL set to 20 or fewer.
URL Context cannot retrieve page Paywall, unsupported type or non-public URL Fetch it yourself only when authorized, or use a publicly accessible supported representation.
Grounding links used to seed a crawler Terms violation risk Maintain your own permitted URL list and do not automate collection of Search grounding links.

Performance, cost and data quality

Fetch once and cache by URL plus relevant request parameters. Send the smallest faithful text to Gemini; repeated navigation and boilerplate increase tokens without improving extraction. Parallelize only within the target site’s limits and your API quota. Store retrieval metadata, model version, prompt version and validation errors so a changed page can be distinguished from a changed model response. For high-value fields, sample outputs for human review and compare extracted values with the source text.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can render a page as PNG, JPEG, WebP or PDF, which is useful when the visual state is the content you need to inspect before extraction. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Its 63 options include full-page capture with lazy images loaded, CSS-selector element capture, device presets and arbitrary viewports, dark mode, retina scale, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Use the ScreenshotNeo documentation for parameters. Example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account to try it.

FAQ

Can Gemini crawl an entire website from one URL?

No. URL Context retrieves supplied URLs and does not automatically traverse nested links. Build a permitted URL queue yourself if you need site-wide coverage.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I always fetch with Python first?

No. Use Python-first fetching when you need custom headers, retries, caching, rendering or an audit trail. Use URL Context when the targets are known, public and within its limits.

Does robots.txt make scraping legal?

No. It expresses crawler preferences; terms, rights, access controls and local law still require separate review.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.