PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Gemini is not a complete web crawler. In Python, treat scraping as two separate jobs: obtain the page (with your HTTP client or Gemini URL Context), then ask Gemini to extract the fields you need. This separation gives you control over URLs, permissions, retries and storage while leaving interpretation of messy text to the model.
The two operations you must keep separate
1. Fetching content
Your program first obtains HTML or another supported representation. A normal HTTP request gives you control over headers, timeouts, rate limits, caching and error handling. You decide which URLs to request and how to respect a site’s access controls.
2. Extracting information
After fetching, send only the relevant page content to Gemini with a precise schema. Ask for fields, types and missing-value behavior instead of a vague “scrape this page” prompt. Extraction is interpretation; it does not make the initial request legal, reliable or complete.
Option A: fetch a page in Python, then ask Gemini to extract it
The following is a package-level pattern, not a guarantee of current client-library syntax. Confirm the current Gemini SDK and HTTP/HTML-parser documentation before deploying it. The example deliberately separates retrieval, status checking, content selection and model output.
#1 Best Overall
- Choose a URL that you are allowed to access and set a clear timeout.
- Fetch it and stop on an unsuccessful HTTP status.
- Extract the useful text (or selected elements) and cap its size.
- Send that text to Gemini with a strict JSON shape.
- Validate the returned object before writing it to your database.
import json
import os
import requests
from bs4 import BeautifulSoup
from google import genai
url = "https://example.com/products/widget"
# 1) Fetch
response = requests.get(
url,
headers={"User-Agent": "MyResearchBot/1.0 ([email protected])"},
timeout=30,
)
response.raise_for_status()
# 2) Select readable content; do not send an entire site blindly
soup = BeautifulSoup(response.text, "html.parser")
for node in soup(["script", "style", "noscript"]):
node.decompose()
text = soup.get_text(" ", strip=True)
text = text[:120_000] # choose a limit for your model and task
# 3) Extract a defined structure
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
prompt = f"""
Extract product data from the page text below.
Return JSON only, matching this shape:
{{"name": string|null, "price": number|null, "currency": string|null,
"availability": string|null, "source_url": "{url}"}}
Use null when a value is not present. Do not infer values.
PAGE TEXT:
{text}
"""
result = client.models.generate_content(
model="gemini-2.5-flash",
contents=prompt,
)
record = json.loads(result.text)
print(json.dumps(record, ensure_ascii=False))
The model name and SDK import can change. Pin and review the version you deploy, handle non-JSON responses, and validate types with your preferred schema library. For large pages, select the article or product container rather than truncating in the middle of important content. Keep the original URL, retrieval time, response status and a content hash alongside the extracted record so you can audit changes.
Option B: let Gemini URL Context retrieve supplied URLs
Google describes URL Context as a way to provide additional context to models in the form of URLs. You give Gemini the full, known URL and ask it to analyze the retrieved content. This is useful when your application already knows its targets and does not need custom crawling logic.
- One request can process up to 20 URLs.
- Retrieved content is limited to 34 MB per URL.
- URLs must be publicly accessible; paywalled pages and some content types are unsupported.
- Google says retrieval may use indexed content first and fall back to a live fetch.
- Responses can include URL citation annotations and retrieval metadata.
- It does not follow nested links from a supplied page, so it is not an unrestricted crawler.
A URL Context request should still state the output schema, missing-value rule and comparison logic. Supplying ten product URLs for a comparison is different from asking the model to discover every link on a site; the latter requires your own link-discovery and queueing code.
Rank #2
URL Context versus Python fetching
| Question | Python fetch, then Gemini | Gemini URL Context |
|---|---|---|
| Who retrieves? | Your application and HTTP stack | Gemini’s URL Context tool |
| URL control | Headers, cookies, retries and pacing are yours | You provide specific public URLs |
| Processing limit | Set by your infrastructure and model request | Up to 20 URLs and 34 MB per URL |
| Nested links | Implement discovery yourself | Not followed automatically |
| Best fit | Repeatable pipelines needing audit and custom access logic | Known public pages that Gemini can retrieve directly |
Do not use Google Search grounding as a crawl-target finder
The Gemini API Additional Terms effective March 23, 2026 prohibit programmatic or automated collection of Grounded Results, Search Suggestions or Links for another purpose, including using links to identify destination pages for crawling or scraping. Fetching a URL your application already knows is a different workflow from collecting search-grounding links to build a crawl queue. Review the terms that apply to your service and geography before shipping.
Gemini CLI is another interface, not a Python crawler
The Gemini CLI web_fetch tool accepts URLs in a prompt and uses Gemini API URL Context. It can be convenient for interactive work, but it is not a Python library and should not be presented as a drop-in replacement for a custom crawler. For scheduled jobs, keep a clear boundary between the CLI workflow and your Python service.
Designing reliable extraction prompts
Specify a contract
Name every field, its type, allowed enum values and the treatment of absent or conflicting text. Tell the model not to infer prices, dates or availability. Include the source URL in the requested object.
Reduce and label input
Remove navigation, scripts and repeated boilerplate when possible. Mark the beginning and end of page text, and identify tables or headings so the model can distinguish labels from values.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Validate and retry safely
Parse JSON, reject unknown or missing required fields, and record the raw response for diagnosis. A retry should use bounded exponential backoff and an idempotent job identifier. Do not silently turn a model failure into an empty record.
Handle dynamic pages
A basic HTTP request may receive an application shell rather than rendered data. If the required content is produced by JavaScript, use a permitted browser-rendering step or a service that renders the page, then pass the resulting content to Gemini. URL Context’s public-access and content-type limits still apply.
Permissions and responsible operation
- Check the site’s
robots.txt; Google documents it as a mechanism for allowing or disallowing crawler access. - Read the site’s terms, authentication requirements and rate limits.
- Obtain permission where required and avoid collecting personal or sensitive data unnecessarily.
- Throttle requests, cache unchanged pages and identify your client honestly.
- Robots.txt alone does not resolve contractual, copyright or jurisdiction-specific questions.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or 429 from your fetch | Access control or excessive rate | Stop, review permission, slow down and use documented authentication; do not evade a block. |
| HTML contains no products | Content rendered after JavaScript | Use an authorized renderer or a page’s documented data endpoint, then extract the rendered text. |
| Gemini invents a value | Prompt permits inference or the field is ambiguous | Require null for absence, provide the exact schema and validate against page evidence. |
| Context or request too large | Page exceeds your selected limit or URL Context’s 34 MB ceiling | Select the relevant element, split work into bounded chunks or reduce the URL set to 20 or fewer. |
| URL Context cannot retrieve page | Paywall, unsupported type or non-public URL | Fetch it yourself only when authorized, or use a publicly accessible supported representation. |
| Grounding links used to seed a crawler | Terms violation risk | Maintain your own permitted URL list and do not automate collection of Search grounding links. |
Performance, cost and data quality
Fetch once and cache by URL plus relevant request parameters. Send the smallest faithful text to Gemini; repeated navigation and boilerplate increase tokens without improving extraction. Parallelize only within the target site’s limits and your API quota. Store retrieval metadata, model version, prompt version and validation errors so a changed page can be distinguished from a changed model response. For high-value fields, sample outputs for human review and compare extracted values with the source text.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can render a page as PNG, JPEG, WebP or PDF, which is useful when the visual state is the content you need to inspect before extraction. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing state.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesIts 63 options include full-page capture with lazy images loaded, CSS-selector element capture, device presets and arbitrary viewports, dark mode, retina scale, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, usage reporting and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Use the ScreenshotNeo documentation for parameters. Example:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account to try it.
FAQ
Can Gemini crawl an entire website from one URL?
No. URL Context retrieves supplied URLs and does not automatically traverse nested links. Build a permitted URL queue yourself if you need site-wide coverage.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Should I always fetch with Python first?
No. Use Python-first fetching when you need custom headers, retries, caching, rendering or an audit trail. Use URL Context when the targets are known, public and within its limits.
Does robots.txt make scraping legal?
No. It expresses crawler preferences; terms, rights, access controls and local law still require separate review.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

