Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To extract structured data from known public web pages with Gemini, pass their URLs to the URL Context tool, describe exactly which fields you need, and request JSON that your application validates before saving. Use Google Search grounding when Gemini must discover pages, not just inspect URLs you already have. These are separate tasks: fetching pages, extracting values, enforcing a schema, and preserving evidence each need their own handling.

Choose the right way to retrieve web information

The best method depends on whether you already know which pages to inspect. Google AI for Developers describes URL Context as a way to extract information such as prices, names, and key findings from multiple URLs. It first tries an internal index cache and can fall back to a live fetch. Search grounding is the better fit when Gemini needs to find relevant pages or answer questions about changing public information.

  • Known URLs: Use URL Context and supply the public URLs in your request. This is the direct choice for extracting specified fields from a list of product pages, articles, or other known pages.
  • Page discovery: Enable Google Search grounding when Gemini needs to find pages. Grounded output includes inline URL citation annotations; keep those annotations or the associated web URI and title records with your results.
  • Discovery plus inspection: Search grounding can find pages, while URL Context can inspect specified pages in depth. Tool availability varies by model and preview status, so check the current Google AI for Developers documentation and the selected model’s support before deploying.

URL Context is retrieval, not a general-purpose browser or a guarantee that every URL will be accessible. Google documents support for examples including HTML, JSON, plain text, XML, CSS, JavaScript, CSV, and RTF; retrieval can still fail safety checks or other URL limitations. Do not assume that a page was fetched successfully or that it contains the requested fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define the extraction contract before calling the API

Decide what a valid record means before writing the prompt. A useful extraction contract specifies the fields, data types, normalization rules, and what to return when information is absent. It should also say whether a field must be quoted from the page or can be summarized. If the result will feed a database, job queue, or other program, make the contract a JSON Schema and validate the response before persisting it.

  • Be precise about fields: Ask for distinct properties such as product name, price, currency, and availability rather than “the important details.”
  • Set missing-value behavior: Require an explicit null or a defined error state when the page does not establish a value. Do not invite the model to fill gaps by guessing.
  • Specify normalization: For example, say whether a price should be represented as a number and currency as a separate code or label. State how to handle variants, sale prices, and ambiguous availability.
  • Keep evidence with the value: For fields where auditability matters, request a short supporting quote or retain the grounding citation records when using Search grounding. A plausible-looking JSON value alone is not proof of its source.

Structured Outputs is for enforcing the final response shape. The Gemini REST API accepts JSON Schema, while Google GenAI SDKs support Pydantic models in Python and Zod in JavaScript. Gemini supports a subset of JSON Schema; keep schemas to supported primitive, object, array, and null forms, and validate at your own application boundary as well.

Python example: extract fields from known URLs

This example uses the Google GenAI Python SDK, Pydantic, URL Context, and a response schema. Install the SDK and Pydantic, set GEMINI_API_KEY, and set GEMINI_MODEL to a model currently available to your account that supports both URL Context and Structured Outputs. Model names and tool availability can change, so the model is deliberately supplied through the environment rather than hard-coded.

pip install google-genai pydantic

export GEMINI_API_KEY="YOUR_API_KEY"
export GEMINI_MODEL="YOUR_SUPPORTED_MODEL"

cat > extract.py <<'PY'
import json
import os
from typing import Optional

from google import genai
from google.genai import types
from pydantic import BaseModel, ConfigDict


class Product(BaseModel):
    model_config = ConfigDict(extra="forbid")
    url: str
    product_name: Optional[str]
    price: Optional[float]
    currency: Optional[str]
    availability: Optional[str]
    evidence: Optional[str]


client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
urls = [
    "https://example.com/product-a",
    "https://example.com/product-b",
]

prompt = f"""
Use URL Context to inspect these public product pages:
{json.dumps(urls)}

Return one record for each URL using the requested schema.
Extract product_name, price, currency, and availability only when supported
by the page. Use null for a field that is missing, ambiguous, or not retrieved;
do not infer it. Price must be a number without a currency symbol. Put a short
verbatim supporting excerpt for the extracted values in evidence, or null if
there is no reliable excerpt. Preserve each page URL in its record.
"""

response = client.models.generate_content(
    model=os.environ["GEMINI_MODEL"],
    contents=prompt,
    config=types.GenerateContentConfig(
        tools=[types.Tool(url_context=types.UrlContext())],
        response_mime_type="application/json",
        response_schema=list[Product],
    ),
)

# Parse and validate before writing or sending records downstream.
records = [Product.model_validate(item) for item in json.loads(response.text)]
print(json.dumps([record.model_dump() for record in records], indent=2))
PY
python extract.py

Replace the example URLs with the pages you are authorized to process. Check the current SDK examples if your installed version uses different names for a tool or configuration field. The response schema constrains the output format, but application-side parsing and validation remain important. This sample stores neither source citations nor a retrieval-status field; for an evidence-sensitive workflow, add appropriate fields and preserve any provenance the API returns instead of treating the model’s summary as a citation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make the result safe to use in an application

Do not send generated text straight into a database as trusted data. Parse it, validate the types and business rules, and decide how to handle missing records before the extraction runs in production.

  1. Validate the input URLs. Allow only expected schemes and destinations, and guard against user-supplied URLs that could expose internal services or data. Keep page content and extracted strings untrusted.
  2. Validate every response. Reject malformed JSON, unexpected fields, invalid types, and values outside your own rules. A schema improves shape consistency; it does not establish factual correctness.
  3. Represent failure separately from absence. A null price can mean the page did not state one, while a retrieval or safety failure means you could not reliably inspect the page. Track an explicit status if downstream users need to distinguish these cases.
  4. Bound the job. Cap the number of URLs, response size, and records accepted per run. Define retries and a failure path rather than retrying indefinitely or silently saving partial output.
  5. Log enough to reproduce a decision. Keep the model identifier, schema version, requested URLs, result status, and citation metadata where available. Avoid logging secrets or unnecessary page data.

Keep discovery, formatting, actions, and provenance separate

Several Gemini capabilities can appear in the same workflow, but they solve different problems. Structured Outputs specifies the final response shape. Function Calling asks your application to execute an application-owned function, such as looking up an internal record or submitting a job; it is not a substitute for choosing a JSON schema. Built-in tools include Google Search, URL Context, File Search, Code Execution, and Google Maps, with model and preview support varying.

For a record that triggers work, first obtain and validate the structured extraction, then make the application decide whether the result meets its rules before it calls an internal function or starts a job. For provenance, retain Search grounding’s inline URL annotations or the returned GroundingChunk web URI/title objects alongside the record. Do not confuse a URL supplied to URL Context with a citation annotation from Search grounding.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup: use ScreenshotNeo for a visual capture

If what you need is a screenshot rather than extracted text fields, ScreenshotNeo offers a one-request website screenshot API. It is not a Gemini extraction endpoint: use the Gemini workflow above for structured page data, and use this when a clean visual capture is the deliverable. Its API can return PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API documentation for request options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
  • Cookie and consent banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. The response includes X-Page-Verdict and X-Billed headers.
  • An MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. All features are available on every plan.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Troubleshoot common extraction failures

  • No useful page content: The URL may be inaccessible, unsupported, blocked by a safety check, or limited by URL Context’s retrieval constraints. Confirm the URL is public and accessible, then inspect the returned response and retrieval metadata if available. Do not present a failed retrieval as an empty-but-successful extraction.
  • JSON parsing or schema validation fails: Confirm the chosen model supports Structured Outputs, reduce the schema to supported JSON Schema forms, and check for optional fields that need explicit null types. Keep application-side validation in place even when the API is configured for JSON.
  • Fields are missing or uncertain: Tighten the prompt so it distinguishes absent information from ambiguity, and specify null behavior. If the page does not contain a value, the correct result is not a plausible guess.
  • Results lack citations: URL Context is the known-URL retrieval path; if you need web discovery with citation annotations, use Search grounding and preserve its annotations or GroundingChunk records. Do not discard provenance before storing the extracted result.
  • A tool or model is unavailable: Support differs across models and preview features. Check the current Google documentation for availability, select a compatible model, and test tool and output-schema support together before running a batch.

Plan for reliability, latency, and cost without guessing

Google’s documentation covered here does not establish a universal extraction accuracy, latency, or cost figure. Those depend on the model, request, retrieval behavior, and current pricing or quota rules, which can change. Check current Google documentation for the model you choose, estimate usage from your actual URL volume and schema, and run representative pages through your validation and failure handling before committing to a production budget.

For reliability, build around observable outcomes rather than assuming every request yields a valid record: record retrieval and validation status, retain provenance when needed, and make retries bounded and safe. For cost control, use a narrow field contract, avoid needlessly large batches, and check current account quotas and model pricing before scaling.

Frequently Asked Questions

Can I use the Gemini API to crawl an entire site from a homepage URL?

The workflow here is for URLs you provide or pages found through Search grounding; it does not establish a general-purpose site crawler. Discover and authorize the pages you need, then submit those URLs under your own limits.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I extract data from a page that requires my login?

The documented URL Context workflow is described for public URLs. Do not assume it can access authenticated or private pages; use an approved retrieval method for content you are authorized to process.

Should I ask Gemini for a quote or a summary?

Choose based on the use case: request a short verbatim excerpt when you need to check the page’s wording, and a summary when you need concise interpretation. Neither replaces retaining citation metadata when source attribution is required.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.