Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
APIs

How to Scrape GraphQL APIs With Python (Queries, Variables, Pagination, and Errors)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To collect records from a GraphQL API with Python, send documented POST requests to the provider’s GraphQL endpoint, put changing values in a separate variables object, inspect both data and errors, and follow that API’s pagination fields until its end-of-results signal. GraphQL “scraping” should mean using an API you are authorized to access—not downloading rendered HTML or bypassing authentication.

What GraphQL scraping actually involves

GraphQL is a strongly typed, self-describing query language and execution system. The service publishes a schema containing the types, fields, arguments, and relationships it permits. Your query selects a shape from that schema; it does not grant arbitrary access to the provider’s database. As the GraphQL Specification Project puts it, “A GraphQL response, on the other hand, contains exactly what a client asks for and no more.”

Before writing code, confirm all of the following in the provider’s official documentation:

  • The exact endpoint URL. /graphql is common, not guaranteed.
  • Authentication, required headers, scopes, and whether your account or plan may use the API.
  • The schema reference, operation names, argument types, and acceptable-use terms.
  • Pagination and rate-limit rules, including any retry headers.
  • Whether introspection is enabled. Some deployments restrict it, so use the published schema when necessary.

Do not copy private browser credentials or replay requests against data you are not allowed to access. A visible browser request is not proof that automated reuse is permitted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a first Python request with requests

A plain HTTP client is usually the clearest starting point. GraphQL-over-HTTP requires servers to support POST with a JSON body. The body contains a query string and can include operationName, variables, and extensions. GET support is optional and must not execute mutations.

Minimal, bounded query

Replace every placeholder with names from the target schema. The example uses a common connection shape; those field names are not universal.

import requests

endpoint = "https://api.example.com/graphql"
query = """
query GetItems($after: String) {
  items(first: 50, after: $after) {
    nodes { id name }
    pageInfo { hasNextPage endCursor }
  }
}
"""

response = requests.post(
    endpoint,
    json={
        "query": query,
        "operationName": "GetItems",
        "variables": {"after": None},
    },
    headers={
        "Accept": "application/graphql-response+json, application/json;q=0.9",
        # "Authorization": "Bearer YOUR_TOKEN",
    },
    timeout=30,
)
response.raise_for_status()          # HTTP delivery/authentication failure
payload = response.json()

if payload.get("errors"):
    raise RuntimeError(payload["errors"])

items = payload["data"]["items"]
print(items["nodes"])

The Accept value is the compatibility-oriented form recommended by the current GraphQL-over-HTTP specification. Follow provider examples if they require a different media type. A timeout prevents a collector from hanging indefinitely; choose a value appropriate to the provider’s documented behavior.

Why variables matter

Declare dynamic IDs, dates, filters, and cursors in the operation signature, then pass values in variables. Do not concatenate user input into the query text. Variables preserve type validation, avoid quoting mistakes, and make operation logging safer.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
query = """
query UserByID($id: ID!, $includeEmail: Boolean!) {
  user(id: $id) {
    id
    name
    email @include(if: $includeEmail)
  }
}
"""

body = {
    "query": query,
    "operationName": "UserByID",
    "variables": {"id": "12345", "includeEmail": False},
}
response = requests.post(endpoint, json=body, timeout=30)
response.raise_for_status()
payload = response.json()
if payload.get("errors"):
    print("GraphQL errors:", payload["errors"])
user = (payload.get("data") or {}).get("user")

Use named operations even when a document contains only one operation. The name appears in provider logs and makes failures easier to identify.

Read HTTP status, GraphQL errors, and partial data

There are two independent layers of failure. An HTTP error can indicate a network problem, rejected authentication, or an unavailable server. A JSON response with a successful HTTP status can still contain GraphQL request or execution errors.

  • Request errors: invalid syntax, unknown fields, validation failures, or variables of the wrong type. Execution does not begin, so data is commonly absent.
  • Execution errors: a resolver failed while other fields succeeded. The response may contain both useful partial data and an errors array.

Always parse the body after delivery succeeds. Decide whether partial records are acceptable for your job; for financial, compliance, or synchronization workloads, fail the page and record the error instead of silently keeping incomplete data.

def graphql_json(response):
    response.raise_for_status()
    payload = response.json()
    errors = payload.get("errors") or []
    if errors:
        for error in errors:
            print("message:", error.get("message"), "path:", error.get("path"))
    if "data" not in payload:
        raise RuntimeError("No data returned")
    return payload["data"], errors

# data, errors = graphql_json(response)

Useful error objects often include a message, a field path, and provider-specific extensions such as a code or rate-limit detail. Treat those extensions as provider-defined rather than portable GraphQL fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Paginate according to the target schema

Pagination is not a universal GraphQL feature with one mandated field layout. Inspect the schema and documentation for cursor or page arguments, the connection’s page-information object, and its terminal signal. Common names include first, after, nodes, pageInfo.hasNextPage, and endCursor, but an API may use different names or offset pagination.

Cursor loop for a connection-shaped schema

import requests

endpoint = "https://api.example.com/graphql"
query = """
query GetItems($after: String) {
  items(first: 50, after: $after) {
    nodes { id name }
    pageInfo { hasNextPage endCursor }
  }
}
"""

session = requests.Session()
session.headers.update({
    "Accept": "application/graphql-response+json, application/json;q=0.9",
    # "Authorization": "Bearer YOUR_TOKEN",
})

after = None
seen_ids = set()
while True:
    response = session.post(
        endpoint,
        json={
            "query": query,
            "operationName": "GetItems",
            "variables": {"after": after},
        },
        timeout=30,
    )
    response.raise_for_status()
    payload = response.json()
    if payload.get("errors"):
        raise RuntimeError(payload["errors"])

    connection = payload["data"]["items"]
    for record in connection["nodes"]:
        # Stable IDs make retries and overlapping pages idempotent.
        if record["id"] not in seen_ids:
            seen_ids.add(record["id"])
            print(record)

    page_info = connection["pageInfo"]
    if not page_info["hasNextPage"]:
        break
    next_cursor = page_info["endCursor"]
    if not next_cursor or next_cursor == after:
        raise RuntimeError("Pagination cursor did not advance")
    after = next_cursor

For long jobs, persist the last successful cursor and the IDs already written. A restart can then resume without duplicating records. If the provider documents snapshot tokens or a consistency mode, use those instead of assuming that a changing dataset will remain stable while you page.

Offset and page-number APIs

Some schemas expose offset/limit, page/perPage, or a total count. Follow that contract exactly; do not send cursor arguments to a field that does not define them. Confirm whether inserts or deletions during a run can shift offsets, and prefer a documented stable sort key when available.

Keep queries small and respect provider limits

Request only fields you will store. Deep, broad nested connections increase response size and execution cost. Use modest page sizes and the provider’s filters rather than downloading a large universe and filtering locally. Do not assume that concurrency is safe: rate limits may penalize parallel requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GitHub’s documented limits are provider-specific

GitHub’s current GraphQL documentation (accessed 2026) requires first or last values from 1 through 100 for connections, limits a single call to 500,000 total nodes, and documents a 10-second request timeout. It also describes possible 502/504 responses and resource exhaustion for very large or deeply nested queries. These numbers apply to GitHub, not to GraphQL in general.

When throttled, honor Retry-After and any rate-limit reset information. Use bounded exponential backoff only where the provider permits it. Do not retry permanent validation or authentication errors, and do not continue hammering an endpoint while rate-limited; GitHub warns that this can result in an integration ban.

Direct HTTP versus the gql Python client

Choice Dependencies and abstraction Execution Schema and subscriptions Best fit
requests or direct HTTPX Small footprint; transport and JSON remain explicit Synchronous with requests; HTTPX also offers async APIs You validate query fields through provider docs or your own tooling One-off calls, simple collectors, and maximum endpoint transparency
gql GraphQL-aware operations and structured transport configuration Documented synchronous RequestsHTTPTransport and synchronous/asynchronous HTTPX transports Can fetch/use schemas for more structured workflows; its HTTP transport does not support subscriptions Multiple operations, reusable clients, and schema-oriented applications

The gql documentation also describes HTTPXAsyncTransport. If you need subscriptions, choose the library’s WebSocket transport rather than its HTTP transport. For an ordinary paginated export, synchronous requests are often easier to operate and less likely to exceed provider limits.

Example using gql

from gql import Client, gql
from gql.transport.requests import RequestsHTTPTransport

transport = RequestsHTTPTransport(
    url="https://api.example.com/graphql",
    headers={"Authorization": "Bearer YOUR_TOKEN"},
    timeout=30,
)
client = Client(transport=transport, fetch_schema_from_transport=False)

operation = gql("""
query GetItems($after: String) {
  items(first: 50, after: $after) {
    nodes { id name }
    pageInfo { hasNextPage endCursor }
  }
}
""")

result = client.execute(operation, variable_values={"after": None})
print(result["items"]["nodes"])

Set schema fetching only when the endpoint permits introspection and your deployment needs it. Otherwise keep a documented schema reference with your code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL and Node.js equivalents for debugging

Even when Python is your production language, a second client helps separate application bugs from endpoint behavior.

curl -X POST https://api.example.com/graphql 
  -H 'Content-Type: application/json' 
  -H 'Accept: application/graphql-response+json, application/json;q=0.9' 
  -H 'Authorization: Bearer YOUR_TOKEN' 
  --data '{"query":"query GetItems($after: String) { items(first: 50, after: $after) { nodes { id name } pageInfo { hasNextPage endCursor } } }","operationName":"GetItems","variables":{"after":null}}'
const endpoint = 'https://api.example.com/graphql';
const query = `query GetItems($after: String) {
  items(first: 50, after: $after) {
    nodes { id name }
    pageInfo { hasNextPage endCursor }
  }
}`;

const res = await fetch(endpoint, {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Accept': 'application/graphql-response+json, application/json;q=0.9',
    'Authorization': 'Bearer YOUR_TOKEN'
  },
  body: JSON.stringify({
    query,
    operationName: 'GetItems',
    variables: { after: null }
  })
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const payload = await res.json();
if (payload.errors) console.error(payload.errors);
console.log(payload.data);
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting checklist

  • 404 or HTML instead of JSON: verify the documented endpoint, API version, and POST route. A website URL is not necessarily the GraphQL endpoint.
  • 401/403: check the token, header spelling, scopes, account status, and whether the operation is allowed. Do not retry unchanged credentials.
  • “Cannot query field” or validation errors: use the exact schema field and argument names for this client; fields visible in another API version may not exist here.
  • Variable type errors: match the declared GraphQL type exactly, including nullability and list notation. Send variables as JSON values, not quoted GraphQL fragments.
  • data plus errors: inspect each error’s path and decide whether the affected records may be retained. Partial data is not an all-clear.
  • Repeated pages: persist and compare cursors; stop if a cursor is empty or unchanged, and investigate unstable sorting.
  • 429, 502, 504, or timeouts: reduce page size and query breadth, obey reset or retry headers, and use bounded backoff where documented.
  • Introspection denied: obtain the provider’s schema reference or generated schema file; do not assume introspection must be enabled.

Or skip the browser setup

If your task is to capture a rendered page rather than collect structured GraphQL records, ScreenshotNeo provides a website screenshot API and MCP server. One request returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the 63 options, including full-page and element capture, device presets, custom CSS/JavaScript, waits, headers, cookies, geolocation, PDF settings, caching, signed links, async webhooks, bulk capture, and the usage API.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can GraphQL return arbitrary database columns?

No. Only fields and relationships exposed by the service’s schema and authorized for your client are queryable.

Is a successful HTTP status proof that the query worked?

No. GraphQL execution errors can accompany partial data, so inspect the JSON body every time.

Should I parallelize page requests?

Not by default. Provider throttling, ordering, and consistency rules may make sequential paging safer.

What if the API has no cursor?

Use its documented offset, page-number, token, or time-window contract and checkpoint progress in the form that contract defines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can GraphQL return arbitrary database columns?

No. Only fields and relationships exposed by the service’s schema and authorized for your client are queryable.

Is a successful HTTP status proof that the query worked?

No. GraphQL execution errors can accompany partial data, so inspect the JSON body every time.

Should I parallelize page requests?

Not by default. Provider throttling, ordering, and consistency rules may make sequential paging safer.

What if the API has no cursor?

Use its documented offset, page-number, token, or time-window contract and checkpoint progress in the form that contract defines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.