Recommended Free Tools
To collect records from a GraphQL API with Python, send documented POST requests to the provider’s GraphQL endpoint, put changing values in a separate variables object, inspect both data and errors, and follow that API’s pagination fields until its end-of-results signal. GraphQL “scraping” should mean using an API you are authorized to access—not downloading rendered HTML or bypassing authentication.
What GraphQL scraping actually involves
GraphQL is a strongly typed, self-describing query language and execution system. The service publishes a schema containing the types, fields, arguments, and relationships it permits. Your query selects a shape from that schema; it does not grant arbitrary access to the provider’s database. As the GraphQL Specification Project puts it, “A GraphQL response, on the other hand, contains exactly what a client asks for and no more.”
Before writing code, confirm all of the following in the provider’s official documentation:
- The exact endpoint URL.
/graphqlis common, not guaranteed. - Authentication, required headers, scopes, and whether your account or plan may use the API.
- The schema reference, operation names, argument types, and acceptable-use terms.
- Pagination and rate-limit rules, including any retry headers.
- Whether introspection is enabled. Some deployments restrict it, so use the published schema when necessary.
Do not copy private browser credentials or replay requests against data you are not allowed to access. A visible browser request is not proof that automated reuse is permitted.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems#1 Best Overall
Build a first Python request with requests
A plain HTTP client is usually the clearest starting point. GraphQL-over-HTTP requires servers to support POST with a JSON body. The body contains a query string and can include operationName, variables, and extensions. GET support is optional and must not execute mutations.
Minimal, bounded query
Replace every placeholder with names from the target schema. The example uses a common connection shape; those field names are not universal.
import requests
endpoint = "https://api.example.com/graphql"
query = """
query GetItems($after: String) {
items(first: 50, after: $after) {
nodes { id name }
pageInfo { hasNextPage endCursor }
}
}
"""
response = requests.post(
endpoint,
json={
"query": query,
"operationName": "GetItems",
"variables": {"after": None},
},
headers={
"Accept": "application/graphql-response+json, application/json;q=0.9",
# "Authorization": "Bearer YOUR_TOKEN",
},
timeout=30,
)
response.raise_for_status() # HTTP delivery/authentication failure
payload = response.json()
if payload.get("errors"):
raise RuntimeError(payload["errors"])
items = payload["data"]["items"]
print(items["nodes"])
The Accept value is the compatibility-oriented form recommended by the current GraphQL-over-HTTP specification. Follow provider examples if they require a different media type. A timeout prevents a collector from hanging indefinitely; choose a value appropriate to the provider’s documented behavior.
Why variables matter
Declare dynamic IDs, dates, filters, and cursors in the operation signature, then pass values in variables. Do not concatenate user input into the query text. Variables preserve type validation, avoid quoting mistakes, and make operation logging safer.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
query = """
query UserByID($id: ID!, $includeEmail: Boolean!) {
user(id: $id) {
id
name
email @include(if: $includeEmail)
}
}
"""
body = {
"query": query,
"operationName": "UserByID",
"variables": {"id": "12345", "includeEmail": False},
}
response = requests.post(endpoint, json=body, timeout=30)
response.raise_for_status()
payload = response.json()
if payload.get("errors"):
print("GraphQL errors:", payload["errors"])
user = (payload.get("data") or {}).get("user")
Use named operations even when a document contains only one operation. The name appears in provider logs and makes failures easier to identify.
Rank #2
Read HTTP status, GraphQL errors, and partial data
There are two independent layers of failure. An HTTP error can indicate a network problem, rejected authentication, or an unavailable server. A JSON response with a successful HTTP status can still contain GraphQL request or execution errors.
- Request errors: invalid syntax, unknown fields, validation failures, or variables of the wrong type. Execution does not begin, so
datais commonly absent. - Execution errors: a resolver failed while other fields succeeded. The response may contain both useful partial
dataand anerrorsarray.
Always parse the body after delivery succeeds. Decide whether partial records are acceptable for your job; for financial, compliance, or synchronization workloads, fail the page and record the error instead of silently keeping incomplete data.
def graphql_json(response):
response.raise_for_status()
payload = response.json()
errors = payload.get("errors") or []
if errors:
for error in errors:
print("message:", error.get("message"), "path:", error.get("path"))
if "data" not in payload:
raise RuntimeError("No data returned")
return payload["data"], errors
# data, errors = graphql_json(response)
Useful error objects often include a message, a field path, and provider-specific extensions such as a code or rate-limit detail. Treat those extensions as provider-defined rather than portable GraphQL fields.
Paginate according to the target schema
Pagination is not a universal GraphQL feature with one mandated field layout. Inspect the schema and documentation for cursor or page arguments, the connection’s page-information object, and its terminal signal. Common names include first, after, nodes, pageInfo.hasNextPage, and endCursor, but an API may use different names or offset pagination.
Cursor loop for a connection-shaped schema
import requests
endpoint = "https://api.example.com/graphql"
query = """
query GetItems($after: String) {
items(first: 50, after: $after) {
nodes { id name }
pageInfo { hasNextPage endCursor }
}
}
"""
session = requests.Session()
session.headers.update({
"Accept": "application/graphql-response+json, application/json;q=0.9",
# "Authorization": "Bearer YOUR_TOKEN",
})
after = None
seen_ids = set()
while True:
response = session.post(
endpoint,
json={
"query": query,
"operationName": "GetItems",
"variables": {"after": after},
},
timeout=30,
)
response.raise_for_status()
payload = response.json()
if payload.get("errors"):
raise RuntimeError(payload["errors"])
connection = payload["data"]["items"]
for record in connection["nodes"]:
# Stable IDs make retries and overlapping pages idempotent.
if record["id"] not in seen_ids:
seen_ids.add(record["id"])
print(record)
page_info = connection["pageInfo"]
if not page_info["hasNextPage"]:
break
next_cursor = page_info["endCursor"]
if not next_cursor or next_cursor == after:
raise RuntimeError("Pagination cursor did not advance")
after = next_cursor
For long jobs, persist the last successful cursor and the IDs already written. A restart can then resume without duplicating records. If the provider documents snapshot tokens or a consistency mode, use those instead of assuming that a changing dataset will remain stable while you page.
Offset and page-number APIs
Some schemas expose offset/limit, page/perPage, or a total count. Follow that contract exactly; do not send cursor arguments to a field that does not define them. Confirm whether inserts or deletions during a run can shift offsets, and prefer a documented stable sort key when available.
Keep queries small and respect provider limits
Request only fields you will store. Deep, broad nested connections increase response size and execution cost. Use modest page sizes and the provider’s filters rather than downloading a large universe and filtering locally. Do not assume that concurrency is safe: rate limits may penalize parallel requests.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →GitHub’s documented limits are provider-specific
GitHub’s current GraphQL documentation (accessed 2026) requires first or last values from 1 through 100 for connections, limits a single call to 500,000 total nodes, and documents a 10-second request timeout. It also describes possible 502/504 responses and resource exhaustion for very large or deeply nested queries. These numbers apply to GitHub, not to GraphQL in general.
When throttled, honor Retry-After and any rate-limit reset information. Use bounded exponential backoff only where the provider permits it. Do not retry permanent validation or authentication errors, and do not continue hammering an endpoint while rate-limited; GitHub warns that this can result in an integration ban.
Direct HTTP versus the gql Python client
| Choice | Dependencies and abstraction | Execution | Schema and subscriptions | Best fit |
|---|---|---|---|---|
requests or direct HTTPX |
Small footprint; transport and JSON remain explicit | Synchronous with requests; HTTPX also offers async APIs |
You validate query fields through provider docs or your own tooling | One-off calls, simple collectors, and maximum endpoint transparency |
gql |
GraphQL-aware operations and structured transport configuration | Documented synchronous RequestsHTTPTransport and synchronous/asynchronous HTTPX transports |
Can fetch/use schemas for more structured workflows; its HTTP transport does not support subscriptions | Multiple operations, reusable clients, and schema-oriented applications |
The gql documentation also describes HTTPXAsyncTransport. If you need subscriptions, choose the library’s WebSocket transport rather than its HTTP transport. For an ordinary paginated export, synchronous requests are often easier to operate and less likely to exceed provider limits.
Example using gql
from gql import Client, gql
from gql.transport.requests import RequestsHTTPTransport
transport = RequestsHTTPTransport(
url="https://api.example.com/graphql",
headers={"Authorization": "Bearer YOUR_TOKEN"},
timeout=30,
)
client = Client(transport=transport, fetch_schema_from_transport=False)
operation = gql("""
query GetItems($after: String) {
items(first: 50, after: $after) {
nodes { id name }
pageInfo { hasNextPage endCursor }
}
}
""")
result = client.execute(operation, variable_values={"after": None})
print(result["items"]["nodes"])
Set schema fetching only when the endpoint permits introspection and your deployment needs it. Otherwise keep a documented schema reference with your code.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11cURL and Node.js equivalents for debugging
Even when Python is your production language, a second client helps separate application bugs from endpoint behavior.
curl -X POST https://api.example.com/graphql
-H 'Content-Type: application/json'
-H 'Accept: application/graphql-response+json, application/json;q=0.9'
-H 'Authorization: Bearer YOUR_TOKEN'
--data '{"query":"query GetItems($after: String) { items(first: 50, after: $after) { nodes { id name } pageInfo { hasNextPage endCursor } } }","operationName":"GetItems","variables":{"after":null}}'
const endpoint = 'https://api.example.com/graphql';
const query = `query GetItems($after: String) {
items(first: 50, after: $after) {
nodes { id name }
pageInfo { hasNextPage endCursor }
}
}`;
const res = await fetch(endpoint, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Accept': 'application/graphql-response+json, application/json;q=0.9',
'Authorization': 'Bearer YOUR_TOKEN'
},
body: JSON.stringify({
query,
operationName: 'GetItems',
variables: { after: null }
})
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const payload = await res.json();
if (payload.errors) console.error(payload.errors);
console.log(payload.data);
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting checklist
- 404 or HTML instead of JSON: verify the documented endpoint, API version, and POST route. A website URL is not necessarily the GraphQL endpoint.
- 401/403: check the token, header spelling, scopes, account status, and whether the operation is allowed. Do not retry unchanged credentials.
- “Cannot query field” or validation errors: use the exact schema field and argument names for this client; fields visible in another API version may not exist here.
- Variable type errors: match the declared GraphQL type exactly, including nullability and list notation. Send variables as JSON values, not quoted GraphQL fragments.
datapluserrors: inspect each error’s path and decide whether the affected records may be retained. Partial data is not an all-clear.- Repeated pages: persist and compare cursors; stop if a cursor is empty or unchanged, and investigate unstable sorting.
- 429, 502, 504, or timeouts: reduce page size and query breadth, obey reset or retry headers, and use bounded backoff where documented.
- Introspection denied: obtain the provider’s schema reference or generated schema file; do not assume introspection must be enabled.
Or skip the browser setup
If your task is to capture a rendered page rather than collect structured GraphQL records, ScreenshotNeo provides a website screenshot API and MCP server. One request returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the 63 options, including full-page and element capture, device presets, custom CSS/JavaScript, waits, headers, cookies, geolocation, PDF settings, caching, signed links, async webhooks, bulk capture, and the usage API.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Can GraphQL return arbitrary database columns?
No. Only fields and relationships exposed by the service’s schema and authorized for your client are queryable.
Best Value
Is a successful HTTP status proof that the query worked?
No. GraphQL execution errors can accompany partial data, so inspect the JSON body every time.
Should I parallelize page requests?
Not by default. Provider throttling, ordering, and consistency rules may make sequential paging safer.
What if the API has no cursor?
Use its documented offset, page-number, token, or time-window contract and checkpoint progress in the form that contract defines.
Frequently Asked Questions
Can GraphQL return arbitrary database columns?
No. Only fields and relationships exposed by the service’s schema and authorized for your client are queryable.
Is a successful HTTP status proof that the query worked?
No. GraphQL execution errors can accompany partial data, so inspect the JSON body every time.
Should I parallelize page requests?
Not by default. Provider throttling, ordering, and consistency rules may make sequential paging safer.
What if the API has no cursor?
Use its documented offset, page-number, token, or time-window contract and checkpoint progress in the form that contract defines.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




