Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use the official Stack Exchange API to collect questions: it is more reliable than parsing page HTML, supports searches and filters, and returns structured data you can paginate. For questions across a site, use /questions; for matching a title or tags, use /search. The API is documented as version 2.3. This guide shows how to request question records, paginate safely, save provenance, and handle rate limits. HTML scraping is a fallback, not the recommended starting point.
Use the Stack Exchange API instead of scraping page HTML
“Scraping Stack Exchange questions” can mean collecting structured question data or saving a visual copy of a page. If you need fields such as title, tags, score, date, and link for analysis or indexing, use the Stack Exchange API. Its documented endpoints and response fields are less vulnerable to page-layout changes than an HTML scraper.
HTML parsing may be useful when you specifically need rendered context that the API does not supply, but it is more fragile and carries additional terms-of-service considerations. Before deploying an HTML scraper or redistributing content, review the current Public Network Terms of Service. The page shows a last-updated date of November 13, 2025.
Free tools Windows power users keep installed
One-click scans. No signup required.
Choose the right endpoint for Stack Exchange questions
Use /questions to collect questions by site and constraints
The /questions endpoint returns a list of questions. Set a site and add the constraints you need. The documented parameters include tagged, fromdate, todate, min, max, sort, order, page, and pagesize. Dates are Unix epoch values. Multiple tags are separated with semicolons; requesting more than five tags returns zero results.
#1 Best Overall
Use this endpoint when the collection is defined by a site and a set of filters—for example, questions tagged python created during a date range. Keep query scope narrow enough that you can resume and audit the collection.
Use /search for title or tag searches
Use /search when the task is to find questions matching a title phrase or tags. At least one of tagged or intitle must be set. A search with several tags uses OR semantics: it can match a question with any of those tags, not necessarily all of them. If you need questions carrying every specified tag, do not assume a single tagged search expresses that condition; verify the results or make narrower requests.
Make a request and choose the fields you need
Register an API application if you need an application key or OAuth access token; the API documentation recommends application registration. For a public, read-only collection, request only data needed for the project. Typical useful fields are question ID, title, link, score, tags, creation date, and—only if necessary—the body. Bodies increase the amount of content you store and process, so leave them out unless your use case requires them.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteThe API supports custom response filters. Create a filter through the API documentation interface for the exact fields you want, then pass its value as filter. Do not guess a custom filter string: use the one generated for your field selection. The example below uses the standard response and writes the returned question objects to JSON Lines. Replace the site and query terms with your own.
Python: paginate and save question records
This script requests up to 100 questions per page, follows has_more, honors the API’s backoff instruction, and saves a checkpoint so an interrupted collection can resume. It uses the documented v2.3 API host and endpoint.
Rank #2
- Used Book in Good Condition
import json
import os
import time
from datetime import datetime, timezone
import requests
API = "https://api.stackexchange.com/2.3/questions"
SITE = "stackoverflow"
OUTPUT = "questions.jsonl"
CHECKPOINT = "questions-page.txt"
# Example: set STACKEXCHANGE_KEY in your environment if you have an API key.
KEY = os.getenv("STACKEXCHANGE_KEY")
params = {
"site": SITE,
"tagged": "python",
"pagesize": 100,
"page": 1,
"order": "desc",
"sort": "creation",
}
if KEY:
params["key"] = KEY
if os.path.exists(CHECKPOINT):
with open(CHECKPOINT, encoding="utf-8") as f:
params["page"] = int(f.read().strip())
session = requests.Session()
while True:
response = session.get(API, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
# Each record keeps enough provenance to reproduce and audit the collection.
retrieved_at = datetime.now(timezone.utc).isoformat()
with open(OUTPUT, "a", encoding="utf-8") as f:
for item in payload.get("items", []):
record = {
"site": SITE,
"question_id": item.get("question_id"),
"title": item.get("title"),
"link": item.get("link"),
"score": item.get("score"),
"tags": item.get("tags"),
"creation_date": item.get("creation_date"),
"retrieved_at": retrieved_at,
"request": {k: v for k, v in params.items() if k != "key"},
}
f.write(json.dumps(record, ensure_ascii=False) + "n")
# Store the next page only after successfully writing this page.
next_page = params["page"] + 1
with open(CHECKPOINT, "w", encoding="utf-8") as f:
f.write(str(next_page))
delay = int(payload.get("backoff", 0))
if delay:
time.sleep(delay)
if not payload.get("has_more", False):
break
params["page"] = next_page
time.sleep(2) # Conservative pacing; do not run near the documented ceiling.
The checkpoint records the next page after the current page has been written. If you intentionally change the query, start a separate output/checkpoint pair or remove the old checkpoint; otherwise the script can continue at a page belonging to a different query. For stronger crash recovery, write to a temporary file and atomically replace the checkpoint after each successful page.
cURL: request one page
For a quick request, use the API’s parameters as query arguments. The following returns one page of questions tagged python; pagination requires changing page and checking the response’s has_more value.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchcurl --get 'https://api.stackexchange.com/2.3/questions'
--data-urlencode 'site=stackoverflow'
--data-urlencode 'tagged=python'
--data-urlencode 'pagesize=100'
--data-urlencode 'page=1'
--data-urlencode 'order=desc'
--data-urlencode 'sort=creation'
Node.js: request and inspect a page
With a modern Node.js runtime that provides fetch, construct the query with URLSearchParams so tag delimiters and other values are encoded correctly.
const q = new URLSearchParams({
site: 'stackoverflow',
tagged: 'python',
pagesize: '100',
page: '1',
order: 'desc',
sort: 'creation'
});
const response = await fetch(`https://api.stackexchange.com/2.3/questions?${q}`);
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const data = await response.json();
for (const question of data.items ?? []) {
console.log(question.question_id, question.title, question.link);
}
console.log({ hasMore: data.has_more, backoffSeconds: data.backoff ?? 0 });
How to paginate without losing or duplicating results
The API uses page numbers starting at 1. The maximum pagesize is 100. Continue until the response wrapper says has_more is false; do not infer completion from a short page alone. Avoid requesting total unless you need a count: the documentation warns that calculating it can cost as much as fetching the items.
- Set a stable query, including site, tags or search phrase, sort order, and any date or score bounds.
- Start with
page=1andpagesize=100or a smaller size appropriate to your processing. - Write each page durably before advancing the checkpoint.
- Honor any returned
backoffvalue before making another request. - Continue while
has_moreis true, and stop when it is false. - On restart, resume the same query from the saved next page; record query parameters so the result set can be audited.
Page-number pagination is straightforward, but a live site can change while a long multi-page job is running. For a reproducible time-bounded dataset, constrain the query with a fixed date range and record the exact parameters and retrieval time. Deduplicate downstream by the pair of site name and question ID rather than assuming repeated retrievals will be identical.
Rank #3
Rate limits, caching, and run reliability
The documented default daily quota is 10,000 requests. The API guidance says that more than 30 requests per second per IP is considered very abusive and can result in requests being cut off harshly. Treat that as a danger threshold, not a target. Stay well below it, honor every backoff response, and do not repeat semantically identical requests more than once per minute.
- Cache responses: save each response keyed by its normalized request parameters so reruns do not repeatedly fetch the same page.
- Use exponential delay on transient failures: for timeouts, server errors, or temporary connectivity issues, wait longer after each failed attempt, with a sensible cap and a limited retry count. Do not rapidly retry a throttled request.
- Checkpoint progress: record completed pages and the query definition. Persist the page only after its records are safely written.
- Separate query runs: different tag, date, or sort parameters should not share a checkpoint.
- Request only necessary fields: use a custom filter when you know the exact fields needed; include bodies only when required.
Preserve provenance and follow attribution rules
Store the site name, question ID, original question link, API request parameters, and retrieval timestamp with every record. Those fields make refreshes, deduplication, and correction of stale records more manageable. Keep the original link rather than treating a copied title or body as a standalone record.
API applications must visibly identify Stack Exchange as the source and comply with the applicable attribution rules. The current requirements depend on how you display or redistribute content, so consult the API attribution guidance and current Public Network Terms before publishing collected material. A dataset used internally and a public product reproducing question content may have different practical obligations; do not assume that fetching through the API alone resolves them.
API collection versus HTML scraping
| Consideration | Official API | HTML parsing |
|---|---|---|
| Coverage and query precision | Documented question, search, tag, date, score, sort, and paging parameters. | Depends on pages fetched and selectors maintained; not a structured query interface. |
| Request cost | Subject to API quota and throttling; cache and page requests deliberately. | Consumes page requests too, with additional work to fetch and parse markup. |
| Freshness | Reflects API responses at retrieval time; record the time and query. | Reflects rendered pages at retrieval time; page changes can alter what the parser sees. |
| Resilience | Documented fields and parameters are preferable to layout-dependent selectors. | More vulnerable to markup and layout changes that break selectors. |
| Compliance | Applications still need source attribution and must follow applicable rules. | Review current Public Network Terms before deploying or redistributing scraped content. |
Choose HTML only for a specific need that the API does not meet, and then parse conservatively, minimize requests, and monitor failures. Do not use a screenshot as a substitute for structured records: a screenshot captures the page’s appearance, not searchable question fields.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common collection problems
The API returns no results for a tag query
Check that the tag spelling and site name are correct. With /questions, more than five semicolon-separated tags yields zero results. With /search, tags use OR semantics; check whether the query is broader or narrower than intended and whether tagged or intitle is present.
Rank #4
The script stops before collecting all pages
Inspect has_more; it, not a guessed page count, indicates whether more pages exist. Verify that the checkpoint belongs to the same query and that it is updated only after writing the current page.
Requests slow down or are cut off
Reduce concurrency and request frequency, honor backoff, and avoid repeating equivalent requests within a minute. Add bounded exponential retries for transient failures, not as a way to bypass throttling.
Dates or score bounds behave unexpectedly
Convert dates to Unix epoch values before passing fromdate or todate. Check which sort is active because min and max apply with the corresponding sortable field; validate the returned items before scaling up the run.
A field you need is missing
Check whether the standard response includes it, then create a custom response filter through the API documentation interface and request that filter. Avoid adding fields indiscriminately, especially full bodies, if the application does not need them.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →An HTML scraper breaks after a site change
That is a common consequence of relying on page structure rather than documented response fields. Reassess whether the API can provide the needed data; if HTML remains necessary, re-check the current terms, test selectors against changed pages, and keep parsing and storage separate so a selector failure does not silently corrupt records.
Best Value
Or skip the browser setup
For structured question data, use the Stack Exchange API examples above; ScreenshotNeo is not a replacement for that API. If your separate task is to capture a rendered page as an image or PDF, ScreenshotNeo offers a one-request screenshot API. Its cleanup can accept cookie or consent banners and remove 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the response identifying the page verdict and billing status. It also provides an MCP server for AI agents with take_screenshot, get_page_info, and capture_pdf.
For example, this captures a rendered Stack Overflow question page to WebP; it does not extract the question as structured data. See the ScreenshotNeo documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stackoverflow.com/questions -o shot.webp
ScreenshotNeo has 1,000 shots per month on its free plan with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Frequently Asked Questions
Can I scrape questions from Stack Exchange sites other than Stack Overflow?
Yes. Set the API’s site parameter to the Stack Exchange site you intend to query; the examples use Stack Overflow.
Can I use the API to retrieve question bodies?
The API supports custom response filters. Use the documentation interface to generate a filter containing the fields your application needs, including a body only when necessary.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

