Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use the official Stack Exchange API to collect questions: it is more reliable than parsing page HTML, supports searches and filters, and returns structured data you can paginate. For questions across a site, use /questions; for matching a title or tags, use /search. The API is documented as version 2.3. This guide shows how to request question records, paginate safely, save provenance, and handle rate limits. HTML scraping is a fallback, not the recommended starting point.

Use the Stack Exchange API instead of scraping page HTML

“Scraping Stack Exchange questions” can mean collecting structured question data or saving a visual copy of a page. If you need fields such as title, tags, score, date, and link for analysis or indexing, use the Stack Exchange API. Its documented endpoints and response fields are less vulnerable to page-layout changes than an HTML scraper.

HTML parsing may be useful when you specifically need rendered context that the API does not supply, but it is more fragile and carries additional terms-of-service considerations. Before deploying an HTML scraper or redistributing content, review the current Public Network Terms of Service. The page shows a last-updated date of November 13, 2025.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right endpoint for Stack Exchange questions

Use /questions to collect questions by site and constraints

The /questions endpoint returns a list of questions. Set a site and add the constraints you need. The documented parameters include tagged, fromdate, todate, min, max, sort, order, page, and pagesize. Dates are Unix epoch values. Multiple tags are separated with semicolons; requesting more than five tags returns zero results.

Use this endpoint when the collection is defined by a site and a set of filters—for example, questions tagged python created during a date range. Keep query scope narrow enough that you can resume and audit the collection.

Use /search for title or tag searches

Use /search when the task is to find questions matching a title phrase or tags. At least one of tagged or intitle must be set. A search with several tags uses OR semantics: it can match a question with any of those tags, not necessarily all of them. If you need questions carrying every specified tag, do not assume a single tagged search expresses that condition; verify the results or make narrower requests.

Make a request and choose the fields you need

Register an API application if you need an application key or OAuth access token; the API documentation recommends application registration. For a public, read-only collection, request only data needed for the project. Typical useful fields are question ID, title, link, score, tags, creation date, and—only if necessary—the body. Bodies increase the amount of content you store and process, so leave them out unless your use case requires them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The API supports custom response filters. Create a filter through the API documentation interface for the exact fields you want, then pass its value as filter. Do not guess a custom filter string: use the one generated for your field selection. The example below uses the standard response and writes the returned question objects to JSON Lines. Replace the site and query terms with your own.

Python: paginate and save question records

This script requests up to 100 questions per page, follows has_more, honors the API’s backoff instruction, and saves a checkpoint so an interrupted collection can resume. It uses the documented v2.3 API host and endpoint.

import json
import os
import time
from datetime import datetime, timezone

import requests

API = "https://api.stackexchange.com/2.3/questions"
SITE = "stackoverflow"
OUTPUT = "questions.jsonl"
CHECKPOINT = "questions-page.txt"

# Example: set STACKEXCHANGE_KEY in your environment if you have an API key.
KEY = os.getenv("STACKEXCHANGE_KEY")
params = {
    "site": SITE,
    "tagged": "python",
    "pagesize": 100,
    "page": 1,
    "order": "desc",
    "sort": "creation",
}
if KEY:
    params["key"] = KEY

if os.path.exists(CHECKPOINT):
    with open(CHECKPOINT, encoding="utf-8") as f:
        params["page"] = int(f.read().strip())

session = requests.Session()
while True:
    response = session.get(API, params=params, timeout=30)
    response.raise_for_status()
    payload = response.json()

    # Each record keeps enough provenance to reproduce and audit the collection.
    retrieved_at = datetime.now(timezone.utc).isoformat()
    with open(OUTPUT, "a", encoding="utf-8") as f:
        for item in payload.get("items", []):
            record = {
                "site": SITE,
                "question_id": item.get("question_id"),
                "title": item.get("title"),
                "link": item.get("link"),
                "score": item.get("score"),
                "tags": item.get("tags"),
                "creation_date": item.get("creation_date"),
                "retrieved_at": retrieved_at,
                "request": {k: v for k, v in params.items() if k != "key"},
            }
            f.write(json.dumps(record, ensure_ascii=False) + "n")

    # Store the next page only after successfully writing this page.
    next_page = params["page"] + 1
    with open(CHECKPOINT, "w", encoding="utf-8") as f:
        f.write(str(next_page))

    delay = int(payload.get("backoff", 0))
    if delay:
        time.sleep(delay)
    if not payload.get("has_more", False):
        break
    params["page"] = next_page
    time.sleep(2)  # Conservative pacing; do not run near the documented ceiling.

The checkpoint records the next page after the current page has been written. If you intentionally change the query, start a separate output/checkpoint pair or remove the old checkpoint; otherwise the script can continue at a page belonging to a different query. For stronger crash recovery, write to a temporary file and atomically replace the checkpoint after each successful page.

cURL: request one page

For a quick request, use the API’s parameters as query arguments. The following returns one page of questions tagged python; pagination requires changing page and checking the response’s has_more value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl --get 'https://api.stackexchange.com/2.3/questions' 
  --data-urlencode 'site=stackoverflow' 
  --data-urlencode 'tagged=python' 
  --data-urlencode 'pagesize=100' 
  --data-urlencode 'page=1' 
  --data-urlencode 'order=desc' 
  --data-urlencode 'sort=creation'

Node.js: request and inspect a page

With a modern Node.js runtime that provides fetch, construct the query with URLSearchParams so tag delimiters and other values are encoded correctly.

const q = new URLSearchParams({
  site: 'stackoverflow',
  tagged: 'python',
  pagesize: '100',
  page: '1',
  order: 'desc',
  sort: 'creation'
});

const response = await fetch(`https://api.stackexchange.com/2.3/questions?${q}`);
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const data = await response.json();
for (const question of data.items ?? []) {
  console.log(question.question_id, question.title, question.link);
}
console.log({ hasMore: data.has_more, backoffSeconds: data.backoff ?? 0 });

How to paginate without losing or duplicating results

The API uses page numbers starting at 1. The maximum pagesize is 100. Continue until the response wrapper says has_more is false; do not infer completion from a short page alone. Avoid requesting total unless you need a count: the documentation warns that calculating it can cost as much as fetching the items.

  1. Set a stable query, including site, tags or search phrase, sort order, and any date or score bounds.
  2. Start with page=1 and pagesize=100 or a smaller size appropriate to your processing.
  3. Write each page durably before advancing the checkpoint.
  4. Honor any returned backoff value before making another request.
  5. Continue while has_more is true, and stop when it is false.
  6. On restart, resume the same query from the saved next page; record query parameters so the result set can be audited.

Page-number pagination is straightforward, but a live site can change while a long multi-page job is running. For a reproducible time-bounded dataset, constrain the query with a fixed date range and record the exact parameters and retrieval time. Deduplicate downstream by the pair of site name and question ID rather than assuming repeated retrievals will be identical.

Rate limits, caching, and run reliability

The documented default daily quota is 10,000 requests. The API guidance says that more than 30 requests per second per IP is considered very abusive and can result in requests being cut off harshly. Treat that as a danger threshold, not a target. Stay well below it, honor every backoff response, and do not repeat semantically identical requests more than once per minute.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Cache responses: save each response keyed by its normalized request parameters so reruns do not repeatedly fetch the same page.
  • Use exponential delay on transient failures: for timeouts, server errors, or temporary connectivity issues, wait longer after each failed attempt, with a sensible cap and a limited retry count. Do not rapidly retry a throttled request.
  • Checkpoint progress: record completed pages and the query definition. Persist the page only after its records are safely written.
  • Separate query runs: different tag, date, or sort parameters should not share a checkpoint.
  • Request only necessary fields: use a custom filter when you know the exact fields needed; include bodies only when required.

Preserve provenance and follow attribution rules

Store the site name, question ID, original question link, API request parameters, and retrieval timestamp with every record. Those fields make refreshes, deduplication, and correction of stale records more manageable. Keep the original link rather than treating a copied title or body as a standalone record.

API applications must visibly identify Stack Exchange as the source and comply with the applicable attribution rules. The current requirements depend on how you display or redistribute content, so consult the API attribution guidance and current Public Network Terms before publishing collected material. A dataset used internally and a public product reproducing question content may have different practical obligations; do not assume that fetching through the API alone resolves them.

API collection versus HTML scraping

Consideration Official API HTML parsing
Coverage and query precision Documented question, search, tag, date, score, sort, and paging parameters. Depends on pages fetched and selectors maintained; not a structured query interface.
Request cost Subject to API quota and throttling; cache and page requests deliberately. Consumes page requests too, with additional work to fetch and parse markup.
Freshness Reflects API responses at retrieval time; record the time and query. Reflects rendered pages at retrieval time; page changes can alter what the parser sees.
Resilience Documented fields and parameters are preferable to layout-dependent selectors. More vulnerable to markup and layout changes that break selectors.
Compliance Applications still need source attribution and must follow applicable rules. Review current Public Network Terms before deploying or redistributing scraped content.

Choose HTML only for a specific need that the API does not meet, and then parse conservatively, minimize requests, and monitor failures. Do not use a screenshot as a substitute for structured records: a screenshot captures the page’s appearance, not searchable question fields.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common collection problems

The API returns no results for a tag query

Check that the tag spelling and site name are correct. With /questions, more than five semicolon-separated tags yields zero results. With /search, tags use OR semantics; check whether the query is broader or narrower than intended and whether tagged or intitle is present.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The script stops before collecting all pages

Inspect has_more; it, not a guessed page count, indicates whether more pages exist. Verify that the checkpoint belongs to the same query and that it is updated only after writing the current page.

Requests slow down or are cut off

Reduce concurrency and request frequency, honor backoff, and avoid repeating equivalent requests within a minute. Add bounded exponential retries for transient failures, not as a way to bypass throttling.

Dates or score bounds behave unexpectedly

Convert dates to Unix epoch values before passing fromdate or todate. Check which sort is active because min and max apply with the corresponding sortable field; validate the returned items before scaling up the run.

A field you need is missing

Check whether the standard response includes it, then create a custom response filter through the API documentation interface and request that filter. Avoid adding fields indiscriminately, especially full bodies, if the application does not need them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An HTML scraper breaks after a site change

That is a common consequence of relying on page structure rather than documented response fields. Reassess whether the API can provide the needed data; if HTML remains necessary, re-check the current terms, test selectors against changed pages, and keep parsing and storage separate so a selector failure does not silently corrupt records.

Or skip the browser setup

For structured question data, use the Stack Exchange API examples above; ScreenshotNeo is not a replacement for that API. If your separate task is to capture a rendered page as an image or PDF, ScreenshotNeo offers a one-request screenshot API. Its cleanup can accept cookie or consent banners and remove 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the response identifying the page verdict and billing status. It also provides an MCP server for AI agents with take_screenshot, get_page_info, and capture_pdf.

For example, this captures a rendered Stack Overflow question page to WebP; it does not extract the question as structured data. See the ScreenshotNeo documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stackoverflow.com/questions -o shot.webp

ScreenshotNeo has 1,000 shots per month on its free plan with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I scrape questions from Stack Exchange sites other than Stack Overflow?

Yes. Set the API’s site parameter to the Stack Exchange site you intend to query; the examples use Stack Overflow.

Can I use the API to retrieve question bodies?

The API supports custom response filters. Use the documentation interface to generate a filter containing the fields your application needs, including a body only when necessary.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.