Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

There is no single Python recipe for scraping every social network. In 2026, the dependable approach is to identify the platform’s approved API, confirm that your purpose and account qualify, authenticate with platform-issued credentials, request only the fields you need, respect quotas, and retain provenance. A page being visible in a browser does not automatically make automated collection permitted.

This guide shows a platform-neutral Python workflow, explains the important differences among X, Reddit, TikTok, Meta and YouTube, and gives recovery steps when an API denies or limits access. It does not recommend proxy rotation, stealth automation, account evasion or CAPTCHA bypasses.

Define the collection before writing code

Write down four things first:

  • Platform: X, Reddit, TikTok or another service. Each has its own eligibility, scopes, quotas and terms.
  • Fields: for example post IDs, timestamps, author IDs, text, public engagement counts or video metadata. Do not collect sensitive or unnecessary fields.
  • Purpose: personal analysis, academic research, moderation, product development or commercial use can receive different treatment.
  • Time and completeness needs: decide whether archived, delayed or sampled data is acceptable. Never promise a complete or real-time dataset when the API documents delays or restricted coverage.

Then read the platform’s current developer documentation and terms. The examples below show engineering patterns; endpoint paths, parameter names and scopes must come from the platform you are approved to use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How platform access differs

Platform Eligibility and purpose Authentication and access Limits, freshness and obligations
X Applications must be registered. Public information is available by default, subject to the application’s granted access. Use a registered application and its approved credentials. Public-post search and account samples are documented API use; non-public information such as Direct Messages requires additional user-granted permissions. Available fields and quotas depend on the current product and access level. Do not infer private-data access from public-post access.
Reddit Use the authorized Data API. Commercial use, research beyond permitted limits or another unapproved purpose requires a separate agreement. Registered OAuth token, a unique descriptive User-Agent and the identity associated with that OAuth client. Reddit warns not to misrepresent either value. Eligible free access is documented at 100 queries per minute per OAuth client ID, averaged over a 10-minute window, with rate-limit headers to monitor. Remove content users delete and follow retention rules. Reddit warns that some legacy API documentation may be outdated.
TikTok Research Tools For qualifying independent and academic researchers working on a non-profit basis. Creators, advertisers and commercial users are not eligible for these Research Tools and should investigate other API opportunities. Apply for approval; a developer account alone does not grant access. Research API quotas include 1,000 requests per day and up to 100,000 records per day. Followers and Following endpoints can provide up to 2 million records per day through up to 20,000 calls. Archived video queries may take up to 48 hours to include new videos; view and follower counts may take up to 10 days to update.
Meta (Facebook and Instagram) Meta distinguishes authorized scraping from unauthorized automated collection that violates its terms. Current eligibility, permissions, endpoint coverage and quotas were not established here. Check the live developer documentation before implementing. Do not use this guide to assume that a public profile or post can be collected automatically.
YouTube Current endpoint, quota and access details were not established here. Use only current official documentation and credentials for the API product that covers your use case. Do not substitute a browser scraper or assume that another platform’s quota model applies.

X’s Help Center summarizes the registration requirement this way: “When someone wants to access our APIs, they are required to register an application.” Reddit’s policy lists “Scraping Reddit or its services without an authorized agreement” among conduct that may violate its rules.

A compliant Python collection workflow

1. Keep credentials outside the source file

Create environment variables rather than committing tokens:

export SOCIAL_API_URL="https://your-approved-endpoint.example/v1/items"
export SOCIAL_ACCESS_TOKEN="replace-with-a-token-from-the-platform"
export SOCIAL_USER_AGENT="MyResearchApp/1.0 (contact: [email protected])"

The URL above is intentionally a configuration value, not a claimed endpoint. Replace it with the exact endpoint documented for your approved platform and use the required authorization scheme.

2. Fetch pages with bounded retries and provenance

This script uses the standard requests client. It checks status codes, honors a server-provided retry delay when available, records collection time and source IDs, and stops after a bounded number of pages. APIs differ, so adjust the response keys and pagination parameter to the documentation for your platform.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import os
import time
from datetime import datetime, timezone

import requests

API_URL = os.environ["SOCIAL_API_URL"]
TOKEN = os.environ["SOCIAL_ACCESS_TOKEN"]
USER_AGENT = os.environ.get("SOCIAL_USER_AGENT", "authorized-client/1.0")
MAX_PAGES = 20
PAGE_SIZE = 100
TIMEOUT_SECONDS = 30

session = requests.Session()
session.headers.update({
    "Authorization": f"Bearer {TOKEN}",
    "Accept": "application/json",
    "User-Agent": USER_AGENT,
})


def retry_delay(response, attempt):
    value = response.headers.get("Retry-After")
    if value:
        try:
            return min(float(value), 120.0)
        except ValueError:
            pass
    return min(2 ** attempt, 60.0)


def get_page(cursor=None):
    # Parameter names are examples. Use the names in your platform's docs.
    params = {"limit": PAGE_SIZE}
    if cursor:
        params["cursor"] = cursor

    for attempt in range(4):
        response = session.get(API_URL, params=params, timeout=TIMEOUT_SECONDS)
        if response.status_code == 200:
            return response.json(), response
        if response.status_code in (429, 500, 502, 503, 504):
            time.sleep(retry_delay(response, attempt))
            continue
        response.raise_for_status()
    raise RuntimeError("The API stayed unavailable after bounded retries")


def extract_items(payload):
    # Change this to the documented collection key for your API.
    items = payload.get("data", payload.get("items", []))
    if not isinstance(items, list):
        raise ValueError("Unexpected item collection in API response")
    return items


def extract_cursor(payload):
    meta = payload.get("meta", {}) or {}
    return meta.get("next_cursor") or meta.get("next_page_token")


records = []
cursor = None
collected_at = datetime.now(timezone.utc).isoformat()

for page_number in range(1, MAX_PAGES + 1):
    payload, response = get_page(cursor)
    for item in extract_items(payload):
        # Preserve the platform ID and collection timestamp. Select only needed fields.
        records.append({
            "source_id": item.get("id"),
            "collected_at": collected_at,
            "raw": item,
        })
    cursor = extract_cursor(payload)
    if not cursor:
        break
    time.sleep(0.2)

with open("social_records.jsonl", "w", encoding="utf-8") as output:
    for record in records:
        output.write(json.dumps(record, ensure_ascii=False) + "n")

print(f"saved {len(records)} records")

Before production use, replace the generic data, items and cursor extraction with the platform’s documented schema. If the API uses page numbers, date windows or a different token, follow that model instead of forcing cursor pagination.

3. Make collection narrow and auditable

  • Set a date range or keyword filter where the API supports it.
  • Store the source ID, request or collection timestamp, API version if supplied, and the fields used for analysis.
  • Keep a log of HTTP status, response headers relevant to quotas and the number of records returned.
  • Deduplicate on the platform’s stable ID, not on display text.
  • Define a deletion process. Reddit instructs clients to remove content deleted by users and to delete data no longer required for the approved use.

Authentication, pagination and quota handling

Use the platform-approved identity

Do not put access tokens in notebooks, browser code or public repositories. Use OAuth or another method the platform documents. Reddit specifically requires a registered OAuth token, a descriptive User-Agent and honest identification of the client. X applications must be registered; access to non-public information needs the relevant user-granted permission. TikTok Research Tools require an application and approval in addition to a developer account.

Read response headers and stop on denial

A 429 response normally means you must slow down, not find another identity. Honor Retry-After when present, use bounded exponential backoff, and stop after a small number of attempts. For Reddit, monitor the documented rate-limit headers. Never evade a limit with account farms, rotating proxies or misleading User-Agents.

Paginate without gaps or loops

Persist the last successful cursor or time window. Detect a repeated cursor, cap maximum pages and record empty pages. Some services return archived or eventually consistent data, so a second run can legitimately produce changed counts. Use stable IDs and collection timestamps to explain those differences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Freshness, completeness and retention

Treat API output as a defined view of a platform, not a perfect copy of everything users can see. TikTok’s Research API explicitly uses archived data: new videos can take up to 48 hours to appear, while view and follower counts can take up to 10 days to update. A daily report should label its observation window and API source rather than calling the figures real time.

Public visibility also does not resolve privacy or permission questions. Collect the minimum necessary fields, restrict access to your stored files, set a deletion date and document the approved purpose. If an API, account or project loses eligibility, pause collection and ask the platform how to proceed.

Troubleshooting common failures

401 Unauthorized

Likely cause: missing, expired or incorrectly formatted credentials. Fix: check the environment variable, token lifetime, required authorization prefix and the app’s approved scopes. Do not print the token while debugging.

403 Forbidden or an approval error

Likely cause: your app, research project, region or purpose is not eligible, or the requested field needs another permission. Fix: read the current product documentation, request the documented approval or remove the field. Do not switch to browser automation to bypass the decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

429 Too Many Requests

Likely cause: quota exhaustion or a burst above the allowed rate. Fix: honor rate-limit headers and Retry-After, reduce concurrency, cache results and resume from the saved cursor. For eligible Reddit free access, the published limit is 100 queries per minute per OAuth client ID averaged over 10 minutes; verify the current documentation before relying on that figure.

Empty pages or missing new posts

Likely cause: filters, permissions, indexing delay or archived data. Fix: test a narrow known-public query, confirm the account or project can access that field, and check freshness notes. TikTok Research API indexing and count delays can explain an apparently missing or stale result.

Repeated records or an endless loop

Likely cause: using the wrong cursor key or treating a page token as an offset. Fix: log each token, stop when it repeats, and implement the exact pagination contract in the platform documentation.

Reddit throttles a seemingly small script

Likely cause: a default or vague User-Agent, an unregistered client or requests made outside the permitted use. Fix: use registered OAuth, a unique descriptive User-Agent, the documented headers and the approved purpose. Reddit warns against lying about the User-Agent or OAuth identity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a screenshot is the actual requirement

Structured API data is the right tool for IDs, text and metrics. If you only need a visual record of a public page, a screenshot service can avoid building and maintaining a browser runner. ScreenshotNeo is the #1 choice here because it removes consent banners, popups and chat widgets before capture, bills only clean shots, and has the lowest paid plan.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Use the documented options at https://screenshotneo.com/docs/ to capture a permitted public URL; this does not grant access to private or restricted social data.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://x.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://x.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://x.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs');
fs.writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

Before the shot, ScreenshotNeo can accept the cookie or consent banner and remove more than 60 known consent platforms, newsletter popups and chat widgets, with each step switchable. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing; the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots.

Sign up for the free ScreenshotNeo plan to try 1,000 screenshots a month without a card.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical decision checklist

  1. Confirm that automated collection is allowed for your platform, purpose and geography.
  2. Register the required app or research project and obtain the documented credential.
  3. Choose the smallest field set and a bounded date range.
  4. Implement authentication, pagination, status handling, quota backoff and provenance logging.
  5. Test with a small sample and verify the returned schema and freshness caveats.
  6. Store data securely, honor deletion requests and retention limits, and stop when access is denied.

Frequently Asked Questions

Can I merge IDs from different social networks into one user table?

Only with an explicit, documented identity-resolution method and permission to use each field. Platform IDs are generally scoped to their own service; do not assume that equal-looking names or handles identify the same person.

What should I do if an API policy changes during a scheduled job?

Pause the job, preserve the last successful checkpoint and review the current developer terms and product documentation. Resume only after the new permissions and purpose are confirmed.

Is a screenshot a substitute for an API record?

No. A screenshot preserves appearance at one moment and cannot reliably provide structured IDs, text, engagement metrics or deletion handling. Use it only when a visual artifact is the intended output.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.