DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
data collection

How to Collect Twitter (X) Data for Sentiment Analysis: An Official API Workflow

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use X’s official API, not HTML scraping, to collect public posts for sentiment analysis. Define a query and time window first, obtain the required developer access, retrieve every paginated response, and preserve collection metadata. Recent search covers roughly the last seven days; full-archive search can reach back to March 2006 but requires the access level specified by X’s current documentation. Your result is the set of posts matching your query and available to your account—not a complete measure of public opinion.

1. Define what your sample is supposed to represent

Before requesting data, write a one-paragraph population definition. Include the topic, languages, dates, account scope, treatment of replies and reposts, and the unit you will classify (normally one post). This prevents a convenient query from silently becoming your research design.

Topic and vocabulary

A keyword query finds the vocabulary you specify. It can miss posts that express the same idea differently and include unrelated uses of an ambiguous term. List synonyms, product names, spelling variants and relevant hashtags, then test the query on a small response before committing to a long collection.

Document inclusion rules

  • Record the exact query string and every revision.
  • State the language or languages and whether multilingual posts are retained.
  • Decide whether replies, reposts and quote posts are included.
  • Specify UTC start and end timestamps, inclusive or exclusive boundaries, and the selected search route.
  • Define deduplication rules and whether the analysis unit is a post, author, conversation or day.

Search operators documented by X include exact phrases, hashtags, mentions, account filters such as from: and to:, language filters such as lang:en, and exclusions including -is:retweet and -is:reply. Operator availability and access requirements can change, so check the current Search Posts documentation when implementing your query.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Choose recent or full-archive search

Route Date coverage Best use Important qualification
Recent search Last seven days Monitoring, rapid-response studies and current events The rolling window means a later run may return a different population.
Full-archive search Back to March 2006 Historical comparisons and long-term studies The quickstart requires Self-serve or Enterprise access; verify your account’s present eligibility.

X’s access tiers, quotas, prices and regional eligibility are volatile. Confirm the current terms for your account before promising a sample size or historical coverage. Do not copy prices from older articles.

3. Create credentials and protect them

Register a developer project with X and create the Bearer Token required for app-only search requests. Follow X’s developer policies, including rules governing storage, redistribution and research use. Keep the token in an environment variable or secret manager; never commit it to source control or include it in a notebook shared publicly.

On macOS or Linux, set a session variable with export X_BEARER_TOKEN='your-token'. In continuous integration, use the platform’s encrypted secret store. Rotate a token that appears in logs, screenshots or error reports.

4. Query the API with Python

The following standard-library example requests recent search pages directly. It requests up to 100 posts per page, follows next_token, writes JSON lines, and records enough metadata to reproduce the run. Replace the query and dates with your study definition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json, os, time, requests

TOKEN = os.environ["X_BEARER_TOKEN"]
QUERY = '("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply'
URL = "https://api.x.com/2/tweets/search/recent"
params = {
    "query": QUERY,
    "max_results": 100,
    "tweet.fields": "id,text,author_id,created_at,lang,conversation_id,public_metrics",
    "expansions": "author_id",
    "user.fields": "username,protected",
}

with open("x_posts.jsonl", "w", encoding="utf-8") as out:
    while True:
        response = requests.get(
            URL,
            headers={"Authorization": f"Bearer {TOKEN}"},
            params=params,
            timeout=60,
        )
        if response.status_code == 429:
            delay = int(response.headers.get("retry-after", "60"))
            time.sleep(min(delay, 900))
            continue
        response.raise_for_status()
        page = response.json()
        out.write(json.dumps(page, ensure_ascii=False) + "n")
        token = page.get("meta", {}).get("next_token")
        if not token:
            break
        params["next_token"] = token

For archive work, use the full-archive endpoint and add ISO 8601 UTC start_time and end_time values accepted by your account. The full-archive quickstart demonstrates this pattern; endpoint names and permissions should be checked against the current documentation.

Python XDK pagination

X’s Python development kit can expose an iterator that handles continuation tokens. It is convenient for notebooks, but still save the original query, timestamps, response metadata and errors. An SDK does not remove rate limits or policy obligations.

5. Equivalent cURL request

curl --get 'https://api.x.com/2/tweets/search/recent' 
  --header "Authorization: Bearer $X_BEARER_TOKEN" 
  --data-urlencode 'query=("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply' 
  --data-urlencode 'max_results=100' 
  --data-urlencode 'tweet.fields=id,text,author_id,created_at,lang,conversation_id,public_metrics' 
  --data-urlencode 'expansions=author_id' 
  --data-urlencode 'user.fields=username,protected'

Save each response before requesting the next page. If a process stops midway, resume from the last recorded token rather than silently replacing the partial collection.

6. Equivalent Node.js request

const token = process.env.X_BEARER_TOKEN;
const params = new URLSearchParams({
  query: '("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply',
  max_results: '100',
  'tweet.fields': 'id,text,author_id,created_at,lang,conversation_id,public_metrics',
  expansions: 'author_id',
  'user.fields': 'username,protected'
});
const res = await fetch(`https://api.x.com/2/tweets/search/recent?${params}`, {
  headers: { Authorization: `Bearer ${token}` }
});
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
console.log(await res.json());

Production Node clients should loop over meta.next_token, apply bounded exponential backoff for 429 responses, and write raw pages to durable storage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Capture metadata and make pagination auditable

A successful first response is not the dataset. Search responses are paginated and can contain a next_token; documentation examples use up to 100 results per request. Store one manifest beside the data containing:

  • query text, endpoint and route (recent or archive);
  • UTC start and end times and the collection start and finish timestamps;
  • requested fields, expansions, page count and result count;
  • account or project access level, software versions and token identity (never the secret itself);
  • HTTP status, error body, retry events and the final page token state.

Keep raw JSON separate from a normalized table. Retaining the original response lets you audit parsing decisions and detect changes in API fields.

8. Understand what the data can and cannot show

Call the result a query-defined sample: posts matching your operators that X returned during your collection period. It is not automatically representative of all X users, all internet users or public opinion.

Missing and changing posts

Protected-account posts, deleted posts and posts withheld in some regions may not be returned. Rate or usage caps can also interrupt collection. Historical studies of the former Twitter Academic API—including a 2022 study that found evidence of “almost complete” samples for many search terms—do not establish completeness for today’s X API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Time comparisons

For trend analysis, keep the query, filters, requested fields and collection procedure constant. Record revisions as separate waves. A change in access, deletion patterns, language use or operator behavior can look like a sentiment change.

9. Prepare posts for sentiment analysis

Normalize without erasing meaning

  • Deduplicate according to your design; reposts may be signal or noise.
  • Preserve the original text, IDs and timestamps, and create a separate analysis column for normalized text.
  • Decide how links, mentions, hashtags, emojis, punctuation and quoted text are represented.
  • Keep replies with conversation identifiers if context matters; do not assume an isolated reply has the same meaning as a standalone post.
  • Route multilingual text to language-appropriate models or analyze languages separately.

Validate labels

Do not call model output ground truth. Sarcasm, negation, slang, coded language and domain terminology can defeat a generic classifier. Sample posts for human annotation, define label instructions, measure agreement, and evaluate candidate models on data resembling your topic and languages. Report uncertainty and ambiguous cases instead of forcing every post into positive, negative or neutral.

10. Troubleshooting

401 or 403 responses

Check that the Bearer Token is present, unexpired and attached as Authorization: Bearer .... A valid token can still lack the product access required for a route, especially full archive. Confirm project permissions and current eligibility.

400 invalid query

Test a simple term, then add one operator at a time. Check balanced quotes, supported operators, date format and URL encoding. Keep the exact accepted query in your manifest.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

429 rate or usage limit

Stop issuing requests, honor retry-after when supplied, and use exponential backoff with jitter. Persist completed pages so a retry does not duplicate work. Review both rate limits and monthly usage caps.

Zero results

Verify the date window, language code and exclusions. A protected author, deleted content, regional withholding or an overly narrow phrase can all produce an empty result. Run a deliberately broad diagnostic query, then restore the study query.

Missing pages or duplicates

Follow every returned next_token until it disappears, save page boundaries, and deduplicate by post ID only when your design calls for it. Never infer completeness from one page.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

11. Performance, reliability and cost planning

Request only fields and expansions needed by the analysis; large expansions increase transfer and parsing time. Use bounded concurrency only where your account’s limits permit it, and prefer a durable queue for long archive jobs. Estimate pages from pilot counts, but treat the estimate as operational planning rather than a promise: query volume, deletions and access caps change the final count. Confirm current X pricing and quotas directly before budgeting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your project also needs visual snapshots of pages or reports, ScreenshotNeo provides a one-call website screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.

Use the API documentation at https://screenshotneo.com/docs/. A complete cURL call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python and Node.js are available when your pipeline already uses those languages:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every feature is on every plan. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I use this workflow for private posts?

No. The documented search workflow concerns public posts, and protected-account content may not be returned. Follow X’s current developer and privacy policies for any research involving people.

How far back can recent search go?

The documented recent-search window is the last seven days. Older dates require full-archive access, whose eligibility must be confirmed for your account.

Should I publish the collected post text?

Check X’s current storage and redistribution rules before sharing text or datasets. A reproducible manifest and post IDs may be preferable to redistributing raw content, subject to those rules.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.