Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Use X’s official API, not HTML scraping, to collect public posts for sentiment analysis. Define a query and time window first, obtain the required developer access, retrieve every paginated response, and preserve collection metadata. Recent search covers roughly the last seven days; full-archive search can reach back to March 2006 but requires the access level specified by X’s current documentation. Your result is the set of posts matching your query and available to your account—not a complete measure of public opinion.
1. Define what your sample is supposed to represent
Before requesting data, write a one-paragraph population definition. Include the topic, languages, dates, account scope, treatment of replies and reposts, and the unit you will classify (normally one post). This prevents a convenient query from silently becoming your research design.
Topic and vocabulary
A keyword query finds the vocabulary you specify. It can miss posts that express the same idea differently and include unrelated uses of an ambiguous term. List synonyms, product names, spelling variants and relevant hashtags, then test the query on a small response before committing to a long collection.
Document inclusion rules
- Record the exact query string and every revision.
- State the language or languages and whether multilingual posts are retained.
- Decide whether replies, reposts and quote posts are included.
- Specify UTC start and end timestamps, inclusive or exclusive boundaries, and the selected search route.
- Define deduplication rules and whether the analysis unit is a post, author, conversation or day.
Search operators documented by X include exact phrases, hashtags, mentions, account filters such as from: and to:, language filters such as lang:en, and exclusions including -is:retweet and -is:reply. Operator availability and access requirements can change, so check the current Search Posts documentation when implementing your query.
#1 Best Overall
2. Choose recent or full-archive search
| Route | Date coverage | Best use | Important qualification |
|---|---|---|---|
| Recent search | Last seven days | Monitoring, rapid-response studies and current events | The rolling window means a later run may return a different population. |
| Full-archive search | Back to March 2006 | Historical comparisons and long-term studies | The quickstart requires Self-serve or Enterprise access; verify your account’s present eligibility. |
X’s access tiers, quotas, prices and regional eligibility are volatile. Confirm the current terms for your account before promising a sample size or historical coverage. Do not copy prices from older articles.
3. Create credentials and protect them
Register a developer project with X and create the Bearer Token required for app-only search requests. Follow X’s developer policies, including rules governing storage, redistribution and research use. Keep the token in an environment variable or secret manager; never commit it to source control or include it in a notebook shared publicly.
On macOS or Linux, set a session variable with export X_BEARER_TOKEN='your-token'. In continuous integration, use the platform’s encrypted secret store. Rotate a token that appears in logs, screenshots or error reports.
4. Query the API with Python
The following standard-library example requests recent search pages directly. It requests up to 100 posts per page, follows next_token, writes JSON lines, and records enough metadata to reproduce the run. Replace the query and dates with your study definition.
import json, os, time, requests
TOKEN = os.environ["X_BEARER_TOKEN"]
QUERY = '("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply'
URL = "https://api.x.com/2/tweets/search/recent"
params = {
"query": QUERY,
"max_results": 100,
"tweet.fields": "id,text,author_id,created_at,lang,conversation_id,public_metrics",
"expansions": "author_id",
"user.fields": "username,protected",
}
with open("x_posts.jsonl", "w", encoding="utf-8") as out:
while True:
response = requests.get(
URL,
headers={"Authorization": f"Bearer {TOKEN}"},
params=params,
timeout=60,
)
if response.status_code == 429:
delay = int(response.headers.get("retry-after", "60"))
time.sleep(min(delay, 900))
continue
response.raise_for_status()
page = response.json()
out.write(json.dumps(page, ensure_ascii=False) + "n")
token = page.get("meta", {}).get("next_token")
if not token:
break
params["next_token"] = token
For archive work, use the full-archive endpoint and add ISO 8601 UTC start_time and end_time values accepted by your account. The full-archive quickstart demonstrates this pattern; endpoint names and permissions should be checked against the current documentation.
Rank #2
Python XDK pagination
X’s Python development kit can expose an iterator that handles continuation tokens. It is convenient for notebooks, but still save the original query, timestamps, response metadata and errors. An SDK does not remove rate limits or policy obligations.
5. Equivalent cURL request
curl --get 'https://api.x.com/2/tweets/search/recent'
--header "Authorization: Bearer $X_BEARER_TOKEN"
--data-urlencode 'query=("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply'
--data-urlencode 'max_results=100'
--data-urlencode 'tweet.fields=id,text,author_id,created_at,lang,conversation_id,public_metrics'
--data-urlencode 'expansions=author_id'
--data-urlencode 'user.fields=username,protected'
Save each response before requesting the next page. If a process stops midway, resume from the last recorded token rather than silently replacing the partial collection.
6. Equivalent Node.js request
const token = process.env.X_BEARER_TOKEN;
const params = new URLSearchParams({
query: '("battery recycling" OR #batteryrecycling) lang:en -is:retweet -is:reply',
max_results: '100',
'tweet.fields': 'id,text,author_id,created_at,lang,conversation_id,public_metrics',
expansions: 'author_id',
'user.fields': 'username,protected'
});
const res = await fetch(`https://api.x.com/2/tweets/search/recent?${params}`, {
headers: { Authorization: `Bearer ${token}` }
});
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
console.log(await res.json());
Production Node clients should loop over meta.next_token, apply bounded exponential backoff for 429 responses, and write raw pages to durable storage.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches7. Capture metadata and make pagination auditable
A successful first response is not the dataset. Search responses are paginated and can contain a next_token; documentation examples use up to 100 results per request. Store one manifest beside the data containing:
- query text, endpoint and route (recent or archive);
- UTC start and end times and the collection start and finish timestamps;
- requested fields, expansions, page count and result count;
- account or project access level, software versions and token identity (never the secret itself);
- HTTP status, error body, retry events and the final page token state.
Keep raw JSON separate from a normalized table. Retaining the original response lets you audit parsing decisions and detect changes in API fields.
8. Understand what the data can and cannot show
Call the result a query-defined sample: posts matching your operators that X returned during your collection period. It is not automatically representative of all X users, all internet users or public opinion.
Missing and changing posts
Protected-account posts, deleted posts and posts withheld in some regions may not be returned. Rate or usage caps can also interrupt collection. Historical studies of the former Twitter Academic API—including a 2022 study that found evidence of “almost complete” samples for many search terms—do not establish completeness for today’s X API.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Time comparisons
For trend analysis, keep the query, filters, requested fields and collection procedure constant. Record revisions as separate waves. A change in access, deletion patterns, language use or operator behavior can look like a sentiment change.
9. Prepare posts for sentiment analysis
Normalize without erasing meaning
- Deduplicate according to your design; reposts may be signal or noise.
- Preserve the original text, IDs and timestamps, and create a separate analysis column for normalized text.
- Decide how links, mentions, hashtags, emojis, punctuation and quoted text are represented.
- Keep replies with conversation identifiers if context matters; do not assume an isolated reply has the same meaning as a standalone post.
- Route multilingual text to language-appropriate models or analyze languages separately.
Validate labels
Do not call model output ground truth. Sarcasm, negation, slang, coded language and domain terminology can defeat a generic classifier. Sample posts for human annotation, define label instructions, measure agreement, and evaluate candidate models on data resembling your topic and languages. Report uncertainty and ambiguous cases instead of forcing every post into positive, negative or neutral.
10. Troubleshooting
401 or 403 responses
Check that the Bearer Token is present, unexpired and attached as Authorization: Bearer .... A valid token can still lack the product access required for a route, especially full archive. Confirm project permissions and current eligibility.
Rank #4
400 invalid query
Test a simple term, then add one operator at a time. Check balanced quotes, supported operators, date format and URL encoding. Keep the exact accepted query in your manifest.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11429 rate or usage limit
Stop issuing requests, honor retry-after when supplied, and use exponential backoff with jitter. Persist completed pages so a retry does not duplicate work. Review both rate limits and monthly usage caps.
Zero results
Verify the date window, language code and exclusions. A protected author, deleted content, regional withholding or an overly narrow phrase can all produce an empty result. Run a deliberately broad diagnostic query, then restore the study query.
Missing pages or duplicates
Follow every returned next_token until it disappears, save page boundaries, and deduplicate by post ID only when your design calls for it. Never infer completeness from one page.
11. Performance, reliability and cost planning
Request only fields and expansions needed by the analysis; large expansions increase transfer and parsing time. Use bounded concurrency only where your account’s limits permit it, and prefer a durable queue for long archive jobs. Estimate pages from pilot counts, but treat the estimate as operational planning rather than a promise: query volume, deletions and access caps change the final count. Confirm current X pricing and quotas directly before budgeting.
Recommended Free Tools
Best Value
Or skip the browser setup
If your project also needs visual snapshots of pages or reports, ScreenshotNeo provides a one-call website screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.
Use the API documentation at https://screenshotneo.com/docs/. A complete cURL call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python and Node.js are available when your pipeline already uses those languages:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is on every plan. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can I use this workflow for private posts?
No. The documented search workflow concerns public posts, and protected-account content may not be returned. Follow X’s current developer and privacy policies for any research involving people.
How far back can recent search go?
The documented recent-search window is the last seven days. Older dates require full-archive access, whose eligibility must be confirmed for your account.
Should I publish the collected post text?
Check X’s current storage and redistribution rules before sharing text or datasets. A reproducible manifest and post IDs may be preferable to redistributing raw content, subject to those rules.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




