Recommended Free Tools
The defensible way to collect YouTube comments is through the YouTube Data API, not by scraping YouTube pages. Use commentThreads.list to collect top-level comments and replies returned with each thread, then call comments.list with a top-level comment’s parentId when you need a complete reply set. Paginate with nextPageToken, record exactly what you included, and treat every finding as a result from your collected sample rather than a census of viewers.
Why the API matters more than page scraping
YouTube’s API Services Developer Policies state: “You and your API Clients must not, and must not encourage, enable, or require others to, directly or indirectly, scrape YouTube Applications or Google Applications, or obtain scraped YouTube data or content.” That makes browser automation, HTML parsing, and unofficial endpoints a poor foundation for a production or research workflow. An API client gives you documented fields, pagination, quota accounting, and a collection method you can describe and reproduce.
Access is still conditional. Your project must have YouTube Data API access enabled, and your use of retrieved data must comply with the API Services Developer Policies. Policy language, endpoint behavior, and quota defaults can change, so check Google’s current developer documentation before deploying a long-running collector.
Choose a collection boundary before writing code
“Scrape the comments” is not a reproducible specification. Write down the population you intend to study and the stopping rule first.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
- Target: one video, a set of video IDs, or channel-related threads.
- Time boundary: for example, comments collected during a stated date range, or all pages available on a collection date.
- Conversation depth: top-level comments only, inline replies, or a second pass that retrieves every reply for selected parents.
- Language handling: retain all languages, detect languages, or analyze a defined subset.
- Exclusions: deleted or unavailable comments, duplicates, spam labels, empty text, and malformed records.
- Pagination cutoff: a maximum number of pages, a maximum number of comments, or exhaustion of
nextPageToken.
Save the query parameters, collection timestamp, video or channel selection logic, page count, and exclusion rules beside the data. Those details are part of the result, not administrative trivia.
Understand the two comment endpoints
commentThreads.list: the starting point
For a video, pass videoId. For channel-related retrieval, the implementation supports channelId and allThreadsRelatedToChannelId. Request snippet for top-level comment data. Adding replies asks for replies that are present in the returned thread.
A thread response is not guaranteed to contain every reply. Treat inline replies as a convenient first page, not proof that the conversation is complete.
comments.list: complete replies for a parent
For a top-level comment, pass its comment ID as parentId. This endpoint returns replies in pages and accepts a maximum page size of 100. Continue while the response contains nextPageToken. A separate pass is necessary whenever reply completeness matters to your question.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesMinimal Python collector for one video
The following example collects top-level threads, stores inline replies when supplied, and keeps requesting pages. Replace the API key and video ID with your values. It deliberately writes raw JSON so you retain the fields needed for later auditing.
import json
import requests
API_KEY = "YOUR_API_KEY"
VIDEO_ID = "VIDEO_ID"
url = "https://www.googleapis.com/youtube/v3/commentThreads"
rows = []
page_token = None
while True:
params = {
"part": "snippet,replies",
"videoId": VIDEO_ID,
"maxResults": 100,
"key": API_KEY,
}
if page_token:
params["pageToken"] = page_token
response = requests.get(url, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
rows.extend(payload.get("items", []))
page_token = payload.get("nextPageToken")
if not page_token:
break
with open("comment_threads.json", "w", encoding="utf-8") as f:
json.dump(rows, f, ensure_ascii=False, indent=2)
print(f"Saved {len(rows)} thread records")
The snippet.topLevelComment.snippet.textDisplay field contains rendered text; textOriginal is preferable when you need the original comment text. Preserve the comment ID, video ID, author metadata that your permitted use allows, publication and update timestamps, like count, and reply count alongside the text.
Fetch every reply for a top-level comment
Use the top-level comment’s ID as parentId. This function follows all pages and returns the API’s reply resources.
def fetch_replies(comment_id, api_key):
url = "https://www.googleapis.com/youtube/v3/comments"
replies = []
page_token = None
while True:
params = {
"part": "snippet",
"parentId": comment_id,
"maxResults": 100,
"key": api_key,
}
if page_token:
params["pageToken"] = page_token
r = requests.get(url, params=params, timeout=30)
r.raise_for_status()
data = r.json()
replies.extend(data.get("items", []))
page_token = data.get("nextPageToken")
if not page_token:
return replies
Do not assume that the reply count in a thread means you already possess that many reply objects. Compare the count with the records returned by your paginated comments.list pass and retain any discrepancy as a data-quality note.
Pagination, quota, and scale planning
comments.list costs one quota unit per call and permits 1–100 results per page. Invalid requests can still cost at least one unit. Google’s API overview describes a default allocation of 10,000 units per day for most endpoints, but says defaults can change and projects can request an extension.
| Workload | Calls you should expect | Planning note |
|---|---|---|
| One video, top-level threads | One call per page | Use the largest practical page size and stop at your documented boundary. |
| Replies for selected parents | One call per reply page, per parent | Reply-heavy videos can multiply calls quickly. |
| Many videos or a channel sample | Thread pages plus reply pages | Estimate worst-case calls before scheduling the run. |
Cache completed pages and make your collector restartable. Persist the last successful page token and the IDs already written. If a request fails after a page is received, retry without duplicating records; use comment IDs as stable deduplication keys.
Clean and prepare the dataset
Preserve raw and derived fields separately
Keep an immutable raw export. Create a second table for normalized text, language, token counts, labels, and analyst notes. Record the normalization choices: lowercasing, URL removal, emoji treatment, repeated-character handling, and whether quoted text was removed.
Handle unavailable and duplicated records
Comments can be deleted, moderated, disabled, or unavailable between pages. Store an explicit status rather than silently dropping an ID. Deduplicate by comment ID, not by text: identical wording can be posted by different people, while an edited comment can retain its identity.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Separate spam from disagreement
Do not equate short, repetitive, or profane text with spam automatically. Create transparent rules, sample the items removed by each rule, and report how many records each exclusion affected.
Analysis that answers real audience questions
Theme and question discovery
Start with an open coding pass on a manageable sample. Label requests for clarification, recurring problems, feature ideas, factual corrections, praise, and off-topic discussion. Merge synonymous labels only after reviewing examples. Then apply the codebook to the larger set and measure agreement or manually validate a random subset.
Aggregate sentiment
YouTube’s derived-metrics policy allows aggregate viewer sentiment analysis based on comment analysis when its conditions are met. Sentiment is a property of the analyzed comments, not a measurement of every viewer. Report the sample size, language coverage, classifier or coding method, validation process, and uncertainty.
Replies and conversation structure
Compare top-level comments with replies to see which questions generate discussion. Useful measures include reply rate, thread length, time to first reply, and recurring terms in high-reply threads. Keep these as descriptive statistics for your selected videos and period.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What you must not infer
Do not infer or estimate sensitive protected attributes from comments. Avoid claims such as “the audience is mostly” a protected group unless you have a lawful, explicitly collected, policy-compliant basis that is outside ordinary comment analysis. A comment sample also cannot establish what silent viewers believe.
How to describe representativeness
Use bounded language: “Among 4,200 comments collected from these 12 videos between these dates, 31% contained a request for setup help.” Do not write “31% of viewers asked for help.” Selection effects are substantial: commenters are self-selected, prolific users contribute multiple records, and moderation or disabled comments alter what is visible.
Published studies illustrate why context belongs next to every number. Shajari, Agarwal, and Alassad’s 2023 study analyzed 20 channels, 7,782 videos, 294,199 commenters, and 596,982 comments; those figures describe that study’s dataset and its investigation of suspicious coordinated behavior, not all YouTube activity. Heydari, Zhang, Appel, Wu, and Ranade’s 2019 work compared discourse measures in particular political and apolitical channel groups, so its rates are not platform-wide baselines.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common failures and fixes
“Comments are disabled” or the response is empty
The video may have comments disabled, be restricted, deleted, or unavailable to the requesting project. Confirm the video ID and permissions, record the empty result, and do not treat it as evidence that nobody commented.
HTTP 403 quota or access errors
Check the API key, enabled API, project quota, and daily usage. Reduce unnecessary fields and calls, resume from saved progress, and request additional quota through Google’s documented process when a legitimate workload requires it.
Replies appear to be missing
Inline replies are not necessarily complete. Call comments.list with the parent ID and follow every page token.
Duplicate rows after a retry
Write pages transactionally and upsert by comment ID. Keep a request log so you can identify which page was retried.
Sentiment looks wrong for slang or irony
Review errors by language, sarcasm, profanity, emoji, and topic vocabulary. Add hand-labeled validation examples and report where the classifier is unreliable instead of presenting a single accuracy number without context.
Best Value
Or skip the browser setup
If your workflow also needs clean screenshots of video pages, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
Use the ScreenshotNeo API documentation for the full option list. A one-call capture looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.youtube.com/watch?v=VIDEO_ID -o shot.webp
There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free. Create a free ScreenshotNeo account to try it.
Other runnable clients
cURL
curl -G "https://www.googleapis.com/youtube/v3/commentThreads"
--data-urlencode part=snippet,replies
--data-urlencode videoId=VIDEO_ID
--data-urlencode maxResults=100
--data-urlencode key=YOUR_API_KEY
Python
import requests
r = requests.get(
"https://www.googleapis.com/youtube/v3/commentThreads",
params={"part": "snippet,replies", "videoId": "VIDEO_ID", "maxResults": 100, "key": "YOUR_API_KEY"},
timeout=30,
)
r.raise_for_status()
print(r.json())
Node.js
const q = new URLSearchParams({
part: 'snippet,replies',
videoId: 'VIDEO_ID',
maxResults: '100',
key: 'YOUR_API_KEY'
});
const res = await fetch(`https://www.googleapis.com/youtube/v3/commentThreads?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = await res.json();
console.log(data.items);
Frequently Asked Questions
Can I collect every YouTube comment with one API request?
No. Results are paginated, and complete replies may require separate comments.list calls for each top-level parent.
Is a comment sample representative of all viewers?
No. Commenters are self-selected and may not reflect silent viewers. Report the videos, dates, exclusions, and sample size that define your result.
How many comments can one page contain?
The comments.list method accepts a page size from 1 to 100.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

