Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The defensible way to collect YouTube comments is through the YouTube Data API, not by scraping YouTube pages. Use commentThreads.list to collect top-level comments and replies returned with each thread, then call comments.list with a top-level comment’s parentId when you need a complete reply set. Paginate with nextPageToken, record exactly what you included, and treat every finding as a result from your collected sample rather than a census of viewers.

Why the API matters more than page scraping

YouTube’s API Services Developer Policies state: “You and your API Clients must not, and must not encourage, enable, or require others to, directly or indirectly, scrape YouTube Applications or Google Applications, or obtain scraped YouTube data or content.” That makes browser automation, HTML parsing, and unofficial endpoints a poor foundation for a production or research workflow. An API client gives you documented fields, pagination, quota accounting, and a collection method you can describe and reproduce.

Access is still conditional. Your project must have YouTube Data API access enabled, and your use of retrieved data must comply with the API Services Developer Policies. Policy language, endpoint behavior, and quota defaults can change, so check Google’s current developer documentation before deploying a long-running collector.

Choose a collection boundary before writing code

“Scrape the comments” is not a reproducible specification. Write down the population you intend to study and the stopping rule first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Target: one video, a set of video IDs, or channel-related threads.
  • Time boundary: for example, comments collected during a stated date range, or all pages available on a collection date.
  • Conversation depth: top-level comments only, inline replies, or a second pass that retrieves every reply for selected parents.
  • Language handling: retain all languages, detect languages, or analyze a defined subset.
  • Exclusions: deleted or unavailable comments, duplicates, spam labels, empty text, and malformed records.
  • Pagination cutoff: a maximum number of pages, a maximum number of comments, or exhaustion of nextPageToken.

Save the query parameters, collection timestamp, video or channel selection logic, page count, and exclusion rules beside the data. Those details are part of the result, not administrative trivia.

Understand the two comment endpoints

commentThreads.list: the starting point

For a video, pass videoId. For channel-related retrieval, the implementation supports channelId and allThreadsRelatedToChannelId. Request snippet for top-level comment data. Adding replies asks for replies that are present in the returned thread.

A thread response is not guaranteed to contain every reply. Treat inline replies as a convenient first page, not proof that the conversation is complete.

comments.list: complete replies for a parent

For a top-level comment, pass its comment ID as parentId. This endpoint returns replies in pages and accepts a maximum page size of 100. Continue while the response contains nextPageToken. A separate pass is necessary whenever reply completeness matters to your question.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal Python collector for one video

The following example collects top-level threads, stores inline replies when supplied, and keeps requesting pages. Replace the API key and video ID with your values. It deliberately writes raw JSON so you retain the fields needed for later auditing.

import json
import requests

API_KEY = "YOUR_API_KEY"
VIDEO_ID = "VIDEO_ID"
url = "https://www.googleapis.com/youtube/v3/commentThreads"

rows = []
page_token = None
while True:
    params = {
        "part": "snippet,replies",
        "videoId": VIDEO_ID,
        "maxResults": 100,
        "key": API_KEY,
    }
    if page_token:
        params["pageToken"] = page_token

    response = requests.get(url, params=params, timeout=30)
    response.raise_for_status()
    payload = response.json()
    rows.extend(payload.get("items", []))
    page_token = payload.get("nextPageToken")
    if not page_token:
        break

with open("comment_threads.json", "w", encoding="utf-8") as f:
    json.dump(rows, f, ensure_ascii=False, indent=2)

print(f"Saved {len(rows)} thread records")

The snippet.topLevelComment.snippet.textDisplay field contains rendered text; textOriginal is preferable when you need the original comment text. Preserve the comment ID, video ID, author metadata that your permitted use allows, publication and update timestamps, like count, and reply count alongside the text.

Fetch every reply for a top-level comment

Use the top-level comment’s ID as parentId. This function follows all pages and returns the API’s reply resources.

def fetch_replies(comment_id, api_key):
    url = "https://www.googleapis.com/youtube/v3/comments"
    replies = []
    page_token = None

    while True:
        params = {
            "part": "snippet",
            "parentId": comment_id,
            "maxResults": 100,
            "key": api_key,
        }
        if page_token:
            params["pageToken"] = page_token
        r = requests.get(url, params=params, timeout=30)
        r.raise_for_status()
        data = r.json()
        replies.extend(data.get("items", []))
        page_token = data.get("nextPageToken")
        if not page_token:
            return replies

Do not assume that the reply count in a thread means you already possess that many reply objects. Compare the count with the records returned by your paginated comments.list pass and retain any discrepancy as a data-quality note.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pagination, quota, and scale planning

comments.list costs one quota unit per call and permits 1–100 results per page. Invalid requests can still cost at least one unit. Google’s API overview describes a default allocation of 10,000 units per day for most endpoints, but says defaults can change and projects can request an extension.

Workload Calls you should expect Planning note
One video, top-level threads One call per page Use the largest practical page size and stop at your documented boundary.
Replies for selected parents One call per reply page, per parent Reply-heavy videos can multiply calls quickly.
Many videos or a channel sample Thread pages plus reply pages Estimate worst-case calls before scheduling the run.

Cache completed pages and make your collector restartable. Persist the last successful page token and the IDs already written. If a request fails after a page is received, retry without duplicating records; use comment IDs as stable deduplication keys.

Clean and prepare the dataset

Preserve raw and derived fields separately

Keep an immutable raw export. Create a second table for normalized text, language, token counts, labels, and analyst notes. Record the normalization choices: lowercasing, URL removal, emoji treatment, repeated-character handling, and whether quoted text was removed.

Handle unavailable and duplicated records

Comments can be deleted, moderated, disabled, or unavailable between pages. Store an explicit status rather than silently dropping an ID. Deduplicate by comment ID, not by text: identical wording can be posted by different people, while an edited comment can retain its identity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Separate spam from disagreement

Do not equate short, repetitive, or profane text with spam automatically. Create transparent rules, sample the items removed by each rule, and report how many records each exclusion affected.

Analysis that answers real audience questions

Theme and question discovery

Start with an open coding pass on a manageable sample. Label requests for clarification, recurring problems, feature ideas, factual corrections, praise, and off-topic discussion. Merge synonymous labels only after reviewing examples. Then apply the codebook to the larger set and measure agreement or manually validate a random subset.

Aggregate sentiment

YouTube’s derived-metrics policy allows aggregate viewer sentiment analysis based on comment analysis when its conditions are met. Sentiment is a property of the analyzed comments, not a measurement of every viewer. Report the sample size, language coverage, classifier or coding method, validation process, and uncertainty.

Replies and conversation structure

Compare top-level comments with replies to see which questions generate discussion. Useful measures include reply rate, thread length, time to first reply, and recurring terms in high-reply threads. Keep these as descriptive statistics for your selected videos and period.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What you must not infer

Do not infer or estimate sensitive protected attributes from comments. Avoid claims such as “the audience is mostly” a protected group unless you have a lawful, explicitly collected, policy-compliant basis that is outside ordinary comment analysis. A comment sample also cannot establish what silent viewers believe.

How to describe representativeness

Use bounded language: “Among 4,200 comments collected from these 12 videos between these dates, 31% contained a request for setup help.” Do not write “31% of viewers asked for help.” Selection effects are substantial: commenters are self-selected, prolific users contribute multiple records, and moderation or disabled comments alter what is visible.

Published studies illustrate why context belongs next to every number. Shajari, Agarwal, and Alassad’s 2023 study analyzed 20 channels, 7,782 videos, 294,199 commenters, and 596,982 comments; those figures describe that study’s dataset and its investigation of suspicious coordinated behavior, not all YouTube activity. Heydari, Zhang, Appel, Wu, and Ranade’s 2019 work compared discourse measures in particular political and apolitical channel groups, so its rates are not platform-wide baselines.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

“Comments are disabled” or the response is empty

The video may have comments disabled, be restricted, deleted, or unavailable to the requesting project. Confirm the video ID and permissions, record the empty result, and do not treat it as evidence that nobody commented.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTTP 403 quota or access errors

Check the API key, enabled API, project quota, and daily usage. Reduce unnecessary fields and calls, resume from saved progress, and request additional quota through Google’s documented process when a legitimate workload requires it.

Replies appear to be missing

Inline replies are not necessarily complete. Call comments.list with the parent ID and follow every page token.

Duplicate rows after a retry

Write pages transactionally and upsert by comment ID. Keep a request log so you can identify which page was retried.

Sentiment looks wrong for slang or irony

Review errors by language, sarcasm, profanity, emoji, and topic vocabulary. Add hand-labeled validation examples and report where the classifier is unreliable instead of presenting a single accuracy number without context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your workflow also needs clean screenshots of video pages, ScreenshotNeo provides a website screenshot API and MCP server. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

Use the ScreenshotNeo API documentation for the full option list. A one-call capture looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.youtube.com/watch?v=VIDEO_ID -o shot.webp

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free. Create a free ScreenshotNeo account to try it.

Other runnable clients

cURL

curl -G "https://www.googleapis.com/youtube/v3/commentThreads" 
  --data-urlencode part=snippet,replies 
  --data-urlencode videoId=VIDEO_ID 
  --data-urlencode maxResults=100 
  --data-urlencode key=YOUR_API_KEY

Python

import requests
r = requests.get(
    "https://www.googleapis.com/youtube/v3/commentThreads",
    params={"part": "snippet,replies", "videoId": "VIDEO_ID", "maxResults": 100, "key": "YOUR_API_KEY"},
    timeout=30,
)
r.raise_for_status()
print(r.json())

Node.js

const q = new URLSearchParams({
  part: 'snippet,replies',
  videoId: 'VIDEO_ID',
  maxResults: '100',
  key: 'YOUR_API_KEY'
});
const res = await fetch(`https://www.googleapis.com/youtube/v3/commentThreads?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = await res.json();
console.log(data.items);

Frequently Asked Questions

Can I collect every YouTube comment with one API request?

No. Results are paginated, and complete replies may require separate comments.list calls for each top-level parent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a comment sample representative of all viewers?

No. Commenters are self-selected and may not reflect silent viewers. Report the videos, dates, exclusions, and sample size that define your result.

How many comments can one page contain?

The comments.list method accepts a page size from 1 to 100.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.