DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Cheerio

How to Scrape IMDb Movie Data with Node.js: Ratings, Metadata, and Legal Access

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use IMDb’s official datasets or a licensed API—not unsolicited page scraping. For a maintainable Node.js pipeline, download the daily-refreshed TSV files from datasets.imdbws.com, stream and decompress them, join title.basics with title.ratings on tconst, and preserve the retrieval date. Choose IMDb’s GraphQL product through AWS Data Exchange when you need real-time, search-based, field-selective responses. Cheerio is suitable only for HTML you are authorized to fetch; Puppeteer or Playwright is needed when permitted pages create fields in the browser.

Choose an access path before writing code

IMDb provides several fundamentally different ways to obtain movie information. They differ in freshness, permission, operating cost, and resistance to site changes.

Path Best for Freshness Important constraint
IMDb Contributor Datasets Permitted non-commercial bulk processing Files are refreshed daily Download, storage, and joins are your responsibility; follow the dataset terms
IMDb GraphQL API on AWS Data Exchange Real-time lookup, search, and selected fields Real time Requires an AWS account, credentials, and a subscription to the product; verify current pricing, limits, retention, and redistribution rights
Cheerio Parsing authorized HTML or embedded structured data As returned by the request It parses markup but does not execute JavaScript
Puppeteer or Playwright Authorized pages whose fields are rendered client-side As rendered by the browser Browser execution costs more and page markup can change; authorization is still required

IMDb’s help guidance says: “The data must be taken only from the datasets made available (see IMDb Contributor Datasets).” Website data mining, robots, screen scraping, or similar extraction requires express written consent. Do not use automation to bypass a CAPTCHA, bot check, robots policy, rate limit, or other access control.

Path A: stream the official TSV files in Node.js

What the files contain

title.basics contains the title identifier, title type, primary and original titles, start and end years, runtime, and genres. title.ratings contains averageRating and numVotes. Other files add crew, principals, episodes, alternative titles, and people. All tables use the alphanumeric tconst (or a corresponding name identifier) as the join key. A missing value is represented by N, not by zero.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install prerequisites

Use a current Node.js release with the built-in fetch API (Node 18 or newer is convenient), then create a project:

mkdir imdb-node && cd imdb-node
npm init -y

The script below uses only Node’s standard library. It scans the compressed files without loading multi-gigabyte datasets into memory and emits records for the IDs supplied on the command line.

Runnable streaming join

import { Readable } from 'node:stream';
import { createGunzip } from 'node:zlib';
import * as readline from 'node:readline';

const BASE = 'https://datasets.imdbws.com/';
const wanted = new Set(process.argv.slice(2));
if (wanted.size === 0) {
  console.error('Usage: node imdb-join.mjs tt0111161 tt0468569');
  process.exit(1);
}

function value(raw) {
  return raw === '\N' ? null : raw;
}
function numberOrNull(raw, integer = false) {
  const v = value(raw);
  if (v === null || v === '') return null;
  const n = integer ? Number.parseInt(v, 10) : Number.parseFloat(v);
  return Number.isFinite(n) ? n : null;
}

async function rows(file) {
  const response = await fetch(BASE + file);
  if (!response.ok || !response.body) {
    throw new Error(`${file}: HTTP ${response.status}`);
  }
  const input = Readable.fromWeb(response.body).pipe(createGunzip());
  const rl = readline.createInterface({ input, crlfDelay: Infinity });
  let headers;
  for await (const line of rl) {
    if (!headers) {
      headers = line.split('t');
      continue;
    }
    const fields = line.split('t');
    yield Object.fromEntries(headers.map((h, i) => [h, fields[i] ?? null]));
  }
}

async function selected(file, transform) {
  const result = new Map();
  for await (const row of rows(file)) {
    if (wanted.has(row.tconst)) result.set(row.tconst, transform(row));
  }
  return result;
}

const basics = await selected('title.basics.tsv.gz', row => ({
  tconst: row.tconst,
  titleType: value(row.titleType),
  primaryTitle: value(row.primaryTitle),
  originalTitle: value(row.originalTitle),
  isAdult: value(row.isAdult),
  startYear: numberOrNull(row.startYear, true),
  endYear: numberOrNull(row.endYear, true),
  runtimeMinutes: numberOrNull(row.runtimeMinutes, true),
  genres: value(row.genres)?.split(',').filter(Boolean) ?? null
}));

const ratings = await selected('title.ratings.tsv.gz', row => ({
  averageRating: numberOrNull(row.averageRating),
  numVotes: numberOrNull(row.numVotes, true)
}));

const retrievedAt = new Date().toISOString();
for (const tconst of wanted) {
  const b = basics.get(tconst);
  if (!b) continue;
  if (b.titleType !== 'movie') continue;
  console.log(JSON.stringify({ ...b, ...(ratings.get(tconst) ?? {
    averageRating: null,
    numVotes: null
  }), retrievedAt }));
}

Save that as imdb-join.mjs and run:

node imdb-join.mjs tt0111161 tt0468569

The program makes two full sequential passes, but memory use is limited to the requested IDs. For a catalog import, stream each row into a database instead of retaining a JavaScript object for every title. Keep the raw row or an import checksum so you can audit a later refresh.

Add related metadata with joins

After the basics-and-ratings join, left-join optional files by tconst. Use title.crew for director and writer identifiers, title.principals for top-billed participants, and title.akas for alternative titles. Resolve person identifiers through name.basics. Episode data belongs to series relationships, so do not treat an episode row as a movie record. Keep nullable fields nullable; an unknown runtime or year is not zero.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ratings are time-stamped snapshots

IMDb publishes a daily-computed average and vote count. Store retrievedAt (and, for a scheduled importer, the dataset retrieval date or revision) with every load. A later import can then explain why a rating changed without claiming that an old value was permanent.

Path B: use the licensed GraphQL API for real-time queries

The IMDb GraphQL product on AWS Data Exchange is appropriate when a user needs current values, title or name search, or a small field-selected response rather than a bulk download. Create an AWS account, subscribe to the product, and obtain the credentials required by that subscription. Authentication headers and endpoint details are product-specific, so use the current AWS Data Exchange documentation for your account.

Minimal Node.js request shape

The following is a complete request once IMDB_GRAPHQL_ENDPOINT and the authentication value required by your subscription are set. Request only fields your application needs, and do not put credentials in source control.

const endpoint = process.env.IMDB_GRAPHQL_ENDPOINT;
const token = process.env.IMDB_API_TOKEN;
if (!endpoint || !token) throw new Error('Set IMDB_GRAPHQL_ENDPOINT and IMDB_API_TOKEN');

const query = `query Movie($id: ID!) {
  title(id: $id) {
    id
    titleText { text }
    releaseYear { year }
    runtime { seconds }
    ratingsSummary { aggregateRating voteCount }
    titleType { id }
  }
}`;

const response = await fetch(endpoint, {
  method: 'POST',
  headers: {
    'content-type': 'application/json',
    'authorization': `Bearer ${token}`
  },
  body: JSON.stringify({ query, variables: { id: 'tt0111161' } })
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const payload = await response.json();
if (payload.errors) throw new Error(JSON.stringify(payload.errors));
console.log(JSON.stringify(payload.data.title, null, 2));

Treat that authorization line as a configuration point: some AWS products expose credentials differently. Add bounded retries with exponential backoff for transient failures, cache responses where the license permits, and check the subscribed product’s current rate limits, price, retention rules, and redistribution rights.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Path C: parse authorized HTML with Cheerio

Cheerio is a lightweight parser with jQuery-like selectors. It works on HTML or XML you have permission to retrieve, but it cannot see elements inserted later by client-side JavaScript.

import * as cheerio from 'cheerio';

const response = await fetch(process.env.AUTHORIZED_URL, {
  headers: { 'user-agent': 'YourAppName/1.0 ([email protected])' }
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const $ = cheerio.load(html);

const title = $('h1').first().text().trim() || null;
const jsonLd = [];
$('script[type="application/ld+json"]').each((_, el) => {
  try { jsonLd.push(JSON.parse($(el).text())); } catch {}
});
console.log(JSON.stringify({ title, jsonLd }, null, 2));

Install it with npm install cheerio. Prefer documented attributes or embedded structured data over brittle positional selectors. Cheerio’s URL helper follows redirects (up to five), rejects non-2xx responses, and accepts request options such as a descriptive user-agent; use it only under an authorization that covers the request.

Path D: browser automation for permitted client-rendered pages

If the authorized target creates its fields with JavaScript, use Puppeteer or Playwright. Puppeteer controls Chrome or Firefox and runs headless by default. Wait for a known selector rather than sleeping for an arbitrary number of seconds, and throttle requests.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(process.env.AUTHORIZED_URL, {
    waitUntil: 'networkidle2',
    timeout: 60000
  });
  await page.waitForSelector('[data-testid="movie-title"]', { timeout: 15000 });
  const result = await page.evaluate(() => ({
    title: document.querySelector('[data-testid="movie-title"]')?.textContent?.trim() ?? null,
    html: document.documentElement.outerHTML
  }));
  console.log(JSON.stringify(result));
} finally {
  await browser.close();
}

A browser does not make an otherwise prohibited request permissible. Do not evade bot controls or present unstable page selectors as an IMDb-supported data contract.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

For a visual snapshot of a page you are allowed to capture, ScreenshotNeo provides a one-request screenshot API and MCP server. It is not a replacement for structured IMDb data, but it can remove browser setup when your deliverable is an image or PDF. Before capture it accepts the cookie or consent banner and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options. A direct call for an authorized IMDb URL is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.imdb.com/title/tt0111161/ -o shot.webp

The same endpoint from Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.imdb.com/title/tt0111161/"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And from Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.imdb.com/title/tt0111161/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('node:fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

Every feature is included on every plan: full-page and element capture, device and viewport settings, custom CSS or JavaScript, waits, blocking rules, headers, cookies, user agents, timezone and geolocation, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. The Free plan includes 1,000 shots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Data-quality checklist for production

  • Keep tconst as a string; it is not a numeric ID.
  • Convert N to null before numeric conversion.
  • Parse averageRating as a decimal and numVotes as an integer, while retaining the original row for auditability.
  • Join ratings to basics on tconst; left-join optional crew, principals, episodes, alternative-title, and name tables.
  • Record retrieved_at, the dataset date or revision, API product revision, or the page URL and authorization basis.
  • Validate titleType === "movie" before presenting movie-only results.
  • Expect missing years, runtimes, genres, and ratings; never coerce missing data to zero.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

HTTP errors or an empty compressed response

Check the file name, network access, and the response status before passing the body to createGunzip(). A proxy or an HTML error page will not decompress as TSV.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every field is null

The dataset uses the literal N marker. Convert it before parsing and confirm that your header and column indexes match the current schema.

A known title does not appear

Verify the exact tconst, then check its titleType. A series or episode is not a movie row, and a title may legitimately lack a rating.

Ratings differ from a previous import

Ratings are daily-computed values. Compare the stored retrieval dates and vote counts; do not overwrite history if you need reproducibility.

Cheerio cannot find a visible field

Inspect the received HTML. If the field is injected by JavaScript, Cheerio cannot produce it; use an authorized browser workflow or an official data source instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

GraphQL returns authorization or throttling errors

Confirm that the AWS Data Exchange subscription is active, credentials are loaded from a secret store, and your request uses the product’s current authentication and rate-limit rules. Retry only transient failures and honor backoff instructions.

Performance, reliability, and cost decisions

  • Bulk files: one download can refresh a large local catalog, but requires disk, decompression time, indexing, and scheduled daily jobs.
  • GraphQL: avoids bulk storage and returns only requested fields, but each request depends on subscription limits, network availability, and current license terms.
  • HTML/browser automation: has the highest maintenance risk because markup, scripts, consent dialogs, and load timing can change. Use queues, concurrency limits, explicit timeouts, and structured logs.
  • Caching: cache API responses or rendered results only when the applicable license permits it. Include the retrieval timestamp in cache keys or records.

For most permitted non-commercial analytics, start with the daily TSVs and a streaming importer. Move to the licensed GraphQL API when real-time search or field-level responses justify its subscription. Use Cheerio or a browser only for an explicitly authorized page-level integration.

Frequently Asked Questions

Can I use the IMDb datasets in a commercial application?

The dataset path described here is for permitted non-commercial use. A commercial application should obtain the appropriate license or use the subscribed IMDb API, then follow that product’s current redistribution terms.

Why should a movie importer keep the original TSV row?

Keeping the source row or a checksum makes a later correction explainable: you can distinguish a parser change from a daily data change without reconstructing the old file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a screenshot API a substitute for IMDb metadata?

No. ScreenshotNeo returns an image or PDF. Use the official datasets or licensed GraphQL fields when your application needs searchable ratings and structured metadata.

The Bottom Line

For Node.js, the durable approach is to stream IMDb’s authorized daily datasets, join basics and ratings by tconst, preserve nulls and retrieval dates, and use the licensed GraphQL API when real-time access is worth the subscription. Treat Cheerio and browser automation as permission-dependent tools, not as a way around IMDb’s access rules.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.