Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A distributed crawler in Node.js is four cooperating systems: a durable URL frontier, BullMQ workers that fetch and parse pages, persistent crawl-state storage, and coordination rules for robots and per-origin politeness. BullMQ distributes jobs and helps recover from worker failures, but it does not canonicalize URLs, deduplicate your frontier, enforce robots.txt, or make database writes exactly once. Those are application responsibilities.

This design scales by adding workers on separate processes or machines while keeping URL identity, crawl policy, and results durable. The examples below use BullMQ with Redis, native Node.js fetch, and a relational database for authoritative crawl state.

1. Define the crawl contract before writing workers

Write down what the crawler is allowed to visit and what ends a crawl. Ambiguity here produces duplicate work and accidental crawling outside the intended site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scope and URL policy

  • Allow only http and https; reject other schemes such as file: or javascript:.
  • Allow-list hostnames (and decide whether subdomains are included). Compare normalized hostnames, not raw link strings.
  • Normalize URLs by lower-casing the host, removing the default port, resolving dot segments, and consistently handling fragments. Fragments usually identify a document location in a browser, not a new HTTP resource, so remove them unless your application explicitly needs them.
  • Choose a query policy. You might retain all parameters, retain an allow-list such as page, or drop tracking parameters. Apply the same policy before deduplication.
  • Set maximum depth, maximum pages, a time budget, and termination conditions. A queue becoming empty is not the only possible stopping rule when retries or delayed jobs remain.

A stable identity function

Create one canonical string and use it everywhere: frontier keys, job IDs, database uniqueness constraints, and result records. A queue job ID prevents duplicate jobs only according to the ID you provide; it is not a crawler-specific canonicalizer.

#1 Best Overall
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
  • Dual band router upgrades to 1200 Mbps high speed internet (300mbps for 2.4GHz plus 900Mbps for 5GHz), reducing buffering and ideal for 4K stream
  • Full Gigabit Ports - Gigabit Router with 4 Gigabit LAN ports, ideal for any internet plan and allow you to directly connect your wired devices
  • Boosted Coverage - Four external antennas equipped with Beamforming technology extend and concentrate the Wi-Fi signals
  • MU-MIMO technology - (5GHz band) allows high speeds for multiple devices simultaneously
  • Access Point Mode - Supports AP Mode to transform your wired connection into wireless network, an ideal wireless router for home
export function canonicalize(raw, base) {
  const u = new URL(raw, base);
  if (!['http:', 'https:'].includes(u.protocol)) return null;
  u.hash = '';
  u.hostname = u.hostname.toLowerCase();
  if ((u.protocol === 'http:' && u.port === '80') ||
      (u.protocol === 'https:' && u.port === '443')) u.port = '';
  for (const key of [...u.searchParams.keys()]) {
    if (/^(utm_|fbclid$|gclid$)/i.test(key)) u.searchParams.delete(key);
  }
  return u.toString();
}

2. Give the frontier durable ownership

Use a database table as the source of truth and BullMQ as the dispatch mechanism. A minimal frontier record needs a unique canonical URL, state, depth, attempt count, next-eligible time, and timestamps. Store fetch metadata separately or in the same record.

State Meaning Transition
queued Eligible for dispatch Inserted into BullMQ with the canonical URL as job identity
leased A worker owns a fetch attempt Set with a lease timeout or heartbeat
fetched Response was classified and links processed Persist status, headers, content and discovered URLs
retry Temporary failure, with a future retry time Requeue only after the backoff time
dead Permanent failure or policy rejection Keep the reason for audit and reporting

Insert the URL under a unique constraint before enqueueing it. If the enqueue operation fails after the insert, a dispatcher can periodically find queued rows without a corresponding job and submit them again. That outbox-style repair closes the gap between database durability and queue delivery.

Redis and BullMQ setup

BullMQ is a Node.js library built on Redis. Its Queue and Worker abstractions allow workers in one process, separate processes, or separate machines to consume the same queue. In production, BullMQ documentation advises Redis persistence, the noeviction max-memory policy, deliberate reconnect behavior, error logging, and graceful worker shutdown.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
TP-Link ER605, Wired Gigabit VPN Router
  • 【Five Gigabit Ports】1 Gigabit WAN Port plus 2 Gigabit WAN/LAN Ports plus 2 Gigabit LAN Port. Up to 3 WAN ports optimize bandwidth usage through one device.
  • 【One USB WAN Port】Mobile broadband via 4G/3G modem is supported for WAN backup by connecting to the USB port. For complete list of compatible 4G/3G modems, please visit TP-Link website.
  • 【Abundant Security Features】Advanced firewall policies, DoS defense, IP/MAC/URL filtering, speed test and more security functions protect your network and data.
  • 【Highly Secure VPN】Supports up to 20× LAN-to-LAN IPsec, 16× OpenVPN, 16× L2TP, and 16× PPTP VPN connections.
  • Security - SPI Firewall, VPN Pass through, FTP/H.323/PPTP/SIP/IPsec ALG, DoS Defence, Ping of Death and Local Management. Standards and Protocols IEEE 802.3, 802.3u, 802.3ab, IEEE 802.3x, IEEE 802.1q
npm install bullmq ioredis cheerio
import { Queue } from 'bullmq';
import IORedis from 'ioredis';

const connection = new IORedis(process.env.REDIS_URL, {
  maxRetriesPerRequest: null,
  enableReadyCheck: true
});

export const crawlQueue = new Queue('crawl-fetch', { connection });

export async function enqueue(url, depth) {
  return crawlQueue.add('fetch', { url, depth }, {
    jobId: url,
    attempts: 4,
    backoff: { type: 'exponential', delay: 2000 },
    removeOnComplete: { count: 10000 },
    removeOnFail: { count: 10000 }
  });
}

The retry settings above are queue-level recovery aids, not proof of exactly-once crawling. A worker can fetch successfully and crash before recording the result, then receive the job again. Make result writes idempotent with unique URL keys and an upsert or compare-and-set update.

3. Implement a worker that classifies every attempt

A worker should claim a lease, check robots policy, enforce an origin limiter, fetch with a bounded timeout, classify the response, extract links, and commit state in one repeat-safe workflow.

import { Worker } from 'bullmq';
import IORedis from 'ioredis';
import * as cheerio from 'cheerio';
import { canonicalize } from './urls.js';
import { allowedByRobots } from './robots.js';
import { claim, saveResult, discover } from './store.js';
import { waitForOrigin } from './politeness.js';

const connection = new IORedis(process.env.REDIS_URL, {
  maxRetriesPerRequest: null
});

const worker = new Worker('crawl-fetch', async job => {
  const { url, depth } = job.data;
  if (!(await claim(url, job.id))) return; // already completed by another attempt

  const target = new URL(url);
  if (!(await allowedByRobots(target))) {
    await saveResult(url, { state: 'dead', reason: 'robots-disallowed' });
    return;
  }
  await waitForOrigin(target.origin); // shared, distributed limiter

  const controller = new AbortController();
  const timer = setTimeout(() => controller.abort(), 30000);
  let response;
  try {
    response = await fetch(url, {
      redirect: 'follow',
      signal: controller.signal,
      headers: { 'user-agent': process.env.CRAWLER_UA ?? 'ExampleCrawler/1.0' }
    });
  } finally { clearTimeout(timer); }

  const contentType = response.headers.get('content-type') ?? '';
  const body = contentType.includes('text/html')
    ? await response.text() : '';
  const links = [];
  if (body) {
    const $ = cheerio.load(body);
    $('a[href]').each((_, el) => {
      const next = canonicalize($(el).attr('href'), url);
      if (next) links.push(next);
    });
  }
  await saveResult(url, {
    state: 'fetched', status: response.status,
    contentType, finalUrl: response.url, fetchedAt: new Date(),
    body, depth
  });
  for (const next of new Set(links)) {
    if (depth + 1 <= Number(process.env.MAX_DEPTH ?? 3))
      await discover(next, depth + 1);
  }
}, { connection, concurrency: Number(process.env.WORKER_CONCURRENCY ?? 8) });

worker.on('error', err => console.error('worker error', err));

async function shutdown() {
  await worker.close();
  await connection.quit();
  process.exit(0);
}
process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);

Classify failures explicitly

  • 2xx responses are successful fetches; persist the final URL after redirects.
  • 3xx responses should be handled by a bounded redirect policy and recorded, not silently followed forever.
  • 4xx responses are usually permanent for that URL, while 429 requires a retry time based on server signals and your policy.
  • 5xx, DNS errors, connection resets and timeouts are normally retryable up to a bounded attempt count.
  • Non-HTML content should be recorded with its media type and skipped by the HTML parser unless your scope includes it.

4. Robots.txt and distributed politeness

RFC 9309 asks crawlers to honor robots.txt rules and states, “These rules are not a form of access authorization.” A successful robots.txt fetch requires following parseable rules; unavailable and unreachable responses have separate handling in the RFC, so do not treat every network error as permission to crawl.

Rank #3
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
  • Dual-band Wi-Fi with 5 GHz speeds up to 867 Mbps and 2.4 GHz speeds up to 300 Mbps, delivering 1200 Mbps of total bandwidth¹. Dual-band routers do not support 6 GHz. Performance varies by conditions, distance to devices, and obstacles such as walls.
  • Covers up to 1,000 sq. ft. with four external antennas for stable wireless connections and optimal coverage.
  • Supports IGMP Proxy/Snooping, Bridge and Tag VLAN to optimize IPTV streaming
  • Access Point Mode - Supports AP Mode to transform your wired connection into wireless network, an ideal wireless router for home
  • Advanced Security with WPA3 - The latest Wi-Fi security protocol, WPA3, brings new capabilities to improve cybersecurity in personal networks

Cache robots decisions for a bounded period, keying them by origin and user-agent. Refresh on expiry or when the policy requires it. A robots parser supplies allow/disallow matching, but it does not impose a universal request interval.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Coordinate limits across machines

A sleep in each Node.js process is insufficient: eight machines can still send eight simultaneous requests to one host. Store the next permitted timestamp or a token bucket in Redis, keyed by origin, and acquire it atomically before every request. Add a maximum in-flight count per origin, honor Retry-After where appropriate, and cap response size and download time.

5. Make state and side effects repeat-safe

Use a unique constraint on canonical URL and write fetch results with an upsert. Include an attempt ID, status, final URL, content type, byte count, error class and timestamps. When discovering links, insert them with ON CONFLICT DO NOTHING; only the transaction that inserted a new row should enqueue a job. If enqueueing is external to the transaction, use the repair scan described earlier.

Rank #4
Sale
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
  • DUAL-BAND WIFI 6 ROUTER: Wi-Fi 6(802.11ax) technology achieves faster speeds, greater capacity and reduced network congestion compared to the previous gen. All WiFi routers require a separate modem. Dual-Band WiFi routers do not support the 6 GHz band.
  • AX1800: Enjoy smoother and more stable streaming, gaming, downloading with 1.8 Gbps total bandwidth (up to 1200 Mbps on 5 GHz and up to 574 Mbps on 2.4 GHz). Performance varies by conditions, distance to devices, and obstacles such as walls.
  • CONNECT MORE DEVICES: Wi-Fi 6 technology communicates more data to more devices simultaneously using revolutionary OFDMA technology
  • EXTENSIVE COVERAGE: Achieve the strong, reliable WiFi coverage with Archer AX1800 as it focuses signal strength to your devices far away using Beamforming technology, 4 high-gain antennas and an advanced front-end module (FEM) chipset
  • OUR CYBERSECURITY COMMITMENT: TP-Link is a signatory of the U.S. Cybersecurity and Infrastructure Security Agency’s (CISA) Secure-by-Design pledge. This device is designed, built, and maintained, with advanced security as a core requirement.

Leases prevent abandoned work from remaining permanently in leased. A periodic sweeper returns expired leases to retry, while a worker heartbeat extends active leases. Design every transition so a late completion cannot overwrite a newer terminal state.

6. Run and scale the system

  1. Start Redis with persistence enabled and configure maxmemory-policy noeviction. Monitor disk, memory, latency and rejected commands.
  2. Run one seeder that normalizes initial URLs, inserts them durably and enqueues only newly inserted rows.
  3. Run workers as separate Node.js processes or machines. Increase concurrency only after measuring origin pressure, memory use and downstream database capacity.
  4. Run a scheduler for delayed retries, expired leases, robots refreshes and the outbox repair scan.
  5. Export queue depth, active jobs, oldest queued age, retry counts, per-origin request rate, response classes and database write failures.
  6. On deployment, stop accepting new work, let active jobs finish or lease-expire, close workers gracefully, then terminate the process.

What to measure

There is no universal throughput number for this architecture. Capacity depends on page size, parsing cost, network latency, origin limits, Redis and database configuration, and worker concurrency. Benchmark with your URL mix and policy; do not infer production capacity from queue depth alone.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

7. Troubleshooting common failures

Symptom Likely cause Fix
Duplicate fetches Different URL spellings or no unique frontier key Centralize canonicalization and enforce a database uniqueness constraint.
Jobs vanish after a restart Redis eviction or non-durable configuration Enable persistence, use noeviction, and run an outbox/lease repair scan.
One host is overloaded Per-process delays are not coordinated Use a Redis-backed per-origin limiter and in-flight cap.
Pages remain leased forever No heartbeat or expiry sweep Add lease TTLs, heartbeat extension and a sweeper.
Retries create conflicting results Non-idempotent result writes Upsert by canonical URL and reject stale attempt updates.
Robots behavior is inconsistent Uncached or incorrectly parsed policy responses Cache by origin/user-agent, distinguish successful, unavailable and unreachable responses, and log the decision.
Memory rises during large pages Unbounded response buffering Enforce a byte limit, abort oversized downloads, and avoid storing bodies unless required.

Or skip the browser setup

If the crawler’s job is to obtain clean visual captures rather than parse HTML, ScreenshotNeo provides a single HTTP endpoint and an MCP server for Claude, Cursor and other MCP clients. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. It supports PNG, JPEG, WebP and PDF, plus full-page and element captures, device and retina settings, custom CSS/JavaScript, waits, blocking rules, cookies, headers, geolocation, caching, signed links, asynchronous jobs, bulk capture and usage reporting.

See the ScreenshotNeo API documentation for parameter details.

Best Value
Sale
TP-Link Dual-Band AX3000 Wi-Fi 6 Wireless Gigabit Internet Router for Home
  • Next-Gen Gigabit Wi-Fi 6 Speeds: 2402 Mbps on 5 GHz and 574 Mbps on 2.4 GHz bands ensure smoother streaming and faster downloads; support VPN server and VPN client¹
  • A More Responsive Experience: Enjoy smooth gaming, video streaming, and live feeds simultaneously. OFDMA makes your Wi-Fi stronger by allowing multiple clients to share one band at the same time, cutting latency and jitter.²
  • Expanded Wi-Fi Coverage: 4 high-gain external antennas and Beamforming technology combine to extend strong, reliable, Wi-Fi throughout your home.
  • Improved Battery Life: Target Wake Time helps your devices to communicate efficiently while consuming less power.
  • Improved Cooling Design: No heat ups, no throttles. A larger heat sink and redefined case design cools the WiFi 6 system and enables your network to stay at top speeds in more versatile environments.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes an MCP server so AI agents can take screenshots directly. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Frequently Asked Questions

Does BullMQ replace a crawler database?

No. It dispatches and retries jobs, while a durable database should remain authoritative for URL identity, crawl state and results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can robots.txt set a universal crawl delay?

No. RFC 9309 defines robots rules and retrieval handling, but your application must choose and coordinate its own per-origin rate policy.

Should every failed fetch be retried?

No. Retry transient network and server failures with bounded backoff; record policy rejections and most permanent client errors without repeated attempts.

Quick Recap

Bestseller No. 1
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
TP-Link AC1200 Gigabit Dual Band WiFi Router (Archer A6)
MU-MIMO technology - (5GHz band) allows high speeds for multiple devices simultaneously
$44.99
SaleBestseller No. 2
Bestseller No. 3
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
TP-Link AC1200 WiFi Router Dual Band Wireless Internet Router (Archer A54)
Supports IGMP Proxy/Snooping, Bridge and Tag VLAN to optimize IPTV streaming
$34.99
SaleBestseller No. 4
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
TP-Link AX1800 WiFi 6 Router (Archer AX21 V5)
VPN SERVER: Archer AX21 Supports both Open VPN Server and PPTP VPN Server
$59.98

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.