October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Automation

How to Build a Fast Scraping Bot with Python Threading (Safely and Reliably)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a scraper that mostly waits for web servers, a modest concurrent.futures.ThreadPoolExecutor can fetch several authorized URLs at once. Give every request a finite timeout, keep a future-to-URL mapping, collect results with as_completed(), and measure throughput and errors before increasing concurrency. Threads do not make unlimited or automatically compliant traffic: the useful worker count depends on your URLs, network, target policies and parser.

When Python threads make a scraper faster

Downloading a page is usually an I/O-bound operation: your program spends much of its time waiting for DNS, a TCP/TLS connection and the server response. While one thread waits, another can process a different request. Python’s concurrency documentation distinguishes this from CPU-bound work, where threads may not provide the same benefit.

Threading is a fit when requests are independent and you have permission to retrieve them. It is not a license to bypass access controls, defeat CAPTCHAs, ignore terms, or send traffic faster than a site can reasonably handle. If parsing, image processing or machine-learning inference dominates the runtime, separate that CPU-heavy stage and benchmark it independently; a process pool or another design may be more appropriate.

What “fast” should mean

Measure more than elapsed seconds. Record total duration, completed pages, HTTP statuses, exceptions, timeouts, retry count, response bytes and the target’s behavior. A faster run that causes many failures or violates a site’s limits is not a successful scraper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A bounded threaded scraper with the standard library

The following complete example uses urllib.request, which is available with Python. It sets a timeout, closes each response with a context manager, preserves the original URL, and continues when one task fails.

from concurrent.futures import ThreadPoolExecutor, as_completed
from dataclasses import dataclass
from time import perf_counter
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen

@dataclass
class FetchResult:
    url: str
    status: int | None
    body: bytes | None
    error: str | None


def fetch(url: str, timeout: float = 20.0) -> FetchResult:
    request = Request(
        url,
        headers={"User-Agent": "AuthorizedResearchBot/1.0 (contact: [email protected])"},
    )
    try:
        with urlopen(request, timeout=timeout) as response:
            body = response.read()
            return FetchResult(url, response.status, body, None)
    except HTTPError as exc:
        return FetchResult(url, exc.code, None, f"HTTP error: {exc.reason}")
    except (URLError, TimeoutError) as exc:
        return FetchResult(url, None, None, f"Network error: {exc}")
    except Exception as exc:
        return FetchResult(url, None, None, f"Unexpected error: {exc}")


def scrape(urls: list[str], max_workers: int = 6) -> list[FetchResult]:
    results: list[FetchResult] = []
    started = perf_counter()
    with ThreadPoolExecutor(max_workers=max_workers) as pool:
        future_to_url = {pool.submit(fetch, url): url for url in urls}
        for future in as_completed(future_to_url):
            url = future_to_url[future]
            try:
                result = future.result()
            except Exception as exc:
                # Protect the rest of the batch if fetch() ever leaks an exception.
                result = FetchResult(url, None, None, f"Worker error: {exc}")
            results.append(result)
            if result.error:
                print(f"FAIL {url}: {result.error}")
            else:
                print(f"OK   {url}: {result.status} ({len(result.body or b'')} bytes)")
    elapsed = perf_counter() - started
    ok = sum(1 for result in results if result.error is None)
    print(f"{ok}/{len(urls)} succeeded in {elapsed:.2f}s")
    return results


if __name__ == "__main__":
    urls = [
        "https://example.com/",
        "https://www.python.org/",
    ]
    scrape(urls, max_workers=6)

max_workers=6 is a conservative starting choice, not a universal optimum. Start lower for a small or sensitive site. The pool bounds the number of simultaneous worker tasks; it does not guarantee a request rate, because response times vary.

Why the future-to-URL mapping matters

as_completed() yields whichever request finishes first, so output is not in input order. Mapping each future to its URL prevents a fast response from being attributed to the wrong page. It also lets one failure be reported while other downloads continue. If you need original ordering, sort the final records by the input list after collection.

Keep downloading separate from parsing

The example returns bytes. Parse HTML after a successful download, or submit parsing to a separate stage. This makes it clear whether a “speedup” came from overlapping network waits or from changing parser work. Decode using the response’s declared charset when available rather than assuming every page is UTF-8.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect robots.txt, permissions and site limits

The standard library includes urllib.robotparser for reading a site’s robots.txt. Use it as one technical input to your access decision, not as a replacement for terms, contracts, authentication requirements or applicable law. Crawl only URLs you are authorized to access, identify your client honestly, and provide contact information where appropriate.

Keep concurrency bounded per target host. Avoid a large burst across many domains, and do not use threads to evade rate limits, bot checks or access controls. For transient failures, a small, bounded backoff may help; do not blindly retry permanent 4xx responses or repeat a request indefinitely. Cache data you are allowed to reuse, and honor a site’s stated request guidance.

Choosing urllib or Requests

urllib.request keeps this tutorial dependency-free and supports timeout-enabled, context-managed responses. Requests offers a higher-level API, sessions, automatic keep-alive and connection pooling. Its documentation identifies Python 3.10+ support for its documented 2.34.2 release; verify the current release and supported Python versions when you deploy.

Consideration urllib.request Requests
Installation Included with Python Third-party package
Timeouts Pass timeout to urlopen Pass timeout arguments to request methods
Connection reuse Lower-level handling Sessions document keep-alive and connection pooling
API ergonomics More explicit request/response objects Convenient methods, sessions and familiar exceptions
Speed No general winner is established; benchmark equivalent code against your authorized workload.

Requests version of the worker

import requests
from concurrent.futures import ThreadPoolExecutor, as_completed


def fetch_with_requests(url: str, timeout: tuple[float, float] = (5, 20)):
    try:
        response = requests.get(
            url,
            timeout=timeout,
            headers={"User-Agent": "AuthorizedResearchBot/1.0"},
        )
        return {
            "url": url,
            "status": response.status_code,
            "body": response.content,
            "error": None,
        }
    except requests.RequestException as exc:
        return {"url": url, "status": None, "body": None, "error": str(exc)}


def run(urls, workers=6):
    with ThreadPoolExecutor(max_workers=workers) as executor:
        futures = {executor.submit(fetch_with_requests, url): url for url in urls}
        for future in as_completed(futures):
            result = future.result()
            print(result["url"], result["status"], result["error"])

A shared session can reuse connections, but concurrent access and session configuration should be tested with your Requests version. Keep timeout values explicit even when using a session.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Benchmark your own worker count

No universal thread count or percentage speedup applies to every scraper. Use the same URL list, headers, timeout, parser and compliance limits for each run:

  1. Run a sequential baseline and save duration, successes, statuses, errors and bytes.
  2. Run small pools, such as 2, 4 and 6 workers, while watching the target’s response behavior.
  3. Increase only while elapsed time improves without unacceptable errors, retries, resource use or policy violations.
  4. Repeat runs at a comparable time and report the environment, date, target set and limits with any numbers you publish.

Do not benchmark against a site you do not control merely to chase a larger number. A local test server or an authorized staging environment gives you reproducible conditions.

Reliability patterns that prevent silent data loss

Use explicit result records

Return URL, status, body (or parsed data), error type and attempt information. Write successful results incrementally if a batch is large, so a process interruption does not discard completed work.

Classify failures

  • Timeouts and connection resets: record them and retry only when the operation is safe and the failure is plausibly transient.
  • HTTP 429: slow down and follow the server’s stated retry guidance rather than adding workers.
  • HTTP 403 or bot checks: do not attempt to evade them; obtain permission or use an approved interface.
  • 4xx responses: check the URL, authentication and requested resource before retrying.
  • 5xx responses: a bounded backoff can be appropriate, with a cap and a final recorded failure.
  • Malformed or unexpected HTML: keep the raw response when permitted and make parsers tolerant of missing fields.

Protect memory and file descriptors

Reading every response fully into memory is simple but may be unsuitable for large files. Stream or impose a maximum size when your application allows it. Always close responses; the context manager in the example does that even when parsing raises an exception.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting

Everything is slow

Check DNS, TLS, server latency, redirects and your timeout. Confirm that the workload is actually network-bound. If parsing consumes most of the elapsed time, optimize or isolate parsing instead of adding threads.

The program exits with incomplete results

Keep the executor in a with block and consume every submitted future. Persist results as they arrive, and catch exceptions around future.result().

Requests time out after adding workers

Reduce max_workers, inspect target rate limits and verify that your network can sustain the connections. More threads can increase contention and trigger defensive controls.

Character decoding is wrong

Use the response’s declared encoding where reliable, inspect the document’s metadata, and preserve raw bytes for later correction. Do not assume a single encoding for every domain.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots rules are unclear

Parse the file, identify the relevant user-agent group and review the site’s terms or contact its operator. A parser can report technical rules; it cannot decide whether your planned collection is lawful or authorized.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than raw HTML, ScreenshotNeo provides a single website-screenshot API call. It accepts the cookie or consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A minimal call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers full-page and selector captures, lazy-image loading, dark mode, device and retina settings, PDF page controls, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every plan includes every feature. The Free plan provides 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

FAQ

Can threads bypass a site’s robots.txt?

No. Threads are an execution technique, not permission. Check robots guidance, terms, authorization and applicable rules before collecting pages.

Should I use one thread per URL?

No. Submit tasks to a bounded executor. An unbounded thread-per-URL design can exhaust local resources and overwhelm a target.

Do I need a browser for every scraper?

No. For pages whose content is delivered in the initial HTTP response, an HTTP client is simpler. Browser automation is needed only when authorized collection genuinely requires client-side rendering or interaction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I preserve input order while using as_completed()?

Yes. Store each result by its original index or URL, then rebuild the output in the order of the input list after all futures finish.

What should I log for a production crawl?

Log the URL, timestamp, attempt, status, elapsed time, response size, exception class, retry decision and final outcome, while excluding credentials and sensitive response data.

When should I stop increasing concurrency?

Stop when latency or success rate stops improving, target errors or rate-limit responses rise, local resources become constrained, or the site’s permitted request behavior would be exceeded.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.