Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Website crawling is a queue-and-parse process: begin with seed URLs, fetch each permitted page, extract the data and links you need, normalize and deduplicate URLs, enforce a domain and page limit, and save structured records. Python’s standard library is enough for a small crawler; Beautiful Soup makes focused HTML extraction easier, while Scrapy is the better choice for reusable spiders, pagination, exports, and large crawls.
What a Python crawler actually does
A crawler is not just an HTML downloader. A useful crawl has a repeatable workflow:
- Seed: Put one or more starting URLs in a queue.
- Fetch: Request a page with an identifying user-agent, timeout, and response checks.
- Parse: Extract fields such as the title, headings, prices, or links.
- Normalize: Resolve relative links, remove URL fragments, and apply your URL policy.
- Deduplicate: Track visited URLs so a navigation loop cannot run forever.
- Enforce scope: Restrict hosts or paths and set a page, depth, byte, and rate budget.
- Persist: Write records incrementally to JSON, CSV, a database, or a feed.
This article demonstrates a bounded, same-host crawl. The example prints page titles; replace that extraction with fields appropriate to your project.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Prepare the Python environment
Python’s urllib.request, urllib.parse, and urllib.robotparser modules can fetch pages, manipulate URLs, and read robots.txt rules. Install Beautiful Soup for convenient HTML parsing:
#1 Best Overall
python -m pip install beautifulsoup4
Use a current Python 3 release, run the crawler from a project environment, and never place credentials in source code. The sample below is an instructional pattern and has not been executed or performance-tested here.
A small, bounded crawler with urllib and Beautiful Soup
This breadth-first crawler starts at one URL, stays on its host, removes fragments, checks robots.txt, identifies itself, and stops after 50 discovered pages. It extracts each page title and queues in-scope links.
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.request import Request, urlopen
from urllib.robotparser import RobotFileParser
from urllib.error import HTTPError, URLError
from bs4 import BeautifulSoup
start_url = "https://example.com/"
user_agent = "ExampleResearchBot/1.0 (+https://example.com/bot-info)"
allowed_host = urlparse(start_url).netloc
queue = deque([start_url])
seen = set()
robots = RobotFileParser(urljoin(start_url, "/robots.txt"))
robots.read()
while queue and len(seen) < 50:
url = queue.popleft()
url, _ = urldefrag(url)
parsed = urlparse(url)
if url in seen or parsed.netloc != allowed_host:
continue
if parsed.scheme not in {"http", "https"}:
continue
if not robots.can_fetch(user_agent, url):
continue
request = Request(url, headers={"User-Agent": user_agent})
try:
with urlopen(request, timeout=20) as response:
content_type = response.headers.get_content_type()
if content_type not in {"text/html", "application/xhtml+xml"}:
continue
html = response.read()
except (HTTPError, URLError, TimeoutError) as exc:
print({"url": url, "error": str(exc)})
continue
seen.add(url)
soup = BeautifulSoup(html, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else ""
print({"url": url, "title": title})
for link in soup.select("a[href]"):
next_url = urljoin(url, link["href"])
next_url, _ = urldefrag(next_url)
next_parsed = urlparse(next_url)
if (next_parsed.scheme in {"http", "https"}
and next_parsed.netloc == allowed_host
and next_url not in seen):
queue.append(next_url)
Save it as crawl.py and run python crawl.py. A successful run prints dictionaries containing URLs and titles until the queue is empty or 50 pages have been processed. Replace start_url and the user-agent contact URL before using it on a real site.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchTurn printed titles into records
For repeatable output, append dictionaries to a list and write JSON or CSV after each successful page. Incremental persistence is safer for long jobs: write each record immediately (or use a database transaction) so a process crash does not discard earlier pages. Store the canonical URL you actually fetched, the HTTP status, retrieval time, and only the fields required for your stated purpose.
Normalize URLs deliberately
urljoin handles relative links and urldefrag removes in-page fragments. Decide how to treat trailing slashes, default ports, case, query parameters, tracking parameters, and redirects. Two syntactically different URLs can represent one page; without a policy, the queue can grow through duplicates. Keep query parameters when they change content, and remove them only when you have established that they are tracking noise.
Rank #2
Robots.txt, terms, and responsible crawling
Read https://target.example/robots.txt for the host you intend to crawl and apply the rules to the exact user-agent you send. Google Search Central describes robots.txt as a way to manage crawler traffic and page paths; a disallowed URL can still be discovered through links. A robots file is a technical signal, not complete legal permission.
- Review the site’s terms of service, privacy obligations, copyright restrictions, and applicable local law.
- Use a useful user-agent string with a project name and contact page or email so an owner can request changes.
- Set conservative delays, timeouts, and retry limits. Retry only transient failures, and stop when a server repeatedly returns errors.
- Use an explicit host and path allowlist. Exclude login, checkout, private, and clearly restricted areas.
- Cap pages, crawl depth, response bytes, and total runtime. Validate content types before parsing.
- Cache responses where appropriate and avoid collecting personal or restricted data without a lawful basis.
Robots rules can change while a crawl is running. For a long job, refresh the policy periodically and record the policy URL and retrieval time with your crawl metadata.
When Beautiful Soup is enough
Beautiful Soup is a Python library for pulling data out of HTML and XML files. It is a practical fit when you have a small page budget, one focused site, and straightforward CSS selectors. For example, soup.select("article h2 a") returns matching links, while element.get_text(" ", strip=True) produces readable text.
It does not provide a complete crawl scheduler, feed exporter, retry middleware, distributed queue, or browser engine. You must add those controls yourself. It also will not execute the JavaScript that builds a page after the initial response.
When to choose Scrapy instead
Scrapy is an application framework for crawling websites and extracting structured data. Choose it when your project needs reusable spiders, recursive following, pagination, CSS and XPath selectors, feed exports, pipelines, crawl-depth controls, caching, or middleware. Its official project site labels version 2.19.0 as the latest release in September 2026; verify the current release before pinning dependencies. Claims on that site about more than 15 years in production and more than 500 contributors are vendor-reported figures, not independent benchmarks.
| Need | urllib + Beautiful Soup | Scrapy |
|---|---|---|
| One site or a small page budget | Good fit | Works, but more setup |
| Recursive crawling and pagination | Manual queue logic | Built-in spider and request pattern |
| CSS/XPath selectors | CSS selectors through Beautiful Soup | Selectors plus XPath |
| Feed exports and pipelines | Implement yourself | Built-in support |
| Depth, caching, and middleware | Implement yourself | Documented framework features |
| JavaScript-rendered pages | Usually insufficient alone | Add a browser-rendering integration |
Scrapy’s tutorial demonstrates a quotes spider, extraction, exports, and recursive following. Start with the small script when you are proving a selector or data model; move to Scrapy when operational controls are becoming application code.
JavaScript pages and browser rendering
A normal HTTP client sees the server response, not necessarily the DOM created by JavaScript. If the fields or links appear only after scripts run, inspect the site’s documented API or add a browser-rendering integration to your crawler. Keep the same robots, scope, rate, privacy, and storage controls around browser requests; rendering does not grant permission to access restricted content.
Production safeguards before expanding the crawl
- Errors: Catch HTTP errors, DNS failures, connection resets, and timeouts separately enough to log useful causes.
- Retries: Use bounded exponential backoff for transient responses; do not hammer a site after repeated failures or rate-limit responses.
- Response limits: Check content type and enforce a maximum body size before calling the parser.
- Rate control: Add a delay or token bucket per host. Concurrency should be a deliberate setting, not an accidental consequence of threads.
- Persistence: Save records and crawl state incrementally, with a unique URL key and a schema version.
- Observability: Log requested URL, final URL after redirects, status, bytes, duration, parser outcome, and skip reason.
- Recovery: Make the queue restartable and record failed URLs for a later, limited retry pass.
- Data minimization: Do not retain fields you do not need, especially identifiers, contact details, or user-generated content.
Troubleshooting common failures
The crawler gets 403 or 429 responses
Check robots.txt and terms first. Slow the request rate, identify the client honestly, reduce concurrency, and stop on repeated blocks. Do not attempt to bypass an access control or CAPTCHA.
Robots parsing fails
Some servers return an unavailable, malformed, or blocked robots file. Treat an inability to establish the site’s policy conservatively: pause and review the host manually rather than assuming permission.
Titles or fields are empty
Inspect the raw response and selector. The content may be injected by JavaScript, delivered in a different content type, or marked up differently across templates. Add template-specific selectors and keep a fallback, or use an approved browser-rendering integration.
The queue grows without finishing
Fragments, tracking parameters, calendars, search pages, and session URLs can generate effectively infinite link spaces. Normalize URLs, allow only required paths, remove known tracking parameters, and enforce both a page budget and a depth limit.
The process runs out of memory
Do not retain every HTML document. Stream or discard bodies after extraction, persist records incrementally, cap response size, and use a database-backed queue for larger jobs.
Pages are slow or time out
Keep a finite timeout, record latency, retry only transient failures, and lower concurrency. A timeout is a failed fetch, not a reason to increase pressure on the server.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your goal is a clean visual capture rather than link discovery and structured field extraction, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the result in X-Page-Verdict and X-Billed headers.
Free tools Windows power users keep installed
One-click scans. No signup required.
For the API parameters, options, signed links, asynchronous jobs, bulk capture, and MCP setup, see the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also exposes an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. It supports full-page and element captures, device and viewport settings, retina scale, PDF paper and page options, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone, geolocation, transparency, resizing, chosen-TTL caching, signed public-image links, webhooks, 100-URL bulk calls, a usage API, and an OpenAPI specification.
Best Value
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000 per month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Every feature is included on every plan; yearly billing gives two months free. You can sign up for the free plan with 1,000 screenshots a month and no card, or start at $5 for 3,000.
FAQ
Is crawling the same as scraping?
Crawling describes discovering and fetching pages; scraping describes extracting useful data from them. A project can crawl without storing extracted fields, or scrape a known set of URLs without following links.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Can I crawl a site that disallows my user-agent?
Do not proceed as though the restriction were permission. Review the site’s terms and contact its owner for an authorized method or data feed.
How do I crawl multiple domains?
Represent allowed hosts as an explicit set, keep robots policy per host, and apply separate rate, error, and page budgets. Never let an extracted external link silently expand the scope.
Frequently Asked Questions
Is crawling the same as scraping?
Crawling describes discovering and fetching pages; scraping describes extracting useful data from them. A project can crawl without storing extracted fields, or scrape a known set of URLs without following links.
Can I crawl a site that disallows my user-agent?
Do not proceed as though the restriction were permission. Review the site’s terms and contact its owner for an authorized method or data feed.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteHow do I crawl multiple domains?
Represent allowed hosts as an explicit set, keep robots policy per host, and apply separate rate, error, and page budgets. Never let an extracted external link silently expand the scope.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

