Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBuild an AI scraper as a pipeline, not as one prompt. Put a policy and permission gate before discovery, use Scrapy for crawl orchestration, add Playwright when pages require JavaScript or authorized interaction, send only the necessary content to an LLM, validate every field against the source, and retain timestamps and provenance with each record. This design produces reproducible JSON while limiting legal, security, cost, and reliability problems.
The reference architecture
A production scraper has separate stages with explicit inputs and outputs. Keeping them separate lets you replace a browser, parser, or model without rewriting the whole system.
- Discovery and policy gate: identify the site owner, purpose, geography, data categories, terms, robots.txt rules, CAPTCHAs, and machine-readable rights reservations.
- Fetch: schedule URLs, enforce concurrency, retry transient failures, and record response status and timing.
- Browser rendering: use Playwright only for pages that need client-side JavaScript, an authorized login flow, or interaction that plain HTTP cannot reproduce.
- Extraction: pass the smallest useful content window to an LLM and require a typed schema.
- Validation: check types, required fields, ranges, duplicates, source evidence, and confidence. Route failures to a re-fetch or a person.
- Storage and monitoring: retain raw hashes or snapshots where lawful, normalized records, model and prompt versions, policy decisions, deletion status, and operational metrics.
Start with permission, purpose, and data minimisation
Check the source before downloading
Prefer an official API or licensed feed when one exists and its licence and limits fit your use case. Canadian privacy commissioners note that an API gives a platform more control over authorized collection and can help detect or mitigate unauthorized scraping.
Read the terms of service, robots.txt, CAPTCHAs, and any rights-reservation signal. CNIL says scraping is not inherently prohibited under GDPR, but recommends excluding sites that oppose scraping through technical or legal means such as CAPTCHAs, robots.txt, or terms. The Italian Garante’s 30 May 2024 guidance also points to reserved areas, anti-scraping clauses, traffic monitoring, and robots.txt as ways to hinder indiscriminate collection.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
Design for personal-data obligations
The European Data Protection Board’s 8 July 2026 guidance describes web scraping as large-scale automated extraction that can create significant risks for people’s personal data. If your dataset contains personal data, document a lawful basis, a specific purpose, transparency information, minimisation rules, retention limits, accuracy checks, and controls for special-category data. The UK’s ICO reported 77 organisational and 16 public responses to its 2024 consultation; 19 respondents (61%) agreed with its initial analysis that legitimate interests was the sole available lawful basis for current web-scraped personal-data training practices, subject to necessity and balancing tests. That is a regulatory position for that context, not a universal permission to scrape.
Keep an audit record
For every source, store the collection date, URL, policy decision, rights signal, lawful-basis analysis, transformation steps, model and version identifiers, and any exclusion or deletion decision. The European Commission’s AI Act obligations for general-purpose AI providers include technical documentation, a copyright-compliance policy, and a sufficiently detailed summary of training content; these records make your own decisions explainable even when you are not a model provider.
Build the crawl with Scrapy, then add Playwright selectively
Install the components
python -m venv .venv
. .venv/bin/activate
pip install scrapy scrapy-playwright pydantic jsonschema
playwright install chromium
Scrapy supplies queues, concurrency, retries, item pipelines, and middleware. Its RobotsTxtMiddleware filters requests forbidden by robots.txt when ROBOTSTXT_OBEY is enabled. Playwright supplies a real browser for client-rendered pages and interactions you are authorized to perform.
A minimal Scrapy project with browser fallback
# settings.py
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 8
DOWNLOAD_TIMEOUT = 45
RETRY_TIMES = 2
DOWNLOAD_HANDLERS = {
"http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
"https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
PLAYWRIGHT_BROWSER_TYPE = "chromium"
PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT = 30000
# spiders/catalog.py
import scrapy
from scrapy_playwright.page import PageMethod
class CatalogSpider(scrapy.Spider):
name = "catalog"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
def parse(self, response):
for card in response.css("article.product"):
yield {
"url": response.url,
"title": card.css("h2::text").get(default="").strip(),
"price_text": card.css(".price::text").get(default="").strip(),
}
for href in response.css("a.next::attr(href)").getall():
yield response.follow(href, callback=self.parse)
def parse_js(self, response):
# Use this callback for a page whose content appears after JavaScript runs.
yield {"url": response.url, "html": response.text}
# Request example for an interactive, authorized page:
# yield scrapy.Request(
# url,
# callback=self.parse_js,
# meta={"playwright": True,
# "playwright_page_methods": [
# PageMethod("wait_for_selector", "article.product"),
# PageMethod("click", "button.load-more"),
# ]})
Do not turn every request into a browser request: browser pages consume more CPU and memory, are slower, and create more opportunities for blocking. Start with HTTP responses and escalate only when a selector, network trace, or page behavior proves it is necessary.
Free tools Windows power users keep installed
One-click scans. No signup required.
Make LLM extraction deterministic enough to operate
Send a narrow, labelled input
Strip navigation, advertisements, repeated footers, and unrelated sections before calling the model. Label the remaining text as untrusted source content and state that instructions inside it are data, not commands. Never let page text override your system or application instructions; this is the basic defence against prompt injection.
SYSTEM = """Extract records from SOURCE_TEXT. Return JSON only, matching this schema:
{"items":[{"name":"string","price": "number|null", "currency":"string|null", "availability":"string|null", "source_quote":"string"}]}
Use null when a value is absent. Never invent values. source_quote must be an exact short span."""
USER = "SOURCE_URL: https://example.com/catalognSOURCE_TEXT:n" + cleaned_text
Validate before writing to your database
from decimal import Decimal
from pydantic import BaseModel, Field, ValidationError
class Item(BaseModel):
name: str = Field(min_length=1)
price: Decimal | None = Field(default=None, ge=0)
currency: str | None = Field(default=None, min_length=3, max_length=3)
availability: str | None = None
source_quote: str = Field(min_length=1)
class Result(BaseModel):
items: list[Item]
def validate_result(model_json: str, source_text: str) -> Result:
result = Result.model_validate_json(model_json)
for item in result.items:
if item.source_quote not in source_text:
raise ValueError(f"Evidence not found for {item.name!r}")
return result
# On ValidationError or ValueError: save the failed payload, re-fetch once,
# then send it to a review queue instead of silently accepting it.
Keep the original URL and capture timestamp beside every item. Add a content hash so a later run can distinguish a changed page from a model variation. Track schema-error rate, evidence failures, duplicate keys, model latency, token usage, and the percentage of pages requiring browser rendering.
JavaScript, pagination, and authenticated flows
Recognise when HTTP is insufficient
- Initial HTML contains an empty application shell while the browser later receives JSON.
- Important content appears only after scrolling, clicking “load more,” or selecting a tab.
- An authorized account is required and the workflow has a documented purpose.
- Content depends on a client-side locale, timezone, or geolocation setting.
Use Playwright’s network events to locate the underlying JSON endpoint; if that endpoint is public and permitted, fetching it directly is usually cheaper and more stable than rendering the page. Never bypass a CAPTCHA or bot check. A block is a policy signal, not an invitation to escalate evasion.
Control pagination and freshness
Persist a canonical URL and an item key, then stop when the next-page link repeats or no new keys appear. Use conditional requests where supported, cache responses with a defined TTL, and schedule high-change sources more often than archival sources. A failed refresh must not overwrite the last known good record; mark it stale and alert instead.
Rank #3
Framework and extraction choices
| Approach | Best use | Strengths | Costs and risks |
|---|---|---|---|
| Scrapy | Large URL sets and recurring crawls | Queues, retries, concurrency, pipelines, and robots.txt middleware | Does not execute page JavaScript by itself; selectors and site changes require maintenance |
| Playwright | Client-rendered pages and authorized interactions | Real browser, waiting, clicking, screenshots, and network inspection | Higher CPU, memory, latency, and operational complexity; browser sessions can trigger anti-bot controls |
| LLM extraction | Messy layouts, varied labels, and multilingual fields | Useful semantic normalization with a single schema | Token cost, latency, nondeterminism, prompt injection, and the need for evidence validation |
For LLMs, compare schema control, validation effort, latency, token cost, multilingual coverage, and resistance to prompt injection. A conventional parser should handle stable fields first; reserve the model for ambiguity that rules cannot resolve.
Storage, privacy, and governance in production
Use layered records
- Raw layer: response hash, status, headers, capture time, and a lawful snapshot or extracted text where retention is permitted.
- Normalized layer: typed fields, canonical URL, source quote, confidence, and deduplication key.
- Lineage layer: parser version, prompt template, model/version, policy decision, and deletion or exclusion status.
Restrict access to raw personal data, encrypt it, set an expiry, and propagate deletion requests through normalized and derived stores. Accuracy checks should compare changed values with the source rather than trusting a confidence number alone.
Monitor failure modes
Alert on robots or policy changes, rising HTTP 403/429 rates, CAPTCHA frequency, blank or unusually small pages, selector misses, schema failures, duplicate spikes, and extraction drift. Keep a small, permissioned fixture set for regression tests. Replay it whenever you change a selector, browser version, prompt, or model.
Hosted proxy and extraction infrastructure
Managed services can provide geographic routing, browser rendering, proxy rotation, or extraction APIs, but coverage, proxy quality, rate limits, retention, contractual permissions, and pricing change frequently. Evaluate each provider against your documented purpose and the target site’s terms. A vendor’s ability to fetch a page does not establish that you are allowed to collect or reuse its content.
Performance, reliability, and cost controls
- Use a bounded queue and per-domain concurrency; polite traffic is less likely to be blocked and is easier to budget.
- Cache successful responses and normalized results with a TTL. Reprocess only changed content.
- Batch LLM inputs by page section, not by unrelated pages, so one malformed page cannot contaminate a whole job.
- Set explicit timeouts for DNS, connection, download, browser navigation, and model calls. Retry only idempotent stages with exponential backoff and a cap.
- Record cost drivers separately: requests, browser minutes, proxy transfer, model input tokens, model output tokens, and human review.
- Use asynchronous workers for large crawls, but preserve per-domain rate limits and a durable job state so a restart cannot duplicate writes.
Troubleshooting guide
Robots.txt blocks requests
Confirm that ROBOTSTXT_OBEY is enabled and inspect the rule for your user agent. If the path is disallowed, stop or obtain permission; do not rotate identities to evade the rule.
The page is blank
Check whether content is client-rendered. Inspect network responses for a permitted data endpoint, or enable Playwright and wait for a stable selector. A timeout, blank page, or bot check should be recorded as a failed fetch, not passed to the model as an empty record.
Playwright times out
Increase the navigation timeout only after checking DNS, blocked resources, and the selector. Wait for a specific selector or network-idle condition rather than an arbitrary long sleep. Close contexts in a finally block to prevent browser leaks.
The model returns invalid or invented JSON
Reduce the input, enforce JSON/schema output, include “use null when absent,” and require an exact source quote. Reject any quote that cannot be found verbatim and send the item for re-fetch or human review.
Best Value
Records duplicate after a restart
Use an idempotency key such as a canonical URL plus source item ID or content hash, and commit through an upsert. Store the crawl job ID separately so retries remain observable.
HTTP 403 or 429 rates rise
Lower concurrency, honor Retry-After, increase spacing, verify your declared user agent, and revisit permission. Do not attempt CAPTCHA bypass or stealth techniques to defeat an explicit control.
Or skip the browser setup
ScreenshotNeo can provide a rendered page image or PDF with one GET request, which is useful when your extraction pipeline needs a stable visual artifact. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, retina scale, PDF page ranges, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account to try it.
Frequently Asked Questions
How should I version selectors and prompts?
Store selectors, prompt templates, and model identifiers in source control, and attach their versions to every record. Run the same permissioned fixture pages before deploying a change.
Can I scrape a page behind a login?
Only when you are authorized and the purpose, account access, and terms permit it. Keep credentials in a secret manager, minimize the fields collected, and never publish session cookies or private page content.
What should happen when a source disappears?
Mark records stale, preserve their last verified timestamp, stop retries after a bounded policy, and apply your documented retention or deletion rule rather than silently treating missing pages as empty data.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

