Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Data parsing turns a response—such as HTML, JSON, XML, plain text, or a downloaded file—into fields and records your program can validate and use. For a static page, a direct HTTP request plus Beautiful Soup or lxml is often enough. For multi-page crawls, Scrapy adds selectors, request scheduling, middleware, and exports. For JavaScript-rendered content, first look for the data request behind the page; use browser automation only when the content depends on browser execution or state.

What parsing does—and where scraping fits

Parsing is the step that interprets a response and selects meaningful values from it. An HTML document might become a record with a product name, price, and canonical URL; a JSON response might already contain those values as typed fields. Scraping is the wider process of obtaining responses, following pages, extracting data, and saving results. A reliable scraper therefore needs more than a parser: it needs request controls, validation, deduplication, error handling, and a plan for markup changes.

Start by identifying what the source actually returns. A browser-visible page may be backed by an HTML response, a JSON endpoint, XML, or a combination. Parsing the underlying permitted data response is usually simpler than extracting text from a rendered page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right input and tool

Input or need Practical starting point When to move up
Static HTML or XML, one or a few pages Python requests with Beautiful Soup or lxml Use Scrapy when link following, retries, crawl limits, or recurring runs become central.
JSON endpoint Request and parse JSON directly, preserving types and pagination metadata Use Scrapy if the endpoint is one part of a larger crawl or needs its request orchestration.
Many linked pages Scrapy spider and selectors Add pipelines, feed exports, caching, bounded concurrency, and a persistence layer as needed.
Content rendered by JavaScript Inspect network requests and reproduce the permitted data request Use Playwright or a Scrapy-Playwright integration if browser execution, cookies, or interaction is truly required.

Beautiful Soup offers a convenient Python interface for navigating parsed markup; lxml is another parser option and supports XPath. Scrapy combines a selector interface with crawl orchestration, downloader middleware, and feed exports. Scrapy selectors support CSS and XPath, and can work with HTML, XML, text, and JSON response types. These tools solve different layers of the problem: a parser extracts from a response, while a crawler manages a stream of requests and results.

#1 Best Overall
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

Parse a static HTML page with Python

For a small, permitted extraction, make a request, check its status, parse the returned HTML, and validate the fields you need. Install the dependencies with python -m pip install requests beautifulsoup4. This example uses semantic selectors; replace the example URL and selectors with ones that match the site and its rules.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

url = "https://example.com/catalog"
response = requests.get(
    url,
    headers={"User-Agent": "ExampleResearchBot/1.0 (contact: [email protected])"},
    timeout=(5, 30),
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
records = []
for card in soup.select("article.product"):
    title = card.select_one("h2")
    link = card.select_one("a[href]")
    price = card.select_one(".price")
    if not title or not link:
        continue

    records.append({
        "title": title.get_text(" ", strip=True),
        "url": urljoin(response.url, link["href"]),
        "price_text": price.get_text(" ", strip=True) if price else None,
        "source_url": response.url,
    })

print(records)

html.parser is Python’s built-in HTML parser, so the example does not require another parser package. Beautiful Soup also lets you choose a parser such as lxml; parser choice can affect how malformed markup is repaired. Install lxml separately with python -m pip install lxml if you choose it, then pass "lxml" to BeautifulSoup. Do not assume invalid source markup will be interpreted identically by every parser.

The example deliberately keeps price as source text. Convert it into a number only after deciding how to handle currency symbols, decimal separators, thousands separators, and missing values. Similarly, normalize whitespace and dates according to an explicit rule rather than silently changing data during extraction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose CSS selectors or XPath

CSS is often easier to read for common class, ID, and descendant selections. XPath is useful when selection depends on relationships such as finding a parent or ancestor from a text node, and it is common in XML navigation. Scrapy supports both, so team familiarity and the target markup can guide the choice.

Question CSS XPath
Typical use Find an element by class, ID, or nesting, such as article.product h2. Navigate ancestors, siblings, and more complex relationships.
Readability Often concise for routine page structure. Can be more expressive for relationship-based selection, but may be less approachable to a new reader.
Stability Both can break when based on unstable generated classes or layout details. Prefer stable semantic attributes and test selectors against representative pages.

For lxml, CSS selection is available through its HTML interfaces, while XPath can be called directly. In Scrapy, a selector can be written as response.css("article.product h2::text").get() or as an XPath expression. A selector that returns no value is not necessarily a parser failure: the response may differ by locale, page type, consent state, or markup version. Treat unexpectedly empty fields as a data-quality signal.

Parse JSON and other response types directly

If the page obtains its content from a JSON endpoint that you are permitted to access, parse that response instead of scraping the rendered text. Direct JSON preserves data types such as numbers, booleans, arrays, and null values; retaining pagination metadata helps you know whether a result set is complete.

import requests

api_url = "https://example.com/api/items"
response = requests.get(api_url, timeout=(5, 30))
response.raise_for_status()
payload = response.json()

items = payload.get("items", [])
next_page = payload.get("next")
for item in items:
    print({
        "id": item.get("id"),
        "name": item.get("name"),
        "source_page": api_url,
    })

Do not assume an endpoint is public or permitted merely because it appears in browser network traffic. Respect authentication boundaries and the site’s terms. For HTML containing embedded JSON, parse the documented or clearly identified data structure carefully; avoid brittle string slicing where a proper JSON parser applies. Scrapy’s dynamic-content guidance recommends reproducing the requests that contain the desired data when that approach is suitable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle JavaScript-rendered pages without defaulting to a browser

First inspect the page’s network activity in your browser’s developer tools. Look for a request whose response contains the records, pagination cursor, or other needed content. If an authorized request can be reproduced directly, this is generally less work and overhead than loading a full browser for every page.

Use browser automation when the required content genuinely depends on JavaScript execution, browser-only state, or an interaction that cannot reasonably be represented as a direct request. Playwright can load and interact with pages; Scrapy-Playwright can connect browser rendering to a Scrapy workflow. Browser automation costs more resources than a simple HTTP request and can complicate request handling. In particular, a directly managed browser may not pass through normal crawler middleware in the same way as Scrapy’s ordinary downloader requests, so plan how you will apply limits, headers, error handling, and observability.

When a screenshot is the actual deliverable

If the task is to inspect or archive what a rendered page looks like—not to extract structured fields—a screenshot or PDF can be a more suitable output than building a browser pipeline. ScreenshotNeo is a website screenshot API and MCP server; it captures images or PDFs and is not a replacement for a parser that must return structured records.

Or skip the browser setup

For a visual capture, one GET request can return a screenshot. See the ScreenshotNeo API documentation for request options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets can be removed; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server exposes take_screenshot, get_page_info, and capture_pdf to AI agents and MCP clients.
  • The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Sign up for ScreenshotNeo free and get 1,000 screenshots a month with no card.

Build a multi-page crawler with Scrapy

Choose Scrapy when following links, controlling request flow, exporting items, or running a recurring crawl matters as much as parsing individual pages. Scrapy spiders yield structured items and follow-up requests; middleware can apply downloader behavior. Its documentation describes feed exports to JSON, XML, or CSV and storage options that include FTP and Amazon S3.

Install Scrapy with python -m pip install scrapy, create a project using scrapy startproject catalog, and add a spider such as this to catalog/spiders/products.py. Replace the domain and selectors with the authorized target’s actual structure.

import scrapy

class ProductsSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/catalog"]

    def parse(self, response):
        for card in response.css("article.product"):
            href = card.css("a[href]::attr(href)").get()
            title = card.css("h2::text").get()
            price = card.css(".price::text").get()
            if href and title:
                yield {
                    "title": " ".join(title.split()),
                    "url": response.urljoin(href),
                    "price_text": " ".join(price.split()) if price else None,
                    "source_url": response.url,
                }

        next_page = response.css("a[rel=next]::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

Run it from the project directory and export records as JSON Lines with scrapy crawl products -O products.jsonl. A feed export is convenient for interchange or inspection; for ongoing production, validate records before writing them into a database or warehouse. Keep extraction logic separate from persistence where possible so failed or rejected records can be replayed without fetching everything again.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make extracted records reliable

Markup changes, missing fields, duplicate links, and inconsistent formatting can turn a successful HTTP response into bad data. Define a schema before the crawl and store provenance alongside values—for example, source URL, fetch time, and the identifier used for deduplication. Track missing or invalid fields rather than quietly emitting incomplete records as if they were sound.

  • Validate: Check required fields, types, allowed ranges, and relationships before accepting a record.
  • Normalize: Set explicit rules for whitespace, encoding, dates, numbers, and missing values. Preserve original text when normalization could discard meaning.
  • Deduplicate: Choose a stable key, such as a permitted source ID or canonical URL, and define whether later observations update or coexist with earlier ones.
  • Log failures: Record HTTP errors, parse exceptions, selector misses, and rejected items with enough provenance to diagnose and replay them.
  • Test selectors: Keep representative pages or fixtures and detect sudden drops in field coverage after a site changes.

Encoding and malformed markup deserve deliberate handling. Beautiful Soup documents that parser choice affects invalid markup behavior; Scrapy documents encoding support and parser alternatives. Do not treat every decode or parse problem as a selector bug.

Scale the crawl without losing control

Scaling is not simply increasing concurrency. A faster request loop can increase load on the target, amplify a selector mistake, or create a larger pile of invalid records. Grow from a measured small crawl and add controls one at a time.

  1. Define the schema and provenance. Decide what constitutes one record and how to identify its source before collecting many pages.
  2. Measure a small run. Begin with direct requests and selectors. Track response outcomes, empty fields, and processing cost to find the actual bottleneck.
  3. Control request volume. Add bounded concurrency, rate limits, pagination safeguards, and retry backoff. Avoid unbounded link following.
  4. Reduce avoidable work. Cache responses where appropriate, deduplicate URLs and records, and avoid requesting data that is already available from an allowed endpoint.
  5. Separate extraction and storage. Use item pipelines or a queue so persistence failures do not force a full recrawl and individual records can be replayed.
  6. Choose an output path. Use JSONL, CSV, or XML for interchange, or write validated records to a database or warehouse. Scrapy feed exports include storage options such as FTP and Amazon S3.
  7. Schedule and monitor. For recurring runs, monitor selector coverage, empty required fields, HTTP errors, and relevant robots.txt changes; review failures before accepting a new batch.

Scrapy’s official overview documents crawl-depth restriction, cookies and sessions, compression, caching, authentication, user-agent controls, robots.txt handling, feed exports, storage backends, and extensibility. Hosted Scrapy API documentation describes synchronous and asynchronous runs, polling, dataset item retrieval, and schedules with JSON, CSV, and JSONL exports. The right deployment depends on whether you want to operate the crawler yourself or use a hosted run workflow; the extraction and compliance requirements still need to be defined either way.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Respect site rules and access boundaries

Compliance belongs in the design, not as a cleanup step after a crawl. Scrapy documents the ROBOTSTXT_OBEY setting and how its parser handles wildcard and path-specific rules. Configure robots.txt handling where applicable to the site’s rules and your legal context, and read the site’s terms and applicable law rather than treating robots.txt as the only permission test.

  • Do not bypass authentication or technical access controls.
  • Rate-limit requests and avoid placing unnecessary load on a service.
  • Minimize personal-data collection; collect sensitive personal information only when you have a documented lawful basis.
  • Recheck relevant site rules when scheduling repeated runs, since they can change.

Troubleshooting common parsing failures

Symptom Likely cause What to check or change
A selector returns no elements The response is a different page, content is rendered later, or markup changed. Inspect the saved response and status code first. Check the selector against current markup, then inspect network requests for the data source before adding a browser.
The browser shows data but the request response does not JavaScript fetches or constructs the content after the initial HTML response. Inspect network traffic for an allowed JSON or data request. Use browser execution only if the required content depends on browser state or interaction.
Characters are garbled or parsing breaks on malformed HTML Encoding assumptions or parser recovery differ from the source. Inspect response encoding and choose the parser deliberately. Test alternatives on representative malformed pages, then normalize text explicitly.
Some records have missing or oddly formatted values Optional fields, locale variants, or changed markup are being treated as uniform. Make fields nullable where appropriate, validate types, normalize with explicit locale rules, and log field coverage by page type.
Records repeat or pages loop Pagination links are duplicated, canonical identities are absent, or crawl following is too broad. Deduplicate requests and records with stable keys; constrain allowed domains and pagination; add a crawl-depth or equivalent boundary.
A crawl is slow or triggers errors Request volume, timeouts, retries, or browser rendering may be excessive. Measure response and failure patterns, bound concurrency, use backoff and suitable timeouts, cache where appropriate, and prefer direct data requests over browser rendering when valid.

Plan cost, performance, and maintenance around the workload

There is no useful universal speed number for parsing or crawling: response size, target behavior, browser use, network conditions, and crawl limits all affect the result. Direct requests generally avoid the extra work of browser rendering, while browser automation is justified when execution or interaction is necessary. Measure your own failure rate, response volume, browser overhead, and downstream storage cost on a representative sample rather than extrapolating from a tiny test.

Retries improve resilience to transient failures but can worsen load or repeat a persistent error if they are unbounded. Use timeouts, bounded retries with backoff, and clear distinction between a failed fetch and a valid response with no matching fields. Cache only when the freshness needs and permission context allow it. For scheduled extraction, selector maintenance and review of changing site rules are recurring operational costs, not one-time setup tasks.

Frequently asked questions

Should I store the raw response as well as parsed records?

When permitted and practical, retaining a limited raw-response sample or fixture can make selector regressions easier to diagnose. Apply the same retention, privacy, and access controls to raw content that apply to the extracted data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I know when a recurring crawl should be paused?

Pause or quarantine a run when required-field coverage drops unexpectedly, errors spike, or the target’s rules or access behavior change. Investigate those signals before publishing a partial or potentially invalid dataset.

Frequently Asked Questions

Should I store the raw response as well as parsed records?

When permitted and practical, retaining a limited raw-response sample or fixture can make selector regressions easier to diagnose. Apply the same retention, privacy, and access controls to raw content that apply to the extracted data.

How do I know when a recurring crawl should be paused?

Pause or quarantine a run when required-field coverage drops unexpectedly, errors spike, or the target’s rules or access behavior change. Investigate those signals before publishing a partial or potentially invalid dataset.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.