Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a website into JSON with CSS selectors, map each output field to a selector and an extraction rule: text, an attribute such as href, or a typed value. For repeated data, select a card or row as a container and define child fields; for a JavaScript-rendered page, wait until the relevant content appears before extracting it. The key is to build and validate the schema against the DOM the scraper actually receives—not just the page source you expect.

How selector-based scraping becomes JSON

CSS selectors identify elements in a page’s DOM. A JSON extraction schema connects those elements to named output fields. For example, a schema might map title to the text of an h1, and url to the href attribute of a link matching a.next. Microlink describes this field-to-rule model in its Scraping API documentation; Ujeebu documents a similar structured extraction pattern in its API documentation.

Selectors do not themselves create JSON. They locate nodes; your scraper or extraction API decides what to read from each node, how to type it, and how to arrange the results. The W3C describes CSS selectors as a way to describe a path to an element in a web page (Selectors Level 4).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Map one field to one value

A field rule typically contains a selector and an extraction instruction. Text is appropriate for a heading or price label; an attribute is appropriate for a link destination, image URL, or data attribute. If the tool supports type conversion, specify it when a downstream system expects a number, URL, or other normalized value.

Use nested objects for related fields

For a product page, an object might contain name, price, and product_url. The fields may come from different elements, but grouping them in one object gives consumers a coherent record. With schema-based hosted extractors, nested rules express this relationship; check the service’s own syntax and type behavior before adopting a schema.

Use repeated containers for arrays

For a list of products, first select the repeated card or row, then define child rules relative to each selected container. Each container becomes one object in an array. This is more robust than independently collecting all titles and all prices and assuming their positions always line up.

Build and validate a useful extraction schema

  1. Fetch and inspect the target. Determine whether you are looking at the original HTML, a rendered DOM, or a browser view after scripts have run. The scraper must operate on the same representation you inspected.
  2. Start with a small schema. Add the fields the consuming application needs, such as a title and canonical link. Microlink’s documentation shows rules with selector, attribute, and type controls; its hosted API returns the requested fields as JSON (Microlink extraction overview).
  3. Identify repeated records. Find a stable common parent for each row or card, then scope child selectors inside it. Add optional fields only after the basic record works.
  4. Set missing-value behavior. Decide whether an absent field should be null, omitted, treated as an error, or supplied by a fallback. Microlink documents null for missing or type-invalid fields; Scrapy’s selector API returns None when a selector has no match.
  5. Validate output types and sample records. Check that prices are numbers if expected, links resolve as intended, and arrays contain the right number of records. Inspect several pages and template variants, not only one successful example.
  6. Export only what the next system needs. Keeping the schema focused makes the output easier to validate and less likely to break when irrelevant parts of a page change.

A selector that returns a value on one page is not proof the extraction is reliable. Monitor empty and null fields: a site redesign may leave the request itself successful while changing the element your selector depended on.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write CSS selectors that survive page changes

Prefer stable identifiers that express meaning: semantic classes, IDs, data attributes, or structured markup. Avoid selectors that depend on a long chain of nested elements or on a child’s position, such as “the third div inside the second section.” Small layout changes can invalidate positional paths without changing the content you want.

  • Scope repeated data to its container. Select a product card first, then find its title and price inside that card.
  • Use specific but not brittle selectors. A meaningful class plus an element type is often more resilient than a full ancestry path.
  • Plan for template variants. If your extraction service supports fallback selectors, provide alternatives only for known variations and test each one.
  • Test missing and malformed fields. Verify how the tool handles absent matches and invalid types so downstream code does not confuse “not present” with a valid empty value.
  • Track null rates. A sudden increase can reveal a changed page template or a selector that stopped matching.

Handle JavaScript-rendered pages

A static page can often be parsed from the HTML returned by a normal HTTP fetch. A client-rendered site may initially return only an application shell; the content appears after JavaScript executes. In that case, selectors aimed at the final content will not match until the page has rendered.

Choose a readiness condition that reflects the target page rather than waiting blindly. Microlink says its extraction rules run on a rendered page when needed (Microlink documentation). Cloudflare Browser Run’s /scrape endpoint accepts a URL or HTML and supports browser navigation waits such as networkidle0 or networkidle2, as well as waitForSelector for a known element (Cloudflare Browser Run scrape endpoint). Browserless says its /scrape request runs selectors against the fully rendered DOM (Browserless scrape API).

Wait for a meaningful signal

When possible, wait for a selector that marks the content you need, such as a product title or results container. Network-idle waits can be useful, but pages with analytics, polling, or persistent connections may never become fully idle. A fixed delay is simple but can be either wasteful or too short under variable network conditions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check what was rendered

If a selector produces no result, inspect the rendered DOM used by the service and verify that the page actually loaded the relevant content. Also check whether the data is inside an iframe, requires a user action, or is withheld until authentication or another interaction. Do not assume that a browser-rendered page and the initial response HTML are equivalent.

Use Scrapy when you want local control

Scrapy is a Python framework for crawling and extracting data. Its selector API supports CSS and XPath, translates CSS queries to XPath internally, and provides ::text and ::attr(name) extraction shortcuts. The selector list’s .get() returns the first match and .getall() returns all matches; no match yields None (Scrapy selectors). Scrapy also supports feed exports, including JSON (Scrapy feed exports).

Here is a minimal spider that extracts a page title and all links into a JSON feed. Create spider.py in a Scrapy project and run it from that project directory:

import scrapy

class PageSpider(scrapy.Spider):
    name = "page"
    start_urls = ["https://example.com/"]

    def parse(self, response):
        yield {
            "title": response.css("h1::text").get(),
            "links": response.css("a::attr(href)").getall(),
        }
scrapy runspider spider.py -O output.json

Replace the example URL and selectors with the target site’s values. -O writes a fresh output file; use Scrapy’s feed export options for your intended destination and format. This example parses the response Scrapy receives; it does not demonstrate browser rendering. For JavaScript-dependent pages, choose a rendering approach or a hosted extraction service that explicitly supports the rendered DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose hosted extraction or a crawler you operate

Hosted APIs can combine fetching, rendering, and selector-based extraction in a request, reducing the browser and parser infrastructure you maintain. Scrapy gives you local control over crawl logic, pipelines, retries, and execution environment. Compare services on JavaScript rendering, nested and repeated schema support, type conversion and null behavior, wait controls, authentication or session support, output formats, quotas and cost, and operational ownership. The cited product documentation establishes different capabilities, not a universal ranking for every workload.

Or skip the browser setup

If the task is to capture a page as an image or PDF rather than extract structured fields, ScreenshotNeo is a website screenshot API and MCP server. Its one-request API returns a PNG, JPEG, WebP, or PDF; screenshots are not a substitute for a selector-to-JSON extraction schema.

For example, this cURL call saves a WebP screenshot of the target URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request parameters and formats. Before capture it accepts the cookie or consent banner like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot missing or incorrect fields

The selector returns no match

First verify that the expected element exists in the DOM received by the scraper. If it appears only after JavaScript runs, add a rendering and readiness step. If it exists, check spelling, casing, scoping, and whether the selector is being applied to the whole document or a parent container. In Scrapy, a missing match produces None from .get().

The request succeeds but JSON values are null

The site may have changed its markup, the field may be optional, or the extraction rule may be reading the wrong attribute or node. Check the rendered element and test the field’s type conversion. Where supported, use a fallback selector for a known alternate template and monitor how often each field is absent.

Text is present but the attribute is empty

Make sure the selector targets the element that owns the attribute. For a link, extract href from the anchor; for an image, inspect whether the page uses src, a lazy-loading attribute, or a CSS background. The appropriate attribute depends on the actual DOM.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Repeated values are misaligned

Do not collect titles and prices as separate global lists and pair by index. Select each repeated card or row and extract its child fields within that container so each array object is assembled from one record.

A JavaScript page times out waiting for network idle

Persistent network activity can prevent an idle condition. If the browser tool supports it, wait for a stable content selector instead. A delay may be a fallback, but test it against slower loads and avoid treating a page as ready merely because a timer elapsed.

Reliability, performance, and responsible use

Keep each extraction focused on fields that are needed, and avoid rendering a browser when the response HTML already contains the data. For pages that require JavaScript, use a targeted wait condition and avoid unnecessary delays. When operating your own crawler, retries, concurrency, caching, and crawl scope are design choices that affect load and resource use; Scrapy’s framework gives control over these concerns, while hosted services handle some of the fetch-and-render work for you.

Check the target site’s terms, robots directives, and applicable law before collecting data. The selector and browser documentation explains technical mechanics; it does not establish permission to access or reuse a particular site’s content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What does a CSS selector extract from a page?

A CSS selector locates DOM elements; the scraper then reads text, an attribute, or a converted value from those elements.

Can CSS selectors scrape data loaded by JavaScript?

Yes, if the extraction process runs selectors against the rendered DOM after the required content has loaded. A parser that sees only the initial HTML may not find it.

When should I use Scrapy instead of a hosted extraction API?

Choose Scrapy when you need to own crawl logic, pipelines, retries, or execution locally; choose a hosted service when you want fetching, rendering, and extraction combined in a managed request.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.