Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To extract a specific field from a website, identify where the value lives, choose a selector or pattern that targets it, map the result to a named output field, and validate it on representative pages. Use CSS or XPath for HTML elements, regex for pattern-based values, and a rendered-page option when JavaScript adds the content after the initial HTML loads.
What a custom extraction rule does
A custom rule tells a crawler or scraping endpoint what to read and where to put the result. For example, a rule might target an article heading and store its text in a field called article_title. The field name is your output label; the selector or pattern determines what value is collected.
Before writing a rule, define the value you need and how the output should be represented: text, an attribute such as a link destination, HTML, or one or more values. These choices affect both the selector and the extractor settings.
Choose a selector or pattern
CSS selectors and XPath for HTML elements
CSS selectors and XPath identify elements in a page’s HTML. Use them when the field is contained in a heading, price element, link, or other identifiable node. A selector that is too broad may match navigation or repeated content instead of the desired value, so inspect the markup and narrow the target to a distinctive element or attribute.
#1 Best Overall
Regex for values defined by a pattern
Regular expressions are useful when the value is better described as a pattern than as an HTML element—for example, date components embedded in a URL. Capture groups can isolate the part you want rather than returning the entire match. Elastic’s Open Web Crawler extraction rules document regex for URL-derived values, including examples that extract year, month, and day components.
Build and validate a rule step by step
- Choose a representative page. Pick a page that actually contains the field and note the expected value and output field name, such as
authororprice. - Inspect the page structure. Use browser developer tools or a crawler’s selector helper to locate the relevant node, attribute, or URL value. Screaming Frog’s custom extraction guide describes using its inbuilt browser to select an element and generate suggested expressions.
- Set the extraction mode. Choose CSS or XPath for an HTML element, or regex for a pattern-based value. Configure the result as text, an attribute, inner HTML, or another supported form appropriate to the field. Screaming Frog’s SEO Spider configuration guide documents extractor modes and output forms.
- Map the result to a named field. In a rules-based crawler, give the output a stable name so downstream processing can distinguish it from other extracted values.
- Test several representative URLs. Compare extracted values with what the pages actually contain. Include pages with different layouts or content states, not only the page used to create the rule.
- Decide what to do with multiple matches. Keep all matches or configure a join behavior if the tool supports it. Elastic documents configurable joining for multiple extracted values.
- Check access permission. Review the target site’s terms and any applicable rules before collecting data. A technically accessible page is not automatically authorized for every use.
When the initial HTML does not contain the field
Some pages insert content with JavaScript after the initial response. A selector may therefore return nothing even though the value appears in a browser. Check whether the field exists in the original HTML; if it does not, use a rendering-enabled crawler or endpoint and compare its extracted result with the browser-visible content.
Screaming Frog documents switching to JavaScript rendering for client-side-only data. Cloudflare’s Browser Rendering /scrape documentation supports selecting page elements with CSS selectors and cautions that a page can count as loaded before JavaScript has finished rendering. When using a rendering path, allow for that timing rather than assuming that the first load event means the field is ready.
Choose an approach for the job
| Approach | Documented capabilities | Useful when |
|---|---|---|
| ScreenshotNeo | Website screenshot API and MCP server; a GET request can return a PNG, JPEG, WebP, or PDF. It captures screenshots rather than extracting arbitrary named fields, so it is not a substitute for a selector-based scraper. | You need a visual capture of a page or PDF as part of a workflow, including one an AI agent can request. |
Cloudflare Browser Rendering /scrape |
Accepts a URL or HTML with CSS selectors for page elements; its documentation includes headings, links, prices, and repeated content. | You want a hosted endpoint to retrieve selected page elements. |
| Screaming Frog SEO Spider | Site crawling with custom extraction using XPath, CSS Path, or regex; supports static or rendered HTML and visual extraction assistance. The custom extraction feature requires a licence. | You want a desktop crawler workflow and configurable extraction across crawled pages. |
| Elastic Open Web Crawler | Rulesets can be scoped to domain entries and URL filters; extraction rules can use CSS or XPath for HTML and regex for URL values, with results assigned to named fields and multiple values joined. | You want a configuration-driven crawler with URL scoping and named output fields. |
These tools document different workflows and capabilities; the cited documentation does not establish which is more accurate, faster, or cheaper overall. For selected-element extraction, Cloudflare documents a hosted endpoint, while Screaming Frog and Elastic describe crawler-based rule configuration.
Rank #3
Troubleshoot missing, wrong, or repeated values
The selected content is missing
Check whether the value is present in the initial HTML or appears only after client-side rendering. If it is inserted by JavaScript, try a rendering-enabled path and account for the page’s rendering timing.
The rule returns the wrong element
Inspect the markup and make the CSS or XPath expression more specific, such as by targeting a distinctive class, parent-child relationship, or attribute. Validate the revised expression on other representative pages because a selector helper’s suggestion is not proof that it works across the site.
The rule works on one URL but not another
Compare the page structures and check any URL filters that determine which pages receive the rule. Elastic documents filters such as begins, ends, contains, and regex; an incorrect scope can leave intended pages out or apply a rule to unintended ones.
Several values are returned
Decide whether each match is meaningful. If the output should be one value, narrow the selector; if all matches matter, retain them or use the crawler’s documented multi-value join behavior.
Best Value
A regex captures too much
Use capture groups to return only the desired substring, then test the pattern against URLs that vary in format. A pattern that matches one example may not cover every date or path structure used by the site.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
For a visual record of a page rather than named-field extraction, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. This does not replace CSS, XPath, or regex extraction rules; it provides a clean page capture or PDF that can support a separate review or processing step.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Sign up for free and try ScreenshotNeo.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteProduct prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

