Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a basic no-code scraper in n8n by fetching a page with an HTTP Request node, extracting fields with HTML Extract and sending the results to a destination such as Google Sheets. This works when the information is present in the HTML returned by the website. If JavaScript must run in a browser before the content appears, use a browser-rendering service such as Browserless instead of expecting a plain HTTP request to render the page.

What this workflow does—and where it stops

An n8n scraper is a visual data pipeline: a trigger starts a run, an HTTP Request retrieves a page, HTML Extract reads selected parts of its markup, and later nodes clean and store the resulting fields. You configure the workflow through node settings rather than writing a scraper program.

The central distinction is how the target serves its content. A standard HTTP request retrieves the server-delivered response; it does not behave like a full browser running the page’s JavaScript. If the target’s initial HTML contains the title, price, or links you need, CSS selectors can extract them. If those values are inserted only after browser JavaScript runs, the response may not contain them, and a browser-rendering layer is needed.

n8n describes HTTP Request as “one of the most versatile nodes in n8n.” It is a general-purpose requester for REST APIs as well as web pages, with configurable methods, URLs, and authentication. See n8n’s HTTP Request node documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build the basic scraper workflow

1. Choose how the run starts

Create a workflow and add either a Manual Trigger for testing and on-demand runs or a Schedule Trigger for recurring collection. Start manually while you are checking the URL and selectors. Add a schedule only after the target, output fields, and request pace are settled.

2. Fetch the page with HTTP Request

Add an HTTP Request node after the trigger. Set Method to GET, enter the page URL, and configure the response to be returned as text or a string so the next node receives the HTML. The exact label for the response-format setting can vary with n8n version; select the option that returns the response body as text rather than attempting to interpret it as JSON.

Use the target page’s direct URL. If the site requires a documented API, authentication, or query parameters, configure those deliberately rather than assuming the public page will expose the same data. The node supports configurable authentication; follow the site’s permitted access method and keep credentials in n8n credentials rather than embedding secrets in text fields.

3. Extract fields with HTML Extract

Add an HTML Extract node and point its HTML input at the property containing the response text from HTTP Request. The property name depends on how the preceding node is configured, so inspect a successful HTTP Request execution and select the actual response-body field rather than guessing it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Add one extraction value for each field. Supply a CSS selector based on the target page’s actual DOM and choose whether the result should be text or an attribute. For a heading, a selector such as h2 can return its text. To extract a link destination, select an anchor and return its href attribute; to get its visible label, return text. Enable array output when the selector matches repeated elements, such as multiple article cards or product rows. The n8n tutorial demonstrates extracting h2 content and then nested anchor text and href values: n8n’s website-scraping tutorial.

For nested structures, scope selectors to the repeated item when possible. For example, extract a card’s title and link relative to each card rather than collecting all titles and all links independently; matching lists in separate arrays can become misaligned if the page has missing fields. Test selectors on more than one representative page, including cases where an expected field is absent.

4. Clean and map the extracted values

Place a cleanup or mapping step between extraction and storage. Normalize whitespace, trim text, convert a price string into a consistent numeric representation only after accounting for the displayed currency and separators, and remove duplicates using a stable key such as a canonical URL or item identifier. Keep the original source URL alongside the normalized fields so you can trace a result back to its page.

Do not treat a missing selector match as a valid empty record without deciding what that means. Depending on the target, a missing value can indicate an optional field, a markup change, a blocked response, or a failed extraction. Preserve enough context to distinguish these cases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Send records to a destination

Connect the cleaned items to Google Sheets, Airtable, a database, or an alerting channel. Choose a destination operation that matches your intended behavior: append for a collection log, or update/upsert where the same record should be refreshed rather than duplicated. n8n’s HTML Extract examples cover uses such as multi-page storage, price tracking, article extraction, and job or product monitoring; see HTML Extract integration examples.

For an auditable run, store the source URL and retrieval time with each record. That makes it easier to identify stale results, diagnose a failed page, and compare changes without confusing a new capture with an earlier one.

Pagination, scheduling, and request reliability

A workflow that extracts one page is not automatically a reliable crawler. Decide how pagination is represented—page-number query parameters, a next-page link, or another pattern—and make each additional request an intentional step. Set a sensible request pace, respect the site’s rate limits, and avoid uncontrolled parallel requests. For each request, handle non-2xx responses explicitly rather than passing an error page into HTML Extract as if it were ordinary content.

  • Record the requested URL and retrieval time with the output.
  • Check that the HTTP response succeeded before extracting fields.
  • Define what to do when a page is empty, a selector stops matching, or a later page cannot be fetched.
  • Throttle requests and keep pagination bounded to the pages you are authorized to collect.
  • Use a representative sample of pages to validate selectors before relying on scheduled runs.

Selectors are coupled to the site’s markup. A redesign can change class names, nesting, or the location of a field without changing the visible page much. Treat extraction as something to monitor and maintain, not a one-time setup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a page needs browser rendering

If HTTP Request retrieves HTML but the desired content is absent because the site fills it in after JavaScript executes, HTML Extract has nothing to select from that response. Confirm this by inspecting the returned body and comparing it with the page content you see in a browser. If the content is missing from the body, changing CSS selectors will not fix the underlying rendering gap.

Use a browser automation or rendering service for those pages. The official Browserless integration for n8n advertises crawling pages and executing JavaScript/Puppeteer server-side; see n8n’s Browserless integration page. That introduces another service and configuration to operate, but can retrieve content that depends on browser execution. It does not make selectors immune to markup changes, remove the need to respect site rules, or guarantee every target will load successfully.

For pages that do not need browser execution, the simpler HTTP Request plus HTML Extract path has fewer moving parts. Browser rendering is specifically for the gap between server-delivered HTML and content produced in a browser; it is not a universal replacement for a well-targeted HTTP request.

Choose an n8n deployment that fits the workflow

n8n documents Cloud, npm, and self-hosted deployment options. The right choice depends on how much infrastructure you want to own, where the target is reachable from, and how you plan to manage credentials and any browser service. See n8n’s deployment options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option Setup and infrastructure Network and credential considerations Browser-rendering consideration
n8n Cloud Hosted n8n; less infrastructure for you to manage. Check that the hosted instance can reach the target and configure credentials within the service. A separate browser-rendering integration or service may still be required for JavaScript-dependent pages.
npm Run n8n through npm in an environment you manage. You control the runtime environment and its network access; handle credentials and operational upkeep accordingly. Plan separately for browser automation when the target requires a rendered page.
Self-hosted You own deployment and supporting infrastructure. You control network placement and credential handling, with the corresponding responsibility for configuration and maintenance. Account for the browser service and its connectivity as an additional component if needed.

The deployment choice does not change what a plain HTTP response contains. If browser execution is necessary, arrange for the rendering service to be available to the workflow regardless of where n8n runs.

Scrape responsibly

Before collecting data, check the target site’s robots.txt and terms of service. n8n’s scraping tutorial recommends looking for robots.txt when no other permission guidance is available. Prefer an official API or RSS feed when one exists, observe authentication requirements and rate limits, and do not scrape private or access-controlled content without authorization.

A page being publicly viewable does not by itself establish permission to collect it at any rate or for any purpose. Keep requests proportionate, avoid bypassing access controls, and use only data you are entitled to process.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

HTML Extract returns no values

First inspect the HTTP Request output. Confirm it contains the expected page HTML and that HTML Extract points to that exact response property. Then inspect the page’s DOM and adjust the selector to match the actual markup. If the response is an error or challenge page, fixing the selector is not the answer.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The browser shows data that the workflow cannot find

The visible page may be populated after JavaScript runs. Compare the HTTP response body with the rendered page. If the fields are absent from the response, use a browser-rendering layer such as Browserless rather than repeatedly changing selectors.

Text appears but links are missing

Check that the extraction is returning the anchor’s href attribute, not only its visible text. Verify the selector reaches the anchor itself and test whether the link is relative or absolute before storing or using it.

Repeated records or mismatched fields

Enable array output for repeated matches and extract related values within the same repeated container where possible. Add a cleanup step that deduplicates on a stable identifier. Avoid independently collecting separate arrays and assuming their positions correspond when some cards may lack a field.

Requests fail or return unexpected pages

Check the status and response body before extraction. Confirm the URL, any required authentication, and whether the target permits automated requests. Handle non-2xx results as failures, throttle the workflow, and log the URL and retrieval time to make intermittent problems diagnosable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
PowerShell for Sysadmins: Workflow Automation Made Easy
  • Book - powershell for sysadmins: workflow automation made easy
  • Language: english
  • Binding: paperback

A scheduled workflow degrades over time

Re-test selectors against current pages and keep an eye on missing fields or sudden empty outputs. Site markup changes are a normal maintenance risk; make the workflow flag extraction anomalies instead of silently writing incomplete rows.

Or skip the browser setup

If your goal is a clean screenshot or a rendered page capture rather than a structured scrape, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. Its API is not a substitute for extracting structured fields with HTML Extract; it is useful when a capture of the rendered page is the output you need.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month with no card.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can n8n scrape a website without code?

Yes. The basic workflow uses HTTP Request to retrieve HTML, HTML Extract to select fields, and later nodes to clean and store them. It works when the target’s response contains the content you need.

Can HTML Extract read data loaded by JavaScript?

Not from a plain HTTP response if that data is absent from the returned HTML. Use a browser-rendering service when the page must execute JavaScript before the content appears.

What should I do if the website offers an API?

Prefer the official API or RSS feed when available, and follow its authentication and rate-limit requirements.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.