Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use rvest to turn permitted HTML into a tidy R data frame: load the document with read_html(), identify the repeated record on the page, select its nodes with CSS selectors or XPath, extract text and attributes, and validate the resulting rows. If the required content is created only after JavaScript runs, switch to read_html_live() or the site’s documented API.

The mental model: document, records, fields

A web page is a hierarchy of HTML elements. A product card, article, search result, or table row is usually a repeated unit. Treat each unit as one record and each value inside it as a field. In a data frame, that means one row per repeated unit and one column per field.

  • Document: the HTML returned by the server or browser.
  • Record selector: a CSS selector such as article.product-card that matches every unit.
  • Field selector: a selector inside one record, such as h2 or a.
  • Value: text, an attribute such as href, or another property you explicitly extract.

Selectors are target-page-specific. Inspect a current sample page, confirm which elements repeat, and record the extraction date when you maintain a real project. A redesign can leave your R code running while silently returning empty or shifted columns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install rvest and inspect a page

install.packages(c("rvest", "dplyr", "tibble", "xml2"))

library(rvest)
library(dplyr)
library(tibble)

page <- read_html("https://example.org/sample-page")
page

example.org/sample-page is a pattern, not a guaranteed data source. Replace it with a page you are allowed to collect and verify its current markup and rules. read_html() parses the HTML returned by a normal request; xml2 supplies the static parser used in this workflow.

Find candidate nodes

# How many article elements are present?
length(html_elements(page, "article"))

# Inspect the first matching element
first_record <- html_element(page, "article")
first_record

# Extract all headings on the page
page |> html_elements("h1, h2, h3") |> html_text2()

html_elements() returns all matches. html_element() returns the first match for each context. The distinction matters: use the plural form for repeated records and the singular form for one field inside each record.

Example project: repeated HTML records to a data frame

The following project assumes the page contains repeated <article> elements, each with an <h2> title and a link. Confirm those selectors in your target page before running it.

library(rvest)
library(dplyr)
library(tibble)

url <- "https://example.org/sample-page"
page <- read_html(url)
records <- html_elements(page, "article")

results <- tibble(
  title = records |> html_element("h2") |> html_text2(),
  link = records |> html_element("a") |> html_attr("href")
)

print(results, n = 10)
write.csv(results, "scraped-results.csv", row.names = FALSE)

html_text2() returns readable text while handling descendant text and whitespace more usefully than a raw text-node extraction. html_attr("href") reads an attribute. To collect another field, add a column using a selector that exists inside every record:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
results <- tibble(
  title = records |> html_element("h2") |> html_text2(),
  link = records |> html_element("a") |> html_attr("href"),
  summary = records |> html_element("p.summary") |> html_text2()
)

Make relative links usable

Sites often store /products/42 instead of a complete URL. Resolve those links against the page URL rather than concatenating strings manually.

results <- results |>
  mutate(link = url_absolute(link, url))

If the target uses a <base> element or an unusual URL scheme, inspect a few outputs before saving them as canonical links.

Handle missing elements deliberately

A missing child element can produce NA or an empty result depending on the selector and context. Keep the row, then inspect missingness instead of dropping records accidentally.

results <- tibble(
  title = records |> html_element("h2") |> html_text2(),
  link = records |> html_element("a") |> html_attr("href"),
  summary = records |> html_element("p.summary") |> html_text2()
) |>
  mutate(
    title = na_if(trimws(title), ""),
    summary = na_if(trimws(summary), "")
  )

colSums(is.na(results))
stopifnot(nrow(results) > 0)

For stricter pipelines, assert expected columns, a sensible row count, and key uniqueness. Save a small sample such as head(results, 5) so a selector failure is visible in code review or scheduled-job logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CSS selectors and XPath you will use most

Goal CSS example What it matches
Element type article Every article element
Class .product-card Any element with that class
Element and class article.product-card Article elements with that class
Descendant article h2 Headings inside an article
Direct child article > a Links directly inside an article
Attribute a[data-id] Links carrying a data-id attribute

Use XPath when the relationship is easier to express that way:

# All links whose visible text contains “Next”
page |> html_elements(xpath = "//a[contains(normalize-space(.), 'Next')]")

Do not select by a generated class name merely because it is convenient. Prefer stable semantic elements, data attributes, or a documented API when available.

Static HTML or JavaScript-rendered content?

A browser may show text that is absent from the HTML returned by a normal request. Check the parsed document first. Search the HTML for a distinctive value, inspect the relevant node, or compare the response source with what the browser displays. Do not assume that a visible element is present in static HTML.

Question Static path Live-browser path
Is the desired value in returned HTML? Use read_html() and selectors. Not needed.
Content appears only after scripts run? Static parsing returns no record or incomplete fields. Use read_html_live() after assessing browser setup and site rules.
Setup and dependencies Usually faster and has fewer external dependencies. Requires a live browser approach and is heavier to operate.
library(rvest)

live_page <- read_html_live("https://example.org/page-with-rendered-data")
rendered_records <- live_page |> html_elements("article")
rendered_titles <- rendered_records |> html_element("h2") |> html_text2()

Use the live method only when it supplies data you cannot obtain from static HTML or an official endpoint. If a site exposes a documented API, it is often a more stable and permission-friendly interface than parsing presentation markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pagination, rate limits, and responsible collection

For one page, a direct read_html() call may be enough. For many pages, use a controlled loop, pause between requests, cache results, and stop when the site signals an error. The rvest maintainers recommend pairing rvest with polite for multi-page scraping because it supports robots.txt awareness and helps avoid hitting a site too aggressively.

install.packages("polite")
library(polite)
library(rvest)

session <- bow("https://example.org")
page <- scrape(session)
html <- read_html(page)

# Follow a discovered next-page link only after checking it exists
next_href <- html |> html_element("a.next") |> html_attr("href")
if (!is.na(next_href)) {
  next_url <- url_absolute(next_href, "https://example.org")
  Sys.sleep(1)
  next_html <- read_html(next_url)
}

Review robots.txt and the site’s terms separately; neither document alone answers every legal or policy question. Consider the site’s API when one is provided. Keep request volume proportionate, identify your application where appropriate, and avoid collecting personal or restricted data without a clear basis.

Production checks and common failure modes

Zero rows

Cause: the selector does not match the current markup, the content is JavaScript-generated, or the request received a block page. Fix: print the page, inspect a distinctive node, verify the response URL and status, then test whether the value exists in static HTML. Use a live browser only if necessary.

Rows exist but columns are empty

Cause: the child selector is wrong or optional on some records. Fix: inspect one record with html_element(), test the selector against several records, and retain explicit NA values for genuinely missing fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every row contains the same value

Cause: a page-level selector was applied to the whole document rather than to each record. Fix: first create records, then run each field extraction on that node set.

Relative links or malformed URLs

Cause: the site uses relative paths, fragments, or protocol-relative links. Fix: use url_absolute(), inspect unusual values, and preserve the original attribute in a separate column if auditability matters.

Requests are slow or blocked

Cause: rate limits, bot defenses, large pages, or unnecessary browser rendering. Fix: prefer static parsing, reduce concurrency, add delays, cache completed pages, follow the site’s rules, and use an official API where available. A live browser is not a bypass for access controls.

Data changes shape after a redesign

Cause: selectors depended on fragile classes or layout nesting. Fix: choose stable attributes, add row-count and missingness assertions, keep a fixture or saved sample for tests, and record when selectors were last verified.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your goal is a clean image or PDF of a page rather than parsing its fields, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

See the parameter reference in the ScreenshotNeo documentation. Replace the example URL with the page you are permitted to capture.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper settings and page ranges, HTML/CSS input, custom JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. Every feature is on every plan: 1,000 shots monthly free with no card, then Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free.

Sign up for the free 1,000-screenshot monthly plan—no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further learning

The free official rvest vignette and project documentation are the best starting points for selectors and extraction. The web-scraping chapter in R for Data Science, 2nd Edition and the University of California, Riverside Data Center tutorial offer supplementary examples, including broader scraping workflows. Treat any book or tutorial example as a pattern: confirm current markup, permissions, and package behavior before adapting it.

Frequently Asked Questions

Can rvest scrape a page behind a login?

Only when you have permission and a supported authentication workflow. Use documented credentials or an API, protect secrets, and do not attempt to bypass access controls.

Should I save the raw HTML?

For maintainable projects, saving a small permitted fixture or response sample helps detect selector changes and makes tests reproducible. Apply the site’s retention and privacy rules.

What is the difference between html_element() and html_elements()?

html_elements() returns all matches in a context; html_element() selects one match per context. Use the plural function for repeated records and the singular function for fields inside each record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.