Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use rvest to turn permitted HTML into a tidy R data frame: load the document with read_html(), identify the repeated record on the page, select its nodes with CSS selectors or XPath, extract text and attributes, and validate the resulting rows. If the required content is created only after JavaScript runs, switch to read_html_live() or the site’s documented API.
The mental model: document, records, fields
A web page is a hierarchy of HTML elements. A product card, article, search result, or table row is usually a repeated unit. Treat each unit as one record and each value inside it as a field. In a data frame, that means one row per repeated unit and one column per field.
- Document: the HTML returned by the server or browser.
- Record selector: a CSS selector such as
article.product-cardthat matches every unit. - Field selector: a selector inside one record, such as
h2ora. - Value: text, an attribute such as
href, or another property you explicitly extract.
Selectors are target-page-specific. Inspect a current sample page, confirm which elements repeat, and record the extraction date when you maintain a real project. A redesign can leave your R code running while silently returning empty or shifted columns.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Install rvest and inspect a page
install.packages(c("rvest", "dplyr", "tibble", "xml2"))
library(rvest)
library(dplyr)
library(tibble)
page <- read_html("https://example.org/sample-page")
page
example.org/sample-page is a pattern, not a guaranteed data source. Replace it with a page you are allowed to collect and verify its current markup and rules. read_html() parses the HTML returned by a normal request; xml2 supplies the static parser used in this workflow.
#1 Best Overall
Find candidate nodes
# How many article elements are present?
length(html_elements(page, "article"))
# Inspect the first matching element
first_record <- html_element(page, "article")
first_record
# Extract all headings on the page
page |> html_elements("h1, h2, h3") |> html_text2()
html_elements() returns all matches. html_element() returns the first match for each context. The distinction matters: use the plural form for repeated records and the singular form for one field inside each record.
Example project: repeated HTML records to a data frame
The following project assumes the page contains repeated <article> elements, each with an <h2> title and a link. Confirm those selectors in your target page before running it.
library(rvest)
library(dplyr)
library(tibble)
url <- "https://example.org/sample-page"
page <- read_html(url)
records <- html_elements(page, "article")
results <- tibble(
title = records |> html_element("h2") |> html_text2(),
link = records |> html_element("a") |> html_attr("href")
)
print(results, n = 10)
write.csv(results, "scraped-results.csv", row.names = FALSE)
html_text2() returns readable text while handling descendant text and whitespace more usefully than a raw text-node extraction. html_attr("href") reads an attribute. To collect another field, add a column using a selector that exists inside every record:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
results <- tibble(
title = records |> html_element("h2") |> html_text2(),
link = records |> html_element("a") |> html_attr("href"),
summary = records |> html_element("p.summary") |> html_text2()
)
Make relative links usable
Sites often store /products/42 instead of a complete URL. Resolve those links against the page URL rather than concatenating strings manually.
results <- results |>
mutate(link = url_absolute(link, url))
If the target uses a <base> element or an unusual URL scheme, inspect a few outputs before saving them as canonical links.
Handle missing elements deliberately
A missing child element can produce NA or an empty result depending on the selector and context. Keep the row, then inspect missingness instead of dropping records accidentally.
results <- tibble(
title = records |> html_element("h2") |> html_text2(),
link = records |> html_element("a") |> html_attr("href"),
summary = records |> html_element("p.summary") |> html_text2()
) |>
mutate(
title = na_if(trimws(title), ""),
summary = na_if(trimws(summary), "")
)
colSums(is.na(results))
stopifnot(nrow(results) > 0)
For stricter pipelines, assert expected columns, a sensible row count, and key uniqueness. Save a small sample such as head(results, 5) so a selector failure is visible in code review or scheduled-job logs.
CSS selectors and XPath you will use most
| Goal | CSS example | What it matches |
|---|---|---|
| Element type | article |
Every article element |
| Class | .product-card |
Any element with that class |
| Element and class | article.product-card |
Article elements with that class |
| Descendant | article h2 |
Headings inside an article |
| Direct child | article > a |
Links directly inside an article |
| Attribute | a[data-id] |
Links carrying a data-id attribute |
Use XPath when the relationship is easier to express that way:
# All links whose visible text contains “Next”
page |> html_elements(xpath = "//a[contains(normalize-space(.), 'Next')]")
Do not select by a generated class name merely because it is convenient. Prefer stable semantic elements, data attributes, or a documented API when available.
Static HTML or JavaScript-rendered content?
A browser may show text that is absent from the HTML returned by a normal request. Check the parsed document first. Search the HTML for a distinctive value, inspect the relevant node, or compare the response source with what the browser displays. Do not assume that a visible element is present in static HTML.
| Question | Static path | Live-browser path |
|---|---|---|
| Is the desired value in returned HTML? | Use read_html() and selectors. |
Not needed. |
| Content appears only after scripts run? | Static parsing returns no record or incomplete fields. | Use read_html_live() after assessing browser setup and site rules. |
| Setup and dependencies | Usually faster and has fewer external dependencies. | Requires a live browser approach and is heavier to operate. |
library(rvest)
live_page <- read_html_live("https://example.org/page-with-rendered-data")
rendered_records <- live_page |> html_elements("article")
rendered_titles <- rendered_records |> html_element("h2") |> html_text2()
Use the live method only when it supplies data you cannot obtain from static HTML or an official endpoint. If a site exposes a documented API, it is often a more stable and permission-friendly interface than parsing presentation markup.
Pagination, rate limits, and responsible collection
For one page, a direct read_html() call may be enough. For many pages, use a controlled loop, pause between requests, cache results, and stop when the site signals an error. The rvest maintainers recommend pairing rvest with polite for multi-page scraping because it supports robots.txt awareness and helps avoid hitting a site too aggressively.
install.packages("polite")
library(polite)
library(rvest)
session <- bow("https://example.org")
page <- scrape(session)
html <- read_html(page)
# Follow a discovered next-page link only after checking it exists
next_href <- html |> html_element("a.next") |> html_attr("href")
if (!is.na(next_href)) {
next_url <- url_absolute(next_href, "https://example.org")
Sys.sleep(1)
next_html <- read_html(next_url)
}
Review robots.txt and the site’s terms separately; neither document alone answers every legal or policy question. Consider the site’s API when one is provided. Keep request volume proportionate, identify your application where appropriate, and avoid collecting personal or restricted data without a clear basis.
Production checks and common failure modes
Zero rows
Cause: the selector does not match the current markup, the content is JavaScript-generated, or the request received a block page. Fix: print the page, inspect a distinctive node, verify the response URL and status, then test whether the value exists in static HTML. Use a live browser only if necessary.
Rows exist but columns are empty
Cause: the child selector is wrong or optional on some records. Fix: inspect one record with html_element(), test the selector against several records, and retain explicit NA values for genuinely missing fields.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #4
Every row contains the same value
Cause: a page-level selector was applied to the whole document rather than to each record. Fix: first create records, then run each field extraction on that node set.
Relative links or malformed URLs
Cause: the site uses relative paths, fragments, or protocol-relative links. Fix: use url_absolute(), inspect unusual values, and preserve the original attribute in a separate column if auditability matters.
Requests are slow or blocked
Cause: rate limits, bot defenses, large pages, or unnecessary browser rendering. Fix: prefer static parsing, reduce concurrency, add delays, cache completed pages, follow the site’s rules, and use an official API where available. A live browser is not a bypass for access controls.
Data changes shape after a redesign
Cause: selectors depended on fragile classes or layout nesting. Fix: choose stable attributes, add row-count and missingness assertions, keep a fixture or saved sample for tests, and record when selectors were last verified.
Recommended Free Tools
Or skip the browser setup
When your goal is a clean image or PDF of a page rather than parsing its fields, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
See the parameter reference in the ScreenshotNeo documentation. Replace the example URL with the page you are permitted to capture.
Best Value
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper settings and page ranges, HTML/CSS input, custom JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. Every feature is on every plan: 1,000 shots monthly free with no card, then Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free.
Sign up for the free 1,000-screenshot monthly plan—no card required.
Further learning
The free official rvest vignette and project documentation are the best starting points for selectors and extraction. The web-scraping chapter in R for Data Science, 2nd Edition and the University of California, Riverside Data Center tutorial offer supplementary examples, including broader scraping workflows. Treat any book or tutorial example as a pattern: confirm current markup, permissions, and package behavior before adapting it.
Frequently Asked Questions
Can rvest scrape a page behind a login?
Only when you have permission and a supported authentication workflow. Use documented credentials or an API, protect secrets, and do not attempt to bypass access controls.
Should I save the raw HTML?
For maintainable projects, saving a small permitted fixture or response sample helps detect selector changes and makes tests reproducible. Apply the site’s retention and privacy rules.
What is the difference between html_element() and html_elements()?
html_elements() returns all matches in a context; html_element() selects one match per context. Use the plural function for repeated records and the singular function for fields inside each record.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

