Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The dependable Ruby workflow is fetch, parse, normalize, and export. Use an HTTP client for pages whose useful markup is in the initial response, Nokogiri for CSS/XPath selection, and a browser such as Selenium only when JavaScript creates the content after load. Start with one URL, verify the response and selectors, then add pacing, retries, and permission checks before crawling at scale.

Choose the right Ruby approach first

Web scraping is not one technique. The page’s rendering model determines the smallest tool that will work:

Page type Recommended Ruby stack Operational cost Typical limitation
Content appears in the first HTML response HTTP client (such as HTTParty) + Nokogiri Low; no browser process Cannot see content that is created later by JavaScript
Content is inserted or changed by JavaScript Selenium WebDriver with Chrome (and Nokogiri when useful) Higher; browser and driver setup, more memory and latency More moving parts and possible bot checks
Large, recurring collection Either stack plus a queue, rate limits, retries, logging, and storage Highest; you own reliability and policy compliance Selectors, layouts, and access rules can change

Nokogiri reads, writes, modifies, and queries HTML and XML. Its documented search interfaces include CSS selectors and XPath, as well as DOM, SAX, and push parsing modes. The current installation documentation lists Ruby 3.2 or newer and JRuby 10.0 or newer; check the live requirements before installing because support changes. The documentation also notes that HTML5 functionality is unavailable on JRuby.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define the fields and check access before coding

Write a small extraction contract

List the fields, their expected types, and what should happen when a field is absent. For example:

#1 Best Overall
  • title: required string
  • price: decimal, or nil when not shown
  • url: absolute URL
  • published_at: ISO date when present

This prevents a scraper from silently producing plausible but wrong data when a page changes.

Confirm permission and site rules

Read the site’s terms, account requirements, applicable law, and any explicit API or export policy for your use case. RFC 9309 describes robots.txt as the Robots Exclusion Protocol and states: “These rules are not a form of access authorization.” Google Search Central likewise explains that robots.txt manages crawler access and traffic, does not keep pages out of search results, and does not enforce crawler behavior. Treat it as an operational signal, never as authentication, a security boundary, or proof that scraping is legally permitted.

Install a minimal static-page stack

Create a project and add the libraries:

bundle init
bundle add httparty nokogiri csv

Use a current Ruby and verify the Nokogiri runtime requirements against its installation documentation, particularly if you use JRuby. Keep the first test to one URL and one request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fetch and inspect one response

Fetching and parsing are separate operations: the HTTP client receives bytes and metadata; Nokogiri turns the response body into a searchable document. Inspect status, content type, and a short body sample before writing selectors.

require "httparty"

url = "https://example.com/articles/1"
response = HTTParty.get(
  url,
  headers: { "User-Agent" => "ResearchBot/1.0 ([email protected])" },
  timeout: 20
)

abort("HTTP #{response.code}") unless response.success?
content_type = response.headers["content-type"].to_s
abort("Not HTML: #{content_type}") unless content_type.include?("html")

puts response.body.byteslice(0, 500)

Do not assume a successful TCP request means useful content. A login page, challenge page, empty shell, or error document can all return HTTP 200. Save status, final URL (if your client exposes it), content type, and a hash or sample of the body in your logs.

Parse with Nokogiri and select stable fields

Prefer semantic attributes, data attributes, and narrowly scoped containers over long positional selectors. Keep selectors in one place so a markup change has one repair point.

require "httparty"
require "nokogiri"
require "uri"

url = "https://example.com/articles/1"
response = HTTParty.get(url, timeout: 20)
raise "HTTP #{response.code}" unless response.success?

doc = Nokogiri::HTML(response.body)

text = ->(node) { node&.text&.strip }
title = text.call(doc.at_css("h1"))
summary = text.call(doc.at_css("[data-summary]"))

canonical = doc.at_css('link[rel="canonical"]')&.[]("href")
absolute_url = canonical ? URI.join(url, canonical).to_s : url

record = {
  title: title,
  summary: summary,
  url: absolute_url
}

p record

at_css returns one node or nil; css returns all matching nodes. XPath is useful when the relationship is more expressive than CSS:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
prices = doc.xpath('//article//span[@data-price]').map { |node| node.text.strip }

Normalize at the boundary: collapse repeated whitespace, remove presentation-only labels, convert decimal values deliberately, and preserve the original URL. Never call text on a missing node without a guard.

Build a complete CSV exporter

The following example fetches a list of article URLs, handles missing fields, and writes UTF-8 CSV. Its selectors are illustrative; inspect the target site’s markup and replace them.

require "httparty"
require "nokogiri"
require "csv"
require "uri"

URLS = [
  "https://example.com/articles/1",
  "https://example.com/articles/2"
].freeze

HEADERS = {
  "User-Agent" => "ResearchBot/1.0 ([email protected])",
  "Accept" => "text/html,application/xhtml+xml"
}.freeze

def clean(node)
  node&.text&.gsub(/s+/, " ")&.strip
end

def scrape(url)
  response = HTTParty.get(url, headers: HEADERS, timeout: 20)
  raise "HTTP #{response.code}" unless response.success?

  doc = Nokogiri::HTML(response.body)
  {
    url: url,
    title: clean(doc.at_css("h1")),
    author: clean(doc.at_css("[rel=author]")),
    published_at: doc.at_css("time")&.[]("datetime"),
    tags: doc.css("a[data-tag]").map { |n| clean(n) }.compact.join("|")
  }
end

CSV.open("articles.csv", "w", write_headers: true,
         headers: %w[url title author published_at tags]) do |csv|
  URLS.each do |url|
    begin
      csv << scrape(url)
    rescue StandardError => e
      warn "#{url}: #{e.message}"
    end
  end
end

For production output, add a schema check that fails when a required field is missing, and retain failed URLs for review rather than discarding them.

Handle JavaScript-rendered pages with Selenium

First verify that the data is absent from the initial response; do not introduce a browser merely because a site uses some JavaScript. When browser rendering is necessary, Selenium WebDriver can load Chrome, wait for a meaningful element, and then expose the rendered DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
bundle add selenium-webdriver nokogiri
require "selenium-webdriver"
require "nokogiri"

options = Selenium::WebDriver::Chrome::Options.new
options.add_argument("--headless=new")
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")

driver = Selenium::WebDriver.for(:chrome, options: options)
begin
  driver.navigate.to("https://example.com/app")
  wait = Selenium::WebDriver::Wait.new(timeout: 15)
  wait.until { driver.find_element(css: "[data-results]").displayed? }

  rendered = Nokogiri::HTML(driver.page_source)
  rows = rendered.css("[data-results] [data-row]").map do |row|
    {
      name: row.at_css("[data-name]")&.text&.strip,
      value: row.at_css("[data-value]")&.text&.strip
    }
  end
  p rows
ensure
  driver.quit
end

Use an explicit wait for a condition that proves the data is ready. A fixed sleep is slower when the page is fast and still unreliable when it is slow. Browser runs also need a compatible Chrome/driver installation, resource limits, cleanup in an ensure block, and a plan for authentication or consent flows that you are authorized to automate.

Scale carefully: pacing, retries, and change detection

Rate-limit deliberately

Use a conservative delay, bounded concurrency, and a per-host budget. A queue makes it possible to pause a host without stopping unrelated work. Identify your client honestly and provide a contact address where appropriate.

Retry only transient failures

  • Retry timeouts, connection resets, and selected 5xx responses with exponential backoff and jitter.
  • Do not blindly retry 401, 403, 404, consent failures, or a bot challenge; investigate the cause.
  • Cap attempts and record the final error with URL, status, and timestamp.

Detect selector and content drift

Track the percentage of records missing required fields. Alert when it changes sharply, store a sample response for diagnosis, and test selectors against fixtures. A tutorial’s selectors and sample site are not guarantees for another site’s structure, stability, accessibility, or permission.

Common failures and fixes

Symptom Likely cause Fix
HTTP 200 but no records Content is JavaScript-rendered, or you received a shell/challenge page Inspect the body; use Selenium only if the data is created after load, and stop on a challenge rather than trying to evade it
Nokogiri::XML::SyntaxError or malformed fields Unexpected markup or wrong parser assumptions Use Nokogiri::HTML for HTML, inspect a saved response, and guard missing nodes
Selectors suddenly return nil Layout or class names changed Prefer semantic/data attributes, add fixture tests, and review a captured page
Timeouts during browser runs Slow resources, an overly broad wait, or constrained host resources Wait for a specific selector, set a bounded timeout, block nonessential resources only when authorized, and monitor memory
403 or a bot-check page Access policy, authentication, rate limit, or anti-bot system Check terms and authorization, slow down, use an official API if available, and do not attempt to bypass the control
JRuby feature mismatch Nokogiri’s documented HTML5 functionality is unavailable on JRuby Use supported parsing features or a CRuby runtime, after checking current documentation

Performance, reliability, and cost decisions

  • Start with HTTP: it avoids browser startup and is usually simpler for static HTML.
  • Use browsers selectively: reserve Selenium for pages whose required data is genuinely rendered client-side.
  • Cache responsibly: avoid refetching unchanged URLs when policy allows, and give your cache a clear freshness rule.
  • Measure the pipeline: record request time, parse time, response size, status, retry count, and missing-field counts.
  • Separate collection from processing: saving raw responses lets you repair selectors without repeatedly requesting the site.

There is no reliable universal speed or success percentage: results depend on the target, network, rendering, policies, and selector quality.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

For a visual capture rather than DOM extraction, one request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Ruby can call the same endpoint:

require "net/http"
require "uri"

uri = URI("https://api.screenshotneo.com/v1/shot")
uri.query = URI.encode_www_form(access_key: "YOUR_API_KEY", url: "https://stripe.com")
response = Net::HTTP.get_response(uri)
raise "HTTP #{response.code}" unless response.is_a?(Net::HTTPSuccess)
File.binwrite("shot.webp", response.body)

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the full parameter set, including full-page and element capture, device and retina settings, PDF controls, custom CSS/JavaScript, waits, request blocking, headers and cookies, timezone and geolocation, transparent backgrounds, resizing, caching TTLs, signed links, asynchronous webhooks, bulk capture, usage, and the OpenAPI specification. Every plan includes all features: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots, with yearly billing giving two months free. Create a free ScreenshotNeo account to try it.

A practical checklist before you crawl more pages

  • Confirm the fields and acceptable missing-value behavior.
  • Inspect one raw response and verify the content type.
  • Choose HTTP plus Nokogiri or Selenium based on where the data appears.
  • Test selectors against saved fixtures and alert on drift.
  • Implement pacing, bounded retries, logging, and duplicate handling.
  • Review terms, authorization, robots.txt, and applicable law separately.
  • Store raw responses or evidence needed to reproduce an extraction.

Frequently Asked Questions

Can Nokogiri scrape a page by itself?

No. Nokogiri parses HTML or XML that you already obtained; an HTTP client or browser must retrieve the page first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I avoid Selenium?

Avoid it when the required fields are present in the initial HTML response. A direct HTTP request is simpler and uses fewer runtime resources.

Does robots.txt make scraping legal?

No. It is a crawler-access protocol, not authorization or a security control. Consider terms, permission, and applicable law independently.

What should I do when a selector returns no nodes?

Save and inspect the actual response, check whether JavaScript supplies the content, verify the selector in browser developer tools, and add a guarded failure rather than emitting an empty record.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.