Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The dependable Ruby workflow is fetch, parse, normalize, and export. Use an HTTP client for pages whose useful markup is in the initial response, Nokogiri for CSS/XPath selection, and a browser such as Selenium only when JavaScript creates the content after load. Start with one URL, verify the response and selectors, then add pacing, retries, and permission checks before crawling at scale.
Choose the right Ruby approach first
Web scraping is not one technique. The page’s rendering model determines the smallest tool that will work:
| Page type | Recommended Ruby stack | Operational cost | Typical limitation |
|---|---|---|---|
| Content appears in the first HTML response | HTTP client (such as HTTParty) + Nokogiri | Low; no browser process | Cannot see content that is created later by JavaScript |
| Content is inserted or changed by JavaScript | Selenium WebDriver with Chrome (and Nokogiri when useful) | Higher; browser and driver setup, more memory and latency | More moving parts and possible bot checks |
| Large, recurring collection | Either stack plus a queue, rate limits, retries, logging, and storage | Highest; you own reliability and policy compliance | Selectors, layouts, and access rules can change |
Nokogiri reads, writes, modifies, and queries HTML and XML. Its documented search interfaces include CSS selectors and XPath, as well as DOM, SAX, and push parsing modes. The current installation documentation lists Ruby 3.2 or newer and JRuby 10.0 or newer; check the live requirements before installing because support changes. The documentation also notes that HTML5 functionality is unavailable on JRuby.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Define the fields and check access before coding
Write a small extraction contract
List the fields, their expected types, and what should happen when a field is absent. For example:
#1 Best Overall
- title: required string
- price: decimal, or
nilwhen not shown - url: absolute URL
- published_at: ISO date when present
This prevents a scraper from silently producing plausible but wrong data when a page changes.
Confirm permission and site rules
Read the site’s terms, account requirements, applicable law, and any explicit API or export policy for your use case. RFC 9309 describes robots.txt as the Robots Exclusion Protocol and states: “These rules are not a form of access authorization.” Google Search Central likewise explains that robots.txt manages crawler access and traffic, does not keep pages out of search results, and does not enforce crawler behavior. Treat it as an operational signal, never as authentication, a security boundary, or proof that scraping is legally permitted.
Install a minimal static-page stack
Create a project and add the libraries:
bundle init
bundle add httparty nokogiri csv
Use a current Ruby and verify the Nokogiri runtime requirements against its installation documentation, particularly if you use JRuby. Keep the first test to one URL and one request.
Fetch and inspect one response
Fetching and parsing are separate operations: the HTTP client receives bytes and metadata; Nokogiri turns the response body into a searchable document. Inspect status, content type, and a short body sample before writing selectors.
Rank #2
require "httparty"
url = "https://example.com/articles/1"
response = HTTParty.get(
url,
headers: { "User-Agent" => "ResearchBot/1.0 ([email protected])" },
timeout: 20
)
abort("HTTP #{response.code}") unless response.success?
content_type = response.headers["content-type"].to_s
abort("Not HTML: #{content_type}") unless content_type.include?("html")
puts response.body.byteslice(0, 500)
Do not assume a successful TCP request means useful content. A login page, challenge page, empty shell, or error document can all return HTTP 200. Save status, final URL (if your client exposes it), content type, and a hash or sample of the body in your logs.
Parse with Nokogiri and select stable fields
Prefer semantic attributes, data attributes, and narrowly scoped containers over long positional selectors. Keep selectors in one place so a markup change has one repair point.
require "httparty"
require "nokogiri"
require "uri"
url = "https://example.com/articles/1"
response = HTTParty.get(url, timeout: 20)
raise "HTTP #{response.code}" unless response.success?
doc = Nokogiri::HTML(response.body)
text = ->(node) { node&.text&.strip }
title = text.call(doc.at_css("h1"))
summary = text.call(doc.at_css("[data-summary]"))
canonical = doc.at_css('link[rel="canonical"]')&.[]("href")
absolute_url = canonical ? URI.join(url, canonical).to_s : url
record = {
title: title,
summary: summary,
url: absolute_url
}
p record
at_css returns one node or nil; css returns all matching nodes. XPath is useful when the relationship is more expressive than CSS:
prices = doc.xpath('//article//span[@data-price]').map { |node| node.text.strip }
Normalize at the boundary: collapse repeated whitespace, remove presentation-only labels, convert decimal values deliberately, and preserve the original URL. Never call text on a missing node without a guard.
Rank #3
Build a complete CSV exporter
The following example fetches a list of article URLs, handles missing fields, and writes UTF-8 CSV. Its selectors are illustrative; inspect the target site’s markup and replace them.
require "httparty"
require "nokogiri"
require "csv"
require "uri"
URLS = [
"https://example.com/articles/1",
"https://example.com/articles/2"
].freeze
HEADERS = {
"User-Agent" => "ResearchBot/1.0 ([email protected])",
"Accept" => "text/html,application/xhtml+xml"
}.freeze
def clean(node)
node&.text&.gsub(/s+/, " ")&.strip
end
def scrape(url)
response = HTTParty.get(url, headers: HEADERS, timeout: 20)
raise "HTTP #{response.code}" unless response.success?
doc = Nokogiri::HTML(response.body)
{
url: url,
title: clean(doc.at_css("h1")),
author: clean(doc.at_css("[rel=author]")),
published_at: doc.at_css("time")&.[]("datetime"),
tags: doc.css("a[data-tag]").map { |n| clean(n) }.compact.join("|")
}
end
CSV.open("articles.csv", "w", write_headers: true,
headers: %w[url title author published_at tags]) do |csv|
URLS.each do |url|
begin
csv << scrape(url)
rescue StandardError => e
warn "#{url}: #{e.message}"
end
end
end
For production output, add a schema check that fails when a required field is missing, and retain failed URLs for review rather than discarding them.
Handle JavaScript-rendered pages with Selenium
First verify that the data is absent from the initial response; do not introduce a browser merely because a site uses some JavaScript. When browser rendering is necessary, Selenium WebDriver can load Chrome, wait for a meaningful element, and then expose the rendered DOM.
bundle add selenium-webdriver nokogiri
require "selenium-webdriver"
require "nokogiri"
options = Selenium::WebDriver::Chrome::Options.new
options.add_argument("--headless=new")
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
driver = Selenium::WebDriver.for(:chrome, options: options)
begin
driver.navigate.to("https://example.com/app")
wait = Selenium::WebDriver::Wait.new(timeout: 15)
wait.until { driver.find_element(css: "[data-results]").displayed? }
rendered = Nokogiri::HTML(driver.page_source)
rows = rendered.css("[data-results] [data-row]").map do |row|
{
name: row.at_css("[data-name]")&.text&.strip,
value: row.at_css("[data-value]")&.text&.strip
}
end
p rows
ensure
driver.quit
end
Use an explicit wait for a condition that proves the data is ready. A fixed sleep is slower when the page is fast and still unreliable when it is slow. Browser runs also need a compatible Chrome/driver installation, resource limits, cleanup in an ensure block, and a plan for authentication or consent flows that you are authorized to automate.
Rank #4
Scale carefully: pacing, retries, and change detection
Rate-limit deliberately
Use a conservative delay, bounded concurrency, and a per-host budget. A queue makes it possible to pause a host without stopping unrelated work. Identify your client honestly and provide a contact address where appropriate.
Retry only transient failures
- Retry timeouts, connection resets, and selected 5xx responses with exponential backoff and jitter.
- Do not blindly retry 401, 403, 404, consent failures, or a bot challenge; investigate the cause.
- Cap attempts and record the final error with URL, status, and timestamp.
Detect selector and content drift
Track the percentage of records missing required fields. Alert when it changes sharply, store a sample response for diagnosis, and test selectors against fixtures. A tutorial’s selectors and sample site are not guarantees for another site’s structure, stability, accessibility, or permission.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 200 but no records | Content is JavaScript-rendered, or you received a shell/challenge page | Inspect the body; use Selenium only if the data is created after load, and stop on a challenge rather than trying to evade it |
Nokogiri::XML::SyntaxError or malformed fields |
Unexpected markup or wrong parser assumptions | Use Nokogiri::HTML for HTML, inspect a saved response, and guard missing nodes |
| Selectors suddenly return nil | Layout or class names changed | Prefer semantic/data attributes, add fixture tests, and review a captured page |
| Timeouts during browser runs | Slow resources, an overly broad wait, or constrained host resources | Wait for a specific selector, set a bounded timeout, block nonessential resources only when authorized, and monitor memory |
| 403 or a bot-check page | Access policy, authentication, rate limit, or anti-bot system | Check terms and authorization, slow down, use an official API if available, and do not attempt to bypass the control |
| JRuby feature mismatch | Nokogiri’s documented HTML5 functionality is unavailable on JRuby | Use supported parsing features or a CRuby runtime, after checking current documentation |
Performance, reliability, and cost decisions
- Start with HTTP: it avoids browser startup and is usually simpler for static HTML.
- Use browsers selectively: reserve Selenium for pages whose required data is genuinely rendered client-side.
- Cache responsibly: avoid refetching unchanged URLs when policy allows, and give your cache a clear freshness rule.
- Measure the pipeline: record request time, parse time, response size, status, retry count, and missing-field counts.
- Separate collection from processing: saving raw responses lets you repair selectors without repeatedly requesting the site.
There is no reliable universal speed or success percentage: results depend on the target, network, rendering, policies, and selector quality.
Recommended Free Tools
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
For a visual capture rather than DOM extraction, one request is enough:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Ruby can call the same endpoint:
require "net/http"
require "uri"
uri = URI("https://api.screenshotneo.com/v1/shot")
uri.query = URI.encode_www_form(access_key: "YOUR_API_KEY", url: "https://stripe.com")
response = Net::HTTP.get_response(uri)
raise "HTTP #{response.code}" unless response.is_a?(Net::HTTPSuccess)
File.binwrite("shot.webp", response.body)
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for the full parameter set, including full-page and element capture, device and retina settings, PDF controls, custom CSS/JavaScript, waits, request blocking, headers and cookies, timezone and geolocation, transparent backgrounds, resizing, caching TTLs, signed links, asynchronous webhooks, bulk capture, usage, and the OpenAPI specification. Every plan includes all features: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots, with yearly billing giving two months free. Create a free ScreenshotNeo account to try it.
A practical checklist before you crawl more pages
- Confirm the fields and acceptable missing-value behavior.
- Inspect one raw response and verify the content type.
- Choose HTTP plus Nokogiri or Selenium based on where the data appears.
- Test selectors against saved fixtures and alert on drift.
- Implement pacing, bounded retries, logging, and duplicate handling.
- Review terms, authorization, robots.txt, and applicable law separately.
- Store raw responses or evidence needed to reproduce an extraction.
Frequently Asked Questions
Can Nokogiri scrape a page by itself?
No. Nokogiri parses HTML or XML that you already obtained; an HTTP client or browser must retrieve the page first.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhen should I avoid Selenium?
Avoid it when the required fields are present in the initial HTML response. A direct HTTP request is simpler and uses fewer runtime resources.
Does robots.txt make scraping legal?
No. It is a crawler-access protocol, not authorization or a security control. Consider terms, permission, and applicable law independently.
What should I do when a selector returns no nodes?
Save and inspect the actual response, check whether JavaScript supplies the content, verify the selector in browser developer tools, and add a guarded failure rather than emitting an empty record.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →

