For most Ruby applications that need to read both HTML and XML, start with Nokogiri. It provides tree-based parsing, CSS and XPath queries, HTML4/HTML5 and XML support, editing, validation, XSLT and builder APIs. Use REXML when you want Ruby’s XML-focused toolkit, and consider Ox or Oga when their streaming, serialization or HTML/XML APIs match a specific workload. The right choice still depends on your Ruby runtime, document size, namespaces, malformed input and security requirements.
Parsing is only half of “getting a website’s source.” An HTTP client first downloads bytes; a parser then turns those bytes into a document your code can query. The examples below keep those steps separate so failures are easier to diagnose.
Choose a parser by the job
| Library | Strong fit | Trade-off to check |
|---|---|---|
| Nokogiri | Combined HTML/XML parsing, CSS and XPath queries, editing, validation, transformation and builders | HTML5 is documented as unavailable on JRuby; native-gem and source-install details vary by platform. |
| REXML | XML documents when Ruby’s tree or stream APIs fit the application | It is XML-focused, and its stream parser does not provide features such as XPath. |
| Ox | XML parsing and writing, object-to-XML serialization and SAX-like streaming | Repository speed claims are project-specific; the reviewed page does not provide enough dated methodology for a neutral benchmark. |
| Oga | Documented HTML/XML, HTML5, DOM, pull/stream, SAX, XPath and CSS functionality | Its README notes limited maintainer spare time, so verify current activity and compatibility before adopting it. |
Compare candidates with representative documents rather than a generic “fastest gem” claim. Include valid and malformed HTML, XML namespaces, large files, your production Ruby version and the exact runtime (CRuby or JRuby).
Install Nokogiri and verify the runtime
Add the dependency
bundle add nokogiri
bundle exec ruby -e 'require "nokogiri"; puts Nokogiri::VERSION'
Supported platforms can normally install Nokogiri’s native gem. A source build may require a C compiler toolchain, Ruby development headers and system dependencies. On CRuby, the implementation uses libxml2 and libxslt; the JRuby implementation uses Java libraries including Xerces and NekoHTML. Read the current installation guide for the deployment image you actually ship.
Recommended Free Tools
#1 Best Overall
HTML5 runtime caveat
Nokogiri’s documented HTML5 API is available from version 1.12.0 onward and is not available on JRuby. Confirm both the gem version and runtime before using Nokogiri.HTML5 in a portable library.
Parse HTML from a URL in Ruby
Use an HTTP client for retrieval and Nokogiri for parsing. The following example checks the response before parsing, follows no redirects implicitly, and extracts links with CSS selectors.
require "net/http"
require "uri"
require "nokogiri"
uri = URI("https://example.com/")
http = Net::HTTP.new(uri.host, uri.port)
http.use_ssl = (uri.scheme == "https")
http.open_timeout = 10
http.read_timeout = 30
request = Net::HTTP::Get.new(uri)
request["User-Agent"] = "MyRubyParser/1.0"
response = http.request(request)
unless response.is_a?(Net::HTTPSuccess)
abort "HTTP #{response.code} #{response.message}"
end
doc = Nokogiri::HTML(response.body, uri.to_s)
puts doc.at_css("title")&.text&.strip
doc.css("a[href]").each do |link|
puts link["href"]
end
The optional base URL helps resolve relative links when you later use Nokogiri’s URL-handling facilities. Treat response bodies as untrusted input, limit download size in production, and do not assume a successful HTTP status means the page contains the markup you expect.
Parse an HTML fragment
fragment = Nokogiri::HTML.fragment('Ready
')
puts fragment.at_css("p.notice").text
Parse HTML5
html5 = Nokogiri.HTML5("")
puts html5.at_css("dialog").text
Use Nokogiri::HTML5.fragment for a fragment. If the application must run on JRuby, use a supported HTML parser path instead of assuming this API exists.
Rank #2
Parse XML with Nokogiri
require "nokogiri"
xml = <<~XML
<catalog xmlns="urn:demo">
<book id="1"><title>Ruby</title></book>
</catalog>
XML
doc = Nokogiri::XML(xml) { |config| config.strict.nonet }
ns = { "d" => "urn:demo" }
book = doc.at_xpath("//d:book", ns)
puts book["id"]
puts book.at_xpath("d:title", ns).text
Namespaces are a common reason an XPath expression appears to return nothing. Bind the document’s namespace URI to a prefix you control, then use that prefix in every XPath step. Keep network access disabled for untrusted XML unless your threat model explicitly requires otherwise; Nokogiri treats documents as untrusted by default, but parser options should still be reviewed for the exact version you deploy.
Query, edit and validate documents
CSS versus XPath
- Use
css("article h2")when selecting by familiar HTML selectors. - Use XPath when you need axes, attributes, namespace-qualified XML or positional logic.
at_cssandat_xpathreturn the first match;cssandxpathreturn collections.
Edit and serialize
doc = Nokogiri::HTML5.fragment('<div class="card">Old</div>')
card = doc.at_css(".card")
card["class"] = "card featured"
card.content = "New"
puts doc.to_html
Nokogiri also documents XML builders, XSLT transformations and XSD validation. Use those APIs when parsing is part of a document-processing pipeline rather than merely extraction; validate against the schema you own and report validation errors instead of silently accepting invalid data.
When REXML, Ox or Oga is a better fit
REXML for Ruby-oriented XML work
require "rexml/document"
include REXML
xml = '<items><item id="7">Ruby</item></items>'
doc = Document.new(xml)
puts doc.elements["items/item"].text
REXML is an XML toolkit for Ruby. It offers tree and stream parsing. Its stream mode can suit sequential processing, but the project documentation notes that stream parsing lacks features such as XPath. Choose tree mode when you need random navigation and stream mode when one-pass event handling matters more.
Ox for XML serialization or event-style processing
Ox documents XML parsing and writing, object-to-XML serialization and SAX-like stream APIs. Those are useful distinctions for services that emit XML or process large inputs incrementally. Do not treat numerical performance claims in a project README as a current, neutral benchmark: versions, callbacks, input shape, runtime and hardware can change the result.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Oga for a broad alternative API
Oga documents HTML/XML parsing, HTML5, DOM, pull and stream parsing, SAX, XPath and CSS. It can be worth evaluating when its API or dependency profile fits your application. Its README also says the maintainer has limited spare time, so check release activity, open issues and compatibility with your Ruby version before committing.
Streaming, memory and performance decisions
Tree parsing
A DOM keeps the document structure in memory and makes repeated CSS/XPath queries and edits convenient. It is usually the simplest choice for pages and XML files whose size is bounded and known.
Event or pull parsing
Stream APIs process records as they arrive and can avoid building a complete tree. They require stateful callbacks or iteration and generally give up random access; REXML’s documented stream limitations include missing XPath. Benchmark your own workload rather than copying a library’s speed number.
Reduce work before parsing
- Set HTTP connect and read timeouts.
- Reject unexpectedly large responses before handing them to a parser.
- Select only the nodes you need and avoid converting every node to a Ruby hash.
- Cache retrieval separately from parsing so repeated queries do not redownload the page.
- Record Ruby version, gem versions, parser mode and input size in performance tests.
Encoding and hostile input
Bytes can be valid in multiple encodings, and perfect automatic detection is impossible. If a feed declares the wrong encoding or you know the source’s encoding, set it explicitly using the parser’s documented options for your installed version. Preserve the original bytes when auditability matters.
Rank #4
For untrusted XML or HTML, apply a resource limit, disable network access and external entity behavior unless required, and avoid evaluating embedded scripts. Sanitize extracted text before displaying it in another context. Parser safety is not the same as application-level HTML sanitization.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
“cannot load such file — nokogiri”
The bundle is not installed in the Ruby environment running the script. Run bundle install, execute through bundle exec, and verify that the deployment image contains the same Gemfile.lock.
Native extension or compiler errors
Your platform may lack a compatible precompiled gem or build tools. Use a supported Ruby/platform combination, install the required compiler and development headers, or choose the platform-supported package documented by Nokogiri.
HTML5 method missing
Check the Nokogiri version and runtime. The documented HTML5 API starts at 1.12.0 and is unavailable on JRuby; use an alternative supported parser path when that combination is required.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
Empty XPath results
Inspect the parsed document, verify the element name and bind namespace URIs. An unprefixed XPath will not match elements in a default XML namespace.
Parser errors on a web page
HTML is often malformed and may be an error page, login page or bot challenge rather than the expected content. Log status, content type, final URL and a bounded response sample, then handle redirects and authentication explicitly.
Encoding corruption
Compare the HTTP charset, XML declaration, HTML metadata and actual bytes. Set the known encoding before parsing and test with non-ASCII fixtures.
Or skip the browser setup
If your goal is a clean visual capture rather than source extraction, ScreenshotNeo returns a PNG, JPEG, WebP or PDF from one request. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for selectors, full-page and lazy-image capture, device and retina settings, PDF options, custom CSS or JavaScript, blocking rules, headers, cookies, geolocation, caching, signed links, asynchronous jobs and bulk capture. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Which parser should you pick?
- Choose Nokogiri for mixed HTML/XML work, CSS or XPath queries, editing, validation or XSLT.
- Choose REXML when an XML-only Ruby toolkit and its tree/stream trade-off are appropriate.
- Evaluate Ox for XML serialization or event-style processing, and Oga for its documented HTML/XML and selector APIs.
- Run a fixture-based test on your real Ruby runtime, including malformed input, namespaces, encodings and the largest expected document.
Frequently Asked Questions
Can Ruby parse a webpage without Nokogiri?
Yes. Ruby’s standard library includes REXML for XML, while other gems such as Ox and Oga provide alternative XML or HTML APIs. You still need an HTTP client to download the page.
Should I use CSS selectors or XPath?
CSS is usually clearer for ordinary HTML selection. XPath is more expressive for XML namespaces, attributes, axes and positional conditions.
Does a parser execute JavaScript on a page?
No. HTTP retrieval gives you the response HTML, not the browser’s post-JavaScript DOM. A browser-rendering service is needed when the content exists only after client-side execution.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




