October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
HTML parsing

Ruby HTML and XML Parsers: Nokogiri, REXML, Ox and Oga Compared

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For most Ruby applications that need to read both HTML and XML, start with Nokogiri. It provides tree-based parsing, CSS and XPath queries, HTML4/HTML5 and XML support, editing, validation, XSLT and builder APIs. Use REXML when you want Ruby’s XML-focused toolkit, and consider Ox or Oga when their streaming, serialization or HTML/XML APIs match a specific workload. The right choice still depends on your Ruby runtime, document size, namespaces, malformed input and security requirements.

Parsing is only half of “getting a website’s source.” An HTTP client first downloads bytes; a parser then turns those bytes into a document your code can query. The examples below keep those steps separate so failures are easier to diagnose.

Choose a parser by the job

Library Strong fit Trade-off to check
Nokogiri Combined HTML/XML parsing, CSS and XPath queries, editing, validation, transformation and builders HTML5 is documented as unavailable on JRuby; native-gem and source-install details vary by platform.
REXML XML documents when Ruby’s tree or stream APIs fit the application It is XML-focused, and its stream parser does not provide features such as XPath.
Ox XML parsing and writing, object-to-XML serialization and SAX-like streaming Repository speed claims are project-specific; the reviewed page does not provide enough dated methodology for a neutral benchmark.
Oga Documented HTML/XML, HTML5, DOM, pull/stream, SAX, XPath and CSS functionality Its README notes limited maintainer spare time, so verify current activity and compatibility before adopting it.

Compare candidates with representative documents rather than a generic “fastest gem” claim. Include valid and malformed HTML, XML namespaces, large files, your production Ruby version and the exact runtime (CRuby or JRuby).

Install Nokogiri and verify the runtime

Add the dependency

bundle add nokogiri
bundle exec ruby -e 'require "nokogiri"; puts Nokogiri::VERSION'

Supported platforms can normally install Nokogiri’s native gem. A source build may require a C compiler toolchain, Ruby development headers and system dependencies. On CRuby, the implementation uses libxml2 and libxslt; the JRuby implementation uses Java libraries including Xerces and NekoHTML. Read the current installation guide for the deployment image you actually ship.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

HTML5 runtime caveat

Nokogiri’s documented HTML5 API is available from version 1.12.0 onward and is not available on JRuby. Confirm both the gem version and runtime before using Nokogiri.HTML5 in a portable library.

Parse HTML from a URL in Ruby

Use an HTTP client for retrieval and Nokogiri for parsing. The following example checks the response before parsing, follows no redirects implicitly, and extracts links with CSS selectors.

require "net/http"
require "uri"
require "nokogiri"

uri = URI("https://example.com/")
http = Net::HTTP.new(uri.host, uri.port)
http.use_ssl = (uri.scheme == "https")
http.open_timeout = 10
http.read_timeout = 30

request = Net::HTTP::Get.new(uri)
request["User-Agent"] = "MyRubyParser/1.0"
response = http.request(request)
unless response.is_a?(Net::HTTPSuccess)
  abort "HTTP #{response.code} #{response.message}"
end

doc = Nokogiri::HTML(response.body, uri.to_s)
puts doc.at_css("title")&.text&.strip

doc.css("a[href]").each do |link|
  puts link["href"]
end

The optional base URL helps resolve relative links when you later use Nokogiri’s URL-handling facilities. Treat response bodies as untrusted input, limit download size in production, and do not assume a successful HTTP status means the page contains the markup you expect.

Parse an HTML fragment

fragment = Nokogiri::HTML.fragment('

Ready

') puts fragment.at_css("p.notice").text

Parse HTML5

html5 = Nokogiri.HTML5("
Hi
") puts html5.at_css("dialog").text

Use Nokogiri::HTML5.fragment for a fragment. If the application must run on JRuby, use a supported HTML parser path instead of assuming this API exists.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse XML with Nokogiri

require "nokogiri"

xml = <<~XML
  <catalog xmlns="urn:demo">
    <book id="1"><title>Ruby</title></book>
  </catalog>
XML

doc = Nokogiri::XML(xml) { |config| config.strict.nonet }
ns = { "d" => "urn:demo" }
book = doc.at_xpath("//d:book", ns)
puts book["id"]
puts book.at_xpath("d:title", ns).text

Namespaces are a common reason an XPath expression appears to return nothing. Bind the document’s namespace URI to a prefix you control, then use that prefix in every XPath step. Keep network access disabled for untrusted XML unless your threat model explicitly requires otherwise; Nokogiri treats documents as untrusted by default, but parser options should still be reviewed for the exact version you deploy.

Query, edit and validate documents

CSS versus XPath

  • Use css("article h2") when selecting by familiar HTML selectors.
  • Use XPath when you need axes, attributes, namespace-qualified XML or positional logic.
  • at_css and at_xpath return the first match; css and xpath return collections.

Edit and serialize

doc = Nokogiri::HTML5.fragment('<div class="card">Old</div>')
card = doc.at_css(".card")
card["class"] = "card featured"
card.content = "New"
puts doc.to_html

Nokogiri also documents XML builders, XSLT transformations and XSD validation. Use those APIs when parsing is part of a document-processing pipeline rather than merely extraction; validate against the schema you own and report validation errors instead of silently accepting invalid data.

When REXML, Ox or Oga is a better fit

REXML for Ruby-oriented XML work

require "rexml/document"
include REXML

xml = '<items><item id="7">Ruby</item></items>'
doc = Document.new(xml)
puts doc.elements["items/item"].text

REXML is an XML toolkit for Ruby. It offers tree and stream parsing. Its stream mode can suit sequential processing, but the project documentation notes that stream parsing lacks features such as XPath. Choose tree mode when you need random navigation and stream mode when one-pass event handling matters more.

Ox for XML serialization or event-style processing

Ox documents XML parsing and writing, object-to-XML serialization and SAX-like stream APIs. Those are useful distinctions for services that emit XML or process large inputs incrementally. Do not treat numerical performance claims in a project README as a current, neutral benchmark: versions, callbacks, input shape, runtime and hardware can change the result.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Oga for a broad alternative API

Oga documents HTML/XML parsing, HTML5, DOM, pull and stream parsing, SAX, XPath and CSS. It can be worth evaluating when its API or dependency profile fits your application. Its README also says the maintainer has limited spare time, so check release activity, open issues and compatibility with your Ruby version before committing.

Streaming, memory and performance decisions

Tree parsing

A DOM keeps the document structure in memory and makes repeated CSS/XPath queries and edits convenient. It is usually the simplest choice for pages and XML files whose size is bounded and known.

Event or pull parsing

Stream APIs process records as they arrive and can avoid building a complete tree. They require stateful callbacks or iteration and generally give up random access; REXML’s documented stream limitations include missing XPath. Benchmark your own workload rather than copying a library’s speed number.

Reduce work before parsing

  • Set HTTP connect and read timeouts.
  • Reject unexpectedly large responses before handing them to a parser.
  • Select only the nodes you need and avoid converting every node to a Ruby hash.
  • Cache retrieval separately from parsing so repeated queries do not redownload the page.
  • Record Ruby version, gem versions, parser mode and input size in performance tests.

Encoding and hostile input

Bytes can be valid in multiple encodings, and perfect automatic detection is impossible. If a feed declares the wrong encoding or you know the source’s encoding, set it explicitly using the parser’s documented options for your installed version. Preserve the original bytes when auditability matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For untrusted XML or HTML, apply a resource limit, disable network access and external entity behavior unless required, and avoid evaluating embedded scripts. Sanitize extracted text before displaying it in another context. Parser safety is not the same as application-level HTML sanitization.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“cannot load such file — nokogiri”

The bundle is not installed in the Ruby environment running the script. Run bundle install, execute through bundle exec, and verify that the deployment image contains the same Gemfile.lock.

Native extension or compiler errors

Your platform may lack a compatible precompiled gem or build tools. Use a supported Ruby/platform combination, install the required compiler and development headers, or choose the platform-supported package documented by Nokogiri.

HTML5 method missing

Check the Nokogiri version and runtime. The documented HTML5 API starts at 1.12.0 and is unavailable on JRuby; use an alternative supported parser path when that combination is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Empty XPath results

Inspect the parsed document, verify the element name and bind namespace URIs. An unprefixed XPath will not match elements in a default XML namespace.

Parser errors on a web page

HTML is often malformed and may be an error page, login page or bot challenge rather than the expected content. Log status, content type, final URL and a bounded response sample, then handle redirects and authentication explicitly.

Encoding corruption

Compare the HTTP charset, XML declaration, HTML metadata and actual bytes. Set the known encoding before parsing and test with non-ASCII fixtures.

Or skip the browser setup

If your goal is a clean visual capture rather than source extraction, ScreenshotNeo returns a PNG, JPEG, WebP or PDF from one request. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for selectors, full-page and lazy-image capture, device and retina settings, PDF options, custom CSS or JavaScript, blocking rules, headers, cookies, geolocation, caching, signed links, asynchronous jobs and bulk capture. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Which parser should you pick?

  1. Choose Nokogiri for mixed HTML/XML work, CSS or XPath queries, editing, validation or XSLT.
  2. Choose REXML when an XML-only Ruby toolkit and its tree/stream trade-off are appropriate.
  3. Evaluate Ox for XML serialization or event-style processing, and Oga for its documented HTML/XML and selector APIs.
  4. Run a fixture-based test on your real Ruby runtime, including malformed input, namespaces, encodings and the largest expected document.

Frequently Asked Questions

Can Ruby parse a webpage without Nokogiri?

Yes. Ruby’s standard library includes REXML for XML, while other gems such as Ox and Oga provide alternative XML or HTML APIs. You still need an HTTP client to download the page.

Should I use CSS selectors or XPath?

CSS is usually clearer for ordinary HTML selection. XPath is more expressive for XML namespaces, attributes, axes and positional conditions.

Does a parser execute JavaScript on a page?

No. HTTP retrieval gives you the response HTML, not the browser’s post-JavaScript DOM. A browser-rendering service is needed when the content exists only after client-side execution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.