Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Nokogiri to turn HTML bytes or a string into a document, then query that document with CSS selectors or XPath. For a normal page, the basic pattern is require "nokogiri", Nokogiri::HTML(html), and a query such as doc.at_css("h1"). Choose HTML5 parsing when browser-style HTML5 tree construction matters, use a fragment parser for snippets, and specify the source encoding when its declaration is unreliable.

Install Nokogiri and parse a complete HTML document

Add Nokogiri to your application’s Gemfile, then install the bundle:

# Gemfile
gem "nokogiri"
bundle install

Here is a complete small example that parses a string and extracts a heading and link:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
require "nokogiri"

html = <<~HTML
  <html>
    <body>
      <article>
        <h1>Example</h1>
        <a href="/next">Next</a>
      </article>
    </body>
  </html>
HTML

doc = Nokogiri::HTML(html)
title = doc.at_css("article h1")&.&text&.&strip
href  = doc.at_xpath("//article//a/@href")&.&value

puts title
puts href

Nokogiri::HTML is the familiar entry point for HTML4-style parsing. It returns a document node you can query; parsing itself does not decide whether the values you extract are valid for your application. Nokogiri documents DOM parsing for HTML4 and HTML5, CSS3 selectors, and XPath 1.0 support at its official documentation.

#1 Best Overall

Parse a file or an IO object

When the HTML is stored in a file, pass its contents to the parser. Reading in binary mode preserves the original bytes in case you need to provide an explicit encoding:

bytes = File.binread("page.html")
doc = Nokogiri::HTML(bytes)

Nokogiri can also parse IO input. For network responses, it is usually clearer to fetch the response yourself, check its status and content type, set timeouts, and then pass the body to Nokogiri. That separation keeps network retries and response limits under your control rather than treating fetching as part of parsing.

Use an appropriate parser name

Nokogiri::HTML is the common convenience interface. The explicit Nokogiri::HTML4 namespace makes the HTML4 parser choice clear, while Nokogiri::HTML5 selects HTML5 parsing behavior. If your code depends on the exact parser mode, name it explicitly rather than relying on an implicit choice.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose CSS selectors or XPath

Both query styles can locate nodes in the parsed document. CSS is often easier to read when the target is described by element names, classes, IDs, or descendant relationships:

cards = doc.css("article.card")
links = doc.css("nav ul.menu li a")
first_heading = doc.at_css("main h1")

XPath is useful when a query depends on structural relationships, predicates, or attributes. For example:

headings = doc.xpath("//article//h2")
https_links = doc.xpath("//a[starts-with(@href, 'https://')]")
first_href = doc.at_xpath("//article//a/@href")&.&value

Use css or xpath when you expect zero or more matches; they return a collection of nodes. Use at_css or at_xpath when one match is sufficient; these return a node or nil. The safe-navigation operators in the first example prevent a missing heading or attribute from raising an error.

Get text and attributes deliberately

For an attribute, a node supports bracket access, such as link["href"]. For text, node.text returns the text content, including text from descendants. Decide how to handle whitespace for the specific field; strip removes leading and trailing whitespace but does not normalize whitespace within the string.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
link = doc.at_css("article a")
if link
  href = link["href"]
  label = link.text.strip
end

If a page has repeated or nested structures, test the selector against representative markup. A query that returns the first matching node can silently select the wrong element if the document contains multiple candidates.

Mix query styles with search

doc.search can accept CSS and XPath expressions when an extraction benefits from both forms. Keep mixed queries readable: split them into named intermediate results if a single expression becomes difficult to understand or maintain.

Decide between HTML4 and HTML5 parsing

Choose the parser based on the kind of tree construction your application needs, not merely on the label of the source document. HTML5 parsing is the better fit when browser-compatible HTML5 parsing behavior matters. HTML4 parsing remains a practical choice for ordinary extraction where that behavior is not needed.

Choice Use it when Important qualification
Nokogiri::HTML or Nokogiri::HTML4 You need the standard HTML parser flow for page extraction and do not require HTML5 tree-construction behavior. HTML4 and HTML5 parsing can construct different trees from malformed or modern markup.
Nokogiri::HTML5 Browser-compatible HTML5 parsing behavior is important to the result. HTML5 functionality is documented as unavailable on JRuby; confirm your runtime before depending on it.

Malformed markup is common on the web. If a selector unexpectedly misses an element, inspect the parsed tree as well as the original source: a parser may repair or rearrange invalid nesting. When tree shape matters to downstream logic, test with the same parser mode and runtime used in production.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set limits for HTML5 parsing when needed

The HTML5 API documents controls including max_errors, max_tree_depth, and max_attributes. These are useful when processing very large or potentially hostile input; choose limits that fit your workload rather than assuming every input should be unrestricted. See the Nokogiri documentation for API details and supported options.

Parse snippets as fragments

A snippet such as a list of <li> elements is not a full page. Use a fragment parser when the input is meant to be a piece of markup rather than a complete document:

fragment = Nokogiri::HTML.fragment("<li>One</li><li>Two</li>")
items = fragment.css("li").map { |item| item.text.strip }

puts items

For HTML5 fragment parsing, use Nokogiri::HTML5.fragment where HTML5 behavior is required and supported by your Ruby runtime:

fragment = Nokogiri::HTML5.fragment("<li>One</li><li>Two</li>")
items = fragment.css("li").map { |item| item.text.strip }

Fragments avoid pretending that a snippet has the context of a whole page. Context can affect how markup is interpreted, so use a full document when page-level structure is part of what you need to examine.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle character encoding explicitly

Nokogiri stores text internally as UTF-8, and methods that return text produce UTF-8 strings. If a document’s declared character set does not match its bytes, do not rely on autodetection; provide the known encoding when parsing.

encoded = File.binread("page.html")
doc = Nokogiri::HTML4.parse(encoded, nil, "EUC-JP")
puts doc.at_css("body")&.&text

Keep the original bytes until parsing so the parser can interpret them using the encoding you specify. When integrating a new source, test representative non-ASCII text, not just ASCII headings: ASCII often looks correct even when a character-set mismatch is present.

Fetch remote HTML safely before parsing

Nokogiri parses markup; it is not a substitute for a network client’s response handling. Fetch the page with an HTTP client you control, then pass the response body to the parser. At minimum, design that fetch step to:

  • Set connection and read timeouts so a slow server cannot hold a worker indefinitely.
  • Limit response size before parsing, especially when the URL or page is not trusted.
  • Check the HTTP status and content type instead of assuming every response is HTML.
  • Handle redirects and retries according to your application’s policy.

The exact Ruby HTTP client and its API depend on your application; no particular client is required by Nokogiri’s parse-then-query workflow. Keeping fetch and parse separate also makes it easier to test parsing against saved response bodies without making a live network request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Treat parsed markup and extracted values as untrusted

Nokogiri’s documented security principle is to treat documents as untrusted by default. Parsing does not validate application data, sanitize markup for display, or make a URL safe to fetch. After extraction, validate values against the rules of your application.

  • Check that required elements exist before using their values.
  • Validate URL schemes and hosts before making requests from extracted links.
  • Validate dates, numbers, and identifiers before storing or acting on them.
  • Do not render extracted HTML without a sanitizer appropriate to the output context.
  • Use HTML5 depth and attribute limits where applicable for hostile or very large input.

For example, finding an <a> element does not establish that its href is a safe destination. Parsing gives you structure to inspect; validation and output-context sanitization remain application responsibilities.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common Nokogiri parsing problems

A CSS or XPath query returns no nodes

  • Cause: The selector does not match the parsed tree, the markup differs from expectations, or the content is rendered dynamically after the server response.
  • Fix: Inspect the response body and parsed structure, confirm the class or attribute spelling, and try a narrow selector against the actual markup. Nokogiri parses the bytes you give it; it does not run a browser’s JavaScript.

A single-node query returns nil

  • Cause: at_css or at_xpath found no match.
  • Fix: Guard the result with an explicit conditional or safe navigation, and decide whether absence is expected or should raise an application-level error.

Text contains unexpected whitespace

  • Cause: text includes whitespace and descendant text from the markup.
  • Fix: Use strip when only leading and trailing whitespace should be removed. Apply a more specific normalization only if the field’s meaning permits it.

Non-ASCII characters appear corrupted

  • Cause: The source bytes and declared or detected encoding do not agree.
  • Fix: Preserve the bytes, identify the actual source encoding, and pass it explicitly to the parser. Verify output with representative characters from that source.

HTML5 constants or methods are unavailable

  • Cause: The runtime may be JRuby, where Nokogiri documents HTML5 functionality as unavailable.
  • Fix: Confirm the runtime and parser support for your deployment target. Use the HTML4 parser if its behavior is acceptable, or select an environment that supports the HTML5 API.

The response is an error page or not HTML

  • Cause: The fetch succeeded at the transport level but returned an unexpected status, content type, or page body.
  • Fix: Check status and content type before parsing, and log a bounded portion of the response for diagnosis without exposing secrets or personal data.

Performance, reliability, and cost considerations

Parsing a local string is usually a different workload from fetching and parsing a large remote page. For reliable extraction, cap response size, select only the nodes you need, and avoid repeatedly parsing the same body. If inputs are untrusted, parser limits and application-level timeouts and size checks matter as much as selector efficiency. Nokogiri itself does not determine the cost of your network requests or validate the meaning of extracted content.

If the source page depends on client-side JavaScript to create the target content, parsing the initial HTML response will not produce that browser-generated DOM. You need a rendered-page capture or browser workflow for that case, rather than expecting an HTML parser to execute scripts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is to obtain a clean screenshot of a URL rather than query its DOM in Ruby, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It removes known consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

For an API request, use the following cURL example, replacing the target URL and supplying your API key:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Frequently Asked Questions

Does Nokogiri execute JavaScript in a page?

No. Nokogiri parses the markup supplied to it; it does not run browser JavaScript to create a rendered DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use HTML5 parsing on JRuby?

Nokogiri documents its HTML5 functionality as unavailable on JRuby. Check runtime support before choosing that API.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.