Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Data extraction in Ruby starts by identifying the input format. Use Ruby’s text and regular-expression tools for genuinely line-oriented text, the standard-library JSON parser for JSON, YAML/Psych for YAML, and Nokogiri for HTML or XML. The examples below target a current Ruby runtime whose documentation matches the Ruby 4.0 documentation index; check the documentation set for the Ruby version you actually run because APIs and defaults can differ between releases and implementations.

Choose the parser from the data format

Do not choose a parser because its name is familiar. A question such as “I’m trying to scrap this json file with nokogiri” mixes two different formats: Nokogiri is for markup, while JSON needs a JSON decoder. Likewise, regular expressions can split a controlled text report, but they are not a general HTML or XML parser.

Input Ruby path Best fit
Line-oriented text String, File, and regular expressions Fixed records with a known delimiter or line shape
JSON Ruby’s JSON standard library Objects and arrays exchanged by APIs or files
YAML YAML/Psych Configuration and human-edited documents, with strict trust controls
HTML/XML Nokogiri DOM queries, XPath, CSS selectors, or streaming SAX/push parsing

Ruby’s official FAQ describes Ruby as good at text processing and demonstrates line-by-line regular-expression parsing. Treat that as guidance for bounded text formats, not permission to parse arbitrary markup with regex.

Prepare a Ruby project

Use the Ruby documentation version that matches your interpreter. Confirm the runtime before debugging parser behavior:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ruby --version

Create a project and add Nokogiri only when you need HTML or XML:

mkdir ruby_extract
cd ruby_extract
bundle init
bundle add nokogiri

The JSON library is part of Ruby’s standard library. YAML is provided through Psych, also documented in Ruby’s standard-library index. Requiring the library explicitly makes dependencies clear:

require "json"
require "yaml"
require "nokogiri"

Extract records from simple text

For a report with one record per line, keep the grammar explicit and reject malformed lines instead of silently producing bad data. This example reads name|email|plan records:

records = []

File.foreach("users.txt", chomp: true).with_index(1) do |line, number|
  next if line.empty?

  match = line.match(/A([^|]+)|([^|]+)|([^|]+)z/)
  unless match
    warn "Skipping malformed line #{number}"
    next
  end

  name, email, plan = match.captures
  records << { "name" => name, "email" => email, "plan" => plan }
end

p records

File.foreach keeps memory use bounded for large files. Add validation for required fields, allowed plan names, and encoding before storing records. If the format gains quoting, escaping, nested fields, or multiline values, move to a format-aware parser instead of extending a regular expression indefinitely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse JSON with Ruby’s JSON library

Decode a file

require "json"

text = File.read("payload.json", encoding: "UTF-8")
data = JSON.parse(text)

case data
when Hash
  puts data.fetch("customer_id")
when Array
  data.each { |row| puts row["id"] }
else
  warn "Expected an object or array"
end

JSON.parse returns Ruby hashes, arrays, strings, numbers, booleans, and nil. Access optional keys with fetch when absence is an error, or row["field"] when absence is acceptable. Never assume an API always returns the same top-level type; validate it before iterating.

Stream newline-delimited JSON

For one JSON object per line, parse each line independently so a large export does not become one large in-memory object:

require "json"

File.foreach("events.ndjson", chomp: true).with_index(1) do |line, number|
  next if line.empty?

  begin
    event = JSON.parse(line)
    puts event.fetch("type")
  rescue JSON::ParserError, KeyError => e
    warn "Line #{number}: #{e.message}"
  end
end

Write extracted data as JSON

File.write("clean.json", JSON.pretty_generate(records), mode: "w", encoding: "UTF-8")

Keep JSON parsing separate from extraction rules. First decode the document, then normalize fields in ordinary Ruby methods; this makes malformed input and schema changes easier to test.

Parse YAML with Psych

YAML can represent richer types than JSON and can contain tags that cause unsafe behavior if handled carelessly. Treat YAML as untrusted unless you control its origin and understand the permitted classes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Safe loading

require "yaml"

text = File.read("settings.yml", encoding: "UTF-8")
settings = YAML.safe_load(text, permitted_classes: [], aliases: false)

host = settings.fetch("host")
port = Integer(settings.fetch("port"))
puts "#{host}:#{port}"

Use the strictest safe-loading options compatible with your document. If aliases or specific classes are genuinely required, permit only those explicitly and document why. Do not deserialize arbitrary YAML from users directly into application objects.

Emit YAML

File.write("export.yml", records.to_yaml, encoding: "UTF-8")

When YAML is only an interchange format, JSON is often easier to constrain and validate. Choose YAML for the properties that require it, not because it looks like a convenient hash literal.

Extract HTML with Nokogiri

Install and parse a document

require "nokogiri"
require "open-uri"

html = URI.open("https://example.com", read_timeout: 20).read
doc = Nokogiri::HTML5(html)

puts doc.at_css("title")&.text&.strip
links = doc.css("a").filter_map do |link|
  href = link["href"]
  next if href.nil? || href.empty?
  { "text" => link.text.strip, "href" => href }
end
p links

Nokogiri documents DOM parsers for HTML4 and HTML5 and supports CSS3 selectors and XPath 1.0 queries. DOM parsing is convenient when you need to navigate repeatedly, inspect ancestors, or extract related fields from each card.

Use CSS selectors for readable queries

doc.css("article.product").each do |article|
  name = article.at_css("h2")&.text&.strip
  price = article.at_css(".price")&.text&.strip
  puts [name, price].compact.join(" - ")
end

Use XPath for structural conditions

doc.xpath("//table//tr[td]").each do |row|
  cells = row.xpath("./td").map { |cell| cell.text.strip }
  puts cells.inspect
end

CSS is usually clearer for classes and simple descendants. XPath is useful when selection depends on position, attributes, or relationships such as “the row whose first cell contains this label.” Test selectors against representative documents; malformed HTML can be repaired differently by parser implementations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract XML

require "nokogiri"

xml = File.read("feed.xml", encoding: "UTF-8")
doc = Nokogiri::XML(xml) { |config| config.strict.nonet }

doc.xpath("//*[local-name()='item']").each do |item|
  title = item.at_xpath("./*[local-name()='title']")&.text&.strip
  id = item.at_xpath("./*[local-name()='id']")&.text&.strip
  puts({ "id" => id, "title" => title }.inspect)
end

Namespaces matter in XML. Prefer a namespace-aware XPath when the document declares stable prefixes; local-name() is a fallback when prefixes vary, but it can match elements from unrelated namespaces.

DOM, SAX, and push parsing: which Nokogiri mode?

Nokogiri also documents SAX and push parsing for XML and HTML4. Choose based on memory and access needs:

  • DOM: loads a tree, making CSS/XPath navigation and related-field extraction straightforward. It uses more memory for very large documents.
  • SAX: calls callbacks as events arrive. It is suitable for large, regular streams when you can process each element without navigating backward.
  • Push parsing: lets your application feed chunks to the parser as they arrive, useful when input is already streamed from a socket or download.

There is no universally best mode in the Nokogiri documentation. Benchmark your real document shape and extraction logic. Native parser behavior can differ between CRuby and JRuby, so do not assume identical edge-case results without checking the runtime and Nokogiri version you deploy.

Security and encoding safeguards

Assume documents are untrusted

Nokogiri’s guiding principles say it should be secure by default by treating all documents as untrusted. Keep network access, entity expansion, and dangerous features disabled unless your use case requires them and you understand the consequence. The nonet option in the XML example prevents network access while parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing safely does not make the extracted content safe to display. Escape values when inserting them into HTML, validate URLs before fetching them, and impose limits on input size and processing time.

Set encoding when it matters

Data is a stream of bytes, and Nokogiri’s documentation notes that 100% accurate encoding detection is impossible. If the source encoding is known, tell the parser explicitly:

bytes = File.binread("legacy.html")
doc = Nokogiri::HTML4.parse(bytes, nil, "Windows-1252")
puts doc.text.encode("UTF-8", invalid: :replace, undef: :replace)

Record the source encoding in your ingestion metadata. Guessing from visible characters can corrupt names, currency symbols, or identifiers without raising an obvious error.

Build a repeatable extraction pipeline

  1. Identify the format. Inspect headers, file extensions, and a small sample; do not infer JSON from a page that merely displays JSON-looking text.
  2. Choose the parser. Use JSON, YAML/Psych, Nokogiri, or line processing according to the table above.
  3. Normalize. Strip presentation whitespace, convert numeric fields deliberately, and preserve original values when auditing matters.
  4. Validate. Check required keys, types, ranges, encodings, and duplicate identifiers.
  5. Emit a stable shape. Return hashes with documented keys or write a versioned JSON schema.
  6. Test failures. Include empty input, malformed syntax, missing fields, unexpected encodings, namespaces, and very large files.

Performance, reliability, and cost considerations

  • Read line-oriented text and NDJSON incrementally with foreach.
  • Use DOM only when tree navigation pays for its memory cost; use SAX or push parsing for large streams.
  • Set network timeouts before downloading HTML. Cache source documents when reproducibility matters.
  • Separate fetching from parsing so a transient HTTP failure cannot be mistaken for an extraction bug.
  • Log parser errors with the source identifier and line or record number, but avoid logging secrets or entire private documents.
  • Pin and review Ruby and Nokogiri versions in deployment. Native parser libraries and CRuby/JRuby differences can affect malformed-input handling.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“undefined method css”

You probably have a string, not a Nokogiri document. Parse first with Nokogiri::HTML5(html), Nokogiri::HTML4(html), or Nokogiri::XML(xml).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JSON raises JSON::ParserError

Log the byte offset and inspect the raw response. Common causes are an HTML error page, a trailing comma, incorrect character encoding, or a JSON object embedded inside another format. Do not pass JSON to Nokogiri.

Selectors return nothing

Print doc.to_html or a small matching node, verify whether the content is generated by JavaScript, and check namespaces for XML. A browser’s developer tools may show a post-JavaScript DOM that was never present in the downloaded source.

Characters are garbled

Determine the source encoding and pass it explicitly to Nokogiri, then transcode to UTF-8 at your application boundary. Check that the HTTP response headers and the document declaration agree.

Memory usage grows on large files

Replace whole-document DOM parsing with line iteration, SAX callbacks, or push parsing. Also avoid collecting every extracted record when you can write or process records incrementally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The page is blocked or incomplete

Respect the site’s access rules, use realistic timeouts, and distinguish a network failure from a valid but empty document. If content requires browser rendering, a plain HTTP fetch may not contain the data you see in a browser.

Or skip the browser setup

When your extraction workflow also needs a clean screenshot of a rendered page, ScreenshotNeo provides a single HTTP request. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Ruby can call the same endpoint without browser automation:

require "net/http"
require "uri"

uri = URI("https://api.screenshotneo.com/v1/shot")
uri.query = URI.encode_www_form(access_key: "YOUR_API_KEY", url: "https://stripe.com")
response = Net::HTTP.get_response(uri)
abort "HTTP #{response.code}" unless response.is_a?(Net::HTTPSuccess)
File.binwrite("shot.webp", response.body)

See the complete option list and authentication details in the ScreenshotNeo documentation. Python and Node.js clients are also straightforward:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is included on every plan. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can Nokogiri parse JSON?

No. Decode JSON with Ruby’s JSON library; use Nokogiri for HTML or XML markup.

Should I use CSS or XPath in Nokogiri?

Use whichever expresses the condition clearly: CSS for common class and descendant queries, XPath for structural and positional conditions.

Is YAML safe to load from users?

Only with restrictive safe-loading settings and explicit permitted classes or aliases. Treat untrusted YAML as data, not executable object state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does browser content differ from Nokogiri output?

Nokogiri parses the response body. JavaScript-rendered content may require a browser-capable fetch or a service that captures the rendered page.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.