Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a static HTML document, use Nokogiri to find the intended <table>, select each row’s <th> and <td> cells, and collect their text. That produces an array of the cells actually present in the document—not automatically a rectangular grid when the table uses rowspan or colspan. This guide shows the basic extraction, CSV export, parser choices, and the extra work needed for irregular tables.

Extract table rows with Nokogiri

Nokogiri parses HTML and supports both CSS and XPath searches. The CSS example below reads a local HTML file, finds a table by ID, and returns one array per row. Replace table#results with a selector matching the table you want.

require "nokogiri"

html = File.read("page.html")
doc = Nokogiri::HTML(html)

table = doc.at_css("table#results")
raise "table not found" unless table

rows = table.css("tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

p rows

For example, a header row containing Name and Score, followed by a row containing Ada and 98, becomes [["Name", "Score"], ["Ada", "98"]]. The output is an array of arrays: each inner array contains the text from the selected cells in that row.

Install Nokogiri

Add Nokogiri to a project with Bundler by declaring it in the Gemfile:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
gem "nokogiri"

Then run bundle install and execute the script through bundle exec ruby your_script.rb. For a one-off script, install the gem with gem install nokogiri and run it with ruby your_script.rb. Nokogiri’s installation requirements can depend on the Ruby runtime and platform; consult the project documentation if installation fails.

Choose a selector that identifies the right table

Pages often contain multiple tables, so avoid assuming the first table is the data you need. A stable ID or class is usually more precise than a broad selector. For example, use doc.at_css("table.prices") for a table with class prices. If a selector matches multiple tables, use doc.css("table.prices") and select the intended match deliberately.

You can express the same target with XPath:

table = doc.at_xpath("//table[@id='results']")

CSS tends to be concise for classes, IDs, and descendant relationships; XPath is useful when the target depends on an attribute or structural relationship. Nokogiri supports both, so choose the form that makes the page-specific target clearest.

Understand what the extracted rows mean

The basic code reads the cells present in the DOM. It does not infer a spreadsheet-like grid, headers, or missing positions. That distinction matters when the visual table combines cells or uses nested markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Header cells and data cells

Selecting th, td includes both header and data cells in document order. If you want only body records, inspect the document structure and scope the selector to tbody, for example:

rows = table.css("tbody tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

Some HTML omits an explicit tbody in source, while an HTML parser may represent table sections in a normalized tree. Inspect the parsed structure and adjust the selector for the actual input. If the table uses multiple header rows or row-header cells, decide explicitly whether those belong in the output.

Rowspan and colspan do not create empty grid cells

A cell with colspan="2" still appears once in the extracted row; a cell with rowspan="2" appears in its source row rather than being copied into the following row. The simple extraction therefore gives unequal row lengths when the source uses spans. This can be correct if you want the DOM cells, but it is not a normalized rectangular table.

If downstream code requires fixed columns, implement a grid-normalization pass: track occupied columns from active row spans, place each cell at the next unoccupied position, and reserve the appropriate number of columns for its colspan. Treat malformed span values carefully and test against representative input. Nokogiri’s cell selection does not perform that transformation for you.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Text extraction and nested elements

cell.text.strip returns the text content of a cell, including text inside nested elements, with leading and trailing whitespace removed. It does not preserve markup or necessarily retain the visual spacing created by line breaks. If a cell contains links, badges, or multiple paragraphs, inspect the extracted value and decide whether you need plain text, a particular descendant’s text, or structured attributes such as link destinations.

Choose HTML parsing behavior for your runtime

Nokogiri::HTML(html) is the straightforward HTML parsing entry point for the pattern above. Nokogiri also documents an HTML5 parser API. Its documentation says HTML5 parsing is available since Nokogiri 1.12.0 and is not available on JRuby; do not assume an HTML5-specific call works across every runtime.

When HTML5 parsing is appropriate

If HTML5 parsing better suits the source document or its error-recovery behavior, use the HTML5 API supported by your installed Nokogiri version and runtime. Check the current API documentation for call signatures and options rather than copying an API intended for another version. The HTML5 parser documents controls including parse-error reporting, maximum tree depth, and maximum attributes per element.

When the runtime is JRuby

Nokogiri’s documented HTML5 functionality is unavailable on JRuby. Use a supported parser API for the installed Nokogiri/runtime combination and verify the selectors against your actual input. Parser implementations and behavior can differ between CRuby and JRuby; for reproducible extraction, record the Ruby version, Nokogiri version, and parser choice alongside your tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Export extracted values as CSV

Extraction and CSV serialization are separate tasks. Do not join cells with commas yourself: values can contain commas, quotation marks, or line breaks, all of which require correct CSV escaping. Use Ruby’s standard CSV library.

require "csv"
require "nokogiri"

html = File.read("page.html")
doc = Nokogiri::HTML(html)
table = doc.at_css("table#results")
raise "table not found" unless table

rows = table.css("tr").map do |row|
  row.css("th, td").map { |cell| cell.text.strip }
end

CSV.open("results.csv", "w", encoding: "UTF-8") do |csv|
  rows.each { |row| csv << row }
end

This writes each extracted row as a CSV record. It does not decide whether the first row is a header or repair rows affected by spans. If the data has a header and you want a table-oriented representation, Ruby’s CSV library also provides header-aware parsing and CSV::Table row and column operations; choose that structure when you need named access rather than only serialized records.

Handle encoding and untrusted HTML safely

Nokogiri documents returned text as UTF-8. Check non-ASCII names, symbols, and punctuation in the resulting file, especially when the source encoding is uncertain. The HTML5 parser documentation describes UTF-8 parsing and an optional encoding parameter, notably relevant when parsing IO. Make the source encoding assumption explicit when it affects the result, and verify the output rather than assuming every upstream document declares or uses the encoding you expect.

Nokogiri’s documented defaults treat input as untrusted: parsing does not load external DTDs or access the network for external resources. Keep those protections enabled. Avoid enabling entity or DTD behavior or disabling network protections for user-supplied or scraped HTML. These parser protections concern processing the supplied document; they do not grant permission to fetch a website or bypass its access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Capture a web page before parsing its table

Nokogiri parses HTML you already have. It does not, by itself, fetch a live page, execute its JavaScript, or wait for a client-rendered table to appear. If you need to obtain a page, first use a permitted retrieval method and confirm that the response contains the table markup. If the table is created only after browser-side scripts run, a static HTML response may not contain it; capture or retrieve the rendered document through an appropriate browser workflow, then parse that HTML.

For browser-based capture, distinguish getting a screenshot from getting DOM data: an image of a table is not HTML for Nokogiri to parse. ScreenshotNeo is a website screenshot API and MCP server, useful when the desired result is a screenshot or PDF rather than extracted cell values.

Or skip the browser setup

For a screenshot or PDF, ScreenshotNeo takes a URL in one request. See the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common extraction problems

The script raises “table not found”

The selector may not match the parsed document, the table may have a different ID or class, or the table may be absent from the HTML you read. Inspect the input and try a broader diagnostic such as doc.css("table").length; if it is zero, the issue is not the row selector. For client-rendered content, obtain HTML that actually includes the rendered table before parsing.

Rows have different numbers of values

Check for rowspan, colspan, header rows, and rows with genuinely missing cells. The basic pattern preserves present cells; it does not fill omitted grid positions. If a rectangular result is required, normalize spans explicitly and decide how to represent genuinely empty cells.

Values contain unexpected whitespace or run together

text.strip trims only the ends of the combined text. Nested markup and line breaks may not map to the visual spacing you expect. Inspect the cell’s child elements and choose a deliberate whitespace-normalization rule, rather than silently changing all internal spacing.

Non-ASCII characters look wrong in the output

Verify the source encoding, parser choice, and output encoding. Nokogiri returns text as UTF-8, but incorrect assumptions about the original bytes or a downstream consumer’s encoding can still produce bad output. Test representative names and symbols through the entire parse-and-write path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The HTML5 parser call fails

Confirm your Nokogiri version and Ruby runtime. HTML5 parsing is documented from v1.12.0 and is unavailable on JRuby. Use an API supported by your installed combination, and consult the current Nokogiri documentation for option names and behavior.

The CSV columns shift or files open incorrectly

CSV escaping prevents commas and quotes in a value from corrupting record boundaries, but it does not repair an irregular table grid. Normalize spans before serialization if fixed columns are required, and confirm the encoding expected by the application opening the CSV.

Performance and reliability considerations

For a single document, parsing once and scoping searches to the target table keeps the code easy to reason about. Avoid repeatedly searching the entire document for each individual cell when you can select rows and cells from the chosen table. For a batch workflow, measure with your own inputs: page size, markup complexity, and parser/runtime can affect cost, and the documentation cited here does not establish a universal speed figure.

HTML parsers are designed to handle imperfect markup, but a successful parse does not prove that the selected cells represent the intended business data. Validate a few expected row counts, headers, and representative values, particularly when the source page changes. Store runtime and parser versions in reproducible jobs, and fail clearly when the target table disappears instead of emitting an apparently valid empty file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently asked questions

Can Nokogiri extract a table directly into a spreadsheet grid?

It can provide the DOM cells, but the basic row-and-cell selection does not expand row and column spans into a rectangular grid. Add normalization logic if you need spreadsheet-style positions.

Does parsing a page with Nokogiri execute JavaScript?

No. Nokogiri parses HTML; JavaScript-rendered table content must be present in the HTML you give it or obtained through a browser-rendering step first.

Can I use XPath instead of CSS?

Yes. Nokogiri supports both. CSS is often concise for common selectors, while XPath can express structural relationships or attribute conditions that suit a particular document.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.