Free tools Windows power users keep installed
One-click scans. No signup required.
For a static HTML document, use Nokogiri to find the intended <table>, select each row’s <th> and <td> cells, and collect their text. That produces an array of the cells actually present in the document—not automatically a rectangular grid when the table uses rowspan or colspan. This guide shows the basic extraction, CSV export, parser choices, and the extra work needed for irregular tables.
Extract table rows with Nokogiri
Nokogiri parses HTML and supports both CSS and XPath searches. The CSS example below reads a local HTML file, finds a table by ID, and returns one array per row. Replace table#results with a selector matching the table you want.
require "nokogiri"
html = File.read("page.html")
doc = Nokogiri::HTML(html)
table = doc.at_css("table#results")
raise "table not found" unless table
rows = table.css("tr").map do |row|
row.css("th, td").map { |cell| cell.text.strip }
end
p rows
For example, a header row containing Name and Score, followed by a row containing Ada and 98, becomes [["Name", "Score"], ["Ada", "98"]]. The output is an array of arrays: each inner array contains the text from the selected cells in that row.
Install Nokogiri
Add Nokogiri to a project with Bundler by declaring it in the Gemfile:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems#1 Best Overall
gem "nokogiri"
Then run bundle install and execute the script through bundle exec ruby your_script.rb. For a one-off script, install the gem with gem install nokogiri and run it with ruby your_script.rb. Nokogiri’s installation requirements can depend on the Ruby runtime and platform; consult the project documentation if installation fails.
Choose a selector that identifies the right table
Pages often contain multiple tables, so avoid assuming the first table is the data you need. A stable ID or class is usually more precise than a broad selector. For example, use doc.at_css("table.prices") for a table with class prices. If a selector matches multiple tables, use doc.css("table.prices") and select the intended match deliberately.
You can express the same target with XPath:
table = doc.at_xpath("//table[@id='results']")
CSS tends to be concise for classes, IDs, and descendant relationships; XPath is useful when the target depends on an attribute or structural relationship. Nokogiri supports both, so choose the form that makes the page-specific target clearest.
Understand what the extracted rows mean
The basic code reads the cells present in the DOM. It does not infer a spreadsheet-like grid, headers, or missing positions. That distinction matters when the visual table combines cells or uses nested markup.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteHeader cells and data cells
Selecting th, td includes both header and data cells in document order. If you want only body records, inspect the document structure and scope the selector to tbody, for example:
rows = table.css("tbody tr").map do |row|
row.css("th, td").map { |cell| cell.text.strip }
end
Some HTML omits an explicit tbody in source, while an HTML parser may represent table sections in a normalized tree. Inspect the parsed structure and adjust the selector for the actual input. If the table uses multiple header rows or row-header cells, decide explicitly whether those belong in the output.
Rank #2
Rowspan and colspan do not create empty grid cells
A cell with colspan="2" still appears once in the extracted row; a cell with rowspan="2" appears in its source row rather than being copied into the following row. The simple extraction therefore gives unequal row lengths when the source uses spans. This can be correct if you want the DOM cells, but it is not a normalized rectangular table.
If downstream code requires fixed columns, implement a grid-normalization pass: track occupied columns from active row spans, place each cell at the next unoccupied position, and reserve the appropriate number of columns for its colspan. Treat malformed span values carefully and test against representative input. Nokogiri’s cell selection does not perform that transformation for you.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Text extraction and nested elements
cell.text.strip returns the text content of a cell, including text inside nested elements, with leading and trailing whitespace removed. It does not preserve markup or necessarily retain the visual spacing created by line breaks. If a cell contains links, badges, or multiple paragraphs, inspect the extracted value and decide whether you need plain text, a particular descendant’s text, or structured attributes such as link destinations.
Choose HTML parsing behavior for your runtime
Nokogiri::HTML(html) is the straightforward HTML parsing entry point for the pattern above. Nokogiri also documents an HTML5 parser API. Its documentation says HTML5 parsing is available since Nokogiri 1.12.0 and is not available on JRuby; do not assume an HTML5-specific call works across every runtime.
When HTML5 parsing is appropriate
If HTML5 parsing better suits the source document or its error-recovery behavior, use the HTML5 API supported by your installed Nokogiri version and runtime. Check the current API documentation for call signatures and options rather than copying an API intended for another version. The HTML5 parser documents controls including parse-error reporting, maximum tree depth, and maximum attributes per element.
When the runtime is JRuby
Nokogiri’s documented HTML5 functionality is unavailable on JRuby. Use a supported parser API for the installed Nokogiri/runtime combination and verify the selectors against your actual input. Parser implementations and behavior can differ between CRuby and JRuby; for reproducible extraction, record the Ruby version, Nokogiri version, and parser choice alongside your tests.
Rank #3
Export extracted values as CSV
Extraction and CSV serialization are separate tasks. Do not join cells with commas yourself: values can contain commas, quotation marks, or line breaks, all of which require correct CSV escaping. Use Ruby’s standard CSV library.
require "csv"
require "nokogiri"
html = File.read("page.html")
doc = Nokogiri::HTML(html)
table = doc.at_css("table#results")
raise "table not found" unless table
rows = table.css("tr").map do |row|
row.css("th, td").map { |cell| cell.text.strip }
end
CSV.open("results.csv", "w", encoding: "UTF-8") do |csv|
rows.each { |row| csv << row }
end
This writes each extracted row as a CSV record. It does not decide whether the first row is a header or repair rows affected by spans. If the data has a header and you want a table-oriented representation, Ruby’s CSV library also provides header-aware parsing and CSV::Table row and column operations; choose that structure when you need named access rather than only serialized records.
Handle encoding and untrusted HTML safely
Nokogiri documents returned text as UTF-8. Check non-ASCII names, symbols, and punctuation in the resulting file, especially when the source encoding is uncertain. The HTML5 parser documentation describes UTF-8 parsing and an optional encoding parameter, notably relevant when parsing IO. Make the source encoding assumption explicit when it affects the result, and verify the output rather than assuming every upstream document declares or uses the encoding you expect.
Nokogiri’s documented defaults treat input as untrusted: parsing does not load external DTDs or access the network for external resources. Keep those protections enabled. Avoid enabling entity or DTD behavior or disabling network protections for user-supplied or scraped HTML. These parser protections concern processing the supplied document; they do not grant permission to fetch a website or bypass its access controls.
Capture a web page before parsing its table
Nokogiri parses HTML you already have. It does not, by itself, fetch a live page, execute its JavaScript, or wait for a client-rendered table to appear. If you need to obtain a page, first use a permitted retrieval method and confirm that the response contains the table markup. If the table is created only after browser-side scripts run, a static HTML response may not contain it; capture or retrieve the rendered document through an appropriate browser workflow, then parse that HTML.
For browser-based capture, distinguish getting a screenshot from getting DOM data: an image of a table is not HTML for Nokogiri to parse. ScreenshotNeo is a website screenshot API and MCP server, useful when the desired result is a screenshot or PDF rather than extracted cell values.
Rank #4
Or skip the browser setup
For a screenshot or PDF, ScreenshotNeo takes a URL in one request. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.
Troubleshoot common extraction problems
The script raises “table not found”
The selector may not match the parsed document, the table may have a different ID or class, or the table may be absent from the HTML you read. Inspect the input and try a broader diagnostic such as doc.css("table").length; if it is zero, the issue is not the row selector. For client-rendered content, obtain HTML that actually includes the rendered table before parsing.
Rows have different numbers of values
Check for rowspan, colspan, header rows, and rows with genuinely missing cells. The basic pattern preserves present cells; it does not fill omitted grid positions. If a rectangular result is required, normalize spans explicitly and decide how to represent genuinely empty cells.
Values contain unexpected whitespace or run together
text.strip trims only the ends of the combined text. Nested markup and line breaks may not map to the visual spacing you expect. Inspect the cell’s child elements and choose a deliberate whitespace-normalization rule, rather than silently changing all internal spacing.
Non-ASCII characters look wrong in the output
Verify the source encoding, parser choice, and output encoding. Nokogiri returns text as UTF-8, but incorrect assumptions about the original bytes or a downstream consumer’s encoding can still produce bad output. Test representative names and symbols through the entire parse-and-write path.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →The HTML5 parser call fails
Confirm your Nokogiri version and Ruby runtime. HTML5 parsing is documented from v1.12.0 and is unavailable on JRuby. Use an API supported by your installed combination, and consult the current Nokogiri documentation for option names and behavior.
Best Value
The CSV columns shift or files open incorrectly
CSV escaping prevents commas and quotes in a value from corrupting record boundaries, but it does not repair an irregular table grid. Normalize spans before serialization if fixed columns are required, and confirm the encoding expected by the application opening the CSV.
Performance and reliability considerations
For a single document, parsing once and scoping searches to the target table keeps the code easy to reason about. Avoid repeatedly searching the entire document for each individual cell when you can select rows and cells from the chosen table. For a batch workflow, measure with your own inputs: page size, markup complexity, and parser/runtime can affect cost, and the documentation cited here does not establish a universal speed figure.
HTML parsers are designed to handle imperfect markup, but a successful parse does not prove that the selected cells represent the intended business data. Validate a few expected row counts, headers, and representative values, particularly when the source page changes. Store runtime and parser versions in reproducible jobs, and fail clearly when the target table disappears instead of emitting an apparently valid empty file.
Recommended Free Tools
Frequently asked questions
Can Nokogiri extract a table directly into a spreadsheet grid?
It can provide the DOM cells, but the basic row-and-cell selection does not expand row and column spans into a rectangular grid. Add normalization logic if you need spreadsheet-style positions.
Does parsing a page with Nokogiri execute JavaScript?
No. Nokogiri parses HTML; JavaScript-rendered table content must be present in the HTML you give it or obtained through a browser-rendering step first.
Can I use XPath instead of CSS?
Yes. Nokogiri supports both. CSS is often concise for common selectors, while XPath can express structural relationships or attribute conditions that suit a particular document.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

