Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Importing an HTML file in Rust is a two-step operation: read the path into text (usually with std::fs::read_to_string), then give that text to an HTML parser such as scraper::Html::parse_document. Use parse_fragment for an HTML snippet, std::fs::read when you need byte-level encoding control, Kuchiki for a mutable DOM-like tree, and html5ever when you need a lower-level HTML5 parser.
The shortest working solution
Create a Rust project and add the parser dependency:
cargo new html_import
cd html_import
cargo add scraper
Put an HTML document at page.html, then replace src/main.rs with:
Recommended Free Tools
use scraper::{Html, Selector};
use std::error::Error;
use std::fs;
fn main() -> Result<(), Box<dyn Error>> {
let html = fs::read_to_string("page.html")?;
let document = Html::parse_document(&html);
let title_selector = Selector::parse("title").expect("valid CSS selector");
if let Some(title) = document.select(&title_selector).next() {
let text = title.text().collect::<Vec<_>>().join(" ");
println!("{text}");
}
Ok(())
}
Run it with cargo run. The ? operator turns a missing file, permission failure, or invalid UTF-8 into an error instead of silently continuing.
#1 Best Overall
Read the file safely
UTF-8 HTML
std::fs::read_to_string("page.html") reads the complete file into a String. It is the most convenient path when the document is valid UTF-8, which is the normal representation for modern HTML.
let html = std::fs::read_to_string("page.html")?;
The path is resolved relative to the process’s current working directory, not necessarily the directory containing your source file. If a test or deployment starts the binary elsewhere, pass an absolute path or construct one from a known configuration directory.
Arbitrary bytes or known legacy encodings
read_to_string rejects invalid UTF-8. To inspect or decode bytes yourself, use std::fs::read:
Rank #2
use std::error::Error;
fn load_html(path: &str) -> Result<String, Box<dyn Error>> {
let bytes = std::fs::read(path)?;
let html = String::from_utf8(bytes)?;
Ok(html)
}
This example still requires UTF-8 at the explicit conversion step, but the decision is now yours: you can detect a byte-order mark, use a separate decoding library for a legacy encoding, or preserve the original bytes for another pipeline. Do not use String::from_utf8_lossy blindly when exact text matters, because replacement characters can change extracted content.
Parse a complete document or a fragment
Complete pages: parse_document
Use Html::parse_document when the input represents a page, including optional or malformed document-level tags such as <html>, <head>, and <body>. The parser builds a read-only tree that you can query with CSS selectors.
let document = scraper::Html::parse_document(&html);
Snippets: parse_fragment
Use Html::parse_fragment for a snippet such as <li>Item</li> that is not intended to be a complete page. Fragment parsing avoids treating the snippet as a full document and is useful for templates, stored fields, and individual components.
Rank #3
use scraper::{Html, Selector};
let fragment = Html::parse_fragment("<li data-id="42">Item</li>");
let item = Selector::parse("li").expect("valid selector");
for node in fragment.select(&item) {
println!("{}", node.text().collect::<Vec<_>>().join(" "));
}
Extract text, attributes, and markup with scraper
The scraper crate is a practical high-level choice when your job is selection rather than in-place editing. Parse selectors once, then reuse them if you process many nodes.
use scraper::{Html, Selector};
use std::error::Error;
use std::fs;
fn main() -> Result<(), Box<dyn Error>> {
let html = fs::read_to_string("page.html")?;
let document = Html::parse_document(&html);
let heading = Selector::parse("h1").expect("valid selector");
let links = Selector::parse("a[href]").expect("valid selector");
if let Some(node) = document.select(&heading).next() {
let text = node.text().collect::<Vec<_>>().join(" ");
println!("heading: {text}");
println!("markup: {}", node.html());
}
for link in document.select(&links) {
let label = link.text().collect::<Vec<_>>().join(" ");
let href = link.value().attr("href").unwrap_or("");
println!("{label} -> {href}");
}
Ok(())
}
text() yields descendant text nodes, so joining the pieces gives you readable output when inline elements divide a sentence. value().attr("name") reads an attribute without assuming it exists. node.html() serializes the selected element’s markup; use the text iterator when you need content without tags.
A reusable importer with clear error messages
Separate file I/O from parsing so callers can distinguish an unreadable path from a document that simply contains no matching element.
use scraper::{Html, Selector};
use std::error::Error;
use std::fmt;
use std::path::Path;
#[derive(Debug)]
struct MissingElement(&'static str);
impl fmt::Display for MissingElement {
fn fmt(&self, f: &mut fmt::Formatter<_>) -> fmt::Result {
write!(f, "required element not found: {}", self.0)
}
}
impl Error for MissingElement {}
fn page_title(path: impl AsRef<Path>) -> Result<String, Box<dyn Error>> {
let path = path.as_ref();
let source = std::fs::read_to_string(path)
.map_err(|e| format!("cannot read {}: {e}", path.display()))?;
let document = Html::parse_document(&source);
let selector = Selector::parse("title").expect("literal selector is valid");
let title = document
.select(&selector)
.next()
.ok_or(MissingElement("title"))?
.text()
.collect::<Vec<_>>()
.join(" ")
.trim()
.to_owned();
Ok(title)
}
fn main() -> Result<(), Box<dyn Error>> {
println!("{}", page_title("page.html")?);
Ok(())
}
For selectors supplied by users, handle Selector::parse as a normal validation error instead of calling expect. Literal selectors in your own source can use expect because a typo is a programming error that should be found during development.
Choose the right Rust HTML crate
| Option | Best fit | Document and fragment support | Tree and selectors | Abstraction level |
|---|---|---|---|---|
scraper |
Extracting text, attributes, and elements with CSS selectors | Html::parse_document and Html::parse_fragment |
Read-only parsed tree with ergonomic CSS selection | High-level application API |
| Kuchiki | DOM-like traversal and mutation | parse_html for documents and parse_fragment for snippets |
Mutable tree with selector-driven traversal | Higher-level tree manipulation |
| html5ever | Custom parsing or serialization infrastructure | HTML5-oriented parsing and serialization | No DOM tree representation by itself; you provide callback handling | Lower-level WHATWG HTML parser |
Use scraper for extraction
If the question is “find every product card,” “read the title,” or “collect all href values,” scraper usually minimizes implementation work. It gives you selectors, descendant text, attributes, and serialization without asking you to build a tree-management layer.
Use Kuchiki for edits
Choose Kuchiki when you need to retain a DOM-like tree and change it: remove nodes, add attributes, replace content, or serialize a modified document. Its HTML parser is built on html5ever and exposes document and fragment parsing through tree-oriented APIs.
cargo add kuchiki
use kuchiki::traits::*;
fn main() {
let source = std::fs::read_to_string("page.html").expect("read page.html");
let document = kuchiki::parse_html().one(source);
if let Ok(title) = document.select_first("title") {
println!("{}", title.text_contents());
}
// Keep `document` as the mutable tree for your edits and serialization.
}
Use html5ever only when you need its lower-level control
html5ever parses and serializes according to WHATWG HTML5 rules, but it uses callbacks and does not supply a DOM tree by itself. That flexibility is valuable for a custom sink, streaming design, or integration with another tree representation; for ordinary application extraction, scraper or Kuchiki is simpler.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common failures and their fixes
- “No such file or directory.” Print
std::env::current_dir()during debugging and verify the runtime working directory. A relative path such aspage.htmlis not relative tosrc/main.rs. - Permission denied. Check the account running the program, directory execute permissions, and container or service restrictions. Handle the returned I/O error; do not substitute an empty document.
- Invalid UTF-8. Replace
read_to_stringwithread, then decode according to the file’s actual encoding. A parser cannot recover the intended characters from an incorrectly decoded string. - A selector returns nothing. Confirm whether the source is a complete document or a fragment, inspect the exact spelling and case of the selector, and print a small serialized section to verify what was parsed. Static HTML will not contain content inserted later by a browser’s JavaScript.
- Text contains unexpected whitespace. HTML text is split across descendant nodes. Collect the iterator and normalize whitespace deliberately rather than assuming one node equals one sentence.
- Malformed markup changes the tree. HTML parsers repair input according to HTML parsing rules. If a tag is omitted or incorrectly nested, inspect the parsed structure instead of relying on the original source layout.
- Fragment selectors behave unexpectedly. Parse snippets with
parse_fragment. If you feed a fragment to document parsing, implied document elements can make tree assumptions fail. - Memory use is too high. Both
read_to_stringand high-level parsers hold substantial in-memory representations. Process files one at a time, avoid cloning the source, and choose a lower-level or streaming design when files are very large.
Reliability and performance practices
- Validate the path before parsing and return separate errors for I/O, decoding, selector validation, and missing required elements.
- Compile each fixed selector once outside a loop. Reusing a selector avoids repeated parsing when extracting many records.
- Keep the original source alive for as long as your parser’s references require it; do not create a temporary string and then try to use references after it is dropped.
- Use document parsing for full pages and fragment parsing for isolated snippets so your assumptions match the input shape.
- Write tests with representative malformed HTML, missing attributes, empty elements, and non-ASCII text. A successful parse does not guarantee that a required field exists.
- If you need to modify and reserialize the tree, use Kuchiki rather than trying to force mutations through scraper’s extraction-oriented API.
Or skip the browser setup
If your actual input is a public web page rather than a local file, ScreenshotNeo can return a rendered screenshot or PDF through one request. It is not a replacement for parsing page.html; it is the simpler route when you need a visual capture of a live URL.
Before capture, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; the response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →See the ScreenshotNeo API documentation for all options. A basic cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Every feature is available on every plan: full-page and element capture, device and viewport controls, custom CSS or JavaScript, waits, request blocking, headers and cookies, PDF settings, caching, signed links, asynchronous jobs, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

