October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
HTML parsing

HTML Parsing in Java with jsoup: A Practical Guide

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use jsoup to turn HTML into a traversable Java document, then select elements with CSS selectors or DOM methods and extract text, attributes, or resolved links. Add the current project version, 1.23.2, through Maven or Gradle; choose the input method that fits your source; and use a safelist when cleaning untrusted markup.

Add jsoup to your Java project

The jsoup project describes the library as an HTML parser and DOM toolkit that also supports URL fetching, CSS and XPath selection, document manipulation, and safelist cleaning. Its documentation says it follows the WHATWG HTML specification and is designed to produce a sensible tree even from malformed “tag-soup” markup.

The official project page lists version 1.23.2. Pin the version rather than relying on an unbounded or floating dependency so builds use a known release; check the project page when updating because versions change.

Maven

<dependency>
  <groupId>org.jsoup</groupId>
  <artifactId>jsoup</artifactId>
  <version>1.23.2</version>
</dependency>

Gradle

implementation 'org.jsoup:jsoup:1.23.2'

After adding the dependency, import the classes you need, commonly org.jsoup.Jsoup, org.jsoup.nodes.Document, org.jsoup.nodes.Element, and org.jsoup.select.Elements.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse HTML from a URL, string, file, or fragment

Choose the parser entry point based on where the markup comes from. A URL fetch returns a parsed Document; a string, file, or stream can be parsed locally. If markup contains relative links, provide its base URI so those links can be resolved.

Fetch a web page

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

public class ParsePage {
    public static void main(String[] args) throws Exception {
        Document doc = Jsoup.connect("https://example.com").get();
        System.out.println("Title: " + doc.title());

        Elements links = doc.select("a[href]");
        for (Element link : links) {
            System.out.println(link.text() + " -> " + link.absUrl("href"));
        }
    }
}

Replace the example URL with the page you are authorized to access. The connection API fetches the response and parses it; get() can fail, so production code should handle network and parsing exceptions rather than assuming every request succeeds. The base URI associated with a fetched document lets absUrl("href") turn a relative href into an absolute address.

Parse a string or fragment

String html = "<article><h2>News</h2><a href='/story'>Read</a></article>";
Document doc = Jsoup.parse(html, "https://example.com/");
Element article = doc.selectFirst("article");
if (article != null) {
    System.out.println(article.select("h2").text());
    Element link = article.selectFirst("a[href]");
    if (link != null) {
        System.out.println(link.absUrl("href"));
    }
}

The second argument supplies the base URI, here the origin against which /story resolves. For a fragment rather than a full document, use jsoup’s fragment parsing API when its fragment-specific behavior is what you need; otherwise parse builds a document around the provided markup.

Parse local files and streams

jsoup provides overloads for files, paths, and input streams. When parsing an input stream, supply a base URI if the document’s relative URLs need resolution. For files, use a path-based overload or a file overload and provide the base URI appropriate to your use case. Select the overload that matches the actual input rather than first reading a large file into a separate string.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Select elements and extract the data you need

A parsed page is a DOM: elements contain child elements and text, and attributes belong to elements. Start by selecting the smallest relevant region, then extract fields from within it. This reduces accidental matches elsewhere on the page.

CSS selectors

select accepts CSS-style selectors and returns matching elements. Common patterns include:

  • article h2 for headings inside articles.
  • .price for elements with the class price.
  • a[href] for links that have an href attribute.
Elements headings = doc.select("article h2");
for (Element heading : headings) {
    System.out.println(heading.text());
}

When you expect one match, use selectFirst and check for null before accessing it. Page structures change, and a missing element should not crash the rest of an extraction job.

Text, HTML, attributes, and URLs

  • element.text() returns the element’s text content in normalized form.
  • element.html() returns its inner HTML; element.outerHtml() includes the element itself.
  • element.attr("name") reads an attribute such as href.
  • element.absUrl("href") resolves a URL-valued attribute against the document’s base URI.

For link extraction, select a[href], then read both text() and absUrl("href"). A plain attr("href") preserves the source value, which may be relative; use the absolute form only when the base URI is meaningful and you want a resolved URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XPath

jsoup also documents XPath selection. It can be useful when a query is more naturally expressed as a path through the tree or when maintaining XPath expressions already used elsewhere. CSS selectors are often easier for common class, tag, and attribute matches. Choose one style consistently for each extraction task, and verify selectors against representative pages because a syntactically valid selector can still match zero elements.

Change markup and sanitize untrusted HTML

For deliberate edits, jsoup exposes methods to set text, HTML, and attributes. Use text-setting methods for text values rather than concatenating untrusted input into markup. To clean HTML from an untrusted source, use the cleaner and a safelist: jsoup parses the input and filters it so only allowed tags and attributes remain.

import org.jsoup.Jsoup;
import org.jsoup.safety.Safelist;

String untrusted = "<p>Hello <script>alert('x')</script><a href='https://example.com'>site</a></p>";
String safeHtml = Jsoup.clean(untrusted, Safelist.basic());
System.out.println(safeHtml);

A safelist is a policy choice, not a universal guarantee that any output is appropriate for every application. Pick the allowed elements, attributes, and protocols to match where the cleaned result will be used, and test the output against your actual trust boundary. Do not assume that parsing alone sanitizes HTML: parsing creates structure; cleaning applies the allow-list.

Choose a parser strategy for document size and format

Ordinary DOM parsing

A regular parse builds a document tree you can traverse, query repeatedly, and modify. It is a natural fit when the page is modest in size or when extraction depends on relationships across the document. Its trade-off is that the tree occupies memory while it is in use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

StreamParser for large documents

The jsoup cookbook includes guidance for StreamParser when working with large documents. Consider it when memory limits matter and you can process content incrementally instead of retaining and querying a complete DOM. If later steps need arbitrary navigation across the whole tree, conventional parsing may be simpler. Assess the real input size, memory budget, and extraction pattern before choosing; streaming is not automatically faster or more convenient for every page.

HTML versus XML parsing

The standard HTML parser is intended for real-world HTML, including malformed markup. jsoup also exposes alternate parser overloads, including an XML parser option. Choose XML parsing when XML-style rules are required; do not switch merely because the input looks tidy. HTML and XML have different parsing expectations, and the desired tree depends on the format and parser you select.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Understand performance claims in context

jsoup 1.23.1 release notes report benchmark improvements on OpenJDK 21: ordinary string parsing was 18% faster on average, InputStream parsing 11% faster, and source-position parsing 70% faster while allocating 64% fewer bytes per document. These are release-note results for stated workloads, not guarantees for every JVM, document, or application. Treat them as a reason to evaluate an upgrade, not as a substitute for measuring your own workload.

The same release notes describe parser changes involving noscript, CDATA, SVG, and MathML, as well as HTTP redirect handling. If your application depends on edge-case markup or redirects, validate its behavior when changing versions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshoot common jsoup parsing problems

  • A selector returns no matches: inspect the fetched or parsed HTML and confirm the selector corresponds to the server-returned markup. Content inserted later by client-side JavaScript may not be present in the HTML response jsoup parses. Check capitalization, classes, attributes, and the selected parent region.
  • A relative link stays relative: parse with the correct base URI or fetch the page through the connection API, then use absUrl("href"). Confirm the source attribute is a URL and the base is the page’s actual location.
  • selectFirst causes a null error: an element was absent. Check for null, handle the missing-data case, and avoid assuming every page variant has the same structure.
  • Fetching throws an exception or returns unexpected content: the request may have failed or the server may have returned a response different from the expected page. Handle exceptions, inspect the response and target URL, and distinguish fetch failures from selectors that simply match nothing.
  • Cleaned output loses an element or attribute: the safelist does not permit it. Review the allow-list and add only what the output context safely requires; do not bypass cleaning for untrusted input just to preserve arbitrary markup.
  • Memory use is too high: avoid retaining multiple full documents unnecessarily, process and release documents as work completes, and evaluate the cookbook’s streaming approach for large inputs.
  • Markup differs from what a browser displays: jsoup parses the HTML it receives; it is not a full browser runtime executing page scripts. Check whether the content exists in the server response before expecting it in the parsed DOM.

Or skip the browser setup

If the goal is to capture a page as an image or PDF rather than extract its DOM, ScreenshotNeo offers a one-request screenshot API. It can remove cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, and cache hits are not billed. It also provides an MCP server for AI agents, and its free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

For a screenshot of a target page, adapt the URL in this cURL call and use your API key. See the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

For HTML extraction and parsing, jsoup remains the Java tool described above; a screenshot API serves a different output need. Learn more at ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Frequently Asked Questions

Does jsoup execute JavaScript on the page?

No. jsoup parses the HTML it receives; it is not a browser runtime that executes client-side scripts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use jsoup for XML?

Yes. Its parser overloads include an XML parser option, which is appropriate when XML parsing rules are required.

Is jsoup open source?

The project repository identifies jsoup as MIT-licensed and maintained by Jonathan Hedley and contributors.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.