Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use the least powerful Java tool that can retrieve the data. If the values are already in the server response, an HTTP client plus an HTML parser is faster, cheaper and easier to debug. If JavaScript, authentication cookies, forms, frames or browser-only behavior creates the page, use a headless browser such as Selenium with Chrome. This guide follows that progression, explains where each approach breaks down, and shows how to move from a local scraper to a responsibly operated cloud job.
What “web scraping with Java” means
Web scraping is the process of fetching a third-party website and parsing its HTML to extract selected data. A reliable scraper is not just a selector loop: it identifies how a page is generated, preserves the session state required to reach it, handles failures, and stores results in a form that can be checked later.
The Java Web Scraping Handbook was originally written in 2018 and republished by ScrapingBee on 17 January 2026. Its official edition covers ordinary HTML, JavaScript-heavy sites, forms, captchas, anti-bot techniques and cloud deployment. Treat its code as an educational pattern: check current JDK, Selenium, browser-driver and library documentation before copying dependency versions or driver setup.
Choose the right scraping architecture
| Approach | JavaScript rendering | Cookies, forms and frames | Complexity and overhead | Best use |
|---|---|---|---|---|
| HTTP client + Jsoup | Does not execute page JavaScript | You implement requests, cookies and form posts | Lowest runtime and resource use | Data present in the initial HTML or a discoverable API |
| HTTP client + underlying API | Not needed when an XHR/fetch endpoint supplies the data | Depends on the API’s authentication and tokens | Usually lighter than a browser, but endpoint-specific | Structured JSON, pagination and stable network calls |
| Selenium with headless Chrome | Executes JavaScript and browser events | Native browser cookies, form filling, frames and waits | Highest memory, startup time and operational complexity | Rendered pages, infinite scroll and browser-only workflows |
| HtmlUnit | GUI-less Java browser; verify compatibility with the target | Browser-like APIs without a visible Chrome process | Less infrastructure than Chrome, but page compatibility varies | Java-only, non-GUI automation where HtmlUnit supports the site |
Start by viewing the raw response and browser developer tools. If the desired text is absent from “View Source” but appears after scripts run, inspect Network requests for the JSON or HTML endpoint before launching a browser. Reproducing that request is generally easier to scale. Use Selenium when the endpoint is difficult to reproduce or the workflow genuinely depends on browser behavior.
Build a static HTML scraper with Java and Jsoup
1. Create a small Maven project
Use a current Jsoup release from its official documentation rather than pinning an old handbook version. A minimal dependency is:
<dependency>
<groupId>org.jsoup</groupId>
<artifactId>jsoup</artifactId>
<version>CURRENT_VERSION</version>
</dependency>
Replace CURRENT_VERSION with the version selected after checking the project’s current release notes.
2. Fetch, parse and validate the document
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
public class StaticScraper {
public static void main(String[] args) throws Exception {
URI uri = URI.create("https://example.com/articles");
HttpClient client = HttpClient.newBuilder()
.connectTimeout(Duration.ofSeconds(15))
.followRedirects(HttpClient.Redirect.NORMAL)
.build();
HttpRequest request = HttpRequest.newBuilder(uri)
.timeout(Duration.ofSeconds(30))
.header("User-Agent", "ResearchBot/1.0 (contact: [email protected])")
.GET().build();
HttpResponse<String> response = client.send(
request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() / 100 != 2) {
throw new IllegalStateException("HTTP " + response.statusCode());
}
Document doc = Jsoup.parse(response.body(), uri.toString());
Elements cards = doc.select("article.card");
for (Element card : cards) {
String title = card.select("h2, h3").text();
String link = card.select("a[href]").attr("abs:href");
if (!title.isBlank() && !link.isBlank()) {
System.out.printf("%st%s%n", title, link);
}
}
}
}
Use CSS selectors that describe the data rather than brittle positional paths. abs:href resolves relative links against the document URL. Check for missing elements before calling text(), and record the source URL and retrieval time with each row so a changed layout can be diagnosed.
Pagination and polite retrieval
Follow a site’s published rules and terms, request only what you need, and add a bounded delay between pages. Stop when the next link is absent or when a maximum-page limit is reached. Deduplicate by a stable key such as a canonical URL. Retry only transient network failures, with exponential backoff and a hard attempt limit; do not hammer a server after a denial or rate-limit response.
Handle forms, cookies and login sessions
A form workflow normally requires a GET to obtain the form and cookies, extraction of hidden fields, then a POST with the same cookie jar. Never hard-code credentials in source code; load them from a secret manager or environment variables, and confirm that automated access is permitted.
Rank #2
import java.net.CookieManager;
import java.net.CookiePolicy;
import java.net.URI;
import java.net.http.*;
import java.nio.charset.StandardCharsets;
import java.net.URLEncoder;
CookieManager cookies = new CookieManager(null, CookiePolicy.ACCEPT_ALL);
HttpClient client = HttpClient.newBuilder()
.cookieHandler(cookies).followRedirects(HttpClient.Redirect.NORMAL).build();
HttpRequest login = HttpRequest.newBuilder(URI.create("https://example.com/login"))
.header("Content-Type", "application/x-www-form-urlencoded")
.POST(HttpRequest.BodyPublishers.ofString(
"username=" + URLEncoder.encode(System.getenv("USER"), StandardCharsets.UTF_8)
+ "&password=" + URLEncoder.encode(System.getenv("PASSWORD"), StandardCharsets.UTF_8)))
.build();
HttpResponse<String> result = client.send(login, HttpResponse.BodyHandlers.ofString());
if (result.statusCode() / 100 != 2 && result.statusCode() != 302) {
throw new IllegalStateException("Login failed: " + result.statusCode());
}
Many sites add CSRF tokens or require a particular referer, origin, or JavaScript-generated value. Parse hidden inputs from the login page and submit them. If the flow includes a challenge, stop and use an authorized integration rather than attempting to defeat it.
Scrape JavaScript-rendered pages with Selenium
When a browser is justified
Choose Selenium when content appears only after script execution, a click triggers the data request, an infinite-scroll list must be advanced, or the workflow depends on browser-managed authentication and frames. The handbook uses Selenium WebDriver with headless Chrome for this class of site.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRunnable headless Chrome example
Use current Selenium Manager and Chrome documentation to align the Selenium library, JDK and installed browser. The example waits for a selector instead of sleeping for an arbitrary period.
import java.time.Duration;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
public class BrowserScraper {
public static void main(String[] args) {
ChromeOptions options = new ChromeOptions();
options.addArguments("--headless=new", "--no-sandbox", "--disable-dev-shm-usage");
options.addArguments("--window-size=1440,1200");
WebDriver driver = new ChromeDriver(options);
try {
driver.get("https://example.com/catalog");
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(30));
wait.until(ExpectedConditions.presenceOfElementLocated(By.cssSelector("article.card")));
driver.findElements(By.cssSelector("article.card")).forEach(card -> {
String title = card.findElement(By.cssSelector("h2, h3")).getText();
System.out.println(title);
});
} finally {
driver.quit();
}
}
}
Browser details that commonly matter
- Frames: switch into the matching iframe before selecting elements, then switch back to the default content.
- Infinite scroll: scroll in a loop, wait for the item count to increase, and stop when it no longer changes or a page limit is reached.
- Downloads and PDFs: configure Chrome’s download preferences or fetch a discovered file URL directly when authorization allows.
- Debugging: save the current URL, page source and a screenshot on failure; these reveal whether a selector, redirect, cookie or challenge caused the error.
- Resource control: reuse a driver for a bounded batch, but always call
quit(). A browser per URL can exhaust memory and file descriptors.
Use the underlying API when the browser exposes one
Open DevTools Network, reload the page, and look for XHR or fetch requests returning JSON. Record the URL, method, query parameters, request body and required headers. Recreate the call with Java’s HTTP client, then validate pagination and authentication expiry. An API response is usually more stable and cheaper to process than rendered markup, but it remains subject to the site’s authorization, terms and rate limits. Do not bypass access controls or reuse private tokens outside their permitted scope.
HtmlUnit as a Java-only alternative
HtmlUnit is a separate GUI-less Java browser project. Its official project information lists version 5.5.0 dated 30 August 2026, and the HtmlUnit 5 line requires JDK 17 or newer. Confirm current Maven or Gradle coordinates and browser-feature compatibility before adopting it. HtmlUnit can reduce the infrastructure of a Chrome process, but a site that relies on browser APIs unsupported by HtmlUnit may still require Selenium.
Anti-bot controls, captchas and responsible operation
Bot checks, captchas, rate limits, fingerprinting and login challenges are signals that the owner wants to control automated access. Verify the target site’s terms, robots guidance and applicable law before deployment. Prefer a documented API, an export, or written permission. Identify your client honestly, limit concurrency, cache results, honor backoff responses and protect personal data. Captcha-solving, proxy rotation and Tor are advanced topics discussed by the handbook, not guarantees of access and not substitutes for authorization.
Deploy a scraper in the cloud
Keep cloud deployment until the local workflow is deterministic. Package configuration separately from code, store secrets in the platform’s secret facility, write structured logs, and make each job idempotent so a retry cannot duplicate records. Set explicit timeouts for connection, page load and total job duration. A browser job needs an image containing Chrome and its libraries, more memory than an HTTP job, and cleanup for temporary profiles. The handbook’s cloud chapter covers serverless and Azure Functions; serverless suitability depends on execution-time, memory and browser-binary limits of the service you select.
- Persist checkpoints and the last successful URL.
- Emit metrics for status codes, empty-result pages, retries and elapsed time.
- Alert on selector drift rather than silently writing empty data.
- Use a queue to cap concurrency and a dead-letter path for repeated failures.
- Retain only the data you need and define deletion periods.
Performance, reliability and cost decisions
| Need | Prefer | Reason |
|---|---|---|
| Thousands of static pages | HTTP client plus parser or API | Low startup and memory overhead; easy concurrency limits |
| One authenticated workflow | Selenium or a permitted API | Preserves browser session behavior or uses structured responses |
| Frequent layout changes | API discovery plus selector tests | Less dependence on visual markup and earlier drift detection |
| Short scheduled jobs | HTTP worker; browser only where required | Lower cloud resource consumption and faster cold starts |
Measure request latency, browser startup time, bytes downloaded, retry rate and records per successful page. Cache immutable pages or API responses within the site’s rules. A timeout is not proof that the page is empty: distinguish DNS, TLS, HTTP status, navigation timeout, selector timeout and parser errors in logs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when your Java job needs a clean visual capture rather than DOM extraction. One GET request returns PNG, JPEG, WebP or PDF. It accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
For a one-call capture, see the ScreenshotNeo API documentation:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also supports full-page and element captures, device and retina settings, custom CSS or JavaScript, waits, request blocking, headers and cookies, geolocation, PDF options, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an MCP server with take_screenshot, get_page_info and capture_pdf for AI clients. Every feature is on every plan: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Rank #4
Troubleshooting common failures
“The selector returns nothing”
Check whether the data exists in the raw response. If not, wait for the correct network or DOM condition, switch into the relevant iframe, or locate the JSON request. Save page source and a screenshot at the failure point.
“ChromeDriver cannot start”
Align the installed Chrome, Selenium and JDK with current Selenium documentation. In containers, verify executable permissions, shared-memory sizing and required libraries; avoid leaving orphaned driver processes.
“Login succeeds but the next page is anonymous”
Keep one cookie-aware client or one WebDriver session, preserve redirects and hidden CSRF fields, and verify the final URL and response status. A new request without the session cookie will normally lose authentication.
“The job times out”
Separate connect, navigation and selector timeouts. Check DNS/TLS and response status, reduce page scope, block unnecessary resources where allowed, and retry only transient failures with backoff.
“Results suddenly become empty”
Treat empty output as an alert. Compare a saved fixture with the current DOM, test selectors in CI, inspect layout or API changes, and stop publishing data until the schema is confirmed.
Best Value
What to take from the handbook
The handbook’s durable lesson is methodological: learn HTTP, HTML and the DOM first; discover an underlying API when practical; add Selenium only for browser behavior; and postpone cloud and anti-bot operations until the local scraper is observable, authorized and repeatable. Its 2018 examples remain useful patterns, but current dependency, browser and platform documentation must determine your production setup.
Frequently Asked Questions
How many pages is The Java Web Scraping Handbook?
The official site lists 120–170 pages depending on format.
Recommended Free Tools
What formats and packages are sold?
The publisher lists PDF, EPUB and MOBI packages, with ebook-only pricing of $29, a standard package at $49 and a complete package at $69; included source code, sandbox access, forum access and updates depend on the package.
Is there a confirmed paperback edition?
A catalog entry reports paperback and ISBN-10/ASIN as not available, so do not assume an Amazon physical edition.
Where does the cloud material begin?
The detailed PDF contents place serverless and Azure Functions material in the cloud chapter beginning on page 102.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →

