Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For pages whose useful content is already in the HTTP response, start with jsoup: it can fetch HTML, parse it into a document, and extract data with DOM traversal, CSS selectors, or XPath. Set explicit request limits and handle missing or changed content. If the content appears only after JavaScript runs or requires browser interaction, use browser automation such as Playwright for Java or Selenium WebDriver instead. For production, add bounded retries, cleanup, data validation, monitoring, and a review of the target site’s rules and your project’s permissions.
Choose the right Java scraping approach
Begin by checking what the server sends, not by assuming you need a browser. If a normal HTTP response contains the content you need, fetching and parsing that response is usually the simpler route. jsoup combines HTTP fetching with HTML parsing and selector-based extraction.
If the data is added only after scripts run, or the task depends on clicking, scrolling, or other browser behavior, evaluate a browser automation framework. A browser can render and interact with a page, but it does not grant permission to access content or bypass a site’s restrictions.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →| Approach | Good fit | Trade-offs |
|---|---|---|
| jsoup direct fetch and parse | Useful content is present in ordinary response HTML. | Lightweight compared with browser automation, but it does not render a JavaScript application as a browser. |
| Playwright for Java | Browser rendering or interactions are required. | Supports Chromium, WebKit, and Firefox, but browser binaries and runtime setup add operational work. |
| Selenium WebDriver | You need browser control and its driver ecosystem, including local or remote sessions. | Requires the Java binding, browser, and driver setup; browser sessions need reliable cleanup. |
This is a qualitative choice, not a performance ranking. Pick the least complex option that returns the data you need, and account for rendering fidelity, interactions, deployment footprint, and browser maintenance.
Set up a small jsoup project
Use Maven or Gradle to declare a dependency and pin the version you intend to run. The jsoup project page lists version 1.23.2; confirm the current version and requirements in the project documentation when updating your build. Avoid copying a jar manually into an application because a build file makes dependency updates and reproducible builds easier.
For Maven, add this dependency to pom.xml:
<dependency>
<groupId>org.jsoup</groupId>
<artifactId>jsoup</artifactId>
<version>1.23.2</version>
</dependency>
The following example is a complete small command-line scraper. It takes the page URL as its first argument, identifies itself with a project-specific user agent, bounds the request and response size, checks the response status, and handles an absent heading. Replace the user-agent name and contact URL with truthful details for your project.
import org.jsoup.Connection;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import java.io.IOException;
public class Main {
public static void main(String[] args) {
if (args.length != 1) {
System.err.println("Usage: java Main <url>");
System.exit(2);
}
String url = args[0];
try {
Connection.Response response = Jsoup.connect(url)
.userAgent("ExampleResearchBot/1.0 (+https://example.org/bot)")
.timeout(10_000)
.maxBodySize(1_000_000)
.followRedirects(true)
.execute();
if (response.statusCode() < 200 || response.statusCode() >= 300) {
System.err.println("Unexpected HTTP status: " + response.statusCode());
System.exit(1);
}
Document doc = response.parse();
Element heading = doc.selectFirst("h1");
String title = doc.title();
String headingText = heading == null ? "" : heading.text();
System.out.println("Title: " + title);
System.out.println("H1: " + headingText);
} catch (IOException e) {
System.err.println("Could not fetch or parse " + url + ": " + e.getMessage());
System.exit(1);
}
}
}
Save this as Main.java in a project with jsoup on the classpath. The sample uses Java language features available in modern Java; compile and run it with the Java version used by your build. It prints the document title and first H1, or an empty H1 value when no matching element exists. For a real extractor, replace the selectors with those that match the target’s HTML and validate the extracted values before storing them.
Extract data without assuming the page stays the same
Use selectors for structure, not certainty
jsoup’s select methods accept CSS selectors; its cookbook also documents DOM traversal and XPath. For a single expected match, selectFirst returns an element or null. Check that result before calling methods such as text() or attr(). For repeated records, select a collection and process each element, while allowing for an empty collection.
Rank #2
for (Element item : doc.select("article.product")) {
Element name = item.selectFirst("h2");
Element price = item.selectFirst(".price");
if (name == null || price == null) {
// Record a structured extraction issue; do not silently invent a value.
continue;
}
String productName = name.text();
String productPrice = price.text();
// Validate and store the values for your application.
}
Prefer selectors tied to meaningful structure over fragile positional assumptions. Page markup can change, and a selector may still return text after the meaning of that text has shifted. Validate formats and required fields, distinguish a missing field from an empty string, and record enough context to investigate extraction failures.
Decide what to do with HTTP responses
The sample inspects the status before parsing. Production code should make response handling explicit: successful responses may still contain an unexpected page, while redirects, access challenges, rate limits, server errors, or missing pages need different handling. Keep status, target, timing, and failure category in logs or metrics, but avoid logging credentials or sensitive page content.
Use a user agent that truthfully identifies your crawler and offers a useful contact route where appropriate. Do not impersonate a browser or another crawler to evade restrictions. If a target expects authentication or a particular session, use only access you are authorized to use and handle credentials securely.
When a browser is necessary
Use browser automation when the data is produced by page scripts or the work requires browser interactions. Playwright’s Java setup is Maven-distributed, supports Chromium, WebKit, and Firefox, and its setup documentation lists Java 8 or higher. Its examples use managed lifecycle patterns such as try-with-resources. Selenium’s Java setup involves the binding, a browser, and a driver; WebDriver can create local or remote browser sessions.
Playwright lifecycle shape
A minimal browser workflow creates Playwright, launches an engine, opens a page, navigates, reads rendered content, and closes the resources. Use the engine and browser-installation steps documented for the Playwright version you choose; browser binaries are a deployment requirement distinct from adding a Maven dependency.
try (Playwright playwright = Playwright.create()) {
Browser browser = playwright.chromium().launch();
try {
Page page = browser.newPage();
page.navigate("https://example.org");
String title = page.title();
System.out.println(title);
} finally {
browser.close();
}
}
For a production project, pin the Playwright dependency and install compatible browser binaries as part of deployment. Set navigation and operation timeouts deliberately, and make selectors wait for the specific content you need rather than relying on an arbitrary long sleep. A rendered page can still be incomplete or inaccessible; treat failed navigation and missing data as outcomes to report, not as a reason to bypass controls.
Selenium lifecycle and remote sessions
Selenium WebDriver provides browser control through browser-specific drivers and can use local or remote sessions. When the browser needs to run on a separate machine or be scaled out, a remote WebDriver or Selenium Grid deployment is an option. Keep driver creation and cleanup in a well-defined scope. Selenium distinguishes closing a window with close from ending the WebDriver session with quit; use quit when the scraping job is finished so the session and its resources are released.
Put limits and cleanup in place
Bound each request
jsoup documents a default total timeout of 30 seconds and a default maximum response body of 2 MB. Both can be configured; setting either limit to zero removes that limit. The example’s 10-second timeout and 1,000,000-byte maximum are deliberate sample values, not universal recommendations. Choose limits suitable for your targets and the size of the documents you expect, then measure outcomes in your own workload.
Rank #4
Keep limits explicit even if the library has defaults. A slow server, oversized response, or unexpected page should not tie up a worker indefinitely or consume unbounded memory. If legitimate pages exceed your configured body limit, revisit the limit deliberately rather than removing it without replacement.
Handle sessions and concurrency intentionally
jsoup sessions keep cookies in memory for the session lifetime. That can help when a sequence of requests needs shared cookies, but it also means session storage and cleanup matter. Avoid one unbounded, long-lived session without a cookie-store plan. The jsoup API advises using a separate request for each concurrent operation when sharing session settings; do not assume a shared mutable session is safe for simultaneous work.
For browser automation, close browsers and drivers in cleanup paths, including when navigation or extraction throws an exception. A job that exits normally but leaks browser processes can degrade later jobs and exhaust machine resources.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Build production scrapers that can be operated
Production readiness is less about adding more scraping libraries than making each job bounded, inspectable, and safe to repeat. The following are engineering practices, not claims of a single prescribed Java architecture.
Best Value
- Separate fetching, extraction, validation, and storage. This makes it easier to distinguish a network failure from a selector change or a persistence error.
- Retry selectively. Use bounded retries with backoff for transient failures; do not endlessly repeat a request or retry an unchanged client error. Respect signals that the service is overloaded.
- Make writes idempotent. If a job runs again after a partial failure, stable record keys and duplicate-handling rules can prevent repeated records.
- Track data quality as well as job status. Record expected-field presence, parse failures, response status, duration, and counts. Alert when a page starts producing missing or implausible values.
- Keep concurrency and request frequency conservative. Set limits based on the target’s published guidance and observed responses rather than assuming parallel requests are harmless.
- Protect secrets and collected data. Keep credentials out of source control and logs, restrict access to stored results, and retain only what the project needs.
There are no comparative throughput, reliability, or cost figures established here for jsoup, Playwright, or Selenium. Benchmark your own representative targets if those figures are important to a capacity decision; account for browser startup and runtime costs in browser-based designs.
Respect crawler guidance and access restrictions
Check a site’s published crawler instructions before collecting pages, identify your client honestly, keep request rates reasonable, and reduce or stop work when the service signals overload. Do not bypass authentication, paywalls, CAPTCHAs, or other explicit access controls.
robots.txt is crawler guidance, not authorization. RFC 9309 states, “These rules are not a form of access authorization.” Google describes robots.txt as a way to tell search engine crawlers which URLs they can access on a site and says it is primarily for crawler traffic management, not page security. A disallowed path is not thereby made private, and a path allowed by robots.txt is not thereby granted to you for every purpose. These points do not determine whether a particular collection project is lawful: contractual terms, privacy, copyright, and regulatory requirements may require separate review.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Common problems and practical fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Expected element is missing | The selector does not match the current HTML, the page has changed, or the content is added after scripts run. | Inspect the response and selector; validate element presence. If rendering is required, assess browser automation rather than assuming a parser can execute scripts. |
| Request times out | The target is slow, unreachable, or the configured timeout is too short for the expected response. | Record the failure, verify the target and network, and adjust a bounded timeout only when justified. Avoid disabling timeouts. |
| Response exceeds the body limit | The response is larger than the configured maximum. | Confirm the response is expected and safe to process, then choose an explicit limit appropriate to the workload. |
| Unexpected status or challenge page | The server returned an error, rate limit, access challenge, or other non-success response. | Handle the outcome explicitly, lower request pressure or stop, and use only authorized access. Do not attempt to evade the challenge. |
| Browser processes remain after a run | A WebDriver or browser was not quit on every code path. | Put cleanup in finally or managed resource scopes and end Selenium sessions with quit. |
| Values are blank despite a successful fetch | The page may not contain the field in its response, markup may have changed, or the selected element may be empty. | Distinguish missing from empty values, inspect the HTML and selector, and route script-rendered content to an appropriate browser workflow. |
Or skip the browser setup
If your goal is a visual capture rather than extracting structured fields, ScreenshotNeo provides a one-request screenshot API; it is not a replacement for a Java HTML parser when you need page data. Its API can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie/consent banners, newsletter popups, and chat widgets are removed before the shot by default, and those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing outcome. An MCP server provides screenshot and PDF tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Learn about ScreenshotNeo, or sign up free for 1,000 screenshots a month with no card.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

