Build a small breadth-first crawler by combining a FIFO queue, a visited set, Java 21’s reusable HttpClient, and Jsoup’s HTML parser. The queue determines breadth-first order: remove the oldest URL, fetch it, discover eligible links, and append unseen links to the tail. Neither HttpClient nor Jsoup supplies this traversal algorithm.
The example below is intentionally bounded to one explicitly selected site scope. It accepts only HTTP(S), follows redirects intentionally, applies request and body limits, identifies itself, spaces requests to the same host, reports page-level failures, and stops after a maximum number of pages.
What the crawler does
For each URL, the crawler performs five operations:
- Canonicalize the URL and check whether it is already known.
- Take the next URL from the queue’s head.
- Fetch the page with one shared
HttpClient. - Parse HTML with Jsoup and resolve links against the page’s URI.
- Apply scope and scheme rules, then enqueue unseen links at the queue’s tail.
A Deque used with removeFirst() and addLast() is the FIFO frontier. A Set prevents duplicate work. A page limit makes the process finite even when a site generates an effectively infinite URL space.
Prerequisites and dependency
Java and Jsoup
Use Java SE 21 as the API baseline. HttpClient has been in the JDK since Java 11; Java 21 documentation describes clients as immutable after construction and reusable for multiple requests. The default redirect policy is NEVER, so this example enables normal redirects explicitly.
The Jsoup project site listed version 1.23.2 on September 29, 2026. Releases can change, so verify the current coordinate before starting a new project. The following Maven entry pins the version shown at that date; it is not presented as a compatibility test.
<dependency>
<groupId>org.jsoup</groupId>
<artifactId>jsoup</artifactId>
<version>1.23.2</version>
</dependency>
Jsoup is open source under the MIT license. On JVM 11 and newer, its connection API uses Java’s HTTP client by default, although this tutorial makes the HTTP request directly so the fetch and parse responsibilities remain visible.
A complete bounded breadth-first crawler
Save this class as BreadthFirstCrawler.java. It crawls pages under one origin and path prefix, accepts HTML responses only, limits response size, observes a simple robots policy, and waits between requests to the same host. The robots parser intentionally handles the common User-agent and Disallow directives; production crawlers should use a complete RFC 9309 implementation.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #2
import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.ArrayDeque;
import java.util.HashSet;
import java.util.Locale;
import java.util.Set;
import java.util.Deque;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
public final class BreadthFirstCrawler {
private static final int MAX_PAGES = 30;
private static final int MAX_BODY_BYTES = 2_000_000;
private static final Duration REQUEST_TIMEOUT = Duration.ofSeconds(15);
private static final Duration HOST_DELAY = Duration.ofMillis(750);
private static final String AGENT = "Freedom251ExampleCrawler/1.0 (contact: [email protected])";
public static void main(String[] args) {
URI start = URI.create(args.length == 0 ? "https://example.com/" : args[0]);
URI scope = originAndPath(start);
Deque<URI> frontier = new ArrayDeque<>();
Set<URI> seen = new HashSet<>();
frontier.addLast(canonicalize(start));
HttpClient client = HttpClient.newBuilder()
.followRedirects(HttpClient.Redirect.NORMAL)
.connectTimeout(Duration.ofSeconds(10))
.build();
Robots robots = Robots.load(client, start, AGENT);
long lastRequestAt = 0;
int processed = 0;
while (!frontier.isEmpty() && processed < MAX_PAGES) {
URI url = frontier.removeFirst();
if (!seen.add(url) || !inScope(url, scope) || !robots.allows(url)) continue;
long wait = HOST_DELAY.toMillis() - (System.currentTimeMillis() - lastRequestAt);
if (lastRequestAt != 0 && wait > 0) sleep(wait);
lastRequestAt = System.currentTimeMillis();
processed++;
try {
HttpRequest request = HttpRequest.newBuilder(url)
.timeout(REQUEST_TIMEOUT)
.header("User-Agent", AGENT)
.header("Accept", "text/html,application/xhtml+xml")
.GET().build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
String type = response.headers().firstValue("content-type").orElse("").toLowerCase(Locale.ROOT);
if (response.statusCode() / 100 != 2 || !type.contains("text/html")) {
System.out.printf("SKIP %s status=%d type=%s%n", url, response.statusCode(), type);
continue;
}
String body = response.body();
if (body.getBytes(java.nio.charset.StandardCharsets.UTF_8).length > MAX_BODY_BYTES) {
System.out.println("SKIP oversized response " + url);
continue;
}
Document document = Jsoup.parse(body, url.toString());
System.out.printf("PAGE %d %s title=%s%n", processed, url, document.title());
for (Element link : document.select("a[href]")) {
try {
URI next = canonicalize(url.resolve(link.attr("href")));
if (inScope(next, scope) && !seen.contains(next) && robots.allows(next)) {
frontier.addLast(next);
}
} catch (IllegalArgumentException ignored) {
System.out.println("BAD LINK on " + url + ": " + link.attr("href"));
}
}
} catch (IOException | InterruptedException | RuntimeException ex) {
System.err.printf("ERROR %s: %s%n", url, ex.getMessage());
if (ex instanceof InterruptedException) Thread.currentThread().interrupt();
}
}
System.out.printf("Finished: processed=%d queued=%d seen=%d%n", processed, frontier.size(), seen.size());
}
static URI canonicalize(URI input) {
if (input.getScheme() == null || input.getHost() == null) throw new IllegalArgumentException("absolute URL required");
String scheme = input.getScheme().toLowerCase(Locale.ROOT);
if (!scheme.equals("http") && !scheme.equals("https")) throw new IllegalArgumentException("HTTP(S) only");
String path = input.getPath().isEmpty() ? "/" : input.getPath();
return URI.create(scheme + "://" + input.getAuthority() + path + (input.getQuery() == null ? "" : "?" + input.getQuery()));
}
static URI originAndPath(URI u) { return URI.create(u.getScheme() + "://" + u.getAuthority() + (u.getPath().isEmpty() ? "/" : u.getPath())); }
static boolean inScope(URI u, URI scope) { return u.getScheme().equals(scope.getScheme()) && u.getAuthority().equalsIgnoreCase(scope.getAuthority()) && u.getPath().startsWith(scope.getPath()); }
static void sleep(long ms) { try { Thread.sleep(ms); } catch (InterruptedException e) { Thread.currentThread().interrupt(); } }
static final class Robots {
final Set<String> disallow;
Robots(Set<String> d) { disallow = d; }
boolean allows(URI u) { return disallow.stream().noneMatch(p -> u.getPath().startsWith(p)); }
static Robots load(HttpClient c, URI page, String agent) {
try {
URI r = URI.create(page.getScheme() + "://" + page.getAuthority() + "/robots.txt");
HttpRequest q = HttpRequest.newBuilder(r).timeout(Duration.ofSeconds(10)).header("User-Agent", agent).GET().build();
HttpResponse<String> x = c.send(q, HttpResponse.BodyHandlers.ofString());
Set<String> blocked = new HashSet<>(); boolean ours = false;
for (String line : x.body().split("\R")) {
String s = line.split("#", 2)[0].trim();
if (s.isEmpty() || !s.contains(":")) continue;
String[] pair = s.split(":", 2); String key = pair[0].trim().toLowerCase(Locale.ROOT); String value = pair[1].trim();
if (key.equals("user-agent")) ours = value.equals("*") || agent.toLowerCase(Locale.ROOT).startsWith(value.toLowerCase(Locale.ROOT));
else if (key.equals("disallow") && ours && !value.isEmpty()) blocked.add(value);
}
return new Robots(blocked);
} catch (Exception e) { System.err.println("robots.txt unavailable: " + e.getMessage()); return new Robots(Set.of()); }
}
}
}
Compile with the Jsoup jar on your class path, then run java BreadthFirstCrawler https://example.com/docs/. Replace the example domain with a site you are authorized to crawl. The scope check requires the same scheme, authority, and path prefix; therefore https://example.com/blog/ does not enter a crawl rooted at https://example.com/docs/.
Why each defensive check matters
Canonical URLs and deduplication
Relative links are resolved against the fetched page before filtering. Empty paths become /, schemes are lower-cased, fragments are discarded by reconstruction, and query strings remain significant. More advanced crawlers may normalize default ports, tracking parameters, encoded paths, and server-specific trailing-slash behavior, but those transformations can change meaning and should be policy decisions.
Status, media type, and size
A successful TCP exchange is not necessarily an HTML page. The example rejects non-2xx responses and content types that do not contain text/html. The two-million-byte guard avoids unbounded memory use; a production implementation should stream to a bounded buffer and enforce a maximum decompressed size.
Timeouts, identity, and pacing
Connection and request timeouts prevent one origin from holding the crawl forever. The user-agent states the crawler’s purpose and supplies a contact address; use a real address in deployment. Sequential requests and a 750-millisecond host delay are conservative defaults, not performance measurements. Crawl-delay guidance is prudent operator behavior where relevant, not a universal RFC 9309 directive.
Recommended Free Tools
Robots policy
RFC 9309 places the file at the origin’s top-level /robots.txt and requires crawlers to follow parseable matching rules after successful retrieval. It also says, “These rules are not a form of access authorization.” Read them as publishing instructions, not a security barrier. A denied URL may still require authentication or other access controls.
Direct HttpClient plus Jsoup versus Jsoup Connection
| Approach | Strength | Trade-off |
|---|---|---|
HttpClient request, then Jsoup.parse |
Separate control over headers, status, limits, redirects, and body handling | More code and explicit error checks |
Jsoup.connect(url).get() |
Short fetch-and-parse path; useful for small scripts | Less visible control over a crawler’s queue, response policy, and bounded body strategy |
Jsoup’s cookbook demonstrates Jsoup.connect(url).get() for HTTP and HTTPS loading. Use that integrated API when its defaults fit; do not make both fetching methods for the same request.
Scaling beyond the starter
Asynchronous work
sendAsync can increase throughput, but submitting every discovered URL at once defeats politeness and can exhaust memory. Add a bounded executor, per-host concurrency limits, per-host rate scheduling, cancellation, and exponential backoff for transient failures. Preserve a durable queue if a crawl must resume.
Persistent state
The sample keeps frontier and visited data in memory. A real crawler needs durable frontier storage, canonicalization rules that remain stable across restarts, retry counts, response metadata, and a checkpoint strategy. For multiple hosts, partition scheduling by origin so one busy domain cannot starve others.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #4
Dynamic pages and non-HTML assets
HttpClient retrieves server responses; it does not execute JavaScript. Pages whose links appear only after browser execution require a rendering system and additional resource controls. Keep binary files out of this HTML crawler unless you deliberately add media-type handling and storage limits.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting
Everything is skipped as non-HTML
Inspect the response’s Content-Type. A missing or unusual header, a redirect landing on a login page, or a server returning an error document can trigger the check. Log status and type before changing the filter.
No links are discovered
Verify that the response actually contains server-rendered anchors. Check the scope path, scheme, and authority; a link outside the selected prefix is intentionally rejected. Fragment-only links collapse to the current document and are already seen.
Requests time out or fail intermittently
Increase the timeout only after confirming the target is allowed to be crawled. Keep retries bounded, back off between attempts, and do not remove the page limit or host delay to compensate.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Redirects leave the intended site
Redirect.NORMAL follows ordinary redirects, but scope is checked before the next queued URL is fetched. Log the final response URI if you need to diagnose a cross-origin redirect, and enforce an allowlist for every redirect target in a production crawler.
Robots behavior looks wrong
Check the origin’s exact /robots.txt, user-agent group, and path matching. The compact parser is educational and does not implement every directive or edge case; replace it with a standards-aware parser for serious use.
Or skip the browser setup
If your actual goal is obtaining clean screenshots rather than traversing links, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return PNG, JPEG, WebP, or PDF, while consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options such as full-page capture, CSS selectors, device presets, custom JavaScript, request blocking, signed links, asynchronous webhooks, bulk capture, and caching. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Frequently Asked Questions
Can this crawler index an entire public internet site?
No. It is deliberately bounded by an origin/path scope and page limit. Broader crawling requires durable state, scheduling, stronger URL normalization, and operational permission from each site.
Does robots.txt make a page legally or technically inaccessible?
No. RFC 9309 describes robots rules as requests to crawlers and explicitly says they are not access authorization.
Why not use Jsoup.connect for every request?
That API is a valid shorter alternative, but direct HttpClient calls make status checks, redirects, headers, timeouts, and body limits explicit.
Will HttpClient execute JavaScript before Jsoup parses a page?
No. HttpClient retrieves the server response; JavaScript-rendered links require a browser-rendering approach.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

