To scrape a website with Kotlin, first fetch its HTML, then parse the response and extract the fields you need. For a Kotlin/JVM project, a practical beginner stack is Ktor Client for HTTP requests and jsoup for HTML parsing and CSS selectors. They handle separate jobs: an HTTP client retrieves the response; an HTML parser turns it into a document you can inspect and query. This approach works when the data is present in the returned HTML. If a page creates the data later with JavaScript, a basic HTTP fetch may not contain it.
Before you scrape: check access and the page
Choose a specific page and a small set of fields, such as a public product title and its displayed price. Look for an official API or data export first; it may provide a more stable and appropriate way to obtain the information. Avoid collecting personal or sensitive information unless you have a clear, lawful reason and have reviewed the relevant requirements.
- Review the site’s terms, applicable rules, and any published access instructions for your intended use.
- Inspect the page in a browser and determine which fields you actually need.
- Fetch the page once and inspect the response HTML. If the text or relevant elements are present, a normal HTTP client and parser may be sufficient.
- If the response lacks the target data, investigate whether the site offers an API or another supported access method before considering browser-based automation.
Robots.txt is relevant to crawler behavior, but it does not grant permission to access a site. IETF RFC 9309 says that parseable rules in a successfully retrieved robots.txt must be followed by crawlers, and explicitly states: “These rules are not a form of access authorization.” Site terms, rate limits, privacy, copyright, and applicable law require context-specific review; robots.txt alone cannot settle those questions.
Choose a Kotlin scraping stack
| Need | Tool | Fit and limitation |
|---|---|---|
| Make HTTP requests from Kotlin | Ktor Client | A Kotlin-oriented HTTP client with multiple platform targets. Select an engine compatible with the target you build for and verify its requirements. |
| Parse static HTML and select elements on the JVM | jsoup | A Java library that provides HTML parsing, DOM traversal, CSS and XPath selectors, and URL-loading features. It does not execute page JavaScript and should not be treated as automatically portable to Kotlin/JS, Kotlin/Native, or Kotlin/Wasm. |
| Build a browser or Node.js web application | Kotlin/JS | Kotlin/JS targets JavaScript environments. That is a web-development target, not the default server-side scraping runtime. |
| Build a Kotlin WebAssembly web application | Kotlin/Wasm | Kotlin’s web documentation describes Wasm web targets and use cases; that does not make it the ordinary choice for a server-side scraper. |
Ktor’s documentation surfaced as version 3.6.0 and lists JVM, Android, Native, JavaScript, and WasmJs client platforms. The jsoup site listed version 1.23.2. These are time-sensitive version observations, not promises about the newest release when you install; check the official project documentation for current dependency coordinates and engine compatibility. The Kotlin/JS and Kotlin/Wasm options are distinct web development approaches, not substitutes for choosing a JVM scraper stack.
#1 Best Overall
Set up a Kotlin/JVM project
The example below uses Ktor Client with the CIO engine to retrieve a page, then jsoup to parse the returned HTML. In a Gradle Kotlin DSL project, add dependencies using versions verified against the current official documentation:
dependencies {
implementation("io.ktor:ktor-client-core:3.6.0")
implementation("io.ktor:ktor-client-cio:3.6.0")
implementation("io.ktor:ktor-client-plugins:3.6.0")
implementation("org.jsoup:jsoup:1.23.2")
}
Before relying on those coordinates, confirm the Ktor artifacts for the version and engine you choose: plugin packaging and setup can change between versions. The example uses the Ktor client, CIO engine, HTTP timeout and user-agent plugins. If your selected release exposes a plugin through a different artifact, add the dependency specified by that release’s setup guide. Keep Ktor modules on a consistent version.
Fetch the page and inspect the response
For a scraper that needs control over status checks, headers, timeouts, or later pagination, make the HTTP request explicitly and pass its response to jsoup. Replace the example URL with a page you are permitted to access. Do not treat this illustrative selector or any particular site’s structure as verified.
import io.ktor.client.HttpClient
import io.ktor.client.engine.cio.CIO
import io.ktor.client.plugins.HttpTimeout
import io.ktor.client.plugins.UserAgent
import io.ktor.client.request.get
import io.ktor.client.statement.bodyAsText
import io.ktor.http.isSuccess
import kotlinx.coroutines.runBlocking
import org.jsoup.Jsoup
import java.net.URI
fun main() = runBlocking {
val pageUrl = "https://example.com/"
val client = HttpClient(CIO) {
install(UserAgent) {
agent = "ExampleResearchBot/1.0 (contact: [email protected])"
}
install(HttpTimeout) {
requestTimeoutMillis = 30_000
connectTimeoutMillis = 10_000
socketTimeoutMillis = 30_000
}
expectSuccess = false
}
try {
val response = client.get(pageUrl)
if (!response.status.isSuccess()) {
error("Page request failed with HTTP ${response.status.value}")
}
val contentType = response.headers["Content-Type"].orEmpty()
if (!contentType.contains("text/html", ignoreCase = true)) {
error("Expected HTML, received Content-Type: $contentType")
}
val html = response.bodyAsText()
val document = Jsoup.parse(html, pageUrl)
println("Page title: ${document.title()}")
} finally {
client.close()
}
}
The user-agent value should identify your application honestly; provide a contact route only if it is real and monitored. Do not impersonate a browser or another organization to bypass site controls. The three timeout values are example settings, not a universal recommendation. Choose limits based on your workload and the target’s behavior.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #2
Checking HTTP status before parsing keeps an error page from being mistaken for a successful result. Checking content type helps catch responses that are not HTML, although servers may omit or mislabel that header. Network exceptions, DNS failures, and timeouts can still occur; handle them at the job boundary and record enough context to diagnose them.
Parse HTML and extract fields
Once you have a jsoup Document, inspect the HTML structure in your browser’s developer tools or by examining the response. Choose selectors that match the page’s actual markup. The following generic example illustrates extracting a title and link from repeated cards; its class names are placeholders for markup you inspect yourself.
data class Listing(
val title: String,
val url: String
)
val listings = document.select(".listing-card").mapNotNull { card ->
val title = card.selectFirst(".listing-title")
?.text()
?.replace(Regex("\s+"), " ")
?.trim()
?.takeIf { it.isNotEmpty() }
?: return@mapNotNull null
val link = card.selectFirst("a[href]")?.absUrl("href")
?.takeIf { it.isNotBlank() }
?: return@mapNotNull null
Listing(title = title, url = link)
}
listings.forEach(::println)
jsoup’s selector and extraction APIs support selecting elements, reading text and attributes, and resolving links. Passing the page URL as the base URI when parsing lets absUrl("href") turn a relative link into an absolute one. If you parse without a base URI, a relative href may not resolve as intended. Use attr("href") if you specifically want the raw attribute rather than a resolved URL.
Normalize values deliberately
- Collapse repeated whitespace and trim surrounding spaces in text fields.
- Parse prices, counts, and dates using explicit rules for decimal separators, currencies, locale, and date format. A displayed value may not be a plain number.
- Represent optional fields as nullable values or skip incomplete records intentionally; do not silently substitute invented values.
- Keep the source URL with a record when it helps you audit where the data came from.
Selectors are coupled to page structure. Prefer selectors tied to meaningful attributes or stable containers over brittle assumptions about element positions. There is no guarantee a site will preserve its markup, so validate the output and make changes visible rather than accepting a run that quietly produces no records.
Recommended Free Tools
Rank #3
Validate and store the extracted data
Convert results into explicit Kotlin data classes, then validate required fields before writing them. For example, reject a record with a blank title or a link that is not an HTTP or HTTPS URL. Decide what should happen to invalid records: log and skip them, or stop the job if missing data would make the whole output misleading.
For a small job, CSV or JSON files may be enough; recurring work may call for a database. Whichever destination you choose, encode values with a proper CSV or JSON library instead of concatenating raw strings, which can break when fields contain commas, quotation marks, or newlines. Store a retrieval timestamp and source URL if you need to distinguish when and where a record was collected. Track useful operational signals such as requested URL, status code, elapsed time, number of extracted records, and validation failures.
Add pagination and scale cautiously
First make a single-page request and extraction reliable. Add pagination only after you understand how the target represents next pages, page numbers, or cursors. Set a clear stopping condition, such as a missing next link or a known page boundary, so a malformed page cannot create an endless loop.
- Use bounded concurrency rather than launching an unbounded number of simultaneous requests.
- Cache responses where appropriate to avoid fetching unchanged pages repeatedly.
- Use retries with backoff only for transient failures; do not retry indefinitely.
- Choose a request rate the site can support. There is no universal safe requests-per-second figure established here.
- Stop on access-denied responses, blocks, or other signs that the site does not want the request to continue. Do not attempt to evade those controls.
For reliability, distinguish a temporary transport problem from a permanent extraction problem. A retry may help with a transient network failure, but it will not fix a selector that no longer matches. Alert on sudden changes in response status, content type, or record count so that a structural change is not mistaken for an empty but successful scrape.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteWhen static HTML is not enough
A normal HTTP client receives the server’s response; jsoup parses that HTML. Neither step runs the page’s JavaScript. If the desired information appears in the browser only after scripts execute, compare the original response with the rendered page and inspect whether the site offers an official API or documented data interface. Review its terms and access rules before using one.
If a browser-based approach appears necessary, select and verify an appropriate automation tool for your platform and current project version. The documentation cited for Ktor and jsoup establishes HTTP and HTML parsing capabilities; it does not establish a particular browser automation library or guarantee that any target can be scraped successfully. Rendering pages also adds operational complexity, so use it only when a simpler permitted route cannot supply the needed data.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If the task is to capture a page as an image or PDF rather than build a custom Kotlin extraction pipeline, ScreenshotNeo is a website screenshot API and MCP server made by Yorker Media. It is not a replacement for parsing structured fields from HTML. Its API accepts one GET request with a URL and can return a PNG, JPEG, WebP, or PDF. Before capture, it can accept the cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
For an image capture from code, the supplied cURL example is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for API details. Failed loads, bot checks or CAPTCHAs, blank pages, timeouts, and cache hits cost nothing; response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Best Value
Sign up for the free plan to try 1,000 screenshots a month with no card.
Troubleshooting Kotlin scraping
| Symptom | Likely cause | What to check |
|---|---|---|
| HTTP error or an error page in the output | The server returned an unsuccessful status, or the request failed. | Check the status before parsing, confirm the URL and access conditions, and inspect logs. Do not treat an error response as extracted content. |
| Parser finds no matching elements | The selector does not match the returned markup, or the content is inserted later by JavaScript. | Inspect the response HTML itself, then verify selectors against its actual structure. If the data is absent, investigate an official API or another permitted route. |
| Links are blank or relative | The raw attribute is relative, or parsing was not given a base URL. | Parse with the page URL as the base URI and use absUrl("href"); check that the element has an href. |
| Timeouts or connection errors | Network conditions, server response time, or a timeout that does not suit the job. | Record the exception and request URL, set sensible timeout limits, and retry only transient failures with backoff. Avoid aggressive retries. |
| Output has missing or malformed values | Optional markup, unexpected formatting, or a page change. | Handle absent fields explicitly, normalize and parse values with known rules, and validate records before storage. |
| It works locally but fails in a different Kotlin target | jsoup is a Java library and the selected Ktor engine or dependency may not support that target. | Confirm platform support and engine requirements for the versions in use. For this tutorial, the intended stack is Kotlin/JVM. |
FAQ
Can I use jsoup with Kotlin?
Yes, in a Kotlin/JVM project. jsoup is a Java library, so its API can be called from JVM Kotlin code. Check compatibility separately before choosing a non-JVM Kotlin target.
Is scraping legal?
There is no blanket answer for every site, dataset, jurisdiction, and use. Review the applicable site terms, access rules, privacy and copyright obligations, and law for your specific situation. Robots.txt is a crawler protocol, not authorization.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Does this example prove a particular website can be scraped?
No. The code demonstrates a general request-and-parse workflow; it does not test a live target or verify its terms, markup, or access rules. Inspect and validate each target independently.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

