Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For straightforward OCaml web scraping, use Cohttp to fetch a page and Lambda Soup to parse its HTML and select elements with CSS selectors. Choose the Cohttp backend that fits your application’s concurrency runtime: Lwt, Async, curl, or Eio. This approach works with HTML returned by the server; the package documentation does not establish that it runs JavaScript or renders a browser page.
How OCaml web scraping fits together
A scraper has two distinct jobs: retrieve a response over HTTP, then interpret the response body. Cohttp provides HTTP client implementations; Lambda Soup provides a document-oriented API for parsing HTML and extracting text or attributes. Keeping those tasks separate makes it easier to diagnose whether a problem is the request, the response content, or the selector.
This stack is a good starting point when a page’s useful content is present in its returned HTML. It is not, by itself, a browser automation system. If a page fills its content with client-side JavaScript, inspect the response you actually receive before assuming that a selector is wrong.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchChoose a Cohttp backend for your runtime
| Backend | When to consider it | Documented qualification |
|---|---|---|
| Cohttp Lwt | Your application already uses Lwt, or you want an asynchronous client in the Lwt ecosystem. | Separate backend package; see Cohttp package documentation. |
| Cohttp Async | Your application is built around Async. | Separate backend package; see Cohttp package documentation. |
| Cohttp curl | You want the curl-backed implementation. | Separate backend package; see Cohttp package documentation. |
| Cohttp Eio | Your application uses Eio and direct-style concurrency. | The Cohttp Eio package describes multicore support for OCaml 5.0+. |
These are implementation choices, not evidence that one backend is faster for scraping. No comparative throughput benchmark is established here. Match the backend to your existing runtime, dependency constraints, and deployment target rather than selecting on an unsupported performance claim.
#1 Best Overall
Package catalog results list Cohttp 6.3.0 and Cohttp Eio 6.3.0, published August 21, 2026. Lambda Soup is listed as version 1.1.1, with a package-page publication date of September 5, 2024; Markup.ml is listed as version 1.0.3. Treat these as catalog observations, not compatibility guarantees. Before pinning, check the current opam constraints for your OCaml version and the exact backend package you plan to install.
Install a practical Lwt and Lambda Soup setup
The example below uses Cohttp’s Lwt Unix client and Lambda Soup. Install the packages in an opam-managed project, then run the program with the dependencies resolved in that switch:
opam install cohttp-lwt-unix lambda-soup
Save this as scrape.ml. It requests one URL, checks the HTTP status, parses the response body, and prints the page title and each link’s visible text and href when present.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
open Lwt.Infix
let scrape url =
let uri = Uri.of_string url in
Cohttp_lwt_unix.Client.get uri >>= fun (response, body) ->
Cohttp_lwt.Body.to_string body >|= fun html ->
let status = Cohttp.Response.status response in
let code = Cohttp.Code.code_of_status status in
if code < 200 || code >= 300 then
failwith (Printf.sprintf "HTTP request failed with status %d" code);
let document = Soup.parse html in
let title =
match Soup.select_one "title" document with
| None -> "(no title element)"
| Some node -> Soup.text node
in
Printf.printf "Title: %sn" title;
Soup.select "a" document
|> Soup.to_list
|> List.iter (fun node ->
let label = Soup.text node in
match Soup.attribute "href" node with
| None -> Printf.printf "Link: %s (no href)n" label
| Some href -> Printf.printf "Link: %s -> %sn" label href)
let () =
try Lwt_main.run (scrape "https://example.com")
with exn ->
prerr_endline (Printexc.to_string exn);
exit 1
The program deliberately reports a non-success status rather than treating an error page as ordinary content. The body is still read before that check; for a production scraper that handles large bodies or needs richer error reporting, consider how and when to consume the response body for your chosen client interface.
Adapt selectors to the page you need
Lambda Soup describes itself as an HTML scraping library inspired by Python’s Beautiful Soup. Its CSS selectors and document traversals are useful when you want to express extraction in terms of page structure. In the example, Soup.select "a" finds links, Soup.text reads a node’s text, and Soup.attribute "href" retrieves an attribute when one exists.
- Target a specific region: replace a broad selector such as
"a"with a selector for the page’s relevant container and descendants. Verify the selector against the HTML you fetched. - Handle missing elements: use the optional result from
select_onerather than assuming every page contains a title or target node. - Handle missing attributes: an element need not have the attribute you expect. The example checks for an absent
hrefinstead of dereferencing it blindly. - Normalize output deliberately: page text may contain whitespace, nested text, or unexpected markup. Decide what normalization your downstream data requires, and test it on representative pages.
CSS selectors describe document structure, not a guarantee that a site will keep that structure stable. If you rely on a particular class, attribute, or nesting pattern, treat changes to it as an expected maintenance case and validate extracted records before storing or acting on them.
Rank #3
When Markup.ml is a better parsing fit
Lambda Soup is the convenient choice when you want a document and selector-based extraction. Its documentation says it is based on Markup.ml. Markup.ml provides HTML5 and XML parsing, error recovery, and lazy, streaming, single-pass processing through parser signals.
Consider Markup.ml directly when streaming input, controlling parser signals, or avoiding a document-oriented extraction workflow is central to the task. It is also relevant when you need to reason explicitly about malformed markup and parser recovery. The right choice depends on the shape of your input and the control you need; the available material does not establish a universal speed or memory advantage for either library.
Know what a plain HTTP scraper cannot tell you
A successful HTTP response is not the same thing as the page a person sees in a browser. The package documentation establishes HTTP clients and HTML parsing, but does not establish that Cohttp plus Lambda Soup executes page JavaScript, handles browser challenges, or reproduces browser rendering. If the HTML body lacks the content you expect, inspect it before changing selectors.
Rank #4
- Used Book in Good Condition
- Check the response status and whether the response body contains the target content.
- Compare the returned markup with the browser-visible page, especially if the page appears to fill in content after loading.
- If the content is absent from the response, investigate whether the site exposes an appropriate data endpoint or whether a browser-rendering workflow is necessary.
- Reassess the extraction selector only after confirming the content is present in the HTML being parsed.
Do not infer permission to automate access from a library’s capabilities. Whether scraping is allowed, what rate limits apply, and which access methods are appropriate depend on the target site and applicable rules. Check the site’s terms and other relevant requirements, and avoid assuming that a package supplies anti-bot handling or permission.
Troubleshooting common scraping failures
| Symptom | Likely explanation | What to check |
|---|---|---|
| The request does not compile or a module is unavailable. | The project may lack the selected backend package, or its opam environment may not be active. | Confirm the switch, installed package names, OCaml constraints, and backend chosen for the code. |
| The program receives an error status. | The server returned a non-success response, or the requested URL is not the endpoint you intended. | Print or log the status and inspect the URL and response behavior; do not parse an error page as a successful record. |
| A selector returns no matching nodes. | The fetched HTML may differ from the browser view, the selector may not match, or the page structure may have changed. | Inspect the response HTML, then test a selector against that document. |
| Text or attributes are missing. | The selected node may not contain the expected text or attribute. | Handle absent nodes and attributes explicitly and validate output before using it. |
| The page looks empty in extracted results. | The relevant content may be added by JavaScript or otherwise unavailable in the response body. | Determine whether the content exists in the downloaded HTML; parsing alone does not establish browser rendering. |
| Results stop matching after a site update. | The markup or selectors may have changed. | Keep representative pages and expected output for validation, and revise selectors when the actual structure changes. |
Performance, reliability, and operating costs
Choose concurrency based on your workload and runtime, but do not assume that more simultaneous requests are always better. The package material establishes the available backend families, not throughput figures, target-site capacity, or a safe request rate. Start conservatively, respect the target’s stated rules, and handle request failures explicitly.
For reliability, distinguish transport and status failures from extraction failures. Record enough context to identify the URL, status, and missing expected fields without exposing credentials or unnecessarily retaining page data. Validate extracted values before downstream use; a syntactically successful parse can still yield incomplete or changed content.
Best Value
These are open-source OCaml libraries, not a hosted scraping service with a price established by the package descriptions. Your operational costs depend on your own hosting, network use, and chosen infrastructure; the available material does not provide a benchmark or total-cost comparison.
Or skip the browser setup
If your task needs a rendered screenshot or PDF rather than HTML extraction, ScreenshotNeo provides a website screenshot API and MCP server. A single request can return an image or PDF, but that output is not a substitute for parsed HTML records when your scraper needs structured text and attributes. Its capture flow accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers.
For example, this cURL request saves a WebP screenshot of Stripe. See the ScreenshotNeo API documentation for request options and response details.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorscurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

