Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use libcurl to download the response and libxml2 to parse the HTML and query it with XPath. The dependable workflow is: initialize libcurl, fetch into a bounded buffer, validate the transfer and HTTP response, parse with htmlReadMemory, extract nodes through an XPath context, and free every curl and libxml2 resource. The complete program below extracts a page title, headings, and links, then adds timeouts, redirect limits, a user agent, and a maximum response size.
What you need
This approach is for pages whose useful data is present in the HTTP response (server-rendered HTML or an endpoint that returns HTML). libcurl is the transfer layer: it handles HTTP, HTTPS, redirects, cookies, authentication, timeouts, and many other protocol details. libxml2 supplies an HTML parser and XPath 1.0 implementation. Neither library is a browser, and neither executes page JavaScript.
Compiler and development packages
Install the development packages for libcurl, libxml2, and a C++ compiler using your operating system’s package manager. Package names and installation paths differ between Linux distributions, macOS, and Windows, so treat the following build command as an example rather than a universal recipe.
pkg-config --cflags --libs libxml-2.0 libcurl
The command should print include and linker flags. If your platform does not provide pkg-config metadata, use the include and library directories supplied by your installation.
#1 Best Overall
A bounded, runnable C++ scraper
Save this as scrape.cpp. It accepts one URL, follows at most five redirects, stops connecting after two seconds, stops the entire transfer after 20 seconds, and refuses to retain more than 10 MiB. The callback deliberately aborts an oversized response instead of allowing unbounded memory growth.
#include <curl/curl.h>
#include <libxml/HTMLparser.h>
#include <libxml/parser.h>
#include <libxml/tree.h>
#include <libxml/xpath.h>
#include <libxml/uri.h>
#include <iostream>
#include <string>
#include <cstdlib>
struct Buffer {
std::string bytes;
std::size_t limit;
bool overflow = false;
};
static std::size_t write_callback(char* ptr, std::size_t size,
std::size_t nmemb, void* userdata) {
auto* out = static_cast<Buffer*>(userdata);
const std::size_t count = size * nmemb;
if (out->bytes.size() + count > out->limit) {
const std::size_t room = out->limit - out->bytes.size();
out->bytes.append(ptr, room);
out->overflow = true;
return 0; // makes curl stop with CURLE_WRITE_ERROR
}
out->bytes.append(ptr, count);
return count;
}
static void print_nodes(xmlXPathObjectPtr result, const char* label) {
if (!result || result->type != XPATH_NODESET || !result->nodesetval) return;
for (int i = 0; i < result->nodesetval->nodeNr; ++i) {
xmlNodePtr node = result->nodesetval->nodeTab[i];
xmlChar* value = xmlNodeGetContent(node);
if (value) {
std::cout << label << ": "
<< reinterpret_cast<const char*>(value) << "n";
xmlFree(value);
}
}
}
int main(int argc, char** argv) {
if (argc != 2) {
std::cerr << "usage: " << argv[0] << " https://example.comn";
return 2;
}
const std::string requested_url = argv[1];
constexpr std::size_t max_bytes = 10 * 1024 * 1024;
Buffer body{"", max_bytes};
if (curl_global_init(CURL_GLOBAL_DEFAULT) != 0) {
std::cerr << "curl_global_init failedn";
return 1;
}
CURL* curl = curl_easy_init();
if (!curl) {
curl_global_cleanup();
return 1;
}
curl_easy_setopt(curl, CURLOPT_URL, requested_url.c_str());
curl_easy_setopt(curl, CURLOPT_FOLLOWLOCATION, 1L);
curl_easy_setopt(curl, CURLOPT_MAXREDIRS, 5L);
curl_easy_setopt(curl, CURLOPT_CONNECTTIMEOUT_MS, 2000L);
curl_easy_setopt(curl, CURLOPT_TIMEOUT_MS, 20000L);
curl_easy_setopt(curl, CURLOPT_USERAGENT, "freedom251-example-scraper/1.0");
curl_easy_setopt(curl, CURLOPT_WRITEFUNCTION, write_callback);
curl_easy_setopt(curl, CURLOPT_WRITEDATA, &body);
const CURLcode transfer = curl_easy_perform(curl);
long status = 0;
char* content_type = nullptr;
curl_easy_getinfo(curl, CURLINFO_RESPONSE_CODE, &status);
curl_easy_getinfo(curl, CURLINFO_CONTENT_TYPE, &content_type);
if (transfer != CURLE_OK) {
std::cerr << "transfer failed: " << curl_easy_strerror(transfer) << "n";
if (body.overflow) std::cerr << "response exceeded the configured limitn";
curl_easy_cleanup(curl);
curl_global_cleanup();
return 1;
}
if (status < 200 || status >= 300 || body.overflow || body.bytes.empty()) {
std::cerr << "unexpected response: HTTP " << status
<< ", bytes=" << body.bytes.size() << "n";
curl_easy_cleanup(curl);
curl_global_cleanup();
return 1;
}
if (content_type)
std::cerr << "content type: " << content_type << "n";
htmlDocPtr doc = htmlReadMemory(
body.bytes.data(), static_cast<int>(body.bytes.size()),
requested_url.c_str(), nullptr,
HTML_PARSE_NONET | HTML_PARSE_RECOVER |
HTML_PARSE_NOERROR | HTML_PARSE_NOWARNING);
if (!doc) {
std::cerr << "libxml2 could not parse the responsen";
curl_easy_cleanup(curl);
curl_global_cleanup();
return 1;
}
xmlXPathContextPtr context = xmlXPathNewContext(doc);
if (!context) {
xmlFreeDoc(doc);
curl_easy_cleanup(curl);
curl_global_cleanup();
return 1;
}
xmlXPathObjectPtr title = xmlXPathEvalExpression(
BAD_CAST "//title", context);
print_nodes(title, "title");
xmlXPathFreeObject(title);
xmlXPathObjectPtr headings = xmlXPathEvalExpression(
BAD_CAST "//h1 | //h2 | //h3", context);
print_nodes(headings, "heading");
xmlXPathFreeObject(headings);
xmlXPathObjectPtr links = xmlXPathEvalExpression(
BAD_CAST "//a[@href]/@href", context);
if (links && links->type == XPATH_NODESET && links->nodesetval) {
for (int i = 0; i < links->nodesetval->nodeNr; ++i) {
xmlNodePtr attr = links->nodesetval->nodeTab[i];
xmlChar* href = xmlNodeGetContent(attr);
if (!href) continue;
xmlChar* absolute = xmlBuildURI(href, BAD_CAST requested_url.c_str());
std::cout << "link: "
<< (absolute ? reinterpret_cast<const char*>(absolute)
: reinterpret_cast<const char*>(href))
<< "n";
if (absolute) xmlFree(absolute);
xmlFree(href);
}
}
xmlXPathFreeObject(links);
xmlXPathFreeContext(context);
xmlFreeDoc(doc);
curl_easy_cleanup(curl);
curl_global_cleanup();
return 0;
}
Build and run it
With packages that expose pkg-config metadata, compile with:
g++ -std=c++17 -Wall -Wextra scrape.cpp -o scrape
$(pkg-config --cflags --libs libxml-2.0 libcurl)
An equivalent command using explicit installation prefixes looks like this:
g++ -Wall -I/opt/curl/include -I/opt/libxml/include/libxml2
scrape.cpp -o scrape -L/opt/curl/lib -L/opt/libxml/lib -lcurl -lxml2
Run it against a page that returns HTML:
./scrape https://example.com/
You should see the response content type on standard error, followed by title, heading, and link records on standard output. The parser receives the requested URL as its base URL; that lets xmlBuildURI resolve relative href values. For provenance, store the final URL, retrieval timestamp, HTTP status, and any parser or transfer error alongside each extracted record.
How the pipeline works
1. Initialize and configure libcurl
Call curl_global_init once for the process, create an easy handle, and set options before curl_easy_perform. A write callback may be called many times and may receive binary data, so append the byte count rather than treating each chunk as a null-terminated string. The example uses a descriptive user agent because libcurl otherwise sends no user-agent by default.
2. Validate the transfer, not just the body
A successful CURLcode means the transfer completed, not that the server returned the document you wanted. Check the HTTP status, whether the body is empty, whether the callback hit its limit, and (when appropriate) the content type. Redirects can lead to a login page, an error document, or another host; record the effective URL with CURLINFO_EFFECTIVE_URL if your application needs the final location.
3. Parse as HTML, safely
htmlReadMemory is designed for HTML and can recover from common malformed markup. HTML_PARSE_NONET prevents the parser from fetching external network resources while parsing downloaded content. Suppressing parser warnings is convenient for a command-line example, but production code should capture diagnostics when malformed input matters. Do not enable external entity or network behavior merely to make a broken page appear complete.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall4. Query with XPath and release objects
Create one XPath context per document, evaluate expressions, and free each xmlXPathObjectPtr. Node text returned by xmlNodeGetContent is allocated by libxml2 and must be released with xmlFree. Free the XPath context and document before cleaning up curl. Test expressions against representative pages: malformed nesting, repeated elements, missing attributes, and namespaces can all change the result.
Useful XPath patterns
| Goal | XPath | Notes |
|---|---|---|
| Document title | string(//title) |
Returns a scalar string; an absent title produces an empty value. |
| All article headings | //article//h1 | //article//h2 | //article//h3 |
Restricting the search to an article avoids navigation headings. |
| Links with a class | //a[contains(concat(' ', normalize-space(@class), ' '), ' product ')] |
The spacing expression matches one class token rather than a substring. |
| Data attribute | //*[@data-id]/@data-id |
Inspect attributes directly when visible text is not a stable identifier. |
| Text normalization | normalize-space(string(.)) |
Collapses runs of whitespace; perform additional encoding conversion in C++ when your storage format requires it. |
XPath 1.0 does not understand CSS selector syntax. Keep expressions small, log the expression used for each field, and treat a missing node as a data-quality condition instead of silently converting it to a valid-looking empty record.
From one page to a polite crawler
The official crawler pattern evaluates //a/@href, resolves URLs with libxml2 URI helpers, and applies explicit bounds. Add these controls before introducing a queue or worker threads.
| Control | Starting policy | Why it matters |
|---|---|---|
| Per-request connect timeout | 2 seconds | Prevents dead hosts from consuming a worker indefinitely. |
| Total transfer timeout | 20 seconds | Bounds slow responses and stalled reads. |
| Redirects | CURLOPT_MAXREDIRS set to 5 |
Stops redirect loops and limits unexpected cross-site hops. |
| Maximum response size | Choose a site-specific cap; the example uses 10 MiB | Protects memory and avoids parsing an accidental large download. |
| Pages per crawl | Set a global maximum | Makes a run predictable and recoverable. |
| Links per page | Set a queueing maximum | Prevents a single index page from exploding the frontier. |
| Concurrency | Use a small, bounded worker pool | Reduces load on the target and limits local sockets, memory, and file descriptors. |
| Cookies and authentication | Enable only when required | Credentials and session state can expose private data or cross trust boundaries. |
Canonicalize and deduplicate URLs
Resolve relative references against the response URL, normalize only transformations that are safe for your target, and maintain a visited set. Decide whether fragments should be removed, whether hostnames are case-normalized, and whether tracking query parameters are meaningful before deduplicating. Keep the original discovered URL as well as the canonical queue key for auditability.
Recommended Free Tools
Retry only transient failures
Retry connection resets, temporary DNS failures, and selected 5xx responses with capped exponential backoff and jitter. Do not blindly retry 4xx responses, authentication failures, parser errors, or a response that exceeded your size limit. Respect the site’s terms, access controls, robots policy, and published rate limits; a technically successful request can still be an unacceptable crawl.
JavaScript-rendered pages: know the boundary
libcurl transfers resources; it does not create a browser DOM or execute JavaScript. If the fields appear only after client-side code runs, first look for an allowed server-rendered page, a documented API, or the same JSON endpoint used by the page. An endpoint is usually cheaper and easier to validate than browser automation.
| Approach | Transfer control | Malformed HTML/XPath | JavaScript execution | Resource profile |
|---|---|---|---|---|
| libcurl + libxml2 | Direct control of timeouts, redirects, cookies, headers, and authentication | HTML recovery plus XPath 1.0 | No | Low overhead; easy to bound |
| Browser automation | Browser-level controls and a rendered session | Browser DOM and JavaScript behavior | Yes | Higher CPU, memory, startup, and operational cost |
| Documented API | Protocol-specific controls | Structured response validation | Not applicable | Usually the smallest and most stable payload |
Do not attempt to solve a JavaScript requirement by increasing parser recovery flags: parsing cannot execute scripts or observe data that was never in the response.
Security and operational safeguards
- Use an honest identifying
CURLOPT_USERAGENT; do not impersonate a browser to evade controls. - Constrain redirects when credentials, cookies, or authorization headers are present. Avoid sending secrets to a different host unless that behavior is explicitly intended.
- Store credentials outside source code and logs. Review whether cookies contain personal or authenticated data before persisting responses.
- Keep response limits, page limits, queue limits, and concurrency limits configurable and visible in run logs.
- Parse downloaded HTML with
HTML_PARSE_NONETunless a narrowly justified external-resource behavior is required. - Treat scraped content as untrusted input. Escape it for the output context, and validate encodings before inserting it into SQL, HTML, JSON, or shell commands.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
CURLE_COULDNT_RESOLVE_HOST |
DNS failure, typo, or blocked resolver | Verify the URL and DNS from the same machine; do not turn this into an unlimited retry loop. |
CURLE_OPERATION_TIMEDOUT |
Slow connection or response | Inspect connect versus total timeout, then increase limits only for a known slow target. |
CURLE_WRITE_ERROR with an overflow message |
The callback reached the maximum response size | Raise the cap deliberately, stream to a file, or reject the page; never remove the bound without a replacement. |
| HTTP 403 or 429 | Access policy or rate limiting | Stop aggressive retries, identify your client honestly, reduce concurrency, and follow the site’s published policy. |
| HTML parses but fields are empty | Wrong XPath, a different template, or client-rendered content | Save a redacted sample, inspect the actual markup, test the expression, and check for an API or server-rendered variant. |
| Only a login page is extracted | Authentication or session cookie is required | Implement the site’s authorized authentication flow, isolate credentials, and validate the final URL and status before parsing. |
| Relative links are malformed | The parser was not given a correct base URL | Pass the effective response URL to htmlReadMemory and resolve with xmlBuildURI. |
| Crashes or leaks during long runs | Unreleased XPath objects, documents, curl handles, or unbounded queues | Use RAII wrappers or cleanup paths, cap the queue, and test with sanitizers. |
Licensing and distribution
curl and libcurl use the permissive curl license, inspired by MIT/X; commercial and closed-source use is allowed when the required copyright and permission notices are retained. libxml2 is distributed under an MIT license. Preserve both notices in your distribution and review the licenses of linked TLS backends and other transitive dependencies separately.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Or skip the browser setup
If your goal is a clean image or PDF rather than parsed fields, ScreenshotNeo is a hosted alternative. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
One request returns PNG, JPEG, WebP, or PDF. The service also supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and arbitrary viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks before capture, hidden selectors, waits for selectors or network idle, request and resource blocking, custom headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and a usage API. Existing integrations can use the parameter names used by other screenshot APIs.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for authentication and output options.
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Sign up for the free ScreenshotNeo plan.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteFAQ
Can I use CSS selectors instead of XPath?
Not directly in libxml2’s XPath engine. Translate the selector into XPath, or add a separate selector library; test the translated expression against the exact HTML templates you crawl.
Should I parse the response before checking its HTTP status?
No. Validate the transfer result, status, size, and content type first. Error pages and login forms are valid HTML but invalid inputs for the record you intended to collect.
Best Value
Is the example safe for arbitrary URLs?
It has useful bounds, but a production service still needs an outbound-network policy, redirect and credential rules, queue limits, logging, and authorization checks appropriate to its environment.
Frequently Asked Questions
Can I use CSS selectors instead of XPath?
Not directly in libxml2’s XPath engine. Translate the selector into XPath, or add a separate selector library; test the translated expression against the exact HTML templates you crawl.
Free tools Windows power users keep installed
One-click scans. No signup required.
Should I parse the response before checking its HTTP status?
No. Validate the transfer result, status, size, and content type first. Error pages and login forms are valid HTML but invalid inputs for the record you intended to collect.
Is the example safe for arbitrary URLs?
It has useful bounds, but a production service still needs an outbound-network policy, redirect and credential rules, queue limits, logging, and authorization checks appropriate to its environment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

