Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a small, known set of pages, fetch HTML with an Elixir HTTP client such as Req or HTTPoison, then use Floki to extract fields with CSS selectors. When you need to discover links, schedule requests, filter domains, control duplicates, or process items through reusable stages, use Crawly. Floki is the parser; Crawly is the crawler framework.

Choose a direct scraper or a crawler framework

The right design depends on whether you already know the URLs and how much orchestration the job needs. A short script is easier to reason about for a few pages. A crawler framework becomes useful when each response can lead to more requests and the crawl needs shared policies.

Need HTTP client and Floki Crawly
One page or a short list of known URLs Usually simpler: request each URL, parse the body, and return data. May add unnecessary overhead.
Follow pagination or discovered links You write and maintain URL traversal. Spider callbacks can return follow-up requests for scheduling.
Domain filtering and duplicate-request control Implement and test these policies yourself. Documented middleware includes domain filtering and duplicate control.
Reusable processing and output stages Add application code for validation and serialization. Documented pipelines provide processing stages.
Content created by browser-side JavaScript Ordinary HTTP fetching does not run page JavaScript; you need a separate rendering solution. Crawly documents configurable browser rendering.

Neither approach is universally faster or suitable for every target. The library documentation does not establish a general throughput comparison; choose according to crawl scope and required controls.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a small scraper with Req and Floki

Req handles the HTTP request, and Floki parses the response HTML and searches it with CSS selectors. The example below extracts product-card titles and prices from a page you are authorized to access. Replace the example URL and selectors with the target site’s actual structure.

1. Add the dependencies

In a Mix project, add Req and Floki to deps in mix.exs:

defp deps do
  [
    {:req, "~> 0.7"},
    {:floki, "~> 0.38"}
  ]
end

Then fetch dependencies:

mix deps.get

These version constraints reflect the documented Req v0.7.4 and Floki v0.38.x material available for this article; check the current package documentation before starting a new project because releases and behavior can change. Req is an extensible HTTP client with documented redirect, retry, response-decoding, and streaming capabilities. Review the current Req documentation at hexdocs.pm/req/Req.html for the options and behavior you plan to use.

2. Fetch, parse, and extract

Place this in a module file such as lib/product_scraper.ex:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
defmodule ProductScraper do
  def fetch_products(url) do
    case Req.get(url, headers: [{"user-agent", "ExampleCatalogBot/1.0 (contact: [email protected])"}], receive_timeout: 15_000) do
      {:ok, %{status: status, body: html}} when status in 200..299 and is_binary(html) ->
        html
        |> Floki.parse_document!()
        |> Floki.find(".product-card")
        |> Enum.map(&product_from_card/1)

      {:ok, %{status: status}} ->
        {:error, {:http_status, status}}

      {:error, exception} ->
        {:error, {:request_failed, exception}}
    end
  end

  defp product_from_card(card) do
    title = card |> Floki.find(".product-title") |> Floki.text() |> String.trim()
    price = card |> Floki.find(".price") |> Floki.text() |> String.trim()

    %{title: present(title), price: present(price)}
  end

  defp present(""), do: nil
  defp present(value), do: value
end

Run it from IEx with a URL and inspect the returned list or error:

iex -S mix
> ProductScraper.fetch_products("https://example.com/catalog")

The selectors are examples, not a promise about any particular site’s markup. Inspect representative pages and verify each selector against the actual HTML before relying on the output. Returning nil for a missing field makes absence explicit; if a missing title or price should invalidate a record, validate it and return an error rather than silently treating incomplete data as complete.

3. Decide how to handle response bodies and failures

A production fetcher should distinguish a successful HTTP response from a transport error and a non-success status. The example returns errors for the latter two rather than trying to parse them as product pages. Add bounded retry behavior only when appropriate for the target and request type; retries can increase load, especially when the server is already returning errors.

Also consider redirects, response size, timeouts, character encoding, and partial results. HTTPoison is another Elixir HTTP client; its request documentation notes that synchronous responses buffer the whole response in memory, so streaming is relevant when bodies are large. See the HTTPoison API documentation and its request documentation for version-specific behavior. Req also documents streaming and response steps; confirm details for the version you install.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract reliably with Floki

Floki’s job is to parse HTML and let your code find nodes, including with CSS selectors. It is not an HTTP client, link scheduler, or browser JavaScript engine. Its overview and API reference describe parsing and node-search functions: Floki overview and Floki v0.38.4 API reference.

  • Select stable targets: Prefer meaningful classes, attributes, or structural relationships over brittle positional assumptions such as “the third span.”
  • Extract attributes when needed: A link’s destination is generally in its href attribute, while visible text comes from node text. Check missing attributes before using them.
  • Normalize deliberately: Trim whitespace and decide how to represent absent, malformed, or repeated values. Preserve raw values when later validation or auditing matters.
  • Test more than one page: Templates can vary across categories, pagination states, or logged-out and logged-in views. A selector that works once can still miss another page type.

Parsing an HTTP response does not execute scripts. If a product list appears only after client-side JavaScript runs, the raw response may not contain it. Confirm that the needed content is present in the downloaded HTML; otherwise use an appropriate rendering solution rather than assuming Floki will evaluate the page.

Follow links without losing control of the crawl

For a few known pages, an explicit URL list may be enough. For discovered pagination or site links, extract candidate URLs, resolve relative paths against the current page URL, and schedule only URLs within the intended scope. Maintain a visited set or equivalent deduplication policy so cycles and repeated links do not trigger unlimited requests.

Keep crawl scope explicit

  • Decide which hostnames and paths are in scope before following links.
  • Normalize URLs consistently before deduplication, while avoiding transformations that change their meaning.
  • Set a crawl boundary, such as a maximum page count or depth, if discovery could expand unexpectedly.
  • Keep the response URL and originating page with extracted records when provenance will help diagnose changed output.

These are application-level responsibilities in a direct scraper. If link discovery, scheduling, duplicate filtering, and reusable request policies are central to the task, Crawly supplies a higher-level spider workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Crawly is the better fit

Crawly organizes a crawl around spiders, requests, middleware, and pipelines. Its documented quickstart uses Floki to parse a page, extract fields, and return items. The README’s sample demonstrates parsing product cards, following a “next” link, validating data, filtering duplicate requests, encoding JSON, and writing output. Those selectors and values are examples only; they do not describe another site’s markup.

In Crawly’s documented request flow, a spider processes a response and can emit items and follow-up requests. Middleware can apply request policies, while pipelines handle item processing. Its v0.17.2 basic concepts documentation lists mechanisms including robots.txt handling, domain filtering, duplicate control, and user-agent behavior. Consult the versioned Crawly v0.17.2 documentation and the Crawly repository README for setup details that match the release you use.

Use the framework when the value of its orchestration exceeds the additional concepts and configuration. For a one-off fetch or a short list, a direct client plus Floki usually keeps the code smaller.

Set responsible request and crawl policies

  • Identify the client honestly. Use a descriptive user agent with a contact route where practical; do not impersonate an ordinary person’s browser to evade controls.
  • Choose conservative concurrency and timeouts. Start with a modest request rate and adjust only in light of target behavior and permission. Crawly’s configuration guide describes per-domain concurrency settings.
  • Respect robots.txt and site rules. Crawly documents robots.txt middleware; use it where appropriate. Its configuration guidance specifically advises against bypassing robots.txt on third-party sites without permission.
  • React to server signals. A 429 or rising 5xx responses should prompt you to reduce pressure, pause, or retry in accordance with the site’s policy. Crawly’s configuration guide notes that aggressive rate limiting and increased 5xx rates can indicate that concurrency should be lowered.
  • Review the actual use. Site terms, access controls, privacy, copyright, and applicable law depend on the target and your purpose. General library documentation cannot settle those site-specific questions.

Framework middleware does not make a crawl automatically appropriate or harmless. Scope, identity, rate, storage, and permitted use remain design decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If what you need is a screenshot rather than extracted HTML fields, ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request can return a PNG, JPEG, WebP, or PDF, and its options include element capture, full-page shots, viewport and device settings, custom CSS and JavaScript, and PDF controls. It does not replace Floki or Crawly for structured extraction and link crawling.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie and consent banners are accepted like a visitor and removed along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the response indicating the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Try it by signing up for ScreenshotNeo free.

Troubleshooting common failures

Symptom Likely cause What to check or change
Request returns a non-2xx status The server rejected the request, rate-limited it, redirected unexpectedly, or returned an error. Inspect status and response headers/body, verify the URL and allowed access, and reduce request pressure for 429 or elevated 5xx responses.
Request times out The host is slow, unreachable, or the timeout is too short for the response. Check connectivity and target behavior; choose a timeout appropriate to the task and avoid unbounded retries.
Floki finds no matching nodes The selector does not match the current markup, the page template differs, or content is inserted by JavaScript. Inspect the actual response HTML, test selectors on representative pages, and use browser rendering if the content exists only after client-side execution.
Fields are blank or inconsistent A selector matches the wrong node, content varies across templates, or text/attributes are missing. Check matching nodes and attributes, handle missing values explicitly, and validate required fields before output.
Memory rises on large responses The complete response is buffered in memory. Consider streaming when appropriate; HTTPoison’s synchronous request documentation notes whole-response buffering.
Crawl repeats pages or leaves the intended site URL normalization, deduplication, or domain scope is missing or misconfigured. Resolve relative links, enforce host/path filters, and enable or implement duplicate control and crawl boundaries.

Performance, reliability, and cost considerations

There is no universal speed figure that can be inferred from the framework documentation. Actual throughput depends on the target, network, response sizes, parsing work, concurrency, and policy constraints. Measure your own permitted workload rather than assuming a framework or client is inherently faster.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For reliability, treat network errors, redirects, missing fields, encoding differences, and partial results as ordinary cases. Log enough context to diagnose failures without collecting unnecessary personal data. A retry policy should be bounded and sensitive to status codes; repeatedly retrying a rate-limited server can worsen the problem. For memory-sensitive work, consider the client’s streaming facilities and the size of the pages being processed.

The cited libraries are open-source software with versioned documentation, but this material does not establish current licensing terms, hosted-service costs, or support arrangements. Check the relevant package and project pages for the release you select. Your operational costs may also include compute, storage, browser rendering infrastructure, and the time needed to maintain selectors.

FAQ

Is there a BeautifulSoup-like library for Elixir?

Floki is the closest match for HTML parsing and CSS-selector extraction. It is not a crawler scheduler or a JavaScript-capable browser.

Should I start with Crawly for a small job?

Usually not if you already have a short list of URLs. A direct HTTP client with Floki is simpler until link discovery and shared crawl policies become meaningful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can Elixir scraping retrieve content behind a login or access restriction?

The general library documentation does not determine whether a particular access method is permitted. Check the site’s terms and permissions before using credentials, cookies, or other access mechanisms.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.