October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
browser automation

Web Scraping With Go in 2026: When to Use Colly, goquery, or Browser Automation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Colly to fetch and coordinate a crawl, goquery to query the HTML you received, and chromedp when the page must actually run in a browser. They are complementary layers rather than three interchangeable scraping libraries. Start with HTTP plus HTML parsing whenever the fields are present in the response; add browser automation only for JavaScript execution, interaction, or browser-only state.

The three layers: crawler, parser, browser

Colly: requests and crawl coordination

The Colly project describes itself as “a Golang framework for building web scrapers.” Its documented responsibilities include making HTTP requests, invoking callbacks for received content, managing cookies, caching, concurrency, robots.txt handling, and distributed crawling. Colly is therefore the orchestration layer: it discovers URLs, applies boundaries, fetches responses, and routes them to your handlers.

goquery: querying a document

goquery provides chainable methods, similar to jQuery, for querying and manipulating an HTML document. It does not fetch URLs, execute JavaScript, click controls, or behave like a browser. Give it HTML (often from a Colly response or a normal Go HTTP request), then use CSS selectors to extract text, attributes, and links.

chromedp: controlling Chrome through CDP

chromedp drives browsers that support the Chrome DevTools Protocol (CDP). Its documented use cases include navigation, DOM queries, scraping, testing, profiling, and headless operation. A browser is the right layer when the useful content appears only after scripts run, when you must click or type, or when page state depends on browser APIs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the smallest tool that can see the data

Requirement Likely fit Decisions to make
Crawl many ordinary HTTP pages Colly Allowed domains, URL filters, depth, request limits, retries, caching, per-domain delay and concurrency
Select fields from returned HTML goquery Selector stability, malformed markup, changing document structure
Interact with rendered content chromedp JavaScript requirement, wait conditions, browser lifecycle, deployment and resource cost

These are engineering criteria, not benchmark results. The available documentation does not provide a controlled, directly comparable speed test for the three tools, so measure your own URLs and workload before claiming that one is faster.

A maintainable Colly and goquery crawler

Keep fetching and extraction separate. Colly can discover and fetch URLs; goquery can parse each response. This example collects article titles and links while limiting the crawl to one domain and two levels.

package main

import (
    "fmt"
    "log"

    "github.com/gocolly/colly/v2"
    "github.com/PuerkitoBio/goquery"
)

func main() {
    c := colly.NewCollector(
        colly.AllowedDomains("example.com"),
        colly.MaxDepth(2),
    )

    c.Limit(&colly.LimitRule{
        DomainGlob:  "example.com/*",
        Parallelism: 2,
        Delay:       500 * time.Millisecond,
    })

    c.OnHTML("article", func(e *colly.HTMLElement) {
        title := e.DOM.Find("h1, h2").First().Text()
        link, _ := e.DOM.Find("a").First().Attr("href")
        fmt.Printf("%st%sn", title, link)
    })

    c.OnResponse(func(r *colly.Response) {
        doc, err := goquery.NewDocumentFromReader(bytes.NewReader(r.Body))
        if err != nil {
            log.Printf("parse %s: %v", r.Request.URL, err)
            return
        }
        _ = doc.Find("title").First().Text()
    })

    c.OnError(func(r *colly.Response, err error) {
        log.Printf("%s: %v", r.Request.URL, err)
    })

    if err := c.Visit("https://example.com/"); err != nil {
        log.Fatal(err)
    }
}

Add bytes and time to the import block. The selector and domain are intentionally examples: replace them with the target site’s structure and your permitted scope. Verify the current canonical goquery module path before running go mod tidy; package documentation has also appeared under the legacy gopkg.in/goquery.v1 path.

Why parse in OnResponse?

OnHTML is convenient for selector callbacks. Parsing in OnResponse gives you an explicit document and makes it easier to pass structured data to another function, retain response metadata, or handle pages whose markup does not match one fixed callback. Do not assume either callback makes JavaScript run: Colly receives HTTP responses, not a rendered browser DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bound the crawl before you start

  • Domain scope: allow only hosts you intend to visit.
  • URL scope: use URL and response filters to exclude logout links, calendars, search explosions, and large downloads.
  • Depth and count: set maximum depth and request limits so a link graph cannot grow without bound.
  • Rate: configure per-domain delay and concurrency appropriate to the site.
  • State: use Colly’s cookie, cache, and retry facilities deliberately; cached responses can hide changes during development.
  • Robots and permission: Colly’s current source checks robots.txt unless configured otherwise. That is an implementation control, not a legal opinion or permission to access a site. Follow the target’s terms, access rules, and applicable law.

Colly documents these controls, including robots support, URL and domain restrictions, concurrency, caching, and distributed patterns. Treat package behavior as version-sensitive and inspect the version you deploy.

When HTTP parsing is not enough

Signs you need a browser

  • The initial response contains an empty application shell and data appears only after JavaScript executes.
  • A button, tab, infinite-scroll trigger, or form must be used before the data exists.
  • The page depends on browser storage, layout events, or other browser APIs.
  • You need to inspect the DOM after rendering rather than the original response.

Use chromedp for these cases. Create a browser context, navigate, wait for a meaningful selector, then extract text or attributes. Prefer a specific readiness condition over a fixed sleep.

package main

import (
    "context"
    "fmt"
    "log"
    "time"

    "github.com/chromedp/chromedp"
)

func main() {
    ctx, cancel := chromedp.NewContext(context.Background())
    defer cancel()

    ctx, cancel = context.WithTimeout(ctx, 45*time.Second)
    defer cancel()

    var title string
    err := chromedp.Run(ctx,
        chromedp.Navigate("https://example.com/"),
        chromedp.WaitVisible("main"),
        chromedp.Text("h1", &title, chromedp.NodeVisible),
    )
    if err != nil {
        log.Fatal(err)
    }
    fmt.Println(title)
}

Install a chromedp version compatible with your Go toolchain and Chrome/Chromium deployment. The package listing available for this article identifies chromedp v0.16.0, published July 14, 2026; verify the current release before pinning it. Browser automation adds a real browser process and its operational overhead, but the reviewed sources do not quantify that overhead against HTTP parsing.

A practical decision sequence

  1. Request one target URL without a browser and inspect the returned HTML.
  2. If the required fields are present, use Colly for crawl control and goquery (or Colly’s HTML callbacks) for extraction.
  3. If fields appear only after scripts or interaction, use chromedp for the affected workflow.
  4. Keep the browser portion narrow: discover URLs or fetch static pages with HTTP, and reserve browser sessions for pages that truly require them.
  5. Record failures, response status, final URL, selector misses, and timing so you can distinguish site changes from network errors.

Common failures and fixes

“The selector returns nothing”

Inspect the actual response body. The selector may be wrong, the markup may have changed, or the data may be injected by JavaScript. Correct the selector for goquery, or move that page to chromedp if the response is only an application shell.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The crawler visits too many URLs”

Add allowed domains, URL and response filters, maximum depth, and a request limit. Exclude query parameters that create unbounded combinations.

“Requests fail or time out”

Log the URL and error, set an explicit timeout, and review retries and per-domain rates. A timeout is not evidence that a browser is required; first check DNS, TLS, redirects, server responses, and your limits.

“Robots rules block a request”

Confirm the target’s robots.txt and access policy. Colly makes robots checking configurable, but disabling it does not grant permission to crawl.

“Chromedp cannot find the element”

Wait for a stable selector, verify the page and frame you navigated to, and use a context deadline. A fixed delay can be useful as a diagnostic, but a readiness selector is generally less brittle.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The browser works locally but not in production”

Check that a compatible Chrome/Chromium binary is installed, that the process can start in the deployment environment, and that your context timeout covers navigation and rendering. Keep browser concurrency bounded.

Performance, reliability, and cost decisions

HTTP fetching plus HTML parsing avoids launching a browser and is usually the simpler operational path, but no controlled evidence here establishes a universal speed ratio. Browser sessions consume additional runtime resources and introduce lifecycle, rendering, and wait-condition failure modes. Benchmark representative pages, concurrency, response sizes, cache behavior, and extraction accuracy on your infrastructure.

Reliability comes from bounded work and observable failure handling: deterministic limits, explicit timeouts, retries that do not amplify load, cache policies suited to freshness requirements, and alerts for selector misses. Version-lock your dependencies, then revalidate behavior when upgrading Colly, goquery, chromedp, or Chrome.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than a custom Go crawler, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It also offers full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for options. The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. An MCP server lets AI agents take screenshots without your maintaining a browser setup. Create a free ScreenshotNeo account.

FAQ

Are Colly and goquery alternatives?

No. Colly coordinates requests and crawling; goquery parses and queries a document. A common design uses both.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do I need a browser for every JavaScript site?

No. Inspect the response first. Some sites expose the needed data in HTML or an underlying request that you can access with permission; use a browser when execution or interaction is actually required.

Is Colly’s advertised request rate a benchmark?

No general comparison should rely on it. The cited project material does not provide a reproducible setup, and no directly comparable test covers these tools.

Which Go browser framework is best for cross-browser work?

The available evidence does not establish a current cross-browser winner. chromedp targets browsers supporting CDP; evaluate other requirements and test your deployment.

Frequently Asked Questions

Can I combine Colly, goquery, and chromedp in one project?

Yes. A typical architecture uses Colly for URL discovery and ordinary responses, goquery for HTML extraction, and chromedp only for routes that require rendering or interaction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does ScreenshotNeo replace a data crawler?

No. It is intended for screenshots, PDFs, page information, and agent-driven capture; use a crawler and parser when you need structured records from many pages.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.