Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Use Colly to fetch and coordinate a crawl, goquery to query the HTML you received, and chromedp when the page must actually run in a browser. They are complementary layers rather than three interchangeable scraping libraries. Start with HTTP plus HTML parsing whenever the fields are present in the response; add browser automation only for JavaScript execution, interaction, or browser-only state.
The three layers: crawler, parser, browser
Colly: requests and crawl coordination
The Colly project describes itself as “a Golang framework for building web scrapers.” Its documented responsibilities include making HTTP requests, invoking callbacks for received content, managing cookies, caching, concurrency, robots.txt handling, and distributed crawling. Colly is therefore the orchestration layer: it discovers URLs, applies boundaries, fetches responses, and routes them to your handlers.
goquery: querying a document
goquery provides chainable methods, similar to jQuery, for querying and manipulating an HTML document. It does not fetch URLs, execute JavaScript, click controls, or behave like a browser. Give it HTML (often from a Colly response or a normal Go HTTP request), then use CSS selectors to extract text, attributes, and links.
chromedp: controlling Chrome through CDP
chromedp drives browsers that support the Chrome DevTools Protocol (CDP). Its documented use cases include navigation, DOM queries, scraping, testing, profiling, and headless operation. A browser is the right layer when the useful content appears only after scripts run, when you must click or type, or when page state depends on browser APIs.
#1 Best Overall
Choose the smallest tool that can see the data
| Requirement | Likely fit | Decisions to make |
|---|---|---|
| Crawl many ordinary HTTP pages | Colly | Allowed domains, URL filters, depth, request limits, retries, caching, per-domain delay and concurrency |
| Select fields from returned HTML | goquery | Selector stability, malformed markup, changing document structure |
| Interact with rendered content | chromedp | JavaScript requirement, wait conditions, browser lifecycle, deployment and resource cost |
These are engineering criteria, not benchmark results. The available documentation does not provide a controlled, directly comparable speed test for the three tools, so measure your own URLs and workload before claiming that one is faster.
A maintainable Colly and goquery crawler
Keep fetching and extraction separate. Colly can discover and fetch URLs; goquery can parse each response. This example collects article titles and links while limiting the crawl to one domain and two levels.
package main
import (
"fmt"
"log"
"github.com/gocolly/colly/v2"
"github.com/PuerkitoBio/goquery"
)
func main() {
c := colly.NewCollector(
colly.AllowedDomains("example.com"),
colly.MaxDepth(2),
)
c.Limit(&colly.LimitRule{
DomainGlob: "example.com/*",
Parallelism: 2,
Delay: 500 * time.Millisecond,
})
c.OnHTML("article", func(e *colly.HTMLElement) {
title := e.DOM.Find("h1, h2").First().Text()
link, _ := e.DOM.Find("a").First().Attr("href")
fmt.Printf("%st%sn", title, link)
})
c.OnResponse(func(r *colly.Response) {
doc, err := goquery.NewDocumentFromReader(bytes.NewReader(r.Body))
if err != nil {
log.Printf("parse %s: %v", r.Request.URL, err)
return
}
_ = doc.Find("title").First().Text()
})
c.OnError(func(r *colly.Response, err error) {
log.Printf("%s: %v", r.Request.URL, err)
})
if err := c.Visit("https://example.com/"); err != nil {
log.Fatal(err)
}
}
Add bytes and time to the import block. The selector and domain are intentionally examples: replace them with the target site’s structure and your permitted scope. Verify the current canonical goquery module path before running go mod tidy; package documentation has also appeared under the legacy gopkg.in/goquery.v1 path.
Why parse in OnResponse?
OnHTML is convenient for selector callbacks. Parsing in OnResponse gives you an explicit document and makes it easier to pass structured data to another function, retain response metadata, or handle pages whose markup does not match one fixed callback. Do not assume either callback makes JavaScript run: Colly receives HTTP responses, not a rendered browser DOM.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Bound the crawl before you start
- Domain scope: allow only hosts you intend to visit.
- URL scope: use URL and response filters to exclude logout links, calendars, search explosions, and large downloads.
- Depth and count: set maximum depth and request limits so a link graph cannot grow without bound.
- Rate: configure per-domain delay and concurrency appropriate to the site.
- State: use Colly’s cookie, cache, and retry facilities deliberately; cached responses can hide changes during development.
- Robots and permission: Colly’s current source checks robots.txt unless configured otherwise. That is an implementation control, not a legal opinion or permission to access a site. Follow the target’s terms, access rules, and applicable law.
Colly documents these controls, including robots support, URL and domain restrictions, concurrency, caching, and distributed patterns. Treat package behavior as version-sensitive and inspect the version you deploy.
When HTTP parsing is not enough
Signs you need a browser
- The initial response contains an empty application shell and data appears only after JavaScript executes.
- A button, tab, infinite-scroll trigger, or form must be used before the data exists.
- The page depends on browser storage, layout events, or other browser APIs.
- You need to inspect the DOM after rendering rather than the original response.
Use chromedp for these cases. Create a browser context, navigate, wait for a meaningful selector, then extract text or attributes. Prefer a specific readiness condition over a fixed sleep.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/chromedp/chromedp"
)
func main() {
ctx, cancel := chromedp.NewContext(context.Background())
defer cancel()
ctx, cancel = context.WithTimeout(ctx, 45*time.Second)
defer cancel()
var title string
err := chromedp.Run(ctx,
chromedp.Navigate("https://example.com/"),
chromedp.WaitVisible("main"),
chromedp.Text("h1", &title, chromedp.NodeVisible),
)
if err != nil {
log.Fatal(err)
}
fmt.Println(title)
}
Install a chromedp version compatible with your Go toolchain and Chrome/Chromium deployment. The package listing available for this article identifies chromedp v0.16.0, published July 14, 2026; verify the current release before pinning it. Browser automation adds a real browser process and its operational overhead, but the reviewed sources do not quantify that overhead against HTTP parsing.
A practical decision sequence
- Request one target URL without a browser and inspect the returned HTML.
- If the required fields are present, use Colly for crawl control and goquery (or Colly’s HTML callbacks) for extraction.
- If fields appear only after scripts or interaction, use chromedp for the affected workflow.
- Keep the browser portion narrow: discover URLs or fetch static pages with HTTP, and reserve browser sessions for pages that truly require them.
- Record failures, response status, final URL, selector misses, and timing so you can distinguish site changes from network errors.
Common failures and fixes
“The selector returns nothing”
Inspect the actual response body. The selector may be wrong, the markup may have changed, or the data may be injected by JavaScript. Correct the selector for goquery, or move that page to chromedp if the response is only an application shell.
“The crawler visits too many URLs”
Add allowed domains, URL and response filters, maximum depth, and a request limit. Exclude query parameters that create unbounded combinations.
“Requests fail or time out”
Log the URL and error, set an explicit timeout, and review retries and per-domain rates. A timeout is not evidence that a browser is required; first check DNS, TLS, redirects, server responses, and your limits.
“Robots rules block a request”
Confirm the target’s robots.txt and access policy. Colly makes robots checking configurable, but disabling it does not grant permission to crawl.
“Chromedp cannot find the element”
Wait for a stable selector, verify the page and frame you navigated to, and use a context deadline. A fixed delay can be useful as a diagnostic, but a readiness selector is generally less brittle.
“The browser works locally but not in production”
Check that a compatible Chrome/Chromium binary is installed, that the process can start in the deployment environment, and that your context timeout covers navigation and rendering. Keep browser concurrency bounded.
Performance, reliability, and cost decisions
HTTP fetching plus HTML parsing avoids launching a browser and is usually the simpler operational path, but no controlled evidence here establishes a universal speed ratio. Browser sessions consume additional runtime resources and introduce lifecycle, rendering, and wait-condition failure modes. Benchmark representative pages, concurrency, response sizes, cache behavior, and extraction accuracy on your infrastructure.
Reliability comes from bounded work and observable failure handling: deterministic limits, explicit timeouts, retries that do not amplify load, cache policies suited to freshness requirements, and alerts for selector misses. Version-lock your dependencies, then revalidate behavior when upgrading Colly, goquery, chromedp, or Chrome.
Rank #4
Or skip the browser setup
If your goal is a clean image or PDF rather than a custom Go crawler, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
It also offers full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options. The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. An MCP server lets AI agents take screenshots without your maintaining a browser setup. Create a free ScreenshotNeo account.
FAQ
Are Colly and goquery alternatives?
No. Colly coordinates requests and crawling; goquery parses and queries a document. A common design uses both.
Free tools Windows power users keep installed
One-click scans. No signup required.
Do I need a browser for every JavaScript site?
No. Inspect the response first. Some sites expose the needed data in HTML or an underlying request that you can access with permission; use a browser when execution or interaction is actually required.
Best Value
Is Colly’s advertised request rate a benchmark?
No general comparison should rely on it. The cited project material does not provide a reproducible setup, and no directly comparable test covers these tools.
Which Go browser framework is best for cross-browser work?
The available evidence does not establish a current cross-browser winner. chromedp targets browsers supporting CDP; evaluate other requirements and test your deployment.
Frequently Asked Questions
Can I combine Colly, goquery, and chromedp in one project?
Yes. A typical architecture uses Colly for URL discovery and ordinary responses, goquery for HTML extraction, and chromedp only for routes that require rendering or interaction.
Does ScreenshotNeo replace a data crawler?
No. It is intended for screenshots, PDFs, page information, and agent-driven capture; use a crawler and parser when you need structured records from many pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




