Free tools Windows power users keep installed
One-click scans. No signup required.
How do you scrape a website in Go? Start with Go’s standard net/http client to download one page, parse the returned HTML with a selector library such as goquery, and adopt Colly when you need link traversal, domain limits, caching, cookies, robots.txt handling, or controlled concurrency. This tutorial builds those layers in order, with runnable code and practical failure handling.
Choose the right Go scraping approach
Fetching and parsing are separate jobs. net/http performs the HTTP request; it does not understand HTML structure. A parser such as goquery turns the response into a document that can be queried with CSS selectors. Colly adds crawler operations around requests and callbacks.
| Approach | What you write | Best fit |
|---|---|---|
net/http plus a parser |
Request lifecycle, URL checks, retries, queueing and extraction | One page, a small script, or a pipeline where every operation should be explicit |
| Colly | Collector configuration and callbacks | Multi-page crawling with repeatable traversal rules |
Neither approach is universally faster. Throughput depends on the target, network, parsing work, limits and configuration; no equivalent independent benchmark establishes a general winner.
Quick start: fetch one page with net/http
Create a module, then save this as main.go:
package main
import (
"fmt"
"io"
"log"
"net/http"
)
func main() {
resp, err := http.Get("https://example.com/")
if err != nil {
log.Fatal(err)
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("unexpected HTTP status: %s", resp.Status)
}
body, err := io.ReadAll(resp.Body)
if err != nil {
log.Fatal(err)
}
fmt.Printf("%s", body)
}
Run it with go run .. The important lifecycle is: create the request, handle a request error, close the body, check the status, and handle read errors. Always close a response body, including when you later replace http.Get with a custom client.
#1 Best Overall
Use a client and timeout in production
The convenience function uses a default client without an overall timeout. A bounded client prevents a stalled server from occupying a worker indefinitely:
client := &http.Client{Timeout: 20 * time.Second}
resp, err := client.Get(targetURL)
Import time when using this form. For more control, build an http.Request so you can set a user agent, cookies, authorization or other headers intentionally. Do not disguise a crawler as a browser to evade access controls.
Parse HTML with goquery
Install the parser in your module:
go get github.com/PuerkitoBio/goquery
Read the response and pass it to goquery. This example prints headings and absolute links:
package main
import (
"fmt"
"log"
"net/http"
"net/url"
"github.com/PuerkitoBio/goquery"
)
func main() {
target := "https://example.com/"
resp, err := http.Get(target)
if err != nil {
log.Fatal(err)
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
log.Fatalf("unexpected HTTP status: %s", resp.Status)
}
doc, err := goquery.NewDocumentFromReader(resp.Body)
if err != nil {
log.Fatal(err)
}
base, err := url.Parse(target)
if err != nil {
log.Fatal(err)
}
doc.Find("h1, h2").Each(func(_ int, s *goquery.Selection) {
fmt.Println("heading:", s.Text())
})
doc.Find("a[href]").Each(func(_ int, s *goquery.Selection) {
href, ok := s.Attr("href")
if !ok {
return
}
link, err := base.Parse(href)
if err == nil {
fmt.Println("link:", link.String())
}
})
}
Selectors such as article h2, a[href] and meta[name='description'] are easier to maintain when the site uses stable semantic elements or classes. Test selectors against representative pages, including pages where a field is missing. Treat Selection.Text() as possibly empty and validate attributes before storing them.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBuild a crawler with Colly
Install Colly with:
go get github.com/gocolly/colly/v2
This crawler restricts requests to one domain, extracts page titles, resolves links, and visits them:
package main
import (
"fmt"
"log"
"github.com/gocolly/colly/v2"
)
func main() {
c := colly.NewCollector(
colly.AllowedDomains("example.com"),
)
c.OnHTML("title", func(e *colly.HTMLElement) {
fmt.Printf("%s: %sn", e.Request.URL, e.Text)
})
c.OnHTML("a[href]", func(e *colly.HTMLElement) {
link := e.Request.AbsoluteURL(e.Attr("href"))
if link != "" {
if err := c.Visit(link); err != nil {
fmt.Println("visit skipped:", err)
}
}
})
c.OnRequest(func(r *colly.Request) {
fmt.Println("visiting", r.URL.String())
})
c.OnError(func(r *colly.Response, err error) {
log.Printf("request failed: %s (%v)", r.Request.URL, err)
})
if err := c.Visit("https://example.com/"); err != nil {
log.Fatal(err)
}
}
AllowedDomains is a safety boundary, not merely an optimization. Add URL-pattern checks when only part of a site is in scope. Colly’s collector model also supports asynchronous operation, caching, cookies and robots.txt controls. Configure those features deliberately rather than enabling concurrency first and discovering that the target is overloaded.
Keep traversal finite
- Set an explicit start URL and allowed domains.
- Reject logout, search, calendar and other unbounded URL patterns unless they are required.
- Use a visited strategy and a maximum page count or depth for experiments.
- Save extracted records incrementally so a later failure does not lose the crawl.
Responsible crawling and request controls
Read the target’s robots.txt and terms before crawling, and keep the request rate low enough not to degrade service. Robots rules are a signal about permitted automated access, not a substitute for authorization where one is required.
- Check every response status; decide explicitly whether redirects and non-2xx responses are data, retries or failures.
- Set request timeouts and bound retries. Retrying a persistent 403 or 404 only creates more load.
- Use bounded concurrency and delays only after observing the target’s behavior.
- Cache responses during development. It reduces duplicate traffic and makes parser debugging repeatable.
- Record URL, status, timestamp and error for each failed request.
Colly documents domain restrictions, asynchronous operation, caching, cookies and robots.txt support. With a hand-built net/http crawler, you must implement equivalent queue, scope, retry and cache behavior yourself.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →JavaScript-rendered and protected pages
A plain HTTP client receives the server’s response; it does not execute page JavaScript. If the data appears only after scripts run, or a target presents a bot check or CAPTCHA, first confirm that an official API or static endpoint is available. A browser-capable or hosted capture service is an advanced branch, not a reason to add a browser dependency to every small scraper. Do not attempt to bypass a CAPTCHA or other access control without authorization.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when your workflow needs a rendered page rather than HTML parsing. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A one-call cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page and element capture, dark mode, device presets, retina scale, PDF settings, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user-agent, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameters used by other screenshot APIs also work for easier migration.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; Growth is $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.
Rank #4
Troubleshooting common failures
Timeouts or a hanging request
Set an overall client timeout, avoid unbounded retries, and log the URL that stalled. For Colly, combine request limits with a conservative delay; a timeout is not a reason to increase concurrency.
403, 429 or CAPTCHA responses
Confirm permission, slow down, honor the site’s terms and check for an official API. Do not rotate identities or automate CAPTCHA solving to evade controls.
Empty selectors
Inspect the actual response body. The content may be rendered by JavaScript, the selector may have changed, or the server may have returned an error page. Add fixture pages and tests for missing elements.
Relative or malformed links
Resolve links against the requesting page’s URL, reject empty or unsupported schemes, and apply the same domain and path restrictions to resolved URLs.
Best Value
Too many duplicate pages
Normalize URLs, remove fragments, constrain query parameters and use Colly’s visit tracking or a persistent cache. Calendar and search parameters commonly create unbounded paths.
Operational checklist
- Define the permitted domain, paths, fields and maximum crawl size.
- Read robots.txt and terms; choose a low request rate.
- Start with
net/httpand a parser for one representative page. - Add status checks, body closure, timeouts, structured logs and bounded retries.
- Move to Colly when traversal, caching, cookies or crawler callbacks justify the dependency.
- Test selectors against changed layouts and missing fields.
- Use a browser-capable or hosted path only for authorized JavaScript-rendered requirements.
FAQ
Is Go good for web scraping?
Yes. The standard library handles HTTP clearly, goquery provides CSS-oriented extraction, and Colly supplies crawler structure. The appropriate choice depends on whether you need one page or controlled multi-page traversal.
Does Colly execute JavaScript?
Colly is an HTTP crawler; it does not provide a general browser runtime. Pages whose data is created only after JavaScript executes require a different, authorized approach.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Should I scrape with an IP rotation service?
Not by default. First establish authorization, reduce request volume, honor site rules and look for an official data interface. Rotation does not make prohibited crawling permissible.
How do I preserve data when a crawl stops?
Write each validated record as it is extracted, and log URL, status and error fields. This allows a later run to resume or reprocess only failed pages.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

