Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Install GoSpider from its Go module, then start with a shallow crawl of a site you are authorized to test: gospider -s "https://example.com/" -o output -c 10 -d 1. The key controls are -s for one site, -S for a file of sites, -d for recursion depth, and -c for concurrent requests to matching domains. Verify the installed binary with gospider --version; version labels differ between distribution sources.

What GoSpider does—and what it does not

GoSpider is an open-source command-line web spider written in Go. It follows and reports URLs it discovers while crawling a site or a list of sites. Its documented capabilities include finding links in JavaScript, trying sitemap and robots files, including subdomains, using third-party URL sources, accepting Burp request input, running crawls in parallel, using random user agents, and producing grep-friendly output. These are discovery options, not guarantees: a crawl can only report material it can reach and identify under the selected settings.

GoSpider is for discovering URLs, not for taking clean page screenshots. If your goal is a visual capture rather than a crawl, ScreenshotNeo is a separate website screenshot API and MCP server; its one-call example appears after the GoSpider workflow below.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use GoSpider only against systems you own or are explicitly authorized to assess. A crawler can make many requests, and depth, concurrency, subdomain inclusion, and external URL sources can materially expand the scope.

Install GoSpider and check the binary

Install with Go

The upstream project documents this Go module installation command:

GO111MODULE=on go install github.com/jaeles-project/gospider@latest

Run it in an environment with Go tooling available. If the shell cannot find gospider afterward, check that Go’s installed-binary directory is included in your executable search path, then open a new shell or invoke the binary by its full path.

Build and run with Docker

The documented Docker route is to clone the GoSpider repository, build the image from the gospider directory, and run its help command:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
docker build -t gospider:latest gospider
docker run -t gospider -h

The build command assumes the repository has been cloned and that your current directory is its parent, so that gospider names the build context. If you are elsewhere, adjust the context path to the cloned directory. The example runs the container’s help output; use the installed CLI’s help to confirm the arguments available in the image you built.

Verify the version and available flags

After either installation path, check the actual binary rather than assuming that all packages carry the same release:

gospider --version
gospider --help

The upstream README usage block shows v1.1.5, while the Kali Linux tools page lists v1.1.6. Those labels demonstrate that sources may differ; they do not establish which version a particular machine has. Use the version command and help output from the binary you will run when confirming an option or diagnosing a difference.

Run a first, bounded crawl

One site, minimal command

To crawl a single authorized site, pass its starting URL with -s:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
gospider -s "https://example.com/"

For a repeatable first pass that saves output, sets a modest request limit, and keeps recursion shallow, use:

gospider -s "https://example.com/" -o output -c 10 -d 1

Replace the example URL with a target in scope. This command selects the output folder, allows up to 10 concurrent requests for matching domains, and limits recursion depth to 1. The project’s documented default concurrency is 5 and default request timeout is 10 seconds; these are program defaults, not speed or completion guarantees.

Understand depth before increasing it

-d or --depth sets the maximum recursion depth. The README says 0 means infinite recursion, so do not use zero casually: an open-ended crawl can keep finding URLs and expand the request volume substantially. For an initial check, -d 1 is a restrained starting point. Increase depth only when the authorization, target capacity, and task require it.

Control request concurrency and timeout

-c or --concurrent sets the maximum concurrent requests for matching domains. It is not the same as the number of sites processed in parallel. -m or --timeout sets request timeout in seconds. A timeout that is too short can fail to capture slower responses; raising it can make a crawl wait longer on unresponsive requests. The right values depend on the target and network, and the documentation does not establish a universal optimal setting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with a modest concurrency such as -c 5 or -c 10, and a shallow depth. If the target is sensitive or the crawl is producing too much traffic, reduce concurrency and use --delay or --random-delay to slow request scheduling. Check the installed binary’s help for accepted syntax and defaults.

Crawl a list of sites and save the results

Put one site per line in a text file, such as sites.txt, then use -S or --sites. For example:

gospider -S sites.txt -o output -c 10 -d 1 -t 20

Here, -S supplies the newline-delimited input list, -o selects the output folder, -c controls concurrent requests for matching domains, and -d controls crawl depth. -t or --threads controls how many sites run in parallel. This separates site-level parallelism from per-domain request concurrency: raising both can increase the total load across the set of targets.

Check that the file contains the intended in-scope sites before starting. If a list includes unrelated hosts or redirects into unexpected areas, stop and review scope rather than assuming that the input filename constrains every discovered URL.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an output format you can use

-o or --output chooses an output folder. Other documented output controls include:

  • --json for JSON output.
  • -q or --quiet to suppress other output and print URLs.
  • -v or --verbose for verbose logs.
  • -l or --length to show response length.
  • -L or --filter-length to filter by lengths.
  • -R or --raw for raw output.

Pick output for the next stage of your workflow: quiet URL output is convenient when another command consumes a URL list, while JSON is suitable when your downstream process expects structured data. The documentation does not specify a JSON schema or establish that all flags combine identically across versions, so validate the help output and a small sample before building automation around a particular format.

Customize requests for an authorized crawl

Headers and cookies

For a permitted authenticated or specialized request, the README documents repeated -H headers and a cookie string. For example:

gospider -s "https://example.com/" -H "Accept: */*" -H "Test: test" --cookie "testA=a; testB=b"

Use --cookie for the cookie header value and repeat -H for additional headers. Treat cookies, authorization headers, and other credentials as secrets: avoid committing them to source control or exposing them in shared command logs. Only send session data to a target for which you have permission.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Proxy and user agent

-p or --proxy configures a proxy. -u or --user-agent accepts the built-in random web or mobile agents, or a custom user-agent string. A proxy is useful when your authorized workflow requires controlled egress; a user-agent setting changes the request identity presented to the site. Neither option grants permission to crawl a target or guarantees that a site will respond in a particular way.

Import a Burp request

Use --burp with a raw Burp request file when you need GoSpider to load headers and cookies from that request:

gospider -s "https://example.com/" --burp burp_req.txt

Keep the request file protected if it contains session credentials. Confirm that the request’s host and authentication context match the authorized target before starting.

Expand URL discovery deliberately

GoSpider’s discovery options can add sources beyond links found during ordinary traversal. Enable only what helps answer the assessment question:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • --js enables JavaScript link finding.
  • --sitemap tries sitemap.xml.
  • --robots tries robots.txt.
  • --subs includes subdomains.
  • --other-source obtains URLs from Archive.org, Common Crawl, VirusTotal, and AlienVault.
  • --include-subs and --include-other-source broaden how those URLs are incorporated.

The README also lists AWS S3 references and link-finder behavior among the project’s features. Discovery can surface URLs that are old, external, or outside the scope you intended; inspect the resulting hosts and paths before feeding them into additional tools. Enabling more sources is not equivalent to proof that a page currently exists or is accessible.

Filter scope and keep request volume responsible

The project examples show --blacklist for URL regular expressions, and note that common static file extensions are filtered by default. A blacklist can help exclude known paths or patterns from a crawl, but regular expressions should be tested carefully: an overly broad expression may discard useful URLs, while a narrow one may not constrain scope as expected.

A cautious tuning sequence is:

  1. Confirm the exact authorized hostnames and whether subdomains are included.
  2. Start from a specific in-scope URL with -d 1.
  3. Use moderate -c concurrency and an appropriate -m timeout.
  4. Add a delay if a slower request pace is appropriate.
  5. Enable JavaScript, sitemap, robots, subdomain, or external-source discovery only when needed.
  6. Review output for unexpected hosts and excessive volume before increasing depth or parallelism.

GoSpider’s documented defaults and examples describe configuration, not independent performance measurements. Crawl time depends on the target, network, response behavior, scope, and selected options; no universal speed estimate follows from the available documentation.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common problems

gospider: command not found

The executable may not be on the shell’s search path, or the installation did not complete in the expected environment. Confirm that the Go install command completed, locate the installed binary, and add its directory to the path or invoke it by full path. With Docker, check that the image build succeeded and that you are running the image name you built.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A flag is rejected or behaves differently

Package versions can differ. Run gospider --version and gospider --help on the exact binary in use. Compare its available arguments with the command you copied, especially if you installed from a distribution package rather than the upstream Go module.

The crawl returns few URLs

Check that the starting URL is reachable from the machine running GoSpider, that the chosen depth is not too shallow for the paths you expect, and that the relevant discovery options are enabled. JavaScript-derived links, sitemap or robots locations, subdomains, and third-party sources are optional; a standard crawl does not imply they were all consulted.

Requests time out or the crawl seems slow

Review -m and the network path. A timeout that is too short can fail on slow responses; a longer timeout can increase waiting on stalled requests. Adjust concurrency and delay in light of target capacity and authorization rather than increasing concurrency automatically. The documented defaults are 10 seconds for timeout and 5 for concurrency, not a promise of a particular completion time.

Too much output or unexpected URLs appear

Use a shallow depth, reduce discovery breadth, inspect the site list, and apply a carefully tested --blacklist pattern where appropriate. If subdomains or third-party sources were enabled, review the discovered hosts against the engagement scope before processing the results further.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

GoSpider discovers URLs; it does not return screenshots. If you need a screenshot of a page rather than a crawl, ScreenshotNeo takes a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. For example, using cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o shot.webp

See the ScreenshotNeo API documentation for setup and options. Cookie banners are accepted and removed along with known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status in headers. An MCP server exposes screenshot and PDF tools to AI agents. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. These are screenshot features, not substitutes for GoSpider’s URL crawling.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

FAQ

Does GoSpider guarantee that every page on a site will be found?

No. It reports URLs it can discover through the selected crawl and discovery methods. Pages that are unlinked, inaccessible, generated in ways the crawler does not identify, or excluded by filtering may not appear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is GoSpider a vulnerability scanner?

The documented purpose here is web crawling and URL discovery. A URL list may support a broader authorized assessment, but crawling alone does not establish that a vulnerability exists.

Can I use it on a list of domains?

Yes. Put one site per line in a file and pass it with -S; use -t to set how many sites run in parallel.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.