Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To add search to a website, build or choose a pipeline that discovers pages, fetches only content you are allowed to crawl, extracts and normalizes text, indexes it, ranks matches, and serves results through an API and search page. For a quick public-site search, a hosted engine may be enough; for private content, custom ranking, or tighter control over recrawling and deletion, you will need a system you operate or a managed crawler you can configure.
“Any website” needs a boundary: decide which domains and paths are in scope, whether the pages are public, and how fresh results must be. A crawler cannot safely treat the whole web—or even every URL on one domain—as an unrestricted collection.
Choose hosted search or a search system you control
Start with the constraints, not the technology. Google’s Programmable Search Engine supports a website, blog, or collection of sites, with options including ranking customization, embedded search, structured-data features, and optional AdSense monetization. Google’s setup tutorial describes adding whole sites, individual URLs, or URL patterns, then using a hosted search page or embedding a search box. That can be the shortest route when its scope, presentation, and data-handling boundaries fit.
A managed crawler and search service is a middle ground: Elastic’s crawler material describes adding a domain to a managed engine, letting its crawler discover pages, and tuning result weights. A self-operated stack gives you more control over crawling, ranking, access checks, and deletion, but you also own each component and its failure modes.
| Approach | Good fit when | Trade-off to assess |
|---|---|---|
| Hosted engine | You need public website search quickly and its inclusion rules, UI, and data handling are acceptable. | Confirm current scope, presentation options, terms, quotas, and pricing with the provider before committing. |
| Managed crawler and search | You want a crawler and index without operating every part yourself, while retaining some tuning control. | Check control over access, freshness, ranking, deletion, analytics, and operational limits. |
| Self-operated | You need private-content access controls, custom ranking or analyzers, specific data residency, or predictable recrawl and deletion behavior. | You must build and operate the crawler, parser, index, query service, monitoring, and security controls. |
Compare options on the same checklist: which URLs are included, how quickly changes appear, expected query latency, privacy and authorization, implementation and operating effort, analytics, and any monetization. Product names, terms, quotas, and prices change; verify them directly before choosing.
#1 Best Overall
Define what your search is allowed to include
Write down scope before the first fetch. Record allowed domains and URL patterns, languages, content types, freshness targets, and who is allowed to see each document. Decide how to treat query parameters, pagination, print views, archives, and URLs that generate effectively unlimited combinations.
- Public content: specify which public pages belong in results and how visitors can filter them.
- Private content: require authentication at crawl time and enforce access again when serving each result. Do not assume that hiding a result in the interface protects its contents.
- Removal and changes: define how an updated, deleted, or newly restricted page is removed or replaced in the index, and how quickly that must happen.
- Freshness: set different recrawl intervals for frequently changing and stable sections, then measure actual index lag.
Build the pipeline in this order
1. Discover approved URLs
Start with URLs you already know are in scope and the site’s XML sitemaps. Fetch and parse robots.txt for each host before following links, then follow only links that pass your domain and path rules. Keep a descriptive user agent and a per-host request limit. A sitemap is a discovery aid, not permission to fetch every listed URL or a guarantee that each URL belongs in search.
2. Fetch carefully and record outcomes
Handle redirects, HTTP status codes, timeouts, retries, compression, and content types deliberately. Store the requested URL, final URL, fetch time, status, and error reason so a failed fetch is distinguishable from a page that was successfully checked and intentionally excluded. Use bounded retries with backoff; retrying every error immediately can overload a host and grow the queue without fixing the cause.
Do not fetch arbitrary URLs supplied by visitors without strict controls. A crawler or query service that can request user-chosen addresses can become a path to internal services or local network resources. Restrict hosts, validate redirects, and keep fetching isolated from sensitive networks.
Rank #2
3. Extract useful content
Parse HTML into a document record instead of indexing raw markup. Preserve the page title, headings, main text, useful metadata, and links. Remove repeated navigation and boilerplate where possible, normalize Unicode and whitespace, and detect language if your search supports more than one language. Treat HTML as untrusted input; do not render extracted markup directly in result snippets.
Some sites put meaningful content in JavaScript-rendered pages. Google says its crawler can render JavaScript, but that does not mean every custom crawler can; rendering adds browser, time, and resource costs. First check whether the site offers server-rendered HTML or an API. Add browser rendering only for the pages that need it, and set explicit limits on scripts, network requests, and wait time.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
4. Canonicalize and deduplicate
Resolve redirects and inspect canonical URLs so alternate addresses for the same content do not become competing results. Normalize only URL details that are genuinely equivalent for your site; do not blindly strip query parameters that distinguish real content. Give each logical document a stable identity and retain versions or update timestamps so a recrawl can replace the old record safely.
5. Create an index
An inverted index maps terms to documents and is a practical base for lexical search. Index separate fields—such as title, headings, and body—so ranking can value a title match more than a common word in the footer. Choose tokenization and stemming or lemmatization for the languages you support. Add phrase and prefix matching, filters, and snippets when the product requires them. Keep enough document state to update or delete a record without leaving stale terms behind.
Rank #3
6. Rank and serve results
Begin with lexical relevance, for example BM25, then test whether field boosts, phrase matches, freshness, popularity or link signals, synonyms, and editorial rules improve real queries. Expose search through a query API with pagination, highlighting, spelling suggestions, and filters as needed. Set timeouts and abuse controls, enforce access checks at query time, and cache only results that are safe to share between users.
7. Build the search experience and operate it
A search box is only the entry point. Results need clear titles, useful snippets, relevant filters, pagination, and an empty state that helps a visitor revise the query. Track which queries lead to clicks, which are reformulated, and which return no results. Operationally, schedule incremental recrawls, back off when hosts fail, monitor queue depth, crawl errors, index lag, and query latency, and process deletions as a first-class job.
A small Python crawler and search prototype
This standard-library example crawls a limited number of pages on one host, consults that host’s robots.txt, excludes pages marked noindex, and offers basic title-and-text matching. It is a learning prototype, not a production search engine: it does not render JavaScript, parse XML sitemaps, provide robust canonical handling, or implement a mature ranking model. Save it as site_search.py, then run python site_search.py https://example.com and enter a query when prompted. Replace the example URL with a site you are authorized to crawl.
import sys
import time
import re
import urllib.parse
import urllib.robotparser
import urllib.request
from html.parser import HTMLParser
USER_AGENT = "ExampleSiteSearchBot/1.0 (contact: [email protected])"
MAX_PAGES = 100
DELAY_SECONDS = 1.0
class PageParser(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
self.links = []
self.title = []
self.in_title = False
self.skip_depth = 0
self.noindex = False
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag in ("script", "style", "noscript", "svg"):
self.skip_depth += 1
if tag == "title":
self.in_title = True
if tag == "meta" and attrs.get("name", "").lower() in ("robots", "googlebot"):
if "noindex" in attrs.get("content", "").lower():
self.noindex = True
if tag == "a" and attrs.get("href"):
self.links.append(attrs["href"])
def handle_endtag(self, tag):
if tag in ("script", "style", "noscript", "svg") and self.skip_depth:
self.skip_depth -= 1
if tag == "title":
self.in_title = False
def handle_data(self, data):
text = data.strip()
if not text:
return
if self.in_title:
self.title.append(text)
if not self.skip_depth:
self.parts.append(text)
def fetch(url, timeout=10):
request = urllib.request.Request(url, headers={"User-Agent": USER_AGENT})
with urllib.request.urlopen(request, timeout=timeout) as response:
content_type = response.headers.get("Content-Type", "")
final_url = response.geturl()
if "text/html" not in content_type.lower():
return final_url, "", ""
raw = response.read(2_000_000)
charset = response.headers.get_content_charset() or "utf-8"
return final_url, raw.decode(charset, errors="replace"), content_type
def main(start):
parsed = urllib.parse.urlsplit(start)
if parsed.scheme not in ("http", "https") or not parsed.hostname:
raise SystemExit("Use an absolute http:// or https:// URL")
host = parsed.hostname.lower()
robots_url = urllib.parse.urlunsplit((parsed.scheme, parsed.netloc, "/robots.txt", "", ""))
robots = urllib.robotparser.RobotFileParser()
robots.set_url(robots_url)
try:
robots.read()
except Exception as exc:
raise SystemExit(f"Could not read robots.txt; stopping rather than assuming permission: {exc}")
queue, seen, docs = [start], set(), []
while queue and len(seen) < MAX_PAGES:
url = queue.pop(0)
clean = urllib.parse.urldefrag(url)[0]
if clean in seen:
continue
seen.add(clean)
candidate = urllib.parse.urlsplit(clean)
if candidate.hostname != host or not robots.can_fetch(USER_AGENT, clean):
continue
try:
final_url, html, _ = fetch(clean)
if urllib.parse.urlsplit(final_url).hostname != host or not html:
continue
parser = PageParser()
parser.feed(html)
if not parser.noindex:
docs.append((final_url, " ".join(parser.title), " ".join(parser.parts)))
for href in parser.links:
target = urllib.parse.urljoin(final_url, href)
target_parts = urllib.parse.urlsplit(target)
if target_parts.hostname == host and target_parts.scheme in ("http", "https"):
queue.append(urllib.parse.urldefrag(target)[0])
except Exception as exc:
print(f"Fetch failed: {clean}: {exc}", file=sys.stderr)
time.sleep(DELAY_SECONDS)
print(f"Indexed {len(docs)} pages. Enter a query; Ctrl-C to quit.")
while True:
query = input("search> ").strip().lower()
terms = re.findall(r"w+", query)
if not terms:
continue
results = []
for url, title, body in docs:
haystack = (title + " " + body).lower()
score = sum(haystack.count(term) + 3 * title.lower().count(term) for term in terms)
if score:
results.append((score, url, title, body))
for score, url, title, body in sorted(results, reverse=True)[:10]:
snippet = re.sub(r"s+", " ", body)[:240]
print(f"n{title or url}n{url}n{snippet}")
if not results:
print("No matches. Try different words.")
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("Usage: python site_search.py https://example.com")
main(sys.argv[1])
The prototype demonstrates the pipeline rather than solving its hard production problems. Its simple occurrence count is not BM25; its document collection is held in memory; it has no durable index, sitemap ingestion, robust redirect/canonical policy, per-user authorization, or search API. Its host check is intentionally narrow, and its conservative behavior when robots cannot be read can stop a crawl that a production policy might handle differently. Add those capabilities deliberately before indexing a real site at scale.
Rank #4
Respect robots.txt, noindex, and private content
These controls answer different questions. robots.txt is a policy for crawler requests: Google explains that it can block crawler access. It is not a secrecy mechanism. A blocked crawler may not see a page’s noindex directive, so do not rely on a robots block to keep a known URL out of search. For material that must not appear, protect it with authentication or make sure the crawler can receive and process a noindex instruction; also remove any existing indexed copy when access rules change.
For Google Search specifically, Google’s technical requirements say a page must be accessible to Googlebot, return HTTP 200, and contain indexable content to be eligible, but eligibility does not guarantee indexing. Those are Google’s public-search rules, not an automatic specification for your own crawler. A site-search system must implement the access and inclusion rules it needs. Google also recommends using sitemaps, checking crawl access, and handling canonical URLs and duplicate content.
Maintain separate crawl and serving checks. A page may have been public when indexed but later become private; a previously stored document must not leak through cached results or snippets. When permissions change, update or remove the indexed content and invalidate any affected result caches.
Test relevance and reliability before launch
Build a test set from tasks real visitors perform, then hand-label which pages should rank for each query. Include exact product or person names, synonyms, misspellings, quoted phrases, filters, pagination, zero-result searches, stale and deleted pages, canonical duplicates, private pages, robots changes, JavaScript-only content, large documents, and hostile input. Tune ranking against these examples rather than assuming a particular algorithm is correct for your audience.
Best Value
Track measures that expose different failures: success rate against expected results, zero-result rate, query reformulation rate, p95 latency, index freshness, crawl error rate, and time to remove a page after deletion or access change. These are useful engineering measures, not universal targets; establish baselines for your site and set thresholds from its needs.
Troubleshooting common failures
- Pages never appear: check the allowed URL rules, robots policy, fetch status, content type, extraction output, noindex directives, and whether the page has actually entered the index.
- Search results are stale: inspect recrawl scheduling, queue backlog, fetch errors, update detection, and whether successful recrawls replace old document versions.
- Duplicate results appear: compare final redirect URLs and canonical signals, then adjust normalization only for URL variants that represent the same content.
- JavaScript pages are empty: inspect the fetched HTML. If the useful content is absent, use a server-rendered source or a bounded rendering step; do not simply add long delays to every fetch.
- Private information appears: disable access to the affected results immediately, remove stored documents and snippets, invalidate caches, then audit both crawl-time and query-time authorization.
- Crawling overloads a host: reduce concurrency and request frequency, honor robots rules, and add backoff for errors and rate limits.
- Good pages rank poorly: examine the actual extracted fields and representative query set before changing boosts. Boilerplate in the index or a poor tokenizer can undermine ranking adjustments.
Or skip the browser setup
A screenshot API can help inspect the rendered appearance of pages during development, but a screenshot is not an extracted text document and does not replace the crawler or search index. ScreenshotNeo accepts a URL and returns a PNG, JPEG, WebP, or PDF. Its optional capture steps accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status.
Example cURL request (replace the URL with the page you want to inspect):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Does building site search automatically make pages appear in Google Search?
No. A site-search index serves queries on your own site; Google Search has its own crawler, eligibility rules, and indexing decisions.
Should a new site-search project start with semantic search?
Not by default. Start with representative queries and a lexical baseline, then add more complex retrieval only when evaluation shows that it improves results.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

