Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
You can build a small web scraper as an MCP server and let Codex call its tools, but “Web MCP” is not an official product name in the OpenAI documentation reviewed. The practical setup is your own server that fetches and extracts public web pages, connected to Codex as an MCP server. OpenAI’s separate Docs MCP is read-only documentation search; it does not scrape arbitrary websites or make API calls on your behalf.
What you are building
Model Context Protocol (MCP) lets an AI client discover and call tools exposed by a server. In this tutorial, the server provides a narrow tool that fetches one public page and returns its title and readable text. Codex can then request that operation as part of a task.
This is different from web search or a browser feature supplied by an AI client: your MCP server owns the fetch-and-extract behavior. It is also different from OpenAI Docs MCP, which serves OpenAI documentation. The official Docs MCP endpoint is https://developers.openai.com/mcp, and OpenAI describes it as documentation-only and read-only.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteThe example deliberately avoids crawling whole sites, signing into accounts, or bypassing access controls. It accepts a single public HTTP or HTTPS URL, fetches a bounded response, extracts text, and returns a small structured result. Before scraping a real site, check its access rules and make requests proportionate; this example does not determine whether a particular site permits your intended use.
#1 Best Overall
Choose the SDK and define the tool boundary
OpenAI’s MCP guide shows the TypeScript SDK package @modelcontextprotocol/sdk with Zod, or the Python package mcp. Use the language your project already runs. There is no established performance ranking between these options in the cited guidance. The example below uses Python and the official mcp package.
Make the tool name describe the action and keep the schema explicit. A one-page extraction operation is easier to reason about than a general-purpose tool with unrelated modes such as login, crawl, download, and submit-form. The server—not the model—must enforce URL restrictions, response limits, timeouts, and authorization.
Build a minimal Python scraper MCP server
Install the package in a virtual environment. The package version is not pinned here; use the current compatible release for your environment and check its SDK documentation if the API has changed.
python -m venv .venv- Activate the environment: on macOS/Linux run
source .venv/bin/activate; on Windows PowerShell run.venvScriptsActivate.ps1. pip install mcp- Save the following as
scraper_server.py.
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urlsplit
from urllib.request import Request, urlopen
import asyncio
import socket
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("public-page-scraper")
MAX_BYTES = 1_000_000
TIMEOUT_SECONDS = 15
USER_AGENT = "ExampleMCPPageReader/1.0"
class PageText(HTMLParser):
"""Extract a title and visible-ish text without executing page code."""
def __init__(self):
super().__init__(convert_charrefs=True)
self.in_title = False
self.skip_depth = 0
self.title_parts = []
self.text_parts = []
def handle_starttag(self, tag, attrs):
if tag.lower() == "title":
self.in_title = True
if tag.lower() in {"script", "style", "noscript", "svg"}:
self.skip_depth += 1
def handle_endtag(self, tag):
if tag.lower() == "title":
self.in_title = False
if tag.lower() in {"script", "style", "noscript", "svg"} and self.skip_depth:
self.skip_depth -= 1
if tag.lower() in {"p", "div", "li", "br", "h1", "h2", "h3"}:
self.text_parts.append(" ")
def handle_data(self, data):
value = " ".join(data.split())
if not value:
return
if self.in_title:
self.title_parts.append(value)
if not self.skip_depth:
self.text_parts.append(value)
@mcp.tool()
async def extract_public_page(url: str) -> dict:
"""Fetch one public HTTP(S) page and return its title and text."""
parts = urlsplit(url)
if parts.scheme not in {"http", "https"} or not parts.hostname:
raise ValueError("url must be an absolute http:// or https:// URL")
if parts.username or parts.password:
raise ValueError("URLs containing credentials are not accepted")
request = Request(url, headers={"User-Agent": USER_AGENT})
try:
with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
content_type = response.headers.get_content_type()
if content_type not in {"text/html", "application/xhtml+xml"}:
raise ValueError(f"Expected HTML, received {content_type}")
raw = response.read(MAX_BYTES + 1)
if len(raw) > MAX_BYTES:
raise ValueError("Page exceeds the 1 MB response limit")
charset = response.headers.get_content_charset() or "utf-8"
html = raw.decode(charset, errors="replace")
final_url = response.geturl()
status = response.status
except (HTTPError, URLError, TimeoutError, socket.timeout) as exc:
raise RuntimeError(f"Page fetch failed: {exc}") from exc
parser = PageText()
parser.feed(html)
title = " ".join(parser.title_parts).strip()
text = " ".join(" ".join(parser.text_parts).split())
return {
"requested_url": url,
"final_url": final_url,
"http_status": status,
"title": title,
"text": text[:12000],
"truncated": len(text) > 12000,
}
if __name__ == "__main__":
mcp.run(transport="streamable-http")
Run it with python scraper_server.py. The server’s transport and local address depend on the installed SDK version and its defaults; consult that version’s documentation for the exact endpoint to enter in a client. The code’s response limit bounds the bytes read, and its returned text is capped separately so a large page does not flood the model context.
Important safety limitation: URL validation is not network authorization
The sample checks the URL scheme and rejects embedded credentials, but that is not sufficient protection if the server will be reachable by untrusted users. A malicious caller could target internal network services through hostnames, redirects, DNS behavior, or unusual address forms. In a real service, resolve and validate destination IPs, restrict allowed hosts where practical, validate every redirect, block private and link-local ranges, and enforce outbound network controls. Apply per-user authorization and rate limits at the server. Do not treat tool annotations or a model’s intent as access control.
The tool reads public web content, so its MCP metadata should accurately communicate that it accesses the open internet, such as setting openWorldHint: true where the SDK supports annotations. Mark it read-only only if the implementation truly makes no changes; keep destructive-action metadata aligned with actual behavior. OpenAI’s guidance emphasizes that every tool input is untrusted.
Test the MCP server locally
Use MCP Inspector, the documented local inspection tool, to connect to and examine an MCP server. For a Streamable HTTP server, point Inspector at the local MCP endpoint reported by your SDK/server configuration.
- Start
python scraper_server.pyin a terminal and leave it running. - Start MCP Inspector according to its current installation instructions, then select the Streamable HTTP connection option and enter the server endpoint.
- Complete initialization and inspect the advertised tool list. Confirm that
extract_public_pagehas one string input namedurland a description explaining the limits. - Call it with a simple public HTML page. Inspect the returned requested and final URLs, status, title, text, and truncation flag.
- Try invalid input such as
file:///etc/passwd, a URL without a host, a credential-bearing URL, a non-HTML resource, a slow page, and an oversized response. Confirm failures are clear and do not disclose secrets or server internals. - Inspect annotations and authorization behavior as well as the happy path. An annotation describes behavior; it does not enforce it.
Useful tests include redirects to a disallowed destination, pages with no title, HTML containing scripts or styles, non-UTF-8 character sets, and responses that close early. Add tests for the restrictions your deployment actually relies on rather than assuming a single successful fetch proves the boundary is safe.
Connect Codex without confusing it with Docs MCP
OpenAI documents this CLI command for its own documentation-only server: codex mcp add openaiDeveloperDocs --url https://developers.openai.com/mcp. Verify configured MCP servers with codex mcp list. The documented configuration is shared by Codex CLI and its IDE extension. That command adds OpenAI Docs MCP; it is not a command for your scraper.
For your own server, use Codex’s MCP server configuration mechanism for the transport and endpoint your installed version supports. For a local stdio server the client normally launches a process; for this example’s Streamable HTTP server the client connects to a reachable HTTP endpoint. The exact setup flow for an arbitrary user-built server is not established by the Docs MCP setup instructions, so do not copy the Docs MCP server name or endpoint and expect it to invoke your code. Confirm the current Codex configuration syntax in the documentation for your installed version, then verify that Codex lists the server and discovers the tool before asking it to call the scraper.
Give Codex a bounded task, for example: “Use extract_public_page on this public product page and summarize the returned description; do not visit other URLs.” The model can choose whether to call the tool, but the server should remain responsible for validation, limits, and access control.
Improve the example before broader use
- Better extraction: the standard-library parser is intentionally basic. It does not render JavaScript, infer the main article region reliably, or load lazy content. A browser-based renderer is appropriate only when the page requires it, and it should retain strict time, resource, and network limits.
- Content limits: tune maximum bytes, extracted characters, and timeout to the pages you need. Return an explicit truncation flag rather than silently implying the complete page was captured.
- Structured outputs: keep fields stable and machine-readable. Add an output schema in the SDK/API version you use so clients can understand the result shape.
- Rate and concurrency: cap requests per user and total concurrent fetches. Add caching only when its freshness and privacy implications are acceptable.
- Robust errors: distinguish invalid input, disallowed destination, timeout, HTTP failure, unsupported content type, and response-size rejection. Avoid returning raw exception details to callers in production.
- Operational visibility: log request IDs, duration, outcome class, and byte counts while excluding credentials, unnecessary personal data, and full page contents unless there is a clear need and a retention policy.
Local versus public deployment
Local execution is useful for development because it avoids exposing a network service, but only the client and process on that machine can use it. For a public deployment or plugin submission, OpenAI’s deployment guidance calls for a stable, publicly reachable HTTPS endpoint using Streamable HTTP. A temporary tunnel or a developer laptop is not a dependable public endpoint.
| Decision | Local development | Public service |
|---|---|---|
| Reachability | Local process or endpoint available to your client | Stable, publicly reachable HTTPS address |
| Transport | Use the transport supported by your local client and SDK | Streamable HTTP is the deployment guidance described by OpenAI |
| Security | Still validate untrusted URLs and bound requests | Preserve authentication and authorization boundaries; limit outbound access |
| Operations | Terminal output may be enough while developing | Plan for service connectivity, logs, metrics, availability, and rollback |
| Data handling | Keep secrets out of source and test data | Decide where requests and logs reside, who can access them, and how long they are retained |
Deployment adds real costs and failure modes: network egress, variable page latency, rate limits, remote-site outages, and service availability. Use bounded timeouts, controlled retries with backoff for transient failures, and monitoring for error rates and latency. Do not retry indefinitely or treat a remote page’s failure as a successful empty result.
Troubleshooting
- Codex does not list the server: check the client’s current MCP configuration syntax, server name, executable or endpoint, and whether the server starts successfully. Restart or reload the client if its configuration requires it; use
codex mcp listto inspect configured servers. - Inspector cannot initialize: confirm the process is running, the URL and transport match, and the endpoint path is correct for the installed SDK. Check the server terminal for startup errors.
- The tool is visible but a call fails: validate that the input is an absolute HTTP(S) URL, the host resolves, the remote server is reachable, and the response is HTML within the byte and time limits.
- The result is empty or incomplete: the page may require JavaScript rendering, use unusual markup, or exceed the text cap. This lightweight parser is not a browser and does not promise full-page semantic extraction.
- Requests hang or consume too many resources: lower the timeout and byte cap, limit concurrency, and enforce network-level restrictions. A Python timeout alone is not a complete defense against every DNS or network-layer behavior.
- A destination is unexpectedly accessible: add host/IP allowlists and redirect validation, plus outbound firewall rules. Scheme checking alone does not prevent server-side request forgery.
Or skip the browser setup
If your actual goal is to capture a page as an image or PDF rather than extract text into a custom MCP tool, ScreenshotNeo provides a screenshot API and MCP server. A one-call capture looks like this; see the ScreenshotNeo API documentation for options and response handling.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
FAQ
Can this scraper read pages that require a login?
Not as written. It sends a basic request without browser login state or account credentials. Do not add authentication to a scraper unless you control the account and have a clear authorization and data-handling design.
Does the example scrape an entire website?
No. Each call fetches one URL. Site-wide crawling needs separate scope, link-following rules, deduplication, rate limits, and stronger controls over which hosts and paths the server may access.
Can I use OpenAI Docs MCP as the scraper server?
No. It searches and returns OpenAI documentation; it is not a general-purpose web scraper.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

