Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

You can build a small web scraper as an MCP server and let Codex call its tools, but “Web MCP” is not an official product name in the OpenAI documentation reviewed. The practical setup is your own server that fetches and extracts public web pages, connected to Codex as an MCP server. OpenAI’s separate Docs MCP is read-only documentation search; it does not scrape arbitrary websites or make API calls on your behalf.

What you are building

Model Context Protocol (MCP) lets an AI client discover and call tools exposed by a server. In this tutorial, the server provides a narrow tool that fetches one public page and returns its title and readable text. Codex can then request that operation as part of a task.

This is different from web search or a browser feature supplied by an AI client: your MCP server owns the fetch-and-extract behavior. It is also different from OpenAI Docs MCP, which serves OpenAI documentation. The official Docs MCP endpoint is https://developers.openai.com/mcp, and OpenAI describes it as documentation-only and read-only.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The example deliberately avoids crawling whole sites, signing into accounts, or bypassing access controls. It accepts a single public HTTP or HTTPS URL, fetches a bounded response, extracts text, and returns a small structured result. Before scraping a real site, check its access rules and make requests proportionate; this example does not determine whether a particular site permits your intended use.

Choose the SDK and define the tool boundary

OpenAI’s MCP guide shows the TypeScript SDK package @modelcontextprotocol/sdk with Zod, or the Python package mcp. Use the language your project already runs. There is no established performance ranking between these options in the cited guidance. The example below uses Python and the official mcp package.

Make the tool name describe the action and keep the schema explicit. A one-page extraction operation is easier to reason about than a general-purpose tool with unrelated modes such as login, crawl, download, and submit-form. The server—not the model—must enforce URL restrictions, response limits, timeouts, and authorization.

Build a minimal Python scraper MCP server

Install the package in a virtual environment. The package version is not pinned here; use the current compatible release for your environment and check its SDK documentation if the API has changed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. python -m venv .venv
  2. Activate the environment: on macOS/Linux run source .venv/bin/activate; on Windows PowerShell run .venvScriptsActivate.ps1.
  3. pip install mcp
  4. Save the following as scraper_server.py.
from html.parser import HTMLParser
from urllib.error import HTTPError, URLError
from urllib.parse import urlsplit
from urllib.request import Request, urlopen
import asyncio
import socket

from mcp.server.fastmcp import FastMCP

mcp = FastMCP("public-page-scraper")

MAX_BYTES = 1_000_000
TIMEOUT_SECONDS = 15
USER_AGENT = "ExampleMCPPageReader/1.0"


class PageText(HTMLParser):
    """Extract a title and visible-ish text without executing page code."""
    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.in_title = False
        self.skip_depth = 0
        self.title_parts = []
        self.text_parts = []

    def handle_starttag(self, tag, attrs):
        if tag.lower() == "title":
            self.in_title = True
        if tag.lower() in {"script", "style", "noscript", "svg"}:
            self.skip_depth += 1

    def handle_endtag(self, tag):
        if tag.lower() == "title":
            self.in_title = False
        if tag.lower() in {"script", "style", "noscript", "svg"} and self.skip_depth:
            self.skip_depth -= 1
        if tag.lower() in {"p", "div", "li", "br", "h1", "h2", "h3"}:
            self.text_parts.append(" ")

    def handle_data(self, data):
        value = " ".join(data.split())
        if not value:
            return
        if self.in_title:
            self.title_parts.append(value)
        if not self.skip_depth:
            self.text_parts.append(value)


@mcp.tool()
async def extract_public_page(url: str) -> dict:
    """Fetch one public HTTP(S) page and return its title and text."""
    parts = urlsplit(url)
    if parts.scheme not in {"http", "https"} or not parts.hostname:
        raise ValueError("url must be an absolute http:// or https:// URL")
    if parts.username or parts.password:
        raise ValueError("URLs containing credentials are not accepted")

    request = Request(url, headers={"User-Agent": USER_AGENT})
    try:
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            content_type = response.headers.get_content_type()
            if content_type not in {"text/html", "application/xhtml+xml"}:
                raise ValueError(f"Expected HTML, received {content_type}")
            raw = response.read(MAX_BYTES + 1)
            if len(raw) > MAX_BYTES:
                raise ValueError("Page exceeds the 1 MB response limit")
            charset = response.headers.get_content_charset() or "utf-8"
            html = raw.decode(charset, errors="replace")
            final_url = response.geturl()
            status = response.status
    except (HTTPError, URLError, TimeoutError, socket.timeout) as exc:
        raise RuntimeError(f"Page fetch failed: {exc}") from exc

    parser = PageText()
    parser.feed(html)
    title = " ".join(parser.title_parts).strip()
    text = " ".join(" ".join(parser.text_parts).split())
    return {
        "requested_url": url,
        "final_url": final_url,
        "http_status": status,
        "title": title,
        "text": text[:12000],
        "truncated": len(text) > 12000,
    }


if __name__ == "__main__":
    mcp.run(transport="streamable-http")

Run it with python scraper_server.py. The server’s transport and local address depend on the installed SDK version and its defaults; consult that version’s documentation for the exact endpoint to enter in a client. The code’s response limit bounds the bytes read, and its returned text is capped separately so a large page does not flood the model context.

Important safety limitation: URL validation is not network authorization

The sample checks the URL scheme and rejects embedded credentials, but that is not sufficient protection if the server will be reachable by untrusted users. A malicious caller could target internal network services through hostnames, redirects, DNS behavior, or unusual address forms. In a real service, resolve and validate destination IPs, restrict allowed hosts where practical, validate every redirect, block private and link-local ranges, and enforce outbound network controls. Apply per-user authorization and rate limits at the server. Do not treat tool annotations or a model’s intent as access control.

The tool reads public web content, so its MCP metadata should accurately communicate that it accesses the open internet, such as setting openWorldHint: true where the SDK supports annotations. Mark it read-only only if the implementation truly makes no changes; keep destructive-action metadata aligned with actual behavior. OpenAI’s guidance emphasizes that every tool input is untrusted.

Test the MCP server locally

Use MCP Inspector, the documented local inspection tool, to connect to and examine an MCP server. For a Streamable HTTP server, point Inspector at the local MCP endpoint reported by your SDK/server configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Start python scraper_server.py in a terminal and leave it running.
  2. Start MCP Inspector according to its current installation instructions, then select the Streamable HTTP connection option and enter the server endpoint.
  3. Complete initialization and inspect the advertised tool list. Confirm that extract_public_page has one string input named url and a description explaining the limits.
  4. Call it with a simple public HTML page. Inspect the returned requested and final URLs, status, title, text, and truncation flag.
  5. Try invalid input such as file:///etc/passwd, a URL without a host, a credential-bearing URL, a non-HTML resource, a slow page, and an oversized response. Confirm failures are clear and do not disclose secrets or server internals.
  6. Inspect annotations and authorization behavior as well as the happy path. An annotation describes behavior; it does not enforce it.

Useful tests include redirects to a disallowed destination, pages with no title, HTML containing scripts or styles, non-UTF-8 character sets, and responses that close early. Add tests for the restrictions your deployment actually relies on rather than assuming a single successful fetch proves the boundary is safe.

Connect Codex without confusing it with Docs MCP

OpenAI documents this CLI command for its own documentation-only server: codex mcp add openaiDeveloperDocs --url https://developers.openai.com/mcp. Verify configured MCP servers with codex mcp list. The documented configuration is shared by Codex CLI and its IDE extension. That command adds OpenAI Docs MCP; it is not a command for your scraper.

For your own server, use Codex’s MCP server configuration mechanism for the transport and endpoint your installed version supports. For a local stdio server the client normally launches a process; for this example’s Streamable HTTP server the client connects to a reachable HTTP endpoint. The exact setup flow for an arbitrary user-built server is not established by the Docs MCP setup instructions, so do not copy the Docs MCP server name or endpoint and expect it to invoke your code. Confirm the current Codex configuration syntax in the documentation for your installed version, then verify that Codex lists the server and discovers the tool before asking it to call the scraper.

Give Codex a bounded task, for example: “Use extract_public_page on this public product page and summarize the returned description; do not visit other URLs.” The model can choose whether to call the tool, but the server should remain responsible for validation, limits, and access control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Improve the example before broader use

  • Better extraction: the standard-library parser is intentionally basic. It does not render JavaScript, infer the main article region reliably, or load lazy content. A browser-based renderer is appropriate only when the page requires it, and it should retain strict time, resource, and network limits.
  • Content limits: tune maximum bytes, extracted characters, and timeout to the pages you need. Return an explicit truncation flag rather than silently implying the complete page was captured.
  • Structured outputs: keep fields stable and machine-readable. Add an output schema in the SDK/API version you use so clients can understand the result shape.
  • Rate and concurrency: cap requests per user and total concurrent fetches. Add caching only when its freshness and privacy implications are acceptable.
  • Robust errors: distinguish invalid input, disallowed destination, timeout, HTTP failure, unsupported content type, and response-size rejection. Avoid returning raw exception details to callers in production.
  • Operational visibility: log request IDs, duration, outcome class, and byte counts while excluding credentials, unnecessary personal data, and full page contents unless there is a clear need and a retention policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Local versus public deployment

Local execution is useful for development because it avoids exposing a network service, but only the client and process on that machine can use it. For a public deployment or plugin submission, OpenAI’s deployment guidance calls for a stable, publicly reachable HTTPS endpoint using Streamable HTTP. A temporary tunnel or a developer laptop is not a dependable public endpoint.

Decision Local development Public service
Reachability Local process or endpoint available to your client Stable, publicly reachable HTTPS address
Transport Use the transport supported by your local client and SDK Streamable HTTP is the deployment guidance described by OpenAI
Security Still validate untrusted URLs and bound requests Preserve authentication and authorization boundaries; limit outbound access
Operations Terminal output may be enough while developing Plan for service connectivity, logs, metrics, availability, and rollback
Data handling Keep secrets out of source and test data Decide where requests and logs reside, who can access them, and how long they are retained

Deployment adds real costs and failure modes: network egress, variable page latency, rate limits, remote-site outages, and service availability. Use bounded timeouts, controlled retries with backoff for transient failures, and monitoring for error rates and latency. Do not retry indefinitely or treat a remote page’s failure as a successful empty result.

Troubleshooting

  • Codex does not list the server: check the client’s current MCP configuration syntax, server name, executable or endpoint, and whether the server starts successfully. Restart or reload the client if its configuration requires it; use codex mcp list to inspect configured servers.
  • Inspector cannot initialize: confirm the process is running, the URL and transport match, and the endpoint path is correct for the installed SDK. Check the server terminal for startup errors.
  • The tool is visible but a call fails: validate that the input is an absolute HTTP(S) URL, the host resolves, the remote server is reachable, and the response is HTML within the byte and time limits.
  • The result is empty or incomplete: the page may require JavaScript rendering, use unusual markup, or exceed the text cap. This lightweight parser is not a browser and does not promise full-page semantic extraction.
  • Requests hang or consume too many resources: lower the timeout and byte cap, limit concurrency, and enforce network-level restrictions. A Python timeout alone is not a complete defense against every DNS or network-layer behavior.
  • A destination is unexpectedly accessible: add host/IP allowlists and redirect validation, plus outbound firewall rules. Scheme checking alone does not prevent server-side request forgery.

Or skip the browser setup

If your actual goal is to capture a page as an image or PDF rather than extract text into a custom MCP tool, ScreenshotNeo provides a screenshot API and MCP server. A one-call capture looks like this; see the ScreenshotNeo API documentation for options and response handling.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can this scraper read pages that require a login?

Not as written. It sends a basic request without browser login state or account credentials. Do not add authentication to a scraper unless you control the account and have a clear authorization and data-handling design.

Does the example scrape an entire website?

No. Each call fetches one URL. Site-wide crawling needs separate scope, link-following rules, deduplication, rate limits, and stronger controls over which hosts and paths the server may access.

Can I use OpenAI Docs MCP as the scraper server?

No. It searches and returns OpenAI documentation; it is not a general-purpose web scraper.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.