Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can use ChatGPT to classify web pages by giving it a clear set of labels and the page content to assess, then asking for one label and supporting evidence per page. For a collection, a well-organized spreadsheet is usually easier to review than a list of URLs. Treat the results as draft labels: check the source content, especially for ambiguous, inaccessible, or consequential pages.

What ChatGPT can—and cannot—classify

ChatGPT can analyze supplied text and supported files, including common spreadsheets and PDFs, and can return structured tables. The exact file types and tools available can vary by model, plan, workspace settings, and account. Check the tools available in your own ChatGPT interface before designing a workflow around a particular feature.

A URL is not the same thing as page content. Pasting a list of URLs does not establish that ChatGPT has fetched and read every page. If you need labels based on the wording or details of a page, provide its text or another representation ChatGPT can actually inspect. For current information, use ChatGPT Search if it is available, and examine the sources it cites.

Neither file analysis nor web search guarantees a correct label. OpenAI cautions that search results and citations may be incomplete, outdated, or incorrect. The official capability information also does not establish a dedicated, universally available webpage-classification tool or a guaranteed classification accuracy rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define the labels before uploading pages

Start by deciding what counts as one page and what distinctions matter to your project. Make the categories mutually distinguishable where possible, and write a short rule for assigning each one. If a page can reasonably fit multiple categories, state which one takes priority—or allow multiple labels if that suits the task.

Include an outcome such as uncertain or needs review. Without it, a model may force a borderline page into a category that does not fit. Do not ask ChatGPT to invent a taxonomy and then treat its choices as your policy; agree on what the labels mean first.

Example label definitions

  • Product page: primarily presents a specific product, its features, or purchase options.
  • Editorial article: primarily explains, reviews, or reports on a topic for readers.
  • Support documentation: primarily gives instructions for using or troubleshooting a product or service.
  • Other: the page does not meet any definition above.
  • Needs review: the supplied content is missing, conflicting, too limited, or difficult to classify confidently.

These are illustrative definitions, not a taxonomy prescribed by OpenAI. Adapt them to your goal. A content audit, accessibility review, and research project may need different labels.

Prepare a spreadsheet for a collection

For batch work, use one row per page and descriptive column headers. OpenAI recommends this general structure for spreadsheet analysis. A useful starting layout is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Column What to put there
page_id A stable identifier you can use to match the result to your records.
url The page address, for reference; do not assume ChatGPT has opened it.
page_title The title shown for the page, if known.
page_text The actual supplied text you want classified. Keep it associated with the correct page.
notes Optional context, such as why the page is included or a known content limitation.

Save the content in a file type supported by the tools on your account. ChatGPT Data Analysis supports common spreadsheets, PDFs, and text or data files, but available types can differ. For exact values or page-by-page traceability, use a structured file rather than relying on a complex, image-heavy, or poorly organized document; such files may not be fully analyzed.

If you only have URLs, first obtain the page content through a method appropriate to your site and permissions. Check that text extraction has not mixed navigation, cookie notices, repeated footers, or content from multiple pages into one record. The data-analysis Python environment described by OpenAI runs in a stateful Jupyter environment for some tasks, but it cannot make external web requests or API calls. It should not be treated as a crawler for a spreadsheet of URLs.

Ask for a consistent, reviewable result

Upload the prepared file, provide the label definitions, and specify the output fields you need. Ask ChatGPT to retain the page identifier so you can join its answers back to the source rows. Requiring a short evidence excerpt makes it easier to notice a label based on irrelevant or misread text.

Prompt template

Classify each page using only the supplied page_title, page_text, and notes. Do not infer that you have visited a URL unless its content is included in the supplied text. Use exactly one label from: Product page, Editorial article, Support documentation, Other, Needs review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Definitions: Product page = primarily presents a specific product, its features, or purchase options. Editorial article = primarily explains, reviews, or reports on a topic for readers. Support documentation = primarily gives instructions for using or troubleshooting a product or service. Other = does not meet those definitions. Needs review = content is missing, conflicting, too limited, or difficult to classify confidently.

Return one result per input page in a table with page_id, label, a brief evidence excerpt from page_text, and a short rationale. Do not invent evidence. If the supplied content is insufficient, choose Needs review and explain what is missing. Preserve the input order and do not omit rows.

ChatGPT can make tables or charts when a structured view is useful, but this prompt is a practical starting point, not a validated schema or a guarantee of a particular response. For high-volume or operational workflows, inspect the output for missing IDs, duplicate rows, labels outside your approved set, and evidence that does not appear in the supplied text.

Use Search when freshness matters

If your classification depends on recent changes—for example, whether a page currently announces a product, event, or policy—ChatGPT Search may help retrieve current material when the feature is available. Ask it to identify the source supporting each decision, and inspect those sources rather than treating citations as automatic proof.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Search is a different input path from classifying a supplied spreadsheet. It can help with current lookup, but the cited material can be incomplete, outdated, or incorrect. A search result that is missing does not prove that the page does not exist or that its content has a particular label. For repeatable audits, preserve the page text and the date or context of collection in your own records.

Review labels and handle exceptions

Use ChatGPT’s output as a first-pass classification, not an unreviewed source of truth. Compare a sample of ordinary labels with the original content, then concentrate additional review on exceptions. This is a quality-control practice, not a claim that sampling establishes a particular accuracy rate.

  • Missing or very short text: check whether extraction failed or the page is genuinely sparse; use Needs review if the evidence cannot support a label.
  • Conflicting page sections: apply your stated rule for the page’s primary purpose, or revise the taxonomy if the categories cannot represent the page.
  • Several pages in one record: separate them so each row represents one page, then rerun those records.
  • Untraceable rationale: verify the excerpt against the supplied text. If it is not there, do not accept the explanation as evidence.
  • Labels outside the allowed set or omitted rows: ask for a correction using the original identifiers and exact label list, then compare row counts again.

For categories with legal, safety, financial, or other material consequences, have a qualified person assess the underlying page and decision rules. A model-generated rationale does not replace that review.

Capture a page when a visual record is useful

Some classification tasks depend on visual presentation—for example, distinguishing a landing page from a long article when the extracted text is incomplete. A screenshot can preserve what a browser rendered, but it is not a substitute for searchable page text: text in an image may be harder to inspect, and a screenshot represents a particular capture rather than a guarantee of current or complete content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If you need a browser-based capture, obtain the page through a browser or capture service, keep the resulting image or PDF tied to its URL and capture context, and provide that artifact only if your ChatGPT account supports analyzing it. Do not assume that a URL-only spreadsheet will be opened just because you have a screenshot workflow.

Or skip the browser setup

For a visual page record, ScreenshotNeo provides a website screenshot API. One GET request can return a PNG, JPEG, WebP, or PDF; its API also accepts the parameter names used by other screenshot APIs. A screenshot is useful as visual evidence, but for text-based classification you should still supply extractable page content to ChatGPT.

cURL example, saving a WebP screenshot of the target page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for request options. The same endpoint can be called from Python or Node.js:

import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server offers screenshot and PDF tools for AI agents. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo for the service and sign up free to start with 1,000 screenshots a month and no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why a page may not appear in ChatGPT Search

OpenAI’s publisher guidance says a website can help its eligibility for ChatGPT Search by allowing OAI-SearchBot to crawl it, but that does not guarantee ranking or placement. This is relevant when using Search to find pages, not a guarantee that a particular page will be indexed or that Search is a comprehensive inventory of a site. If completeness matters, work from your own page list and provide the content you need classified.

Practical troubleshooting

ChatGPT ignores some rows

Check that the file has one page per row, descriptive headers, and a clear page identifier. Ask for one output row per input identifier and compare counts. Split a particularly large or complex job into manageable batches while retaining the same definitions and prompt.

The answer labels URLs without evidence

Do not treat the address as proof that the page was read. Include the page text, or use an available Search workflow and inspect its cited sources. Ask ChatGPT to mark insufficient input as Needs review rather than guessing.

Uploaded content is incomplete or garbled

Check the source file’s structure and whether its important content is selectable text or embedded in complex images. Re-export to a supported, well-structured format where possible, or provide clean text with one record per page. OpenAI notes that complex, image-heavy, or poorly structured files may not be fully analyzed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Features or file options are missing

Tool and file availability can depend on account, plan, model, and workspace configuration. Confirm what is visible in your ChatGPT interface; do not assume another account’s options apply to yours.

Search sources do not support the label

Open the cited pages and check that they substantiate the relevant page content and classification. Search citations can be incomplete, outdated, or incorrect. If evidence remains unclear, use Needs review or provide the source text directly.

Frequently Asked Questions

Can ChatGPT classify a list of URLs without page text?

A URL list alone does not establish that ChatGPT has fetched and read the pages. For a content-based classification, provide the page content or verify the sources retrieved through an available Search workflow.

Does ChatGPT guarantee correct page labels?

No. The cited OpenAI documentation describes capabilities and caveats, not a guaranteed accuracy rate or a universally available dedicated classification tool.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.