Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The right extraction method depends on where the words live. For a short visible passage, select it and copy it. For a readable article, use Reader Mode. For code running in an already loaded browser, read the rendered DOM with innerText. For repeatable requests, fetch and parse the HTML—but expect client-side JavaScript to add content that a plain request cannot see. If the words are pixels in an image, use text recognition (OCR), not ordinary DOM extraction.

Choose the method before you start

Decide what you need, how often you will do it, and whether the page has finished rendering. The following choices prevent most wasted effort.

Method Best for Needs an article-like page? Handles JavaScript-rendered content? Typical effort
Select and copy One visible passage No Yes, if it is visible Lowest
Reader Mode Cleaning up an article for reading or copying Usually Only after the browser has made it available Low
Rendered DOM Scripts running in a loaded browser No Yes, after the relevant nodes exist Medium
Fetch and parse Repeatable server-side or command-line extraction No No, unless you add a browser renderer Medium
Image text recognition Scans, screenshots and text embedded in images No Not applicable Medium

Copy a visible passage without code

  1. Open the page and wait until the passage is visible.
  2. Drag across only the text you need. On a touch device, press and hold, then adjust the selection handles.
  3. Use the browser or operating system’s Copy command, then paste into your destination.

This is the most reliable choice for a one-off excerpt because you can see exactly what is being collected. It also avoids requesting permissions or interpreting a site’s HTML. If selection starts too early or ends too late, begin inside the first word and finish after the last punctuation mark; selecting the containing heading or paragraph can be easier than dragging across a complex layout.

Use Reader Mode for article cleanup

Reader Mode presents an article-like page as a simplified reading view. It can hide sidebars, footers and advertisements and let you change text size, contrast and layout. Copy from that view when the page is cluttered and your goal is the central reading text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When Reader Mode will not appear

A browser cannot assume every page is an article. Dashboards, search results, home pages, stores and pages without a recognizable article structure may be ineligible. In that case, return to manual selection or use a DOM-based method.

What Reader Mode does not guarantee

It is a presentation and extraction aid, not a universal scraper. It may omit captions, interactive widgets, comments or content that the site does not mark as part of the article. Check the result against the original page before treating it as complete.

Extract rendered text with JavaScript

If developer tools, a browser extension or your own script is running in the page context, select the smallest useful container and read its rendered text:

const articleText = document.querySelector("article")?.innerText ?? "";
console.log(articleText);

HTMLElement.innerText approximates the text a person could select and copy. It reflects rendered appearance, including line breaks and visibility decisions. textContent is different: it reads the node’s text regardless of whether the browser is visually displaying it in the same way.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Choose a precise selector

Using document.body.innerText often captures navigation, cookie notices, footers and unrelated controls. Prefer an article element or a site-specific content selector. If the selector is absent, fail clearly instead of silently returning an empty string:

const container = document.querySelector("article, main, .post-content");
if (!container) throw new Error("Content container not found");
const text = container.innerText.trim();
if (!text) throw new Error("Container has no rendered text");
console.log(text);

Wait for dynamic content

The live DOM can change after the initial page load. A script that reads immediately may miss comments, infinite-scroll items or content inserted by a framework. In a browser automation tool, wait for a known selector or for the application’s loading indicator to disappear, then read innerText. Do not use an arbitrary short delay as proof that rendering is finished; a slow connection can take longer.

Fetch the HTML and parse it

Fetching is appropriate when the desired text is present in the server response and you need a repeatable workflow. A fetch promise does not reject merely because the server returned 404 or 500, so check the status before parsing.

async function extract(url) {
  const response = await fetch(url);
  if (!response.ok) {
    throw new Error(`HTTP ${response.status} for ${url}`);
  }
  const html = await response.text();
  const doc = new DOMParser().parseFromString(html, "text/html");
  const node = doc.querySelector("article, main");
  return (node?.textContent ?? doc.body?.textContent ?? "")
    .replace(/s+/g, " ")
    .trim();
}

extract("https://example.com/article")
  .then(console.log)
  .catch(console.error);

Response.text() reads the response body as text. DOMParser creates an in-memory document so you can query it without inserting the markup into the current page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

Fetch versus the page you see

The original response can differ substantially from the rendered page. JavaScript may request article data, expand a component, personalize the text or replace a loading shell after the response arrives. A fetch-only extractor will not execute those changes. When the required words appear only after scripts run, use a real browser context and the rendered-DOM approach instead.

Handle untrusted markup safely

Parsing into a detached document is safer than injecting unknown HTML into your application, but treat all fetched content as untrusted data. Extract text nodes or sanitized fields; do not copy arbitrary markup into innerHTML and execute it. Validate URLs and apply timeouts and size limits in production code.

Read clipboard text in a web app

A user-facing tool can let someone copy text normally and then press an explicit “Paste” button. The asynchronous Clipboard API provides readText():

async function pasteText() {
  try {
    const value = await navigator.clipboard.readText();
    document.querySelector("#output").value = value;
  } catch (error) {
    console.error("Clipboard read denied or unavailable", error);
  }
}

document.querySelector("#paste").addEventListener("click", pasteText);

Clipboard reads require a secure context and can be denied by the user, browser permissions or embedding policy. They are not guaranteed to work from an insecure HTTP page or without a user gesture. Rich formats use navigator.clipboard.read(), whose supported types and policy restrictions vary. Design a fallback—such as a focused text box where the user can paste with the keyboard—and explain why the permission is needed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract words from screenshots and other images

If letters are pixels, there is no HTML text node to read. Use OCR or a browser’s image-text feature. Mozilla documents a Firefox “Copy Text from Image” option for supported macOS configurations; that platform scope is not a promise of universal availability. For scans, diagrams and low-resolution screenshots, expect recognition errors and verify names, numbers and punctuation manually.

Firefox’s page extraction behavior

Firefox documentation describes a Page Extractor that can work with live-DOM text, Reader Mode and PDF text, with limited site-specific handling. It may return no result and does not necessarily wait for every later dynamic update. Treat an empty result as a signal to wait, choose a more specific element, open Reader Mode, or use a browser automation flow that waits for the application state you need.

Build a dependable extraction workflow

For a one-time task

  1. Try selecting and copying the visible passage.
  2. If navigation and ads interfere, try Reader Mode.
  3. If the words are inside an image, use image text recognition.

For a script running in a browser

  1. Wait for the target selector or a clear application-ready condition.
  2. Select the article or content container, not the entire body.
  3. Read innerText when you want user-visible text; use textContent only when hidden or formatting-insensitive text is intentional.
  4. Record the URL, timestamp and selector so an empty result is diagnosable.

For repeated server-side jobs

  1. Fetch the URL with an explicit timeout.
  2. Check response.ok or the numeric status before consuming the body.
  3. Parse the returned HTML and target stable selectors.
  4. Detect pages that contain only a loading shell; route those URLs to a browser renderer.
  5. Normalize whitespace only after preserving meaningful headings, lists and paragraph boundaries.

Troubleshooting common failures

Symptom Likely cause Fix
The copied result contains menus and cookie text The selection included the whole page Use an article selector or Reader Mode and select only the content region.
Fetch returns almost no article text The page fills its content with JavaScript Use a browser-rendered DOM and wait for the target node.
fetch did not throw for a 404 HTTP errors do not automatically reject the promise Check response.ok or response.status.
innerText is empty The selector is wrong or the content has not rendered Inspect the live DOM, correct the selector and wait for the element.
Clipboard read raises a permission error Insecure context, denied permission or missing user gesture Serve over HTTPS, request the read from an explicit click and provide manual paste.
OCR misreads a word Blur, unusual fonts, compression or low contrast Use a higher-resolution source, crop tightly and proofread critical text.
Reader Mode is unavailable The page is not recognized as an article Copy manually or extract a selected DOM container.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your real requirement is a clean visual record of a rendered webpage rather than a text string, ScreenshotNeo makes one GET request and returns a PNG, JPEG, WebP or PDF. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the 63 capture options, including full-page lazy-image loading, element capture, custom CSS and JavaScript, waits, headers, cookies, user agents, geolocation, PDF settings, caching, signed links, asynchronous jobs, webhooks and bulk capture. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for the free ScreenshotNeo plan when a rendered screenshot or PDF is the appropriate output.

Privacy, permission and reliability checks

  • Do not extract private or access-controlled material unless you are authorized.
  • Respect a site’s terms, authentication boundaries and rate limits when automating requests.
  • Store the source URL and extraction time with your output so changing pages can be audited.
  • Expect browser APIs and feature availability to change; test clipboard, Reader Mode and image recognition on the browsers and operating systems you support.

Frequently Asked Questions

Can I extract text from a page that requires login?

Only when you are authorized and your tooling is allowed to use that authenticated session. A public fetch will not reproduce private session content.

Why does copied text contain unexpected line breaks?

Rendered text follows the page’s visual layout. Normalize whitespace for plain search or storage, but preserve paragraph and heading boundaries when the structure matters.

Is a screenshot an alternative to extracted text?

It is an alternative output, not a text parser. Use it when visual fidelity, evidence or a PDF is the goal; use DOM extraction or OCR when you need searchable characters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.