Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Beautiful Soup when you want a practical one-liner: parse the HTML with an explicit parser, then call get_text(" ", strip=True). For zero dependencies, subclass Python’s built-in HTMLParser and collect data callbacks. If you want readable plain ASCII that preserves more document structure, consider html2text.

Choose the right conversion method

HTML-to-text conversion processes markup you already have in a string, file, or response body. It does not download or execute a web page by itself. Your choice depends on the output and dependency constraints:

Approach Best for Trade-off
Beautiful Soup Convenient extraction with control over separators, whitespace, and selected elements Requires an external package
html.parser.HTMLParser Dependency-free parsing and custom formatting You implement boundaries, filtering, and cleanup
html2text Readable plain ASCII with Markdown-like structure Output is formatted text rather than only concatenated text nodes

Method 1: Beautiful Soup’s get_text()

Install and parse an HTML string

Install Beautiful Soup 4 with pip install beautifulsoup4. Always name the parser explicitly; parser implementations can build different trees from malformed markup. The project documents Beautiful Soup at crummy.com/software/BeautifulSoup/bs4/doc/.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

html = "<p>Hello <b>world</b>.</p><p>Next paragraph.</p>"
soup = BeautifulSoup(html, "html.parser")
text = soup.get_text(" ", strip=True)
print(text)
# Hello world. Next paragraph.

get_text() returns a Unicode string containing text beneath the document or tag. Its first argument is a separator inserted between text fragments; strip=True trims whitespace around each fragment. A space is useful for inline elements, while a newline is often better when you intentionally select block elements.

Keep paragraph boundaries

Flattening an entire document can lose the distinction between paragraphs, headings, and list items. Select the blocks you need and join each block separately:

from bs4 import BeautifulSoup

html = """
<article>
  <h1>Release notes</h1>
  <p>First paragraph.</p>
  <p>Second <strong>important</strong> paragraph.</p>
</article>
"""
soup = BeautifulSoup(html, "html.parser")
blocks = soup.select("h1, p")
text = "nn".join(block.get_text(" ", strip=True) for block in blocks)
print(text)

For lower-level control, soup.stripped_strings yields already-trimmed text fragments that you can classify or join yourself.

Remove unwanted elements first

When navigation, advertisements, or metadata should not appear, remove those elements before extraction:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
for tag in soup.select("script, style, template, nav, aside"):
    tag.decompose()
text = soup.get_text(" ", strip=True)

With Beautiful Soup 4.9.0 and later, when using html.parser or lxml, the contents of script, style, and template elements are generally not treated as human-visible text. This behavior is parser- and version-qualified, so explicit removal is safer when your output must exclude them.

Method 2: A dependency-free extractor with HTMLParser

Collect text callbacks

Python’s standard library includes html.parser.HTMLParser. The Python documentation describes it as a parser able to parse invalid markup, but it is not a one-call tag stripper: you decide how to join fragments and where to insert boundaries.

from html.parser import HTMLParser

class TextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_data(self, data):
        self.parts.append(data)

html = "<p>Hello <b>world</b>.</p>"
parser = TextExtractor()
parser.feed(html)
parser.close()
text = " ".join(" ".join(parser.parts).split())
print(text)
# Hello world .

The final normalization collapses runs of whitespace, but it cannot know that punctuation should touch the preceding word. For production formatting, insert boundaries at block tags instead of blindly joining every callback.

Preserve block structure

from html.parser import HTMLParser

class BlockTextExtractor(HTMLParser):
    BLOCKS = {"p", "div", "section", "article", "h1", "h2", "h3", "li", "br"}

    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag in self.BLOCKS and self.parts and not self.parts[-1].endswith("n"):
            self.parts.append("n")

    def handle_endtag(self, tag):
        if tag in self.BLOCKS:
            self.parts.append("n")

    def handle_data(self, data):
        self.parts.append(data)

    def text(self):
        lines = [" ".join(line.split()) for line in "".join(self.parts).splitlines()]
        return "n".join(line for line in lines if line)

html = "<h1>Title</h1><p>Hello <b>world</b>.</p>"
parser = BlockTextExtractor()
parser.feed(html)
parser.close()
print(parser.text())

convert_charrefs=True is the default and converts character references in normal data. The parser’s scripting option also affects how noscript content is handled. See the Python references for the html module and structured markup tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decode HTML entities safely

Named and numeric references such as &amp;, &nbsp;, and &#169; need deliberate handling. Python’s html.unescape() converts references according to HTML5 rules:

from html import unescape

value = "Rock &amp; Roll &#169;"
print(unescape(value))
# Rock & Roll ©

Beautiful Soup converts entities while parsing. Do not unescape twice unless your input genuinely contains text that is still escaped after parsing; otherwise characters can be transformed unexpectedly.

Read HTML from files and HTTP responses

Local files

from pathlib import Path
from bs4 import BeautifulSoup

raw = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(raw, "html.parser")
text = soup.get_text("n", strip=True)
Path("page.txt").write_text(text, encoding="utf-8")

When starting with bytes, decode them with the correct character encoding before parsing. Beautiful Soup documents Unicode conversion and encoding detection, but a reliable source encoding is still preferable.

HTTP response bodies

import requests
from bs4 import BeautifulSoup

response = requests.get("https://example.com", timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
text = soup.get_text(" ", strip=True)

This extracts the server response. It does not run JavaScript, click consent controls, or include content injected after page load. For dynamically rendered text, obtain rendered HTML through a browser workflow or another rendering service first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use html2text when readable structure matters

The html2text package page describes the project as a Python script that converts HTML into clean, easy-to-read plain ASCII text. It is useful when links, headings, lists, and emphasis should remain readable rather than becoming one concatenated string.

import html2text

html = "<h1>Guide</h1><p>Read <a href='https://example.com'>the docs</a>.</p>"
converter = html2text.HTML2Text()
converter.ignore_links = False
text = converter.handle(html)
print(text)

The available evidence does not establish a feature-by-feature comparison, maintenance level, or suitability for every HTML dialect. Choose it for its output style, then test representative documents from your own pipeline.

Handle malformed markup, scripts, and dynamic content

Malformed HTML

Real-world markup may contain unclosed tags, nested elements in unusual orders, or invalid attributes. Python’s parser accepts invalid markup, while Beautiful Soup lets you choose a parser. Keep the parser name in code and tests so a deployment change does not silently alter the tree.

Scripts and styles

Neither JavaScript nor CSS should be treated as visible article text. Explicitly decompose script and style tags when using Beautiful Soup, and ignore those sections in a custom HTMLParser implementation if they enter your callbacks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Rendered versus source HTML

A parser sees only the markup supplied to it. Client-side frameworks may inject the article, comments, or navigation after load. Fetching the initial response and parsing it cannot recover that later DOM. Use a browser-rendered capture or an endpoint that returns the rendered content when those nodes are required.

Performance, reproducibility, and output quality

  • Memory: Beautiful Soup builds a parse tree, which makes selection and removal convenient. A custom streaming-style callback collector can keep your own data structure smaller, but formatting remains your responsibility.
  • Whitespace: Pick separators based on the consumer. Search indexing may prefer normalized spaces; email or document export usually needs paragraph newlines.
  • Encoding: Decode bytes once, using the response or file’s declared encoding. Keep the resulting text as Unicode throughout processing.
  • Testing: Include inline tags, nested lists, entities, empty blocks, malformed markup, and script/style elements in fixtures. Assert the exact whitespace your downstream system expects.
  • Security: Treat extracted text as untrusted input. Escape it when inserting into HTML again, and do not execute scripts from the source document.

Troubleshooting common failures

“No module named bs4”

Install the package in the same environment running your script: python -m pip install beautifulsoup4. If your project uses a virtual environment, activate it before installing.

Words run together

Use get_text(" ", strip=True) for inline content, or select block elements and join them with "nn". A tag stripper cannot infer every visual boundary.

Navigation or cookie text pollutes the result

Target the article container, such as soup.select_one("article"), or remove known selectors before calling get_text(). Site-specific selectors are more reliable than a universal rule.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Entities appear as literal codes

Check whether the string was escaped twice. Parse once with Beautiful Soup or call html.unescape() on the still-escaped value; avoid applying both blindly.

JavaScript content is missing

The original response did not contain the content. Obtain the rendered DOM first; HTML parsing alone does not execute scripts.

Different machines produce different output

Confirm the Beautiful Soup parser name, library versions, input encoding, and cleanup selectors. Parser differences on invalid markup can change which text nodes are visited.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual goal is to obtain a clean page capture before processing its content, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can use its MCP tools take_screenshot, get_page_info, and capture_pdf.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One request returns an image or PDF; it does not replace parsing HTML when you need the page’s text. Use it when a rendered visual or PDF is the input to your workflow.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all options. The same request from Python is:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);

ScreenshotNeo includes full-page and element capture, device and viewport settings, custom CSS and JavaScript, waits, request blocking, authentication headers and cookies, PDF controls, caching, signed links, asynchronous webhooks, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Sign up for the free plan.

Frequently asked questions

Does Beautiful Soup fetch a URL?

No. Give it HTML text or bytes; use an HTTP client or browser workflow to obtain the page first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which parser should I use for reproducible results?

Name the parser explicitly, commonly html.parser, and pin or record your library versions. Different parsers can interpret invalid markup differently.

Can HTML-to-text conversion preserve formatting?

Only the formatting you model. get_text() returns text, while block-aware joins or html2text can retain readable paragraph, list, heading, and link structure.

Frequently Asked Questions

Does Beautiful Soup fetch a URL?

No. Give it HTML text or bytes; use an HTTP client or browser workflow to obtain the page first.

Which parser should I use for reproducible results?

Name the parser explicitly, commonly html.parser, and record your library versions because parsers can interpret invalid markup differently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can HTML-to-text conversion preserve formatting?

Only the formatting you model. Use block-aware joins or html2text when headings, lists, paragraphs, or links must remain readable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.