Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Beautiful Soup when you want a practical one-liner: parse the HTML with an explicit parser, then call get_text(" ", strip=True). For zero dependencies, subclass Python’s built-in HTMLParser and collect data callbacks. If you want readable plain ASCII that preserves more document structure, consider html2text.
Choose the right conversion method
HTML-to-text conversion processes markup you already have in a string, file, or response body. It does not download or execute a web page by itself. Your choice depends on the output and dependency constraints:
| Approach | Best for | Trade-off |
|---|---|---|
| Beautiful Soup | Convenient extraction with control over separators, whitespace, and selected elements | Requires an external package |
html.parser.HTMLParser |
Dependency-free parsing and custom formatting | You implement boundaries, filtering, and cleanup |
html2text |
Readable plain ASCII with Markdown-like structure | Output is formatted text rather than only concatenated text nodes |
Method 1: Beautiful Soup’s get_text()
Install and parse an HTML string
Install Beautiful Soup 4 with pip install beautifulsoup4. Always name the parser explicitly; parser implementations can build different trees from malformed markup. The project documents Beautiful Soup at crummy.com/software/BeautifulSoup/bs4/doc/.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchfrom bs4 import BeautifulSoup
html = "<p>Hello <b>world</b>.</p><p>Next paragraph.</p>"
soup = BeautifulSoup(html, "html.parser")
text = soup.get_text(" ", strip=True)
print(text)
# Hello world. Next paragraph.
get_text() returns a Unicode string containing text beneath the document or tag. Its first argument is a separator inserted between text fragments; strip=True trims whitespace around each fragment. A space is useful for inline elements, while a newline is often better when you intentionally select block elements.
#1 Best Overall
Keep paragraph boundaries
Flattening an entire document can lose the distinction between paragraphs, headings, and list items. Select the blocks you need and join each block separately:
from bs4 import BeautifulSoup
html = """
<article>
<h1>Release notes</h1>
<p>First paragraph.</p>
<p>Second <strong>important</strong> paragraph.</p>
</article>
"""
soup = BeautifulSoup(html, "html.parser")
blocks = soup.select("h1, p")
text = "nn".join(block.get_text(" ", strip=True) for block in blocks)
print(text)
For lower-level control, soup.stripped_strings yields already-trimmed text fragments that you can classify or join yourself.
Remove unwanted elements first
When navigation, advertisements, or metadata should not appear, remove those elements before extraction:
Recommended Free Tools
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
for tag in soup.select("script, style, template, nav, aside"):
tag.decompose()
text = soup.get_text(" ", strip=True)
With Beautiful Soup 4.9.0 and later, when using html.parser or lxml, the contents of script, style, and template elements are generally not treated as human-visible text. This behavior is parser- and version-qualified, so explicit removal is safer when your output must exclude them.
Method 2: A dependency-free extractor with HTMLParser
Collect text callbacks
Python’s standard library includes html.parser.HTMLParser. The Python documentation describes it as a parser able to parse invalid markup, but it is not a one-call tag stripper: you decide how to join fragments and where to insert boundaries.
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
def handle_data(self, data):
self.parts.append(data)
html = "<p>Hello <b>world</b>.</p>"
parser = TextExtractor()
parser.feed(html)
parser.close()
text = " ".join(" ".join(parser.parts).split())
print(text)
# Hello world .
The final normalization collapses runs of whitespace, but it cannot know that punctuation should touch the preceding word. For production formatting, insert boundaries at block tags instead of blindly joining every callback.
Rank #2
Preserve block structure
from html.parser import HTMLParser
class BlockTextExtractor(HTMLParser):
BLOCKS = {"p", "div", "section", "article", "h1", "h2", "h3", "li", "br"}
def __init__(self):
super().__init__(convert_charrefs=True)
self.parts = []
def handle_starttag(self, tag, attrs):
if tag in self.BLOCKS and self.parts and not self.parts[-1].endswith("n"):
self.parts.append("n")
def handle_endtag(self, tag):
if tag in self.BLOCKS:
self.parts.append("n")
def handle_data(self, data):
self.parts.append(data)
def text(self):
lines = [" ".join(line.split()) for line in "".join(self.parts).splitlines()]
return "n".join(line for line in lines if line)
html = "<h1>Title</h1><p>Hello <b>world</b>.</p>"
parser = BlockTextExtractor()
parser.feed(html)
parser.close()
print(parser.text())
convert_charrefs=True is the default and converts character references in normal data. The parser’s scripting option also affects how noscript content is handled. See the Python references for the html module and structured markup tools.
Decode HTML entities safely
Named and numeric references such as &, , and © need deliberate handling. Python’s html.unescape() converts references according to HTML5 rules:
from html import unescape
value = "Rock & Roll ©"
print(unescape(value))
# Rock & Roll ©
Beautiful Soup converts entities while parsing. Do not unescape twice unless your input genuinely contains text that is still escaped after parsing; otherwise characters can be transformed unexpectedly.
Read HTML from files and HTTP responses
Local files
from pathlib import Path
from bs4 import BeautifulSoup
raw = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(raw, "html.parser")
text = soup.get_text("n", strip=True)
Path("page.txt").write_text(text, encoding="utf-8")
When starting with bytes, decode them with the correct character encoding before parsing. Beautiful Soup documents Unicode conversion and encoding detection, but a reliable source encoding is still preferable.
HTTP response bodies
import requests
from bs4 import BeautifulSoup
response = requests.get("https://example.com", timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
text = soup.get_text(" ", strip=True)
This extracts the server response. It does not run JavaScript, click consent controls, or include content injected after page load. For dynamically rendered text, obtain rendered HTML through a browser workflow or another rendering service first.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteUse html2text when readable structure matters
The html2text package page describes the project as a Python script that converts HTML into clean, easy-to-read plain ASCII text. It is useful when links, headings, lists, and emphasis should remain readable rather than becoming one concatenated string.
import html2text
html = "<h1>Guide</h1><p>Read <a href='https://example.com'>the docs</a>.</p>"
converter = html2text.HTML2Text()
converter.ignore_links = False
text = converter.handle(html)
print(text)
The available evidence does not establish a feature-by-feature comparison, maintenance level, or suitability for every HTML dialect. Choose it for its output style, then test representative documents from your own pipeline.
Handle malformed markup, scripts, and dynamic content
Malformed HTML
Real-world markup may contain unclosed tags, nested elements in unusual orders, or invalid attributes. Python’s parser accepts invalid markup, while Beautiful Soup lets you choose a parser. Keep the parser name in code and tests so a deployment change does not silently alter the tree.
Scripts and styles
Neither JavaScript nor CSS should be treated as visible article text. Explicitly decompose script and style tags when using Beautiful Soup, and ignore those sections in a custom HTMLParser implementation if they enter your callbacks.
Rendered versus source HTML
A parser sees only the markup supplied to it. Client-side frameworks may inject the article, comments, or navigation after load. Fetching the initial response and parsing it cannot recover that later DOM. Use a browser-rendered capture or an endpoint that returns the rendered content when those nodes are required.
Performance, reproducibility, and output quality
- Memory: Beautiful Soup builds a parse tree, which makes selection and removal convenient. A custom streaming-style callback collector can keep your own data structure smaller, but formatting remains your responsibility.
- Whitespace: Pick separators based on the consumer. Search indexing may prefer normalized spaces; email or document export usually needs paragraph newlines.
- Encoding: Decode bytes once, using the response or file’s declared encoding. Keep the resulting text as Unicode throughout processing.
- Testing: Include inline tags, nested lists, entities, empty blocks, malformed markup, and script/style elements in fixtures. Assert the exact whitespace your downstream system expects.
- Security: Treat extracted text as untrusted input. Escape it when inserting into HTML again, and do not execute scripts from the source document.
Troubleshooting common failures
“No module named bs4”
Install the package in the same environment running your script: python -m pip install beautifulsoup4. If your project uses a virtual environment, activate it before installing.
Words run together
Use get_text(" ", strip=True) for inline content, or select block elements and join them with "nn". A tag stripper cannot infer every visual boundary.
Navigation or cookie text pollutes the result
Target the article container, such as soup.select_one("article"), or remove known selectors before calling get_text(). Site-specific selectors are more reliable than a universal rule.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Entities appear as literal codes
Check whether the string was escaped twice. Parse once with Beautiful Soup or call html.unescape() on the still-escaped value; avoid applying both blindly.
JavaScript content is missing
The original response did not contain the content. Obtain the rendered DOM first; HTML parsing alone does not execute scripts.
Different machines produce different output
Confirm the Beautiful Soup parser name, library versions, input encoding, and cleanup selectors. Parser differences on invalid markup can change which text nodes are visited.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your actual goal is to obtain a clean page capture before processing its content, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can use its MCP tools take_screenshot, get_page_info, and capture_pdf.
One request returns an image or PDF; it does not replace parsing HTML when you need the page’s text. Use it when a rendered visual or PDF is the input to your workflow.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all options. The same request from Python is:
Best Value
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);
ScreenshotNeo includes full-page and element capture, device and viewport settings, custom CSS and JavaScript, waits, request blocking, authentication headers and cookies, PDF controls, caching, signed links, asynchronous webhooks, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Sign up for the free plan.
Frequently asked questions
Does Beautiful Soup fetch a URL?
No. Give it HTML text or bytes; use an HTTP client or browser workflow to obtain the page first.
Which parser should I use for reproducible results?
Name the parser explicitly, commonly html.parser, and pin or record your library versions. Different parsers can interpret invalid markup differently.
Can HTML-to-text conversion preserve formatting?
Only the formatting you model. get_text() returns text, while block-aware joins or html2text can retain readable paragraph, list, heading, and link structure.
Frequently Asked Questions
Does Beautiful Soup fetch a URL?
No. Give it HTML text or bytes; use an HTTP client or browser workflow to obtain the page first.
Which parser should I use for reproducible results?
Name the parser explicitly, commonly html.parser, and record your library versions because parsers can interpret invalid markup differently.
Can HTML-to-text conversion preserve formatting?
Only the formatting you model. Use block-aware joins or html2text when headings, lists, paragraphs, or links must remain readable.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

