For most readable-text jobs, parse the markup with Beautiful Soup, choose a parser explicitly, and call get_text(" ", strip=True). The separator keeps words apart when tags sit between them, while strip=True removes surrounding whitespace. Start with this complete example:
from bs4 import BeautifulSoup
html = "<article><h1>Hello</h1><p>Python makes parsing practical.</p></article>"
soup = BeautifulSoup(html, "lxml")
text = soup.get_text(" ", strip=True)
print(text)
# Hello Python makes parsing practical.
This guide explains when to use Beautiful Soup, when the standard-library HTMLParser is a better fit, how parser choice changes results, and why removing tags is not the same as finding an article’s main content.
The shortest reliable solution
Install Beautiful Soup and an explicit parser backend:
python -m pip install beautifulsoup4 lxml
Then parse the HTML and extract text from either the whole document or a selected element:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
from bs4 import BeautifulSoup
with open("page.html", encoding="utf-8") as f:
html = f.read()
soup = BeautifulSoup(html, "lxml")
all_text = soup.get_text(" ", strip=True)
print(all_text)
main = soup.select_one("main")
if main is None:
raise ValueError("No main element found")
print(main.get_text(" ", strip=True))
get_text() returns the text beneath a document or tag. Passing a separator is important: without it, text in adjacent tags can run together. Selecting main first prevents menus, footers, and unrelated page chrome from entering the result.
Choose a parser backend deliberately
Beautiful Soup provides one tree API over several parsers. The same malformed HTML can produce different trees, so the parser is part of your program’s behavior rather than an incidental installation detail.
| Approach | Strength | Trade-off | Best fit |
|---|---|---|---|
Beautiful Soup + lxml |
Friendly tree API with a robust parser backend | Requires third-party dependencies | General extraction from messy pages |
Beautiful Soup + html5lib |
HTML5-style error recovery | Usually slower and adds a dependency | Input where browser-like recovery matters |
Beautiful Soup + html.parser |
Simple installation and familiar API | Different recovery behavior on invalid markup | Small scripts and controlled input |
html.parser.HTMLParser |
Python standard library and callback control | You implement collection and cleanup | Dependency-free, event-driven processing |
Name the parser in code (for example, BeautifulSoup(html, "lxml")) and pin it in your project requirements. That makes deployments reproducible and lets tests detect a parser change instead of silently changing extracted text.
Beautiful Soup extraction patterns
Extract the complete document
from bs4 import BeautifulSoup
def visible_text(html: str) -> str:
soup = BeautifulSoup(html, "lxml")
return soup.get_text(" ", strip=True)
print(visible_text("<p>One</p><p>Two</p>"))
# One Two
The result is one normalized string. This is useful for search indexing, logging, deduplication, or a quick inspection of a page.
Extract one known region
from bs4 import BeautifulSoup
def article_text(html: str) -> str:
soup = BeautifulSoup(html, "lxml")
node = soup.select_one("article, main, [role='main']")
if node is None:
return ""
return node.get_text(" ", strip=True)
Use the selectors that match your site. A selector is usually more dependable than trying to infer “the article” from every tag on an arbitrary page. Check for None; calling get_text() on a missing node raises an error.
Process fragments with stripped_strings
When you need to inspect, filter, or transform each fragment separately, iterate over stripped_strings:
Rank #2
from bs4 import BeautifulSoup
soup = BeautifulSoup("<p> First <strong>important</strong> point. </p>", "lxml")
fragments = list(soup.stripped_strings)
print(fragments)
# ['First', 'important', 'point.']
cleaned = " | ".join(fragments)
print(cleaned)
This gives you control over the joining rule instead of asking Beautiful Soup to produce one final string immediately.
Remove known page chrome before extraction
Tag removal and content selection solve different problems. If a cookie banner, navigation menu, comments area, or duplicate mobile markup is inside your selected region, remove those nodes first:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutefrom bs4 import BeautifulSoup
def clean_article(html: str) -> str:
soup = BeautifulSoup(html, "lxml")
for selector in ("script", "style", "template", "nav", "footer", ".cookie-banner", ".comments"):
for node in soup.select(selector):
node.decompose()
article = soup.select_one("article, main, [role='main']") or soup
return article.get_text(" ", strip=True)
Beautiful Soup’s documentation notes that script, style, and template contents are generally not treated as human-readable text when lxml or html.parser is used. Explicit removal is still useful when you want the same behavior across parser choices or when unwanted visible elements such as navigation and comments are present.
Use the standard library when dependencies are not wanted
Python’s html.parser is an event-driven parser. Subclass it, collect data in handle_data, and normalize the collected fragments yourself:
from html.parser import HTMLParser
class TextExtractor(HTMLParser):
def __init__(self):
super().__init__()
self.parts = []
def handle_data(self, data):
self.parts.append(data)
def extract_text(html: str) -> str:
extractor = TextExtractor()
extractor.feed(html)
extractor.close()
return " ".join(" ".join(extractor.parts).split())
html = "<article><h1>Title</h1><p>Body</p></article>"
print(extract_text(html))
# Title Body
The callback receives text as the parser encounters it. This gives you low-level control, but it does not provide CSS selection, tree navigation, or automatic article identification. Add your own state if you need to ignore data inside selected tags:
from html.parser import HTMLParser
class ArticleParser(HTMLParser):
def __init__(self):
super().__init__()
self.depth = 0
self.parts = []
def handle_starttag(self, tag, attrs):
if tag == "article":
self.depth += 1
def handle_endtag(self, tag):
if tag == "article" and self.depth:
self.depth -= 1
def handle_data(self, data):
if self.depth:
self.parts.append(data)
parser = ArticleParser()
parser.feed(html)
text = " ".join(" ".join(parser.parts).split())
For anything beyond simple streaming collection, Beautiful Soup is usually less code and easier to maintain.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsWhitespace, entities, and readable output
Prevent words from colliding
Prefer get_text(" ", strip=True) over get_text(strip=True) when inline elements can touch. The explicit space separates text fragments created by tags such as <span>, <em>, and links.
Choose your line-break policy
Readable text is not always one paragraph. If downstream code needs paragraph boundaries, iterate over block elements and join them with newlines:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "lxml")
paragraphs = [p.get_text(" ", strip=True) for p in soup.select("article p")]
text = "nn".join(p for p in paragraphs if p)
This preserves a deliberate paragraph structure while still normalizing spaces inside each paragraph. Do not assume every <br> represents a paragraph; treat it according to your output format.
Normalize after extraction, not before selection
Whitespace collapsing is useful for a final string, but applying it before selecting elements can hide boundaries that your selectors or tests rely on. Select and remove nodes first, then normalize the result.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Static HTML is not the same as rendered page text
These libraries parse the HTML string you give them. If a page inserts its article with JavaScript after the initial response, the article will not exist in that string. You need to obtain rendered HTML with a browser-capable capture step, or use an endpoint that returns the content directly, and then pass the resulting HTML to Beautiful Soup.
Even with a rendered snapshot, extraction still requires a content decision. Navigation, cookie notices, chat widgets, comments, and repeated responsive components may be legitimate HTML but not part of the article. Combine a rendered capture with selectors, explicit removal rules, or a dedicated content-extraction stage.
Build a repeatable extraction pipeline
- Acquire bytes and decode them correctly. Preserve the response or saved HTML used for a failing case.
- Parse with a named backend. Use one of
lxml,html5lib, orhtml.parserexplicitly. - Remove predictable noise. Delete selectors for scripts, navigation, banners, comments, or duplicate containers that do not belong in the output.
- Select the content root. Prefer a known
article,main, or site-specific selector; fall back to the document only when that is acceptable. - Extract with a separator. Call
get_text(" ", strip=True)or processstripped_strings. - Normalize for the consumer. Choose one-line text, paragraph-separated text, or a list of fragments deliberately.
- Test representative fixtures. Include valid pages, malformed markup, missing selectors, empty elements, and pages containing navigation or consent UI.
For performance, parse once and reuse the resulting tree for all selectors. If you only need a small fragment, select it before extracting instead of converting the entire document to text. For reliability, log the parser name, selector used, and whether the fallback path ran; those details explain most changes in output.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
“No module named bs4” or parser-not-found errors
Install the package in the same environment that runs the script: python -m pip install beautifulsoup4 lxml. Virtual environments and system Python installations can have separate package directories.
Words are joined together
You likely omitted the separator. Replace get_text(strip=True) with get_text(" ", strip=True), or join stripped_strings with a space.
The output contains menus or a footer
You extracted the whole document. Select the article container first and remove known navigation, footer, comments, and banner selectors before calling get_text().
The expected element is missing
Inspect the HTML string you actually parsed. The content may be injected by JavaScript, the selector may differ for another template, or the page may have returned an error document. Check the selected node for None and keep a fallback policy explicit.
Results change after deployment
Different parser backends recover malformed markup differently. Pin the backend, keep dependency versions controlled, and run the same fixture tests in development and production.
Text includes an unwanted widget
Add that widget’s stable class, ID, or structural selector to your removal list before selecting the content root. Avoid broad rules that could delete legitimate article text.
The page is blank even though a browser shows content
The initial HTML likely lacks client-rendered content, or access is blocked by a bot check. Obtain a browser-rendered snapshot first, then parse its HTML; do not expect an HTML parser to execute JavaScript.
Best Value
Or skip the browser setup
If obtaining clean, rendered HTML is the difficult part, ScreenshotNeo can capture a URL before you run your own extraction. Its API returns a screenshot or PDF rather than article text, so use it when you need a dependable visual/rendered capture alongside your parser workflow.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for request options. The same call from Python:
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
- Cookie banners, newsletter popups, and chat widgets are removed before the shot.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.
FAQ
Frequently Asked Questions
Can Beautiful Soup parse an HTML fragment instead of a complete page?
Yes. Pass the fragment string to BeautifulSoup with the same explicit parser. You can then select elements and call get_text() exactly as you would for a full document.
Should extracted text be stored as one string or a list?
Use one string for search, display, or simple export. Keep a list of stripped fragments or paragraphs when later processing needs boundaries, headings, or per-element metadata.
Does HTMLParser provide CSS selectors?
No. HTMLParser reports parsing events through callbacks. If you need CSS selection or tree navigation, use Beautiful Soup or build those capabilities yourself.
How can I verify that a selector still matches after a site redesign?
Treat selectors as configuration: test them against saved representative HTML, alert when the expected node is absent, and retain a documented fallback rather than silently returning an unrelated page-wide string.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

