Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For most readable-text jobs, parse the markup with Beautiful Soup, choose a parser explicitly, and call get_text(" ", strip=True). The separator keeps words apart when tags sit between them, while strip=True removes surrounding whitespace. Start with this complete example:

from bs4 import BeautifulSoup

html = "<article><h1>Hello</h1><p>Python makes parsing practical.</p></article>"
soup = BeautifulSoup(html, "lxml")
text = soup.get_text(" ", strip=True)
print(text)
# Hello Python makes parsing practical.

This guide explains when to use Beautiful Soup, when the standard-library HTMLParser is a better fit, how parser choice changes results, and why removing tags is not the same as finding an article’s main content.

The shortest reliable solution

Install Beautiful Soup and an explicit parser backend:

python -m pip install beautifulsoup4 lxml

Then parse the HTML and extract text from either the whole document or a selected element:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

with open("page.html", encoding="utf-8") as f:
    html = f.read()

soup = BeautifulSoup(html, "lxml")
all_text = soup.get_text(" ", strip=True)
print(all_text)

main = soup.select_one("main")
if main is None:
    raise ValueError("No main element found")
print(main.get_text(" ", strip=True))

get_text() returns the text beneath a document or tag. Passing a separator is important: without it, text in adjacent tags can run together. Selecting main first prevents menus, footers, and unrelated page chrome from entering the result.

Choose a parser backend deliberately

Beautiful Soup provides one tree API over several parsers. The same malformed HTML can produce different trees, so the parser is part of your program’s behavior rather than an incidental installation detail.

Approach Strength Trade-off Best fit
Beautiful Soup + lxml Friendly tree API with a robust parser backend Requires third-party dependencies General extraction from messy pages
Beautiful Soup + html5lib HTML5-style error recovery Usually slower and adds a dependency Input where browser-like recovery matters
Beautiful Soup + html.parser Simple installation and familiar API Different recovery behavior on invalid markup Small scripts and controlled input
html.parser.HTMLParser Python standard library and callback control You implement collection and cleanup Dependency-free, event-driven processing

Name the parser in code (for example, BeautifulSoup(html, "lxml")) and pin it in your project requirements. That makes deployments reproducible and lets tests detect a parser change instead of silently changing extracted text.

Beautiful Soup extraction patterns

Extract the complete document

from bs4 import BeautifulSoup

def visible_text(html: str) -> str:
    soup = BeautifulSoup(html, "lxml")
    return soup.get_text(" ", strip=True)

print(visible_text("<p>One</p><p>Two</p>"))
# One Two

The result is one normalized string. This is useful for search indexing, logging, deduplication, or a quick inspection of a page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract one known region

from bs4 import BeautifulSoup

def article_text(html: str) -> str:
    soup = BeautifulSoup(html, "lxml")
    node = soup.select_one("article, main, [role='main']")
    if node is None:
        return ""
    return node.get_text(" ", strip=True)

Use the selectors that match your site. A selector is usually more dependable than trying to infer “the article” from every tag on an arbitrary page. Check for None; calling get_text() on a missing node raises an error.

Process fragments with stripped_strings

When you need to inspect, filter, or transform each fragment separately, iterate over stripped_strings:

from bs4 import BeautifulSoup

soup = BeautifulSoup("<p> First <strong>important</strong> point. </p>", "lxml")
fragments = list(soup.stripped_strings)
print(fragments)
# ['First', 'important', 'point.']

cleaned = " | ".join(fragments)
print(cleaned)

This gives you control over the joining rule instead of asking Beautiful Soup to produce one final string immediately.

Remove known page chrome before extraction

Tag removal and content selection solve different problems. If a cookie banner, navigation menu, comments area, or duplicate mobile markup is inside your selected region, remove those nodes first:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup

def clean_article(html: str) -> str:
    soup = BeautifulSoup(html, "lxml")

    for selector in ("script", "style", "template", "nav", "footer", ".cookie-banner", ".comments"):
        for node in soup.select(selector):
            node.decompose()

    article = soup.select_one("article, main, [role='main']") or soup
    return article.get_text(" ", strip=True)

Beautiful Soup’s documentation notes that script, style, and template contents are generally not treated as human-readable text when lxml or html.parser is used. Explicit removal is still useful when you want the same behavior across parser choices or when unwanted visible elements such as navigation and comments are present.

Use the standard library when dependencies are not wanted

Python’s html.parser is an event-driven parser. Subclass it, collect data in handle_data, and normalize the collected fragments yourself:

from html.parser import HTMLParser

class TextExtractor(HTMLParser):
    def __init__(self):
        super().__init__()
        self.parts = []

    def handle_data(self, data):
        self.parts.append(data)


def extract_text(html: str) -> str:
    extractor = TextExtractor()
    extractor.feed(html)
    extractor.close()
    return " ".join(" ".join(extractor.parts).split())

html = "<article><h1>Title</h1><p>Body</p></article>"
print(extract_text(html))
# Title Body

The callback receives text as the parser encounters it. This gives you low-level control, but it does not provide CSS selection, tree navigation, or automatic article identification. Add your own state if you need to ignore data inside selected tags:

from html.parser import HTMLParser

class ArticleParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.depth = 0
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag == "article":
            self.depth += 1

    def handle_endtag(self, tag):
        if tag == "article" and self.depth:
            self.depth -= 1

    def handle_data(self, data):
        if self.depth:
            self.parts.append(data)

parser = ArticleParser()
parser.feed(html)
text = " ".join(" ".join(parser.parts).split())

For anything beyond simple streaming collection, Beautiful Soup is usually less code and easier to maintain.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Whitespace, entities, and readable output

Prevent words from colliding

Prefer get_text(" ", strip=True) over get_text(strip=True) when inline elements can touch. The explicit space separates text fragments created by tags such as <span>, <em>, and links.

Choose your line-break policy

Readable text is not always one paragraph. If downstream code needs paragraph boundaries, iterate over block elements and join them with newlines:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "lxml")
paragraphs = [p.get_text(" ", strip=True) for p in soup.select("article p")]
text = "nn".join(p for p in paragraphs if p)

This preserves a deliberate paragraph structure while still normalizing spaces inside each paragraph. Do not assume every <br> represents a paragraph; treat it according to your output format.

Normalize after extraction, not before selection

Whitespace collapsing is useful for a final string, but applying it before selecting elements can hide boundaries that your selectors or tests rely on. Select and remove nodes first, then normalize the result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Static HTML is not the same as rendered page text

These libraries parse the HTML string you give them. If a page inserts its article with JavaScript after the initial response, the article will not exist in that string. You need to obtain rendered HTML with a browser-capable capture step, or use an endpoint that returns the content directly, and then pass the resulting HTML to Beautiful Soup.

Even with a rendered snapshot, extraction still requires a content decision. Navigation, cookie notices, chat widgets, comments, and repeated responsive components may be legitimate HTML but not part of the article. Combine a rendered capture with selectors, explicit removal rules, or a dedicated content-extraction stage.

Build a repeatable extraction pipeline

  1. Acquire bytes and decode them correctly. Preserve the response or saved HTML used for a failing case.
  2. Parse with a named backend. Use one of lxml, html5lib, or html.parser explicitly.
  3. Remove predictable noise. Delete selectors for scripts, navigation, banners, comments, or duplicate containers that do not belong in the output.
  4. Select the content root. Prefer a known article, main, or site-specific selector; fall back to the document only when that is acceptable.
  5. Extract with a separator. Call get_text(" ", strip=True) or process stripped_strings.
  6. Normalize for the consumer. Choose one-line text, paragraph-separated text, or a list of fragments deliberately.
  7. Test representative fixtures. Include valid pages, malformed markup, missing selectors, empty elements, and pages containing navigation or consent UI.

For performance, parse once and reuse the resulting tree for all selectors. If you only need a small fragment, select it before extracting instead of converting the entire document to text. For reliability, log the parser name, selector used, and whether the fallback path ran; those details explain most changes in output.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“No module named bs4” or parser-not-found errors

Install the package in the same environment that runs the script: python -m pip install beautifulsoup4 lxml. Virtual environments and system Python installations can have separate package directories.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Words are joined together

You likely omitted the separator. Replace get_text(strip=True) with get_text(" ", strip=True), or join stripped_strings with a space.

The output contains menus or a footer

You extracted the whole document. Select the article container first and remove known navigation, footer, comments, and banner selectors before calling get_text().

The expected element is missing

Inspect the HTML string you actually parsed. The content may be injected by JavaScript, the selector may differ for another template, or the page may have returned an error document. Check the selected node for None and keep a fallback policy explicit.

Results change after deployment

Different parser backends recover malformed markup differently. Pin the backend, keep dependency versions controlled, and run the same fixture tests in development and production.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Text includes an unwanted widget

Add that widget’s stable class, ID, or structural selector to your removal list before selecting the content root. Avoid broad rules that could delete legitimate article text.

The page is blank even though a browser shows content

The initial HTML likely lacks client-rendered content, or access is blocked by a bot check. Obtain a browser-rendered snapshot first, then parse its HTML; do not expect an HTML parser to execute JavaScript.

Or skip the browser setup

If obtaining clean, rendered HTML is the difficult part, ScreenshotNeo can capture a URL before you run your own extraction. Its API returns a screenshot or PDF rather than article text, so use it when you need a dependable visual/rendered capture alongside your parser workflow.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. The same call from Python:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.

FAQ

Frequently Asked Questions

Can Beautiful Soup parse an HTML fragment instead of a complete page?

Yes. Pass the fragment string to BeautifulSoup with the same explicit parser. You can then select elements and call get_text() exactly as you would for a full document.

Should extracted text be stored as one string or a list?

Use one string for search, display, or simple export. Keep a list of stripped fragments or paragraphs when later processing needs boundaries, headings, or per-element metadata.

Does HTMLParser provide CSS selectors?

No. HTMLParser reports parsing events through callbacks. If you need CSS selection or tree navigation, use Beautiful Soup or build those capabilities yourself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How can I verify that a selector still matches after a site redesign?

Treat selectors as configuration: test them against saved representative HTML, alert when the expected node is absent, and retain a documented fallback rather than silently returning an unrelated page-wide string.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.