The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Beautiful Soup parses HTML or XML that you give it; it does not download web pages or run their JavaScript. A basic scraper therefore has two jobs: retrieve the page with an HTTP client such as Requests, then parse the returned markup with Beautiful Soup. This guide shows the complete Python 3 workflow, how to choose a parser, find and extract data, and diagnose common failures.
What Beautiful Soup does—and what it does not do
Beautiful Soup turns supplied HTML or XML into a navigable tree of Python objects. It can search that tree, extract text and attributes, and modify the parsed structure. The project describes it as “a Python library for pulling data out of HTML and XML files.” Beautiful Soup documentation
As an Amazon Associate I earn from qualifying purchases.
It does not send HTTP requests, crawl a site, or execute browser JavaScript. Use an HTTP client to retrieve markup, check what came back, and then pass the response content to Beautiful Soup. If a page is built in the browser after scripts run, the initial HTTP response may not contain the data you see on screen.
Free tools Windows power users keep installed
One-click scans. No signup required.
Install Beautiful Soup and Requests
Use Python 3. Install the Beautiful Soup package and Requests, which the example uses to fetch a page:
#1 Best Overall
python -m pip install beautifulsoup4 requests
The package is named beautifulsoup4, but you import it from the bs4 namespace. The older Python 2-specific installation path is not relevant to this Python 3 example. Beautiful Soup on PyPI
Fetch a page, parse it, and extract data
Save this as scrape.py and run it with Python. It checks the HTTP response before parsing, explicitly selects a parser, handles a missing title, and extracts links that actually have an href attribute.
import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
response = requests.get(
url,
headers={"User-Agent": "Example scraper contact: [email protected]"},
timeout=20,
)
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")
title = soup.title.get_text(strip=True) if soup.title else "(no title found)"
print("Title:", title)
for link in soup.find_all("a", href=True):
text = link.get_text(" ", strip=True)
href = link.get("href")
print(text, href)
Replace the example URL and contact text with values appropriate for your use. Requests’ quickstart documents its response object and the process of making requests. Requests Quickstart
Why pass response.content?
response.content is the response body as bytes. Beautiful Soup can use that input while accounting for encoding information in the markup. Requests also provides response.text, a decoded string whose encoding is selected from the response information and detection. If characters look corrupted, inspect the response headers and encoding before changing your selectors.
Rank #2
Check what the server returned
raise_for_status() stops the script when the response has an unsuccessful HTTP status instead of silently parsing an error page as if it were the intended document. For unexpected results, inspect response.status_code, response.headers, and a portion of response.text or response.content before debugging the parse.
Choose a parser deliberately
Beautiful Soup supports Python’s built-in html.parser and optional parsers including lxml and html5lib. The parser interprets markup and constructs the tree Beautiful Soup searches. Malformed HTML can produce different trees with different parsers, so specify the parser instead of relying on an environment-dependent default.
| Parser | When to choose it | Consideration |
|---|---|---|
html.parser |
A built-in option for HTML with no extra parser package. | Explicitly name it in the constructor for consistent intent. |
lxml |
When you have installed and chosen the lxml parser; use its XML mode for XML. | It is an optional dependency. Parser choice can affect the resulting tree. |
html5lib |
When you have installed and chosen html5lib for parsing HTML. | It is optional, and can produce a different tree from other parsers. |
Install an optional parser separately if you choose it, then name it in the constructor, for example BeautifulSoup(response.content, "lxml"). For XML, use BeautifulSoup(xml_bytes, "xml") with lxml installed, as directed in the Beautiful Soup documentation. The documentation options establish different parser choices; they do not establish current speed rankings, so choose for input compatibility and predictable output rather than an assumed benchmark.
Find elements and extract their values
Use find() for one match
find() returns the first matching element or None. Check for a match before accessing its text or attributes:
heading = soup.find("h1")
if heading is not None:
print(heading.get_text(" ", strip=True))
Use find_all() for repeated matches
find_all() returns all matching elements. You can filter by tag name and attributes, then extract text or an attribute from each tag:
for item in soup.find_all("a", class_="product-link"):
label = item.get_text(" ", strip=True)
destination = item.get("href")
print(label, destination)
In Python, class_ is used because class is a reserved word. The attribute lookup with get() safely returns None if the attribute is absent.
Use CSS selectors when they clarify relationships
select() accepts CSS selectors and is useful when a relationship or combination of attributes is clearer as a selector:
for item in soup.select("article.product a.product-link[href]"):
print(item.get_text(" ", strip=True), item.get("href"))
Use whichever form makes the target structure easiest to understand and maintain. Verify the selector against the markup you actually fetched; do not assume a page’s visual arrangement or an element’s position will remain stable.
Extract text without assuming a match exists
For a tag, get_text(" ", strip=True) combines descendant text, separates pieces with spaces, and strips surrounding whitespace. A missing element is not an empty tag, so test the lookup result before calling methods on it. Avoid brittle assumptions such as treating the third paragraph as a price unless the target page explicitly guarantees that structure.
When the returned HTML does not match the browser
A simple Requests-and-Beautiful-Soup script parses the HTTP response; it does not render the page as a browser or execute page scripts. A site may insert content after JavaScript runs, so the browser’s visible text can be absent from the response. Inspect the response body first. If the required content is not present there, changing Beautiful Soup selectors will not make it appear: use a retrieval method appropriate to the page, subject to the site’s rules.
Troubleshoot common scraping failures
| Symptom | Likely cause | What to check or do |
|---|---|---|
Lookup returns None or an empty list |
The selector does not match the returned markup, the page structure changed, or the data is added by JavaScript. | Inspect the response body and parsed tree; verify tag names, attributes, and nesting. If the content is not in the response, the simple request-and-parse approach cannot extract it. |
| The script parses an error page or unexpected document | The HTTP request returned an unsuccessful response or different content than expected. | Check the status code, headers, and response body; call raise_for_status() before parsing. |
| Text contains broken or unexpected characters | The response encoding or decoding may not match the content. | Inspect response headers and encoding, and compare response.text with the raw response.content bytes before altering selectors. |
| Results differ across computers | A different parser may have constructed a different tree from imperfect markup. | Specify the parser explicitly and ensure the chosen optional parser is installed in each environment. |
| A tag exists but the expected attribute is missing | The matching element may not have that attribute. | Use tag.get("attribute") and handle a None result instead of assuming the attribute exists. |
Keep scraping responsible and maintainable
- Check the target site’s current terms, access controls, robots directives, and other applicable requirements before collecting data; obligations can differ by site and jurisdiction.
- Obtain authorization where needed and avoid sending requests at a rate that burdens the service.
- Keep selectors tied to meaningful tags and attributes, and revisit them when the site’s markup changes.
- Log enough response and extraction information to distinguish an HTTP problem from a parsing mismatch.
Beautiful Soup and Requests documentation describe software behavior, not whether scraping a particular site is permitted. Make that determination for the specific site and your circumstances.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchOr skip the browser setup
If your task is to capture a rendered page rather than parse HTML yourself, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF; its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Each step can be turned off.
For example, this saves a WebP screenshot of the target URL. See the ScreenshotNeo API documentation for request options and response details.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo bills only clean shots: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents, including Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, with no card required.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Frequently Asked Questions
Does Beautiful Soup scrape a page on its own?
No. An HTTP client retrieves the page; Beautiful Soup parses the markup it receives.
Can Beautiful Soup extract content that appears only after JavaScript runs?
Not from a response that does not contain that content. Beautiful Soup parses markup and does not execute page scripts.
Should I use find_all() or select()?
Use the method that expresses the target structure most clearly and is easiest for you to maintain; verify it against the fetched markup.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




