What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use Beautiful Soup to parse the response into a tree, select the correct <table>, then walk each <tr> and its <th>/<td> cells. Normalize whitespace, preserve links or other attributes separately when needed, and validate row widths before exporting. If your goal is a DataFrame from a conventional table, pandas.read_html() is usually shorter.
Install the libraries
Install Requests and Beautiful Soup for fetching and parsing. Add pandas when you want DataFrame output. The examples use Python’s built-in html.parser; Beautiful Soup also supports lxml and html5lib.
python -m pip install requests beautifulsoup4 pandas lxml html5lib
Choose a parser explicitly. Different parsers can build different trees from malformed HTML, so an explicit choice makes results more reproducible. Beautiful Soup documents html.parser, lxml and html5lib; it describes lxml as faster for performance-sensitive parsing (Beautiful Soup documentation).
Fetch the HTML safely
Check the HTTP response before parsing. Requests uses the server’s declared encoding for response.text; set response.encoding yourself when the declaration is wrong (Requests Quickstart).
#1 Best Overall
import requests
url = "https://example.com/results"
response = requests.get(url, timeout=30)
response.raise_for_status()
# If you know the page's actual encoding, set it before reading text:
# response.encoding = "utf-8"
html = response.text
A timeout prevents a stalled request from hanging indefinitely. raise_for_status() turns HTTP error responses into an exception you can handle instead of parsing an error page as if it were data.
Locate the intended table
Do not assume the first table is the one you need. Target a stable id, class, or other attribute whenever possible.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
table = soup.find("table", id="results")
if table is None:
raise ValueError("The results table was not found")
CSS selectors are useful when the markup requires a more specific match:
table = soup.select_one("table.results-table[data-kind='public']")
if table is None:
raise ValueError("No matching table")
Use find_all() when you expect several tables and inspect their captions, ids, or surrounding headings before choosing one.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
Extract headers and rows
This complete example fetches a table, reads headers, extracts cell text, and rejects inconsistent rows. It searches descendants by default, which handles cells nested inside <thead> and <tbody>.
import csv
import requests
from bs4 import BeautifulSoup
url = "https://example.com/results"
response = requests.get(url, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
table = soup.find("table", id="results")
if table is None:
raise RuntimeError("Could not find table#results")
# Prefer an explicit header row. Fall back to the first row if needed.
header_row = table.find("tr")
if header_row is None:
raise RuntimeError("Table contains no rows")
header_cells = header_row.find_all("th")
if not header_cells:
header_cells = header_row.find_all(["th", "td"])
headers = [cell.get_text(" ", strip=True) for cell in header_cells]
records = []
for row_number, row in enumerate(table.find_all("tr"), start=1):
cells = row.find_all(["th", "td"])
values = [cell.get_text(" ", strip=True) for cell in cells]
if not values or row is header_row:
continue
if len(values) != len(headers):
raise ValueError(
f"row {row_number} has {len(values)} cells; expected {len(headers)}"
)
records.append(dict(zip(headers, values)))
with open("results.csv", "w", newline="", encoding="utf-8") as output:
writer = csv.DictWriter(output, fieldnames=headers)
writer.writeheader()
writer.writerows(records)
print(f"Wrote {len(records)} records")
get_text(" ", strip=True) joins text from nested elements with spaces and trims surrounding whitespace. It does not retain an anchor’s URL, an image’s alt value, or the distinction between nested semantic elements.
Keep links and other cell attributes
Extract structured content deliberately instead of relying only on text.
for row in table.find_all("tr"):
for cell in row.find_all(["th", "td"]):
label = cell.get_text(" ", strip=True)
links = [a.get("href") for a in cell.find_all("a", href=True)]
print({"text": label, "links": links})
Use recursive=False when you need only direct child rows or cells and must avoid nested tables. Inspect empty rows, nested tables, and cells with rowspan or colspan before assuming a rectangular shape.
Recommended Free Tools
Handle headers, missing values and irregular markup
Separate thead and tbody when present
thead = table.find("thead")
tbody = table.find("tbody")
header_cells = thead.find_all(["th", "td"]) if thead else []
body_rows = tbody.find_all("tr") if tbody else table.find_all("tr")
Some pages place header cells in the first body row; detect that case rather than hard-coding a location.
Validate widths and types
Log or reject rows whose cell count differs from the header count. Convert numeric and date strings only after extraction, with explicit handling for thousands separators, currency symbols, blanks, and locale-specific formats. Keep the raw text if the original representation matters.
Account for rowspan and colspan
A visual table with merged cells is not automatically a rectangular list. A cell with rowspan applies to subsequent rows and colspan occupies multiple columns. Either implement a grid-expansion routine that tracks occupied coordinates or use a parser that understands the structure, then inspect representative output. Never silently zip short rows to headers.
Use pandas for a DataFrame
For an ordinary HTML table whose desired result is tabular data, pandas.read_html() can avoid manual traversal. The pandas API describes it as reading HTML tables into a list of DataFrame objects (pandas.read_html).
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →import pandas as pd
frames = pd.read_html(
"https://example.com/results",
match="Product",
attrs={"id": "results"},
header=0,
)
if not frames:
raise ValueError("No matching table")
df = frames[0]
df.to_csv("results.csv", index=False)
It always returns a list, even when one table matches. Use match to require text and attrs for valid table attributes; other useful options include header, index_col, skiprows, converters and missing-value handling. Inspect and clean the result: pandas attempts to assume little about the source structure and may require manually assigned column names. It attempts to handle rowspan and colspan. The documentation notes that an empty list is a rare possible result.
| Need | Better starting point |
|---|---|
| Exact control over every cell, links, attributes or irregular markup | Beautiful Soup traversal |
| Quick rectangular DataFrame from a conventional table | pandas.read_html() |
| Several tables and one must be selected by text or attributes | read_html(match=..., attrs=...), then inspect the returned list |
When no table is found
- Wrong document: print or save
response.textand confirm the expected table markup is actually present. - Wrong selector: inspect ids, classes and capitalization; test
soup.find_all("table"). - Malformed HTML: try an explicitly installed parser such as
lxmlorhtml5lib. Parser libraries can construct different trees from invalid markup. - Content loaded after the response: the initial HTML may not contain rows that a browser later receives through JavaScript. Beautiful Soup parses what you fetched; it does not execute page scripts. Identify an authorized data endpoint or use a browser-rendering capture workflow.
- Access or error page: check status, redirects, response headers and a short prefix of the body before parsing.
For pandas specifically, its HTML-parsing guide notes that lxml is fast but does not guarantee results for strictly invalid markup. It documents fallback behavior involving BeautifulSoup and html5lib and recommends installing those dependencies alongside lxml when that fallback is needed (pandas HTML table parsing gotchas).
Performance, reliability and responsible output
- Fetch once and parse once; reuse the parsed table rather than repeatedly searching the document.
- Use a stable selector and assertions for required headers so a site redesign fails loudly.
- Choose lxml when parsing speed matters and it is available; keep the parser choice fixed in production.
- Set connection and read timeouts, catch request exceptions, and preserve the source URL and retrieval timestamp with exported data.
- Write UTF-8 output and test blank cells, duplicate headers, nested tags and merged cells.
- Respect the site’s terms, access controls and applicable laws; collect only data you are authorized to use.
Or skip the browser setup
If the table is on a page that needs browser rendering, you can first capture a clean page image or PDF with ScreenshotNeo, a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.
The direct API call returns an image or PDF; it does not replace HTML parsing when you need cell values. For a visual record or a page whose layout must be rendered, use:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options such as full-page capture with lazy images, CSS-selector element capture, device and viewport settings, custom JavaScript or CSS, waits, request blocking, cookies and headers, PDFs, caching, signed links, asynchronous jobs, bulk capture and usage reporting.
Best Value
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Should I use Beautiful Soup or pandas?
Use Beautiful Soup when cell-level control or non-tabular content matters. Use pandas when the intended result is a cleaned DataFrame from a conventional table.
Why does read_html() return a list?
A page can contain multiple tables, so pandas returns one DataFrame per match. Select and validate the expected frame instead of assuming index zero is always correct.
Can Beautiful Soup scrape a JavaScript-generated table?
Only if the rows are present in the HTML you provide. Beautiful Soup does not execute JavaScript; obtain rendered HTML or an authorized underlying endpoint first.
Frequently Asked Questions
How do I preserve a table cell’s hyperlink?
Extract the cell text and call cell.find_all("a", href=True) to store each URL separately; get_text() returns text only.
What does an empty pandas result mean?
Verify the fetched document, your match text and attrs. Pandas documents an empty list as a rare outcome when no table is recognized.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches

