What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Beautiful Soup to parse the response into a tree, select the correct <table>, then walk each <tr> and its <th>/<td> cells. Normalize whitespace, preserve links or other attributes separately when needed, and validate row widths before exporting. If your goal is a DataFrame from a conventional table, pandas.read_html() is usually shorter.

Install the libraries

Install Requests and Beautiful Soup for fetching and parsing. Add pandas when you want DataFrame output. The examples use Python’s built-in html.parser; Beautiful Soup also supports lxml and html5lib.

python -m pip install requests beautifulsoup4 pandas lxml html5lib

Choose a parser explicitly. Different parsers can build different trees from malformed HTML, so an explicit choice makes results more reproducible. Beautiful Soup documents html.parser, lxml and html5lib; it describes lxml as faster for performance-sensitive parsing (Beautiful Soup documentation).

Fetch the HTML safely

Check the HTTP response before parsing. Requests uses the server’s declared encoding for response.text; set response.encoding yourself when the declaration is wrong (Requests Quickstart).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

url = "https://example.com/results"
response = requests.get(url, timeout=30)
response.raise_for_status()

# If you know the page's actual encoding, set it before reading text:
# response.encoding = "utf-8"
html = response.text

A timeout prevents a stalled request from hanging indefinitely. raise_for_status() turns HTTP error responses into an exception you can handle instead of parsing an error page as if it were data.

Locate the intended table

Do not assume the first table is the one you need. Target a stable id, class, or other attribute whenever possible.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
table = soup.find("table", id="results")

if table is None:
    raise ValueError("The results table was not found")

CSS selectors are useful when the markup requires a more specific match:

table = soup.select_one("table.results-table[data-kind='public']")
if table is None:
    raise ValueError("No matching table")

Use find_all() when you expect several tables and inspect their captions, ids, or surrounding headings before choosing one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract headers and rows

This complete example fetches a table, reads headers, extracts cell text, and rejects inconsistent rows. It searches descendants by default, which handles cells nested inside <thead> and <tbody>.

import csv
import requests
from bs4 import BeautifulSoup

url = "https://example.com/results"
response = requests.get(url, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

table = soup.find("table", id="results")
if table is None:
    raise RuntimeError("Could not find table#results")

# Prefer an explicit header row. Fall back to the first row if needed.
header_row = table.find("tr")
if header_row is None:
    raise RuntimeError("Table contains no rows")

header_cells = header_row.find_all("th")
if not header_cells:
    header_cells = header_row.find_all(["th", "td"])
headers = [cell.get_text(" ", strip=True) for cell in header_cells]

records = []
for row_number, row in enumerate(table.find_all("tr"), start=1):
    cells = row.find_all(["th", "td"])
    values = [cell.get_text(" ", strip=True) for cell in cells]
    if not values or row is header_row:
        continue
    if len(values) != len(headers):
        raise ValueError(
            f"row {row_number} has {len(values)} cells; expected {len(headers)}"
        )
    records.append(dict(zip(headers, values)))

with open("results.csv", "w", newline="", encoding="utf-8") as output:
    writer = csv.DictWriter(output, fieldnames=headers)
    writer.writeheader()
    writer.writerows(records)

print(f"Wrote {len(records)} records")

get_text(" ", strip=True) joins text from nested elements with spaces and trims surrounding whitespace. It does not retain an anchor’s URL, an image’s alt value, or the distinction between nested semantic elements.

Keep links and other cell attributes

Extract structured content deliberately instead of relying only on text.

for row in table.find_all("tr"):
    for cell in row.find_all(["th", "td"]):
        label = cell.get_text(" ", strip=True)
        links = [a.get("href") for a in cell.find_all("a", href=True)]
        print({"text": label, "links": links})

Use recursive=False when you need only direct child rows or cells and must avoid nested tables. Inspect empty rows, nested tables, and cells with rowspan or colspan before assuming a rectangular shape.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle headers, missing values and irregular markup

Separate thead and tbody when present

thead = table.find("thead")
tbody = table.find("tbody")
header_cells = thead.find_all(["th", "td"]) if thead else []
body_rows = tbody.find_all("tr") if tbody else table.find_all("tr")

Some pages place header cells in the first body row; detect that case rather than hard-coding a location.

Validate widths and types

Log or reject rows whose cell count differs from the header count. Convert numeric and date strings only after extraction, with explicit handling for thousands separators, currency symbols, blanks, and locale-specific formats. Keep the raw text if the original representation matters.

Account for rowspan and colspan

A visual table with merged cells is not automatically a rectangular list. A cell with rowspan applies to subsequent rows and colspan occupies multiple columns. Either implement a grid-expansion routine that tracks occupied coordinates or use a parser that understands the structure, then inspect representative output. Never silently zip short rows to headers.

Use pandas for a DataFrame

For an ordinary HTML table whose desired result is tabular data, pandas.read_html() can avoid manual traversal. The pandas API describes it as reading HTML tables into a list of DataFrame objects (pandas.read_html).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import pandas as pd

frames = pd.read_html(
    "https://example.com/results",
    match="Product",
    attrs={"id": "results"},
    header=0,
)
if not frames:
    raise ValueError("No matching table")
df = frames[0]
df.to_csv("results.csv", index=False)

It always returns a list, even when one table matches. Use match to require text and attrs for valid table attributes; other useful options include header, index_col, skiprows, converters and missing-value handling. Inspect and clean the result: pandas attempts to assume little about the source structure and may require manually assigned column names. It attempts to handle rowspan and colspan. The documentation notes that an empty list is a rare possible result.

Need Better starting point
Exact control over every cell, links, attributes or irregular markup Beautiful Soup traversal
Quick rectangular DataFrame from a conventional table pandas.read_html()
Several tables and one must be selected by text or attributes read_html(match=..., attrs=...), then inspect the returned list

When no table is found

  • Wrong document: print or save response.text and confirm the expected table markup is actually present.
  • Wrong selector: inspect ids, classes and capitalization; test soup.find_all("table").
  • Malformed HTML: try an explicitly installed parser such as lxml or html5lib. Parser libraries can construct different trees from invalid markup.
  • Content loaded after the response: the initial HTML may not contain rows that a browser later receives through JavaScript. Beautiful Soup parses what you fetched; it does not execute page scripts. Identify an authorized data endpoint or use a browser-rendering capture workflow.
  • Access or error page: check status, redirects, response headers and a short prefix of the body before parsing.

For pandas specifically, its HTML-parsing guide notes that lxml is fast but does not guarantee results for strictly invalid markup. It documents fallback behavior involving BeautifulSoup and html5lib and recommends installing those dependencies alongside lxml when that fallback is needed (pandas HTML table parsing gotchas).

Performance, reliability and responsible output

  • Fetch once and parse once; reuse the parsed table rather than repeatedly searching the document.
  • Use a stable selector and assertions for required headers so a site redesign fails loudly.
  • Choose lxml when parsing speed matters and it is available; keep the parser choice fixed in production.
  • Set connection and read timeouts, catch request exceptions, and preserve the source URL and retrieval timestamp with exported data.
  • Write UTF-8 output and test blank cells, duplicate headers, nested tags and merged cells.
  • Respect the site’s terms, access controls and applicable laws; collect only data you are authorized to use.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If the table is on a page that needs browser rendering, you can first capture a clean page image or PDF with ScreenshotNeo, a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.

The direct API call returns an image or PDF; it does not replace HTML parsing when you need cell values. For a visual record or a page whose layout must be rendered, use:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options such as full-page capture with lazy images, CSS-selector element capture, device and viewport settings, custom JavaScript or CSS, waits, request blocking, cookies and headers, PDFs, caching, signed links, asynchronous jobs, bulk capture and usage reporting.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Should I use Beautiful Soup or pandas?

Use Beautiful Soup when cell-level control or non-tabular content matters. Use pandas when the intended result is a cleaned DataFrame from a conventional table.

Why does read_html() return a list?

A page can contain multiple tables, so pandas returns one DataFrame per match. Select and validate the expected frame instead of assuming index zero is always correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can Beautiful Soup scrape a JavaScript-generated table?

Only if the rows are present in the HTML you provide. Beautiful Soup does not execute JavaScript; obtain rendered HTML or an authorized underlying endpoint first.

Frequently Asked Questions

How do I preserve a table cell’s hyperlink?

Extract the cell text and call cell.find_all("a", href=True) to store each URL separately; get_text() returns text only.

What does an empty pandas result mean?

Verify the fetched document, your match text and attrs. Pandas documents an empty list as a rare outcome when no table is recognized.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.