Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a conventional HTML table already present in a page’s markup, use pandas.read_html(): it returns a list of DataFrames, so inspect the list and select the table you actually need. Use Beautiful Soup when you need custom table selection, cell-by-cell control, links, or attributes that a DataFrame conversion does not handle the way you want.

Choose the right extraction method

Approach Best for What you get Trade-off
pandas.read_html() Ordinary rendered HTML tables you want to analyze as tabular data A list of DataFrames Selection and cleanup may be needed when a page has several tables or irregular headers
Beautiful Soup Custom selection, nested structures, cell-by-cell processing, or extracting links and attributes Elements from the parsed HTML tree, which you shape into the data you need You write and maintain the row and cell traversal logic

Neither method by itself guarantees that the output matches the page’s intended meaning. Verify the selected table, headers, row count, representative values, and missing data before using the result.

Read a table into a DataFrame with pandas

The pandas API describes read_html as reading HTML tables into a list of DataFrame objects. Pass it a URL, a file path, or file-like HTML input. The returned list matters: a page with one table still produces a list, and a page with several tables produces multiple candidates.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the packages

For a straightforward setup, install pandas and the parser libraries explicitly:

python -m pip install pandas lxml

pandas tries lxml by default, and may fall back to Beautiful Soup with html5lib if it cannot parse with lxml. If your environment needs that fallback, install the dependencies:

python -m pip install beautifulsoup4 html5lib

Parser behavior and package availability vary by environment; keeping the parser dependency explicit helps make setup and troubleshooting clearer.

Fetch and inspect all tables

import pandas as pd

url = "https://example.com/data"
tables = pd.read_html(url)

print(f"Found {len(tables)} tables")
for index, table in enumerate(tables):
    print(f"nTable {index}")
    print(table.head())

Replace the example URL with the page you are authorized to retrieve. Do not assume tables[0] is the desired result simply because it is first: navigation, summaries, or other page content may produce earlier tables. Inspect dimensions and column names before selecting one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Select a table by text or attributes

When the page has multiple tables, match can select tables containing distinctive text. Use a phrase that is actually present in the target table:

import pandas as pd

url = "https://example.com/data"
tables = pd.read_html(url, match="Quarterly revenue")

if not tables:
    raise RuntimeError("No table matched the requested text")

revenue = tables[0]
print(revenue.head())

You can also select by a table attribute, such as an HTML id:

tables = pd.read_html(url, attrs={"id": "results-table"})
if not tables:
    raise RuntimeError("No table with id='results-table' was found")
results = tables[0]

Use a valid attribute and value from the page’s markup. Selection filters which tables are returned; it does not verify that their columns or values are semantically correct.

Configure headers and value parsing

Tables often need small adjustments after selection. pandas exposes options for headers, skipped rows, numeric converters, thousands separators, decimal marks, encoding, and link extraction. For example, if a page uses commas for thousands and periods for decimals, specify those conventions where appropriate:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
tables = pd.read_html(
    url,
    match="Quarterly revenue",
    thousands=",",
    decimal="."
)
frame = tables[0]

If the real column names appear on a row other than the one pandas inferred, inspect the page’s header rows and adjust the header option. If a title or note row precedes the data, consider skiprows. For columns with unusual formats, use converters to define parsing explicitly. These options are controls, not automatic validation: check a few known values against the source.

Use Beautiful Soup for custom extraction

Beautiful Soup is a better fit when you need to choose a table using surrounding structure, walk individual rows, preserve particular cell logic, or read anchor URLs and other attributes. Its documentation supports the constructor pattern shown below; this example explicitly chooses Python’s built-in html.parser.

Parse markup and walk table rows

from bs4 import BeautifulSoup

html = """<table id="results-table">
  <tr><th>Name</th><th>Score</th></tr>
  <tr><td>Ava</td><td>97</td></tr>
  <tr><td>Noah</td><td>91</td></tr>
</table>"""

soup = BeautifulSoup(html, "html.parser")
table = soup.find("table", id="results-table")
if table is None:
    raise ValueError("Target table was not found")

rows = []
for tr in table.find_all("tr"):
    cells = tr.find_all(["th", "td"])
    rows.append([cell.get_text(" ", strip=True) for cell in cells])

for row in rows:
    print(row)

The output is a list of lists, not a DataFrame. The first row in this example contains headings; decide explicitly whether to keep it as data, separate it into column names, or handle multi-row headings differently.

Extract links or attributes

If a cell contains a link, text extraction alone loses its destination. Read the element attribute as well:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
for tr in table.find_all("tr"):
    cells = tr.find_all(["th", "td"])
    row = []
    for cell in cells:
        link = cell.find("a", href=True)
        row.append({
            "text": cell.get_text(" ", strip=True),
            "href": link["href"] if link else None,
        })
    print(row)

Adapt the record shape to your use case; for cells without links, href is None. If you only need plain text, the simpler traversal avoids carrying unnecessary fields.

Choose a parser for the actual HTML

Parser Documented trade-off Use when
html.parser Built into Python, so it needs no separate parser package; parser choice can affect the tree You want a simple dependency setup and the markup parses as expected
lxml Fast, but requires an external C dependency; pandas describes it as less predictable on invalid markup It is available in your environment and the input is sufficiently well-formed
html5lib More lenient with malformed markup, but slower You need tolerant parsing and accept the performance trade-off

Beautiful Soup can use each of these backends. Invalid markup may yield different trees under different parsers, so compare the parsed result with the target page rather than assuming every backend will identify the same rows and cells.

Validate and clean the extracted data

Parsing turns markup into data structures; it does not establish that the data is complete or correct. Before analysis or storage, check the following:

  • Table identity: confirm the selected table is the one you intended, particularly when using an index into pandas’ returned list.
  • Column names and header rows: check whether headers became missing values, were read from the wrong row, or span multiple rows.
  • Row count and sample cells: compare the number of extracted records and several representative values with the page.
  • Spanned cells: inspect how rowspan and colspan affect the resulting shape.
  • Missing values and numeric conventions: verify blanks, thousands separators, decimal marks, and converted types.
  • Links and attributes: explicitly extract these when they matter; visible text alone does not preserve them.

Keep a small validation check in your script, such as asserting expected columns or a plausible row count, so a page redesign is less likely to silently change downstream output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to do when a table is not in the HTML response

read_html and Beautiful Soup parse HTML markup that is available to them. If the table is not present in the input markup, these parsers cannot extract it from that markup. First confirm the HTML you are passing contains the table. If a browser displays content that is absent from your HTML input, the page may require a different retrieval or rendering workflow; do not treat an empty parse as proof that the site has no table.

Or skip the browser setup

If your task is to capture a rendered page for inspection rather than turn its table directly into a DataFrame, ScreenshotNeo provides a screenshot API and MCP server. Its API accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. For example, save a screenshot of a rendered page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request parameters. ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client. The free plan includes 1,000 shots a month without a card; paid plans start at $5 for 3,000 shots. This is for visual capture, not a replacement for extracting table cells into structured Python data. Learn about ScreenshotNeo, or sign up free for 1,000 screenshots a month, with no card required.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common extraction problems

No tables returned or no table matched

Check that the input is HTML containing a table and that your match text or attrs value occurs on the intended table. Print the number of tables without a filter first, then inspect their headers and first rows.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parser dependency or parsing errors

Confirm that the parser package your environment needs is installed. Try an explicitly installed parser backend; if the markup is malformed, compare output using a more lenient parser such as html5lib. A different parser can produce a different tree, so verify the result rather than only checking that an exception disappeared.

Unexpected headers, columns, or missing values

Inspect the table’s header rows and any title or notes embedded in the table. Adjust pandas’ header or skipped-row settings if needed, and check whether cells spanning rows or columns changed the shape. Validate column names and sample values after applying the change.

Numbers parse incorrectly

Determine the source’s thousands and decimal conventions, then pass the matching thousands and decimal settings or a column converter. Check missing values and representative numeric cells after conversion.

Beautiful Soup finds the wrong table or loses link destinations

Narrow the lookup with an appropriate table attribute or surrounding structure. Use find_all deliberately when nested tables are present, and retrieve each relevant anchor’s href; get_text() returns text, not link targets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Does pandas return one DataFrame or several?

read_html returns a list of DataFrames, including when the HTML contains only one table. Inspect the list and select the intended entry.

Can Beautiful Soup create a DataFrame?

Yes. Its traversal gives you values to organize, but the row and column structure is your code to define. Use pandas directly when ordinary table-to-DataFrame conversion is sufficient.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.