Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a conventional HTML table already present in a page’s markup, use pandas.read_html(): it returns a list of DataFrames, so inspect the list and select the table you actually need. Use Beautiful Soup when you need custom table selection, cell-by-cell control, links, or attributes that a DataFrame conversion does not handle the way you want.
Choose the right extraction method
| Approach | Best for | What you get | Trade-off |
|---|---|---|---|
pandas.read_html() |
Ordinary rendered HTML tables you want to analyze as tabular data | A list of DataFrames | Selection and cleanup may be needed when a page has several tables or irregular headers |
| Beautiful Soup | Custom selection, nested structures, cell-by-cell processing, or extracting links and attributes | Elements from the parsed HTML tree, which you shape into the data you need | You write and maintain the row and cell traversal logic |
Neither method by itself guarantees that the output matches the page’s intended meaning. Verify the selected table, headers, row count, representative values, and missing data before using the result.
Read a table into a DataFrame with pandas
The pandas API describes read_html as reading HTML tables into a list of DataFrame objects. Pass it a URL, a file path, or file-like HTML input. The returned list matters: a page with one table still produces a list, and a page with several tables produces multiple candidates.
Free tools Windows power users keep installed
One-click scans. No signup required.
Install the packages
For a straightforward setup, install pandas and the parser libraries explicitly:
#1 Best Overall
python -m pip install pandas lxml
pandas tries lxml by default, and may fall back to Beautiful Soup with html5lib if it cannot parse with lxml. If your environment needs that fallback, install the dependencies:
python -m pip install beautifulsoup4 html5lib
Parser behavior and package availability vary by environment; keeping the parser dependency explicit helps make setup and troubleshooting clearer.
Fetch and inspect all tables
import pandas as pd
url = "https://example.com/data"
tables = pd.read_html(url)
print(f"Found {len(tables)} tables")
for index, table in enumerate(tables):
print(f"nTable {index}")
print(table.head())
Replace the example URL with the page you are authorized to retrieve. Do not assume tables[0] is the desired result simply because it is first: navigation, summaries, or other page content may produce earlier tables. Inspect dimensions and column names before selecting one.
Select a table by text or attributes
When the page has multiple tables, match can select tables containing distinctive text. Use a phrase that is actually present in the target table:
import pandas as pd
url = "https://example.com/data"
tables = pd.read_html(url, match="Quarterly revenue")
if not tables:
raise RuntimeError("No table matched the requested text")
revenue = tables[0]
print(revenue.head())
You can also select by a table attribute, such as an HTML id:
Rank #2
tables = pd.read_html(url, attrs={"id": "results-table"})
if not tables:
raise RuntimeError("No table with id='results-table' was found")
results = tables[0]
Use a valid attribute and value from the page’s markup. Selection filters which tables are returned; it does not verify that their columns or values are semantically correct.
Configure headers and value parsing
Tables often need small adjustments after selection. pandas exposes options for headers, skipped rows, numeric converters, thousands separators, decimal marks, encoding, and link extraction. For example, if a page uses commas for thousands and periods for decimals, specify those conventions where appropriate:
Recommended Free Tools
tables = pd.read_html(
url,
match="Quarterly revenue",
thousands=",",
decimal="."
)
frame = tables[0]
If the real column names appear on a row other than the one pandas inferred, inspect the page’s header rows and adjust the header option. If a title or note row precedes the data, consider skiprows. For columns with unusual formats, use converters to define parsing explicitly. These options are controls, not automatic validation: check a few known values against the source.
Use Beautiful Soup for custom extraction
Beautiful Soup is a better fit when you need to choose a table using surrounding structure, walk individual rows, preserve particular cell logic, or read anchor URLs and other attributes. Its documentation supports the constructor pattern shown below; this example explicitly chooses Python’s built-in html.parser.
Parse markup and walk table rows
from bs4 import BeautifulSoup
html = """<table id="results-table">
<tr><th>Name</th><th>Score</th></tr>
<tr><td>Ava</td><td>97</td></tr>
<tr><td>Noah</td><td>91</td></tr>
</table>"""
soup = BeautifulSoup(html, "html.parser")
table = soup.find("table", id="results-table")
if table is None:
raise ValueError("Target table was not found")
rows = []
for tr in table.find_all("tr"):
cells = tr.find_all(["th", "td"])
rows.append([cell.get_text(" ", strip=True) for cell in cells])
for row in rows:
print(row)
The output is a list of lists, not a DataFrame. The first row in this example contains headings; decide explicitly whether to keep it as data, separate it into column names, or handle multi-row headings differently.
Extract links or attributes
If a cell contains a link, text extraction alone loses its destination. Read the element attribute as well:
for tr in table.find_all("tr"):
cells = tr.find_all(["th", "td"])
row = []
for cell in cells:
link = cell.find("a", href=True)
row.append({
"text": cell.get_text(" ", strip=True),
"href": link["href"] if link else None,
})
print(row)
Adapt the record shape to your use case; for cells without links, href is None. If you only need plain text, the simpler traversal avoids carrying unnecessary fields.
Choose a parser for the actual HTML
| Parser | Documented trade-off | Use when |
|---|---|---|
html.parser |
Built into Python, so it needs no separate parser package; parser choice can affect the tree | You want a simple dependency setup and the markup parses as expected |
lxml |
Fast, but requires an external C dependency; pandas describes it as less predictable on invalid markup | It is available in your environment and the input is sufficiently well-formed |
html5lib |
More lenient with malformed markup, but slower | You need tolerant parsing and accept the performance trade-off |
Beautiful Soup can use each of these backends. Invalid markup may yield different trees under different parsers, so compare the parsed result with the target page rather than assuming every backend will identify the same rows and cells.
Validate and clean the extracted data
Parsing turns markup into data structures; it does not establish that the data is complete or correct. Before analysis or storage, check the following:
- Table identity: confirm the selected table is the one you intended, particularly when using an index into pandas’ returned list.
- Column names and header rows: check whether headers became missing values, were read from the wrong row, or span multiple rows.
- Row count and sample cells: compare the number of extracted records and several representative values with the page.
- Spanned cells: inspect how
rowspanandcolspanaffect the resulting shape. - Missing values and numeric conventions: verify blanks, thousands separators, decimal marks, and converted types.
- Links and attributes: explicitly extract these when they matter; visible text alone does not preserve them.
Keep a small validation check in your script, such as asserting expected columns or a plausible row count, so a page redesign is less likely to silently change downstream output.
What to do when a table is not in the HTML response
read_html and Beautiful Soup parse HTML markup that is available to them. If the table is not present in the input markup, these parsers cannot extract it from that markup. First confirm the HTML you are passing contains the table. If a browser displays content that is absent from your HTML input, the page may require a different retrieval or rendering workflow; do not treat an empty parse as proof that the site has no table.
Or skip the browser setup
If your task is to capture a rendered page for inspection rather than turn its table directly into a DataFrame, ScreenshotNeo provides a screenshot API and MCP server. Its API accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. For example, save a screenshot of a rendered page:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request parameters. ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client. The free plan includes 1,000 shots a month without a card; paid plans start at $5 for 3,000 shots. This is for visual capture, not a replacement for extracting table cells into structured Python data. Learn about ScreenshotNeo, or sign up free for 1,000 screenshots a month, with no card required.
Troubleshoot common extraction problems
No tables returned or no table matched
Check that the input is HTML containing a table and that your match text or attrs value occurs on the intended table. Print the number of tables without a filter first, then inspect their headers and first rows.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Parser dependency or parsing errors
Confirm that the parser package your environment needs is installed. Try an explicitly installed parser backend; if the markup is malformed, compare output using a more lenient parser such as html5lib. A different parser can produce a different tree, so verify the result rather than only checking that an exception disappeared.
Unexpected headers, columns, or missing values
Inspect the table’s header rows and any title or notes embedded in the table. Adjust pandas’ header or skipped-row settings if needed, and check whether cells spanning rows or columns changed the shape. Validate column names and sample values after applying the change.
Best Value
Numbers parse incorrectly
Determine the source’s thousands and decimal conventions, then pass the matching thousands and decimal settings or a column converter. Check missing values and representative numeric cells after conversion.
Beautiful Soup finds the wrong table or loses link destinations
Narrow the lookup with an appropriate table attribute or surrounding structure. Use find_all deliberately when nested tables are present, and retrieve each relevant anchor’s href; get_text() returns text, not link targets.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFAQ
Does pandas return one DataFrame or several?
read_html returns a list of DataFrames, including when the HTML contains only one table. Inspect the list and select the intended entry.
Can Beautiful Soup create a DataFrame?
Yes. Its traversal gives you values to organize, but the row and column structure is your code to define. Use pandas directly when ordinary table-to-DataFrame conversion is sufficient.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

