Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use pandas’ read_html() to turn a Wikipedia page’s HTML tables into pandas DataFrames. It returns a list—not a single DataFrame—so inspect the results and deliberately select the table you need. For repeatable work, clean and validate the selected table, and consider MediaWiki’s REST API when the data is available there and rendered page markup is too unstable.

Load a Wikipedia page’s tables with pandas

Install pandas and an HTML parser before running the example. For instance, in a terminal:

python -m pip install pandas lxml

Then save this as a Python script or run it in a notebook. Replace the example URL with the Wikipedia article you want to inspect.

import pandas as pd

url = "https://en.wikipedia.org/wiki/List_of..."
tables = pd.read_html(url)

print(f"Found {len(tables)} tables")
for i, table in enumerate(tables):
    print(f"nTable {i}: {table.shape}")
    print(table.head())
    print("Columns:", table.columns.tolist())

read_html() searches the page’s HTML <table> elements and parses their rows and cells. Its return value is a list of DataFrames, even if only one table is found. The official pandas API reference documents that behavior, and the pandas I/O guide shows URL-based HTML-table reads.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does pd.read_html() return a list?

A web page can contain multiple tables, and pandas cannot know which one represents your target data. The list preserves the parsed tables so you can inspect and select the right result. tables[0] means “the first table in page order,” not “the table I intended.” Check each candidate’s preview, shape, and column labels before analysis.

Select the intended Wikipedia table

For a page with several tables, filter by visible table text with match, by a valid HTML attribute with attrs, or both. This example asks for tables whose text matches “Population” and whose class is wikitable:

tables = pd.read_html(
    url,
    match="Population",
    attrs={"class": "wikitable"},
    header=0,
)

print(f"Matched {len(tables)} table(s)")
for i, table in enumerate(tables):
    print(i, table.shape, table.columns.tolist())
    print(table.head())

if not tables:
    raise ValueError("No table matched; inspect the page and adjust match or attrs")

df = tables[0]

The filters narrow the candidates; still inspect the returned list. A match can occur in more than one table, and a page’s markup or visible wording may change. Use an attribute only if it is actually present on the target table. The pandas reference describes match and attrs as selectors for table text and valid HTML attributes.

Inspect before assigning meaning

  • Print len(tables) to see how many tables were parsed.
  • For each candidate, inspect tables[i].head(), tables[i].shape, and tables[i].columns.
  • Check whether the first row is data or a header, and whether labels are duplicated or hierarchical.
  • Confirm that the expected records and fields are present before using a DataFrame in calculations.

Configure parsing and clean the DataFrame

Wikipedia tables are presentation markup, not a guaranteed data schema. Header rows can span columns, footnotes can be attached to values, and cells can contain links or missing-value markers. Parse with the relevant options, then inspect and clean the result rather than assuming the first parse is analysis-ready.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set headers and skipped rows deliberately

header identifies the row or rows to use as column labels; skiprows skips rows when a table has introductory or otherwise unwanted rows. If the column labels look like data, are unexpectedly blank, or contain multiple levels, inspect the table’s header rows and spans, then adjust these settings. index_col can designate a column as the index when that fits the data you need.

tables = pd.read_html(url, match="Population", header=0, skiprows=None)
df = tables[0]
print(df.columns)
print(df.head())

This is a starting point, not a universal header recipe: the correct row depends on the particular table. If headers span rows or columns, a parsed multi-level column index may be appropriate; normalize it only after checking what each level means.

Normalize column labels

Once you have inspected the parsed labels, standardize them for downstream code. For example, flatten a multi-level header into readable strings and trim incidental whitespace:

if isinstance(df.columns, pd.MultiIndex):
    df.columns = [
        " ".join(str(part) for part in column if str(part) != "nan").strip()
        for column in df.columns
    ]
else:
    df.columns = [str(column).strip() for column in df.columns]

print(df.columns.tolist())

Review the resulting names for duplicates or lost distinctions before referring to columns by name. Do not flatten a multi-level header blindly if the levels encode information your analysis needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert numeric values and dates

Footnote markers and separators can leave numeric-looking cells as text. Remove only formatting you have verified, then use pd.to_numeric() with errors="coerce" when invalid values should become missing values rather than stop the conversion:

column = "Population"
cleaned = df[column].astype("string").str.replace(",", "", regex=False)
cleaned = cleaned.str.replace(r"[.*?]", "", regex=True).str.strip()
df[column] = pd.to_numeric(cleaned, errors="coerce")

That example removes commas and bracketed footnote text; adapt it to the actual cells. Check the results for newly missing values so unexpected formats do not silently disappear from your analysis.

For dates, inspect the page’s displayed format first. You can use parse_dates or a column-specific converters function, both documented by pandas, when the source representation is clear. Avoid assuming that ambiguous date strings use a particular day-month order.

Control separators and missing values

Options such as thousands, decimal, na_values, and keep_default_na let you describe how the source represents numbers and missing data. For example, set na_values to the source’s actual missing-value marker; change keep_default_na only when you have a reason to alter pandas’ default missing-value handling. Confirm the parsed types and null counts after reading:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
print(df.dtypes)
print(df.isna().sum())

Keep hyperlinks when needed

By default, the table is parsed as data rather than as a preservation of every link in its presentation. Use extract_links="all" if you need link information from table cells, then inspect the resulting values because the representation differs from plain text:

linked_tables = pd.read_html(url, match="Population", extract_links="all")
linked_df = linked_tables[0]
print(linked_df.head())

Useful read_html() controls

Option What it controls When to consider it
match Filters tables by matching text When the page has many tables and the desired table has distinctive visible wording
attrs Targets valid HTML table attributes When inspection confirms a useful attribute such as a class or id
header, skiprows Which rows become headers and which are skipped When labels appear on an unexpected row or introductory rows interfere
index_col Which column becomes the DataFrame index When a column is the intended row identifier
parse_dates, converters Date parsing or custom per-column conversion When source formats are understood and need explicit interpretation
thousands, decimal Numeric separators When the table’s number formatting differs from defaults
na_values, keep_default_na How missing values are recognized When the source has known missing-value markers or defaults need adjustment
displayed_only Whether to consider only displayed table content When hidden markup affects what should be parsed
extract_links Whether links are extracted from cells When linked labels or destinations are part of the data you need
flavor Which supported HTML parser is used When choosing among installed parser options or diagnosing parser differences

These options and their accepted values are documented in the pandas API reference. Use only the controls relevant to the actual table, and inspect the resulting DataFrame after changes.

Make the workflow repeatable and auditable

Wikipedia page markup can change, so a script that works today may select a different table or parse different headers later. Keep the source URL and retrieval time with the output, and validate the columns and basic shape expected by your analysis.

from datetime import datetime, timezone

retrieved_at = datetime.now(timezone.utc).isoformat()
expected_columns = {"Country", "Population"}
actual_columns = set(df.columns)
missing = expected_columns - actual_columns
if missing:
    raise ValueError(f"Expected columns missing: {sorted(missing)}")

metadata = {
    "source_url": url,
    "retrieved_at_utc": retrieved_at,
}
df.to_csv("wikipedia_table.csv", index=False)
print(metadata)

Store the metadata alongside the saved data in your own pipeline, for example in a JSON file or a database record. A successful HTTP read does not prove that you captured the intended table; assertions make structural changes visible instead of letting them quietly flow into later calculations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose between read_html(), targeted parsing, and the API

For an ordinary rendered table, read_html() is the quickest route from a page URL to DataFrames. If a page has complex markup or you need precise control beyond table extraction, targeted parsing may be more suitable but requires you to handle the page structure. When the data is available as structured Wikimedia data, or the rendered HTML is unstable for your workflow, evaluate the official MediaWiki REST API instead. An API can avoid dependence on presentation markup, but first confirm that it exposes the specific data and fields you need.

Parser choice is another practical trade-off. pandas documents support for lxml, html5lib, and bs4; parser availability and behavior depend on your installed environment. The pandas HTML parsing guidance covers setup and parser gotchas. Try an installed supported flavor when parsing fails, and validate the output because switching parsers does not resolve ambiguity in the page’s table structure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Too many tables or the wrong table

Cause: The page contains several HTML tables, and taking the first DataFrame assumes page order matches your intent. Fix: Narrow candidates with match or a verified attrs value, then print each candidate’s preview, columns, and shape before selecting it.

Parser or dependency error

Cause: The selected parser may not be installed, or the HTML may expose a parser-specific issue. Fix: Install the parser you intend to use and try another supported flavor—lxml, html5lib, or bs4—as appropriate. Consult pandas’ HTML parsing guidance for setup details, then check that the parsed rows and columns remain correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unexpected columns, blank labels, or NaN headers

Cause: The table may have multiple header rows, spans, or markup that does not map cleanly to simple column names. Fix: Inspect the top rows and parsed column index; adjust header or skiprows to match the table, and normalize labels only after establishing what they mean.

Numbers or dates remain text

Cause: Footnotes, thousands separators, decimal conventions, or ambiguous date formats can prevent the values from parsing as intended. Fix: Inspect representative raw cells, use relevant parsing options or a converter, and verify converted values and missing-value counts before analysis.

The page changes and the result no longer fits

Cause: A rendered table’s labels or markup may have changed. Fix: Keep schema checks in the pipeline and re-inspect the page. If the needed structured data is available through MediaWiki’s REST API, evaluate it rather than relying on presentation HTML.

Or skip the browser setup

If you need a screenshot or PDF of a Wikipedia page rather than its table as structured data, ScreenshotNeo provides a one-request screenshot API. This does not replace pandas when your goal is a DataFrame; it captures a visual page output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL example, with the target URL set to a Wikipedia page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://en.wikipedia.org/wiki/List_of... -o shot.webp

See the ScreenshotNeo documentation for API details. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for free.

Frequently Asked Questions

Can I use `read_html()` directly on a Wikipedia URL?

Yes. The examples use a page URL as the input. Inspect the returned DataFrames and confirm the intended table before relying on the result.

Does `read_html()` preserve links from table cells?

Use `extract_links=”all”` when link information matters, then inspect the resulting values because they are not plain-text cells.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use a Wikipedia table or the MediaWiki REST API?

Use `read_html()` for a straightforward rendered table. Evaluate the REST API when the required data is available in structured form or page markup is too unstable for your workflow.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.