Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk7 min

Pandas Python Tutorial: Load, Clean, and Analyze Data

A practical introduction to pandas in Python, from Series and DataFrames to loading, inspecting, cleaning, selecting, grouping, and saving data.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pandas is a Python library for working with structured data: it helps you load tables, inspect and clean them, select records, calculate summaries, combine datasets, and save results. Its two core structures are the one-dimensional Series and the two-dimensional, labeled DataFrame. This guide builds a small example and follows the common steps of a data-analysis workflow.

What is pandas in Python?

Pandas is an open-source library for data analysis and manipulation. It is designed for tabular and heterogeneous data: a DataFrame can have named columns containing different kinds of values, such as dates, text, and numbers. That differs from a typical NumPy array, which is generally organized around a single data type. [O’Reilly sample chapter]

As an Amazon Associate I earn from qualifying purchases.

Series and DataFrame

  • Series is a one-dimensional labeled sequence of values, such as one column of measurements.
  • DataFrame is a two-dimensional table with labeled rows and columns. Each column is a Series, and columns can use different data types.

Labels matter: rows have an index, and columns have names. Pandas operations often use those labels to align values when selecting, combining, or calculating.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I install and import pandas?

Install pandas into the same Python environment that will run your script or notebook. The standard options include pip and Conda; exact compatibility can depend on your Python environment, so use the installation guidance for your chosen package manager if an install fails. The current release and compatibility requirements are listed in the official pandas installation documentation.

Using pip

In a terminal, with the intended environment activated, run:

python -m pip install pandas

On systems where the Python 3 command is python3, use python3 -m pip install pandas. In a notebook, a package installed into a different Python environment may not be available to its kernel.

Using Conda

With the desired Conda environment active, install pandas using Conda’s package command:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
conda install pandas

After installation, import the library. The conventional alias pd keeps examples concise:

import pandas as pd

How do I create a DataFrame and read a CSV?

You can create a DataFrame directly from Python data. Here, each dictionary key becomes a column name, and values at matching positions form a row.

import pandas as pd

data = {
    "name": ["Ari", "Bo", "Chen"],
    "team": ["North", "South", "North"],
    "sales": [12, 9, 15],
}
df = pd.DataFrame(data)
print(df)

For a CSV file, use read_csv and pass the file path:

df = pd.read_csv("sales.csv")

If the file is not in the script’s current working directory, provide the correct relative or absolute path. Pandas also has I/O functions for formats such as Excel and can work with SQL databases and URLs; particular formats or database connections may require additional packages or drivers. See the pandas I/O guide for format-specific options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I inspect a DataFrame?

Inspect data before transforming it. These methods answer different questions:

  • df.head() shows the first five rows by default; pass a number to request a different count.
  • df.tail() shows the last five rows by default.
  • df.shape returns a pair: number of rows, then number of columns.
  • df.info() prints column names, non-null counts, and data types.
  • df.describe() summarizes numeric columns by default, including count, mean, and quartiles.
print(df.head())
print(df.shape)
df.info()
print(df.describe())

Use info() to spot columns with missing values or unexpected types; use describe() to notice implausible ranges or distributions. These summaries are clues, not proof that data is valid.

How do I select columns and rows with loc and iloc?

Use a column name to select a column, or a list of names to select several. In the example below, df["sales"] returns a Series and df[["name", "sales"]] returns a DataFrame.

sales = df["sales"]
name_and_sales = df[["name", "sales"]]

Choose by label with loc

loc selects by row and column labels. With the default index, the row labels are 0, 1, and 2:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
first_row = df.loc[0]
selected = df.loc[0:1, ["name", "sales"]]

Label slices with loc include the stop label when it is present. If you assign a custom index, use those index labels rather than assuming row numbers.

Choose by position with iloc

iloc selects by zero-based integer position. Its slice stop is excluded, following normal Python slicing:

first_row = df.iloc[0]
first_two_rows = df.iloc[0:2, [0, 2]]

In the second example, 0:2 means positions 0 and 1, while column positions 0 and 2 correspond to name and sales.

Filter rows with a condition

Boolean filtering keeps rows whose condition evaluates to true. This selects records with sales greater than 10:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
high_sales = df[df["sales"] > 10]

For multiple conditions, put each comparison in parentheses and combine them with & for AND or | for OR. Use these operators rather than Python’s and and or for element-wise pandas conditions.

How do I handle missing values?

Missing data can affect calculations and conclusions. First inspect where values are absent:

df.isna().sum()

Choose a treatment based on what the field means and how much data is missing. Dropping rows removes records; filling values keeps them but introduces a replacement. Neither choice is automatically correct.

Drop incomplete rows

Use dropna when rows with missing values should not be included in the analysis:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
complete_rows = df.dropna()

By default this drops rows containing at least one missing value. A dataset with many columns can therefore lose more records than expected.

Fill missing values

Use fillna when a defensible replacement exists. For example, filling a missing numeric measurement with zero asserts that the value was zero, which may not be true:

df["sales"] = df["sales"].fillna(0)

For other data, a median, category such as "Unknown", or a method based on surrounding observations may be more appropriate. Document the choice, especially when it changes the meaning of the data.

How do I group, combine, and reshape data?

Summarize groups

groupby splits rows into groups, applies a calculation, and combines the results. For the sample table, this calculates total sales by team:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
totals = df.groupby("team")["sales"].sum()
print(totals)

Use multiple grouping columns when the question calls for summaries by more than one category. The result is typically indexed by the grouping labels.

Concatenate or merge tables

Concatenation stacks or places objects alongside one another; merging matches records using one or more key columns. These operations answer different questions.

# Stack rows with the same columns
all_sales = pd.concat([january, february], ignore_index=True)

# Match records that share a customer_id
combined = orders.merge(customers, on="customer_id", how="left")

A merge can change the number of rows: duplicate keys on either side may produce multiple matches. Check key uniqueness and inspect the output when row counts matter. For merge options and join behavior, consult the pandas merging guide.

Pivot data for a summary view

A pivot reorganizes values into a table keyed by categories. For example, if a table has team, month, and sales columns, this creates a team-by-month summary using summed sales:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
summary = df.pivot_table(
    index="team",
    columns="month",
    values="sales",
    aggfunc="sum",
)

A pivot table aggregates duplicate combinations according to aggfunc; a plain pivot expects each index/column combination to identify a single value. Choose based on whether the source has repeated combinations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do I save results, work with dates, or plot data?

Write a result to a file

To save a DataFrame as CSV without writing its index as an extra column, use:

df.to_csv("cleaned_sales.csv", index=False)

For Excel, use to_excel; support can depend on an installed Excel-writing engine. Check the output path and open the file or read it back to confirm that the intended columns and rows were written.

Parse and analyze dates

When a CSV column contains dates, ask read_csv to parse it, then use the datetime accessors or a datetime index for date-based operations:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
df = pd.read_csv("sales.csv", parse_dates=["date"])
monthly = df.set_index("date").resample("ME")["sales"].sum()

Date parsing depends on the input format and pandas version. If values are ambiguous or parsing fails, inspect the original strings and specify a format appropriate to the file. See the time series guide.

Make a quick plot

Pandas provides plotting methods that use Matplotlib. A simple line plot of a numeric column can be created with:

df["sales"].plot()

For a chart with labeled axes, titles, or more control, use Matplotlib directly. Plotting is most useful after checking that the selected data and ordering represent the question you want to answer.

What should I check when pandas code misbehaves?

  • Unexpected row count: Check filters, missing-value drops, and whether a merge key appears more than once.
  • Unexpected values after selection: Confirm whether you intended labels (loc) or positions (iloc), and verify the index.
  • Wrong calculations: Inspect data types, missing values, and whether numeric-looking text was parsed as text.
  • Slow work on a large table: Avoid row-by-row Python loops when a column operation, boolean filter, or groupby expresses the same calculation. Read only needed columns when loading files and consider appropriate data types.
  • Import or installation errors: Verify that pandas was installed in the same environment as the interpreter or notebook kernel, and consult the official installation instructions for supported configurations.

Where can I learn more?

The official pandas documentation is the best place to check release-specific APIs and detailed behavior. Python Guides also publishes a free pandas course covering installation and common operations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a longer book-based course, Wes McKinney’s Python for Data Analysis, 3rd Edition covers data loading, cleaning, merging, groupby, visualization, and time series. O’Reilly says this edition is updated for Python 3.10 and pandas 1.4, so treat it as a structured learning resource rather than a reference for current pandas releases. [O’Reilly book page]

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.