Recommended Free Tools
Pandas is a Python library for working with structured data: it helps you load tables, inspect and clean them, select records, calculate summaries, combine datasets, and save results. Its two core structures are the one-dimensional Series and the two-dimensional, labeled DataFrame. This guide builds a small example and follows the common steps of a data-analysis workflow.
What is pandas in Python?
Pandas is an open-source library for data analysis and manipulation. It is designed for tabular and heterogeneous data: a DataFrame can have named columns containing different kinds of values, such as dates, text, and numbers. That differs from a typical NumPy array, which is generally organized around a single data type. [O’Reilly sample chapter]
As an Amazon Associate I earn from qualifying purchases.
Series and DataFrame
Seriesis a one-dimensional labeled sequence of values, such as one column of measurements.DataFrameis a two-dimensional table with labeled rows and columns. Each column is a Series, and columns can use different data types.
Labels matter: rows have an index, and columns have names. Pandas operations often use those labels to align values when selecting, combining, or calculating.
Free tools Windows power users keep installed
One-click scans. No signup required.
How do I install and import pandas?
Install pandas into the same Python environment that will run your script or notebook. The standard options include pip and Conda; exact compatibility can depend on your Python environment, so use the installation guidance for your chosen package manager if an install fails. The current release and compatibility requirements are listed in the official pandas installation documentation.
#1 Best Overall
Using pip
In a terminal, with the intended environment activated, run:
python -m pip install pandas
On systems where the Python 3 command is python3, use python3 -m pip install pandas. In a notebook, a package installed into a different Python environment may not be available to its kernel.
Using Conda
With the desired Conda environment active, install pandas using Conda’s package command:
conda install pandas
After installation, import the library. The conventional alias pd keeps examples concise:
import pandas as pd
How do I create a DataFrame and read a CSV?
You can create a DataFrame directly from Python data. Here, each dictionary key becomes a column name, and values at matching positions form a row.
import pandas as pd
data = {
"name": ["Ari", "Bo", "Chen"],
"team": ["North", "South", "North"],
"sales": [12, 9, 15],
}
df = pd.DataFrame(data)
print(df)
For a CSV file, use read_csv and pass the file path:
df = pd.read_csv("sales.csv")
If the file is not in the script’s current working directory, provide the correct relative or absolute path. Pandas also has I/O functions for formats such as Excel and can work with SQL databases and URLs; particular formats or database connections may require additional packages or drivers. See the pandas I/O guide for format-specific options.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
How do I inspect a DataFrame?
Inspect data before transforming it. These methods answer different questions:
df.head()shows the first five rows by default; pass a number to request a different count.df.tail()shows the last five rows by default.df.shapereturns a pair: number of rows, then number of columns.df.info()prints column names, non-null counts, and data types.df.describe()summarizes numeric columns by default, including count, mean, and quartiles.
print(df.head())
print(df.shape)
df.info()
print(df.describe())
Use info() to spot columns with missing values or unexpected types; use describe() to notice implausible ranges or distributions. These summaries are clues, not proof that data is valid.
How do I select columns and rows with loc and iloc?
Use a column name to select a column, or a list of names to select several. In the example below, df["sales"] returns a Series and df[["name", "sales"]] returns a DataFrame.
sales = df["sales"]
name_and_sales = df[["name", "sales"]]
Choose by label with loc
loc selects by row and column labels. With the default index, the row labels are 0, 1, and 2:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
first_row = df.loc[0]
selected = df.loc[0:1, ["name", "sales"]]
Label slices with loc include the stop label when it is present. If you assign a custom index, use those index labels rather than assuming row numbers.
Choose by position with iloc
iloc selects by zero-based integer position. Its slice stop is excluded, following normal Python slicing:
first_row = df.iloc[0]
first_two_rows = df.iloc[0:2, [0, 2]]
In the second example, 0:2 means positions 0 and 1, while column positions 0 and 2 correspond to name and sales.
Filter rows with a condition
Boolean filtering keeps rows whose condition evaluates to true. This selects records with sales greater than 10:
high_sales = df[df["sales"] > 10]
For multiple conditions, put each comparison in parentheses and combine them with & for AND or | for OR. Use these operators rather than Python’s and and or for element-wise pandas conditions.
How do I handle missing values?
Missing data can affect calculations and conclusions. First inspect where values are absent:
df.isna().sum()
Choose a treatment based on what the field means and how much data is missing. Dropping rows removes records; filling values keeps them but introduces a replacement. Neither choice is automatically correct.
Drop incomplete rows
Use dropna when rows with missing values should not be included in the analysis:
complete_rows = df.dropna()
By default this drops rows containing at least one missing value. A dataset with many columns can therefore lose more records than expected.
Fill missing values
Use fillna when a defensible replacement exists. For example, filling a missing numeric measurement with zero asserts that the value was zero, which may not be true:
df["sales"] = df["sales"].fillna(0)
For other data, a median, category such as "Unknown", or a method based on surrounding observations may be more appropriate. Document the choice, especially when it changes the meaning of the data.
How do I group, combine, and reshape data?
Summarize groups
groupby splits rows into groups, applies a calculation, and combines the results. For the sample table, this calculates total sales by team:
totals = df.groupby("team")["sales"].sum()
print(totals)
Use multiple grouping columns when the question calls for summaries by more than one category. The result is typically indexed by the grouping labels.
Concatenate or merge tables
Concatenation stacks or places objects alongside one another; merging matches records using one or more key columns. These operations answer different questions.
# Stack rows with the same columns
all_sales = pd.concat([january, february], ignore_index=True)
# Match records that share a customer_id
combined = orders.merge(customers, on="customer_id", how="left")
A merge can change the number of rows: duplicate keys on either side may produce multiple matches. Check key uniqueness and inspect the output when row counts matter. For merge options and join behavior, consult the pandas merging guide.
Pivot data for a summary view
A pivot reorganizes values into a table keyed by categories. For example, if a table has team, month, and sales columns, this creates a team-by-month summary using summed sales:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →summary = df.pivot_table(
index="team",
columns="month",
values="sales",
aggfunc="sum",
)
A pivot table aggregates duplicate combinations according to aggfunc; a plain pivot expects each index/column combination to identify a single value. Choose based on whether the source has repeated combinations.
Best Value
How do I save results, work with dates, or plot data?
Write a result to a file
To save a DataFrame as CSV without writing its index as an extra column, use:
df.to_csv("cleaned_sales.csv", index=False)
For Excel, use to_excel; support can depend on an installed Excel-writing engine. Check the output path and open the file or read it back to confirm that the intended columns and rows were written.
Parse and analyze dates
When a CSV column contains dates, ask read_csv to parse it, then use the datetime accessors or a datetime index for date-based operations:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minutedf = pd.read_csv("sales.csv", parse_dates=["date"])
monthly = df.set_index("date").resample("ME")["sales"].sum()
Date parsing depends on the input format and pandas version. If values are ambiguous or parsing fails, inspect the original strings and specify a format appropriate to the file. See the time series guide.
Make a quick plot
Pandas provides plotting methods that use Matplotlib. A simple line plot of a numeric column can be created with:
df["sales"].plot()
For a chart with labeled axes, titles, or more control, use Matplotlib directly. Plotting is most useful after checking that the selected data and ordering represent the question you want to answer.
What should I check when pandas code misbehaves?
- Unexpected row count: Check filters, missing-value drops, and whether a merge key appears more than once.
- Unexpected values after selection: Confirm whether you intended labels (
loc) or positions (iloc), and verify the index. - Wrong calculations: Inspect data types, missing values, and whether numeric-looking text was parsed as text.
- Slow work on a large table: Avoid row-by-row Python loops when a column operation, boolean filter, or groupby expresses the same calculation. Read only needed columns when loading files and consider appropriate data types.
- Import or installation errors: Verify that pandas was installed in the same environment as the interpreter or notebook kernel, and consult the official installation instructions for supported configurations.
Where can I learn more?
The official pandas documentation is the best place to check release-specific APIs and detailed behavior. Python Guides also publishes a free pandas course covering installation and common operations.
For a longer book-based course, Wes McKinney’s Python for Data Analysis, 3rd Edition covers data loading, cleaning, merging, groupby, visualization, and time series. O’Reilly says this edition is updated for Python 3.10 and pandas 1.4, so treat it as a structured learning resource rather than a reference for current pandas releases. [O’Reilly book page]
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




