Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pandas and NumPy solve different data problems: pandas adds labels, alignment, and tabular operations, while NumPy provides direct operations on homogeneous multidimensional arrays. Reliable advanced work depends on understanding which object you are selecting, whether an operation returns a view or a copy, and whether a result still aligns with the original rows.

This guide uses APIs documented for pandas 3.0.6 and the NumPy 2.3 stable manual. It covers label-versus-position selection, NumPy indexing semantics, hierarchical indexes, and row-aligned group transformations.

Choose pandas or NumPy by data model

Pandas is built around labeled Series and DataFrame objects. Labels identify rows and columns, participate in alignment during assignment and arithmetic, and make heterogeneous tabular data practical. NumPy is centered on homogeneous, multidimensional arrays whose operations are primarily governed by shape, dtype, and position.

Wes McKinney describes the distinction this way: “While pandas adopts many coding idioms from NumPy, the biggest difference is that pandas is designed for working with tabular or heterogeneous data. NumPy, by contrast, is best suited for working with homogeneously typed numerical array data.” (The quotation appears in the publisher-hosted sample of the third edition of Python for Data Analysis.)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Question Pandas NumPy
Primary data model Labeled Series and DataFrame tables Homogeneous numerical arrays
Selection basis Labels or positions, chosen explicitly Positions, slices, masks, and index arrays
Combining objects Indexes align labels automatically Shapes and broadcasting determine compatibility
Typical strength Filtering, joining, missing data, grouping, and reshaping tables Numerical array computation and multidimensional math

Use pandas when the identity of a row or column matters. Use NumPy when the data are naturally a uniform numeric block or when an array-oriented algorithm is the clearest expression. Converting between them is straightforward, but converting a labeled table to an array discards index and column labels unless you preserve them separately.

Use .loc for labels and .iloc for positions

The central pandas selection distinction is simple: .loc interprets keys as labels, while .iloc interprets them as zero-based integer positions. The difference remains important when an index contains integers that do not describe row positions.

A label is not necessarily a row number

import pandas as pd

sales = pd.DataFrame(
    {"item": ["keyboard", "mouse", "monitor"], "units": [12, 30, 7]},
    index=[101, 205, 410],
)

sales.loc[205, "item"]   # "mouse": label 205
sales.iloc[1, 0]          # "mouse": second row, first column

Here, label 205 happens to identify the second row, but that is a coincidence. sales.loc[1] asks for a label of 1 and raises KeyError because that label is absent. sales.iloc[1] asks for the second row and succeeds.

Slices and missing labels

Label selection follows the index’s labels; positional selection follows Python-style integer positions. A label-based slice can include both endpoint labels when the index supports the operation, whereas positional slicing uses the usual half-open convention and excludes the stop position. Test the exact index type and slice you use rather than assuming labels are consecutive integers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
sales.loc[101:410, "units"]  # label-based range
sales.iloc[0:2, 1]             # rows at positions 0 and 1

A scalar label lookup with an absent label raises KeyError. Positional access outside the valid range raises IndexError. For a condition, use a boolean mask and select with .loc when you want the resulting labels retained:

large_orders = sales.loc[sales["units"] >= 10, ["item", "units"]]

Alignment changes assignment results

Pandas aligns a Series by index labels during assignment and many arithmetic operations. That is useful when labels represent entities, but surprising if you expected values to be assigned by row order.

scores = pd.Series([90, 80], index=[205, 101])
sales["score"] = scores

The score 90 goes to row label 205 and 80 to label 101; pandas does not simply place them in the Series’ displayed order. Inspect .index before combining objects. If positional assignment is intentional, make that intent explicit with an array after checking that lengths match.

Understand NumPy view and copy behavior

NumPy has two broad indexing categories. Basic slicing generally produces a view into the original array, while advanced indexing—integer arrays or boolean arrays—produces a copy. The distinction controls whether later mutation affects the source and how much additional memory selection may require.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Basic slicing returns a view

import numpy as np

a = np.array([10, 20, 30, 40])
part = a[1:3]
part[0] = 99
# a is now [10, 99, 30, 40]

part refers to the same underlying data for this basic slice. Use part.base or np.shares_memory(a, part) when you need to verify sharing in a particular expression.

Advanced indexing returns a copy

a = np.array([10, 20, 30, 40])
chosen = a[[1, 3]]
chosen[0] = 999
# chosen is [999, 40], but a remains [10, 20, 30, 40]

mask = a % 20 == 0
selected = a[mask]

Integer-array selection and boolean-mask selection create independent arrays. This is useful when you need a safely editable subset, but it also means that modifying the result will not update the source. If you need to write selected values back, assign through the original array with the mask or integer index expression:

a[a % 20 == 0] = -1

Do not infer a universal speed or memory ranking from these rules. Advanced indexing must gather arbitrary elements, and the practical cost depends on array shape, dtype, and workload. Measure a real workload when performance matters.

Use a MultiIndex for hierarchical labels

A pandas MultiIndex stores multiple levels of labels on a one- or two-dimensional object. It represents relationships such as region and quarter, customer and invoice, or sensor and timestamp without forcing the data into a higher-dimensional container. Hierarchical labels support grouped selection, aggregation, and reshaping.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Create and select hierarchical data

regional = pd.DataFrame(
    {"revenue": [120, 135, 98, 110]},
    index=pd.MultiIndex.from_tuples(
        [("West", "Q1"), ("West", "Q2"), ("East", "Q1"), ("East", "Q2")],
        names=["region", "quarter"],
    ),
)

regional.loc["West"]              # all West rows, retaining the inner level
regional.loc[("East", "Q2"), "revenue"]

Because the levels are labels, selection remains meaningful even when rows are not arranged as a rectangular three-dimensional array. You can reshape a hierarchical result with operations such as unstack:

by_quarter = regional["revenue"].unstack("quarter")

The result has regions as its row index and quarters as columns. Reshaping can introduce missing combinations; handle those explicitly rather than assuming every level pair exists.

Sort before repeated hierarchical lookups

Hierarchical selection is most predictable when the index is sorted by its levels. An unsorted MultiIndex can make lookups inefficient and may produce a performance warning. Sort it when you perform repeated partial-key selection:

regional = regional.sort_index()
west = regional.loc["West"]

Sorting changes row order, so do it at a deliberate point in the pipeline and preserve a stable identifier if downstream code depends on the original order. The pandas advanced-indexing guide documents the selection, reshaping, and sortedness behavior.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep row alignment with GroupBy.transform()

Aggregation answers “what is one value per group?” Transformation answers “what value should each original row receive based on its group?” A transform result has the grouped object’s index, so it can be assigned back to the source table without a manual merge.

Center each value on its group mean

orders = pd.DataFrame({
    "team": ["A", "A", "B", "B"],
    "hours": [10, 14, 8, 12],
})

team_mean = orders.groupby("team")["hours"].transform("mean")
orders["hours_from_team_mean"] = orders["hours"] - team_mean

The transformed Series contains four values, in the same index order as orders. Team A receives its two-row mean in both rows, and team B receives its own mean in both rows. Subtracting it is therefore an element-by-element, label-aligned operation.

Standardize within each group

grouped = orders.groupby("team")["hours"]
orders["z_within_team"] = grouped.transform(
    lambda values: (values - values.mean()) / values.std(ddof=0)
)

The lambda is applied separately to each group, but the outputs are reassembled to the original index. This is appropriate for features such as within-customer deviation, relative rank inputs, or group-level imputation features where every source row must remain present.

Aggregation and transform are not interchangeable

Operation Shape of result Use it when
groupby(...).agg(...) Usually one row (or one value) per group You need a summary table for reporting or a later join
groupby(...).transform(...) Same number of rows and index as the input selection Each original row needs a group-derived value

For example, orders.groupby("team")["hours"].mean() returns two group means, not four row-level values. Use transform("mean") when those means must be broadcast to all contributing rows. If you need a reduced result, aggregation is clearer and avoids manufacturing repeated values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A dependable workflow for mixed pandas and NumPy code

  1. Identify the data model. Keep labels and heterogeneous columns in pandas; convert to NumPy only when an array algorithm genuinely needs it.
  2. State the selection coordinate system. Use .loc for labels and .iloc for positions. Do not rely on an integer-looking index to imply row positions.
  3. Inspect indexes before combining objects. Check obj.index, names, and ordering before assignment, arithmetic, joins, or concatenation.
  4. Check mutation semantics. Treat NumPy basic slices as potentially shared views and advanced selections as copies. Make copies explicitly when ownership needs to be clear.
  5. Choose the groupwise shape first. Use aggregation for one result per group and transformation for an output aligned to every source row.
  6. Sort hierarchical indexes for recurring partial-key access. Confirm that the resulting row order is acceptable to later steps.

Further reading

Pandas indexing and selecting data covers labels, positions, and alignment. The NumPy indexing guide details basic and advanced indexing. For hierarchical selection and reshaping, see the pandas MultiIndex guide; for row-aligned group operations, see the pandas groupby guide.

For a broader, worked treatment, Python for Data Analysis, 3rd Edition by Wes McKinney (published August 2022) covers NumPy, pandas, data cleaning, merging, reshaping, and groupby. O’Reilly describes that edition as updated for Python 3.10 and pandas 1.4, so use it for concepts and examples while checking the current pandas 3.x documentation for API details. See the book listing and its publisher-hosted chapter sample.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.