October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

DuckDB vs pandas: How One SQL Switch Can Speed Up Analytics

DuckDB can run SQL against pandas DataFrames and query Parquet directly, but a 10× speedup depends on workload. Here’s how to decide and benchmark fairly.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DuckDB can make some pandas analytics workloads much faster, but a 10× speedup is not guaranteed by switching libraries. It is most promising when a script performs large aggregations, joins, or scans of columnar files: DuckDB can run SQL directly against a pandas DataFrame or read supported files such as Parquet without first building a full pandas copy. Whether it wins depends on the query, data layout, memory, and the cost of loading and converting results.

What changes when you use DuckDB with pandas?

DuckDB is an analytical database engine with a Python package. Instead of expressing every operation through pandas methods, you can issue SQL against a DataFrame that already exists in Python. For example:

As an Amazon Associate I earn from qualifying purchases.

pip install duckdb
import duckdb

result = duckdb.query("SELECT sum(a) FROM mydf").to_df()

Here, mydf is a pandas DataFrame variable. DuckDB’s replacement-scan behavior resolves that name, reads the DataFrame’s columns and types, and returns the result as another DataFrame. You do not need to import it into a separate database table first. This is useful for SQL-shaped work, but it does not translate arbitrary pandas syntax or make every pandas operation interchangeable. See DuckDB’s SQL on pandas example.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When can the switch pay off?

Large aggregations and joins

Grouping, joining, sorting, and window calculations over substantial datasets are natural analytical workloads for DuckDB. It can parallelize work across threads, while pandas operations may have different performance and memory characteristics depending on the specific operation. A change is most plausible when profiling shows that such a step dominates your script—not when the time is mostly spent elsewhere.

Queries against Parquet files

DuckDB can query Parquet directly, so you can avoid first loading every column into a pandas DataFrame. When a query needs only a few columns, projection can reduce the data read. Filtering, row-group layout, and the number of times you reuse the data also matter. DuckDB’s file-format guide reports that, in its own TPC-H microbenchmark, queries on Parquet files ran approximately 1.1–5.0× slower than on a DuckDB database. That figure compares Parquet scans with DuckDB database storage; it is not a DuckDB-versus-pandas benchmark. The guide recommends loading data first when storage is available and the workload is join-heavy or repeatedly queries the same data. See the file-format performance guide.

Work that strains available memory

DuckDB can spill some grouping, joining, sorting, and windowing work to disk, which can help with datasets that do not fit comfortably in memory. It is not a guarantee against out-of-memory errors: queries with multiple blocking operators can still exceed available memory, and aggregates such as list() and string_agg() do not support disk offload. Scratch storage and the temporary-directory configuration can matter for these workloads. Check the workload tuning guide before relying on spill behavior.

What the published speed claims actually show

DuckDB’s 2021 comparison used the TPC-H lineitem and orders tables, around 1 GB of uncompressed CSV data at scale factor 1, in Google Colab. It tested selected aggregations and a join, and compared DuckDB configured for one and two threads because that environment supported two. The article also considered direct Parquet querying against reading Parquet into pandas. These are useful examples of potential gains, not a current universal benchmark for all pandas scripts. Dataset size and format, query shape, thread count, software versions, memory, and whether loading and output conversion are timed can change the outcome. DuckDB describes the comparison and its boundaries in its 2021 article.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A later DuckDB benchmark-history article makes another important distinction: it reports raw query performance separately from import/export performance across pandas, Arrow, and Parquet. Its replacement-scan benchmark reads one column from a 100-million-row dataset at the 5 GB scale and calculates a single aggregate, focusing on scan speed rather than aggregation or output conversion. A fast central query does not necessarily mean the whole script is faster by the same factor. See DuckDB’s benchmark history.

Rank #3
Thank You Data Analyst Humor Gift for Data Scientists Analysts, Office Décor for Business Intelligence Experts, Analytics Professional Appreciation Gift, Office Pencil Holder Desk for Desk SD278
  • Perfect Gift for Data Analysts – A fun and unique desk sign for business intelligence experts, data scientists, and analytics professionals.
  • Bold & Readable Design – High-contrast lettering ensures visibility on any desk, making it an instant conversation starter.
  • Compact & Lightweight – Small enough to fit any workspace without taking up too much room but big enough to make an impact.
  • Durable & Long-Lasting Material – Made with premium materials to withstand daily office use while maintaining its sleek look.
  • Great for Any Occasion – Ideal for birthdays, work anniversaries, promotions, or just a fun appreciation gift for number crunchers

How to test whether your script gets faster

Benchmark the operation you actually need and compare equivalent results. A query-only timing can be useful for diagnosing engine speed, but it cannot stand in for end-to-end runtime if one path has additional reading, conversion, or export work.

  1. Choose a representative input and output. Use the same dataset, filters, columns, and expected result for both implementations. Confirm results are equivalent before comparing times.
  2. Record the setup. Note data dimensions and format, CPU and memory, Python, pandas, and DuckDB versions, DuckDB thread count, and whether the cache is warm or cold.
  3. Time the stages separately. Measure file reading, conversion, the central query or transformation, and output conversion. Also report total end-to-end runtime and peak memory.
  4. Repeat runs and state the statistic. Run each path more than once under comparable conditions, and say whether the reported result is a median, mean, or another statistic.
  5. Inspect a disappointing query. Use EXPLAIN to inspect the plan and EXPLAIN ANALYZE to profile it. DuckDB notes that multithreaded step CPU-time totals can exceed total wall-clock time, so do not interpret those totals as elapsed time. Details are in the tuning guide.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When pandas may remain the better choice

A library switch is not automatically a performance improvement. Simple vectorized transformations, small datasets, and code that depends heavily on pandas APIs may not benefit enough to justify rewriting. A 2025 academic evaluation of single-machine dataframe libraries found pandas consistently best for small datasets in that study; its abstract does not establish a universal DuckDB-versus-pandas ranking. Tool choice also depends on whether the data fits in RAM, available hardware, and the workload.

DuckDB is designed for larger, less frequent analytical queries rather than many small concurrent queries. More threads are not always faster, either; the tuning guide advises limiting threads when appropriate. Consider the change when the work is SQL-shaped and measured, and account for both the performance difference and migration effort.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.