Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk5 min

Best Alternatives to CSV for Large-Scale Data Benchmarks

Parquet, ORC, and Arrow IPC solve different large-data problems. Choose by query mix, storage, conversion cost, and streaming needs—not a universal leaderboard.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no format that wins every large-scale data benchmark. Start with Parquet for compressed on-disk analytics, test ORC when selective scans or a Hadoop-oriented stack matter, and include Arrow IPC/Feather when the workload benefits from processing data in Arrow’s in-memory layout. Keep CSV as a baseline for interoperability, inspection, or sequential streaming. The useful winner is the format that performs best on the queries, software, and hardware your benchmark actually uses.

Which format should you test first?

For a general analytical benchmark, use Parquet as the on-disk starting point, then add ORC and Arrow IPC/Feather if their strengths match the workload. Retain CSV when readers need a straightforward text representation or a sequential stream. A benchmark that compares only whole-file read time can miss the reasons these formats differ: projection, filtering, compression, conversion, and startup behavior all affect results.

As an Amazon Associate I earn from qualifying purchases.

Format Best fit to evaluate Trade-off to account for
Parquet Compressed columnar files for analytical storage and scans. Often smaller than Arrow IPC, but data must be decoded for reading.
ORC Hadoop-oriented workloads and selective scans that can use indexes or predicate pushdown. Benefits depend on reader support, data layout, and query; do not assume it beats Parquet universally.
Arrow IPC / Feather V2 Arrow-aware in-memory processing or interchange where avoiding deserialization and extra copies matters. Files may be larger than Parquet, so storage and network costs can offset faster access.
CSV Interoperability, inspection, and sequential streaming. Text must be scanned and types inferred, adding parsing work and potential ambiguity.

Apache Arrow describes Parquet and Arrow as complementary: “Therefore, Arrow and Parquet complement each other and are commonly used together in applications.” Parquet can be the durable, compact file format while Arrow is used for in-memory processing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the formats behave in a benchmark

Parquet: a practical storage baseline

Parquet is columnar and compressed, making it a strong first candidate when the benchmark stores analytical data on disk and reads selected columns. Its compact files do not make reads free: the reader decodes the stored representation, and the cost depends on the data, codecs, and query.

#1 Best Overall
Sandisk 2TB Extreme Portable SSD, Up to 1050MB/s, USB-C, USB 3.2 Gen 2, IP65 Water and Dust Resistance, Updated Firmware, External Solid State Drive, SDSSDE61-2T00-G25
  • Get NVMe solid state performance with up to 1050MB/s read and 1000MB/s write speeds in a portable, high-capacity drive(1) (Based on internal testing; performance may be lower depending on host device & other factors. 1MB=1,000,000 bytes.)
  • Up to 3-meter drop protection and IP65 water and dust resistance mean this tough drive can take a beating(3) (Previously rated for 2-meter drop protection and IP55 rating. Now qualified for the higher, stated specs.)
  • Use the handy carabiner loop to secure it to your belt loop or backpack for extra peace of mind.
  • Help keep private content private with the included password protection featuring 256‐bit AES hardware encryption.(3)
  • Easily manage files and automatically free up space with the SanDisk Memory Zone app.(5). Non-Operating Temperature -20°C to 85°C

ORC: test it for selective scans

ORC is self-describing and type-aware, and supports indexes and predicate pushdown that can help readers skip stripes or narrow work to row ranges. The ORC documentation describes a default stripe size of roughly 64 MB. This is a format capability, not a guarantee of faster queries: the engine, predicates, and file organization determine whether the reader can exploit it.

Arrow IPC and Feather V2: reduce conversion work

Arrow IPC stores Arrow’s columnar in-memory layout on disk and can be memory-mapped. If the consuming system works natively with Arrow, this can avoid deserialization and extra copies. The trade-off is file size: Arrow IPC may take more space than Parquet. Feather V2 is the Arrow IPC file format under a retained name and API.

Rank #2
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
  • Solid state performance with up to 800MB/s read speeds in a portable drive. (Based on internal testing; performance may be lower depending on host device, interface, usage conditions and other factors. 1MB=1,000,000 bytes.)
  • Back up your content and memories on a storage solution that fits seamlessly into your mobile lifestyle.
  • Take it with you on your adventures—up to two-meter drop protection means this durable drive can take a beating. (Based on internal testing.)
  • Secure it to your belt loop or backpack for extra peace of mind thanks to the tough rubber hook.
  • From Sandisk, a brand professional photographers trust to take on assignments.

Arrow streams and CSV: incremental processing

An Arrow stream places the schema before record batches, so a receiver can process batches as they arrive. CSV can also be consumed sequentially, and remains convenient to inspect and exchange. Unlike typed, self-describing formats, CSV requires parsing text and inferring or separately specifying types.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What published size comparisons do—and do not—show

Microsoft Research’s 2024 paper, A Deep Dive into Common Open Formats for Analytical DBMSs, reports selected real-world column data totals of 489.7 GB in raw CSV, 64.7 GB in Parquet, and 133.9 GB in ORC. In those selected data, Parquet occupied about 13% of the raw CSV total and ORC about 27%. The paper also reports 522.5 GB for Arrow with default settings and 237.4 GB for Arrow with dictionary encoding. These are totals for the paper’s selected data, not general compression ratios; results varied by dataset, column type, and encoding, including integer columns with different distinct-value distributions.

Rank #3
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

A broader study by Chunwei Liu, Anna Pavlenko, Matteo Interlandi, and Brandon Haynes in The VLDB Journal (November 2024) evaluated Arrow, Parquet, and ORC using TPC-DS scale 10, the Join Order Benchmark, the Public BI Benchmark, and real-world GIS, machine-learning, financial, RAG, and embedding datasets. It tested Arrow 5.0.0, ORC 1.7.2, Parquet Java API 1.9.0, and PyArrow 17.0.0. The authors find different trade-offs among formats and report that none is optimal for certain popular machine-learning tasks.

One query comparison in that study found ORC outperforming Parquet and Arrow Feather; in that experiment, compressed Arrow Feather was 3–4 times slower than Parquet, and uncompressed Feather was more than 7 times slower. That result belongs to its specific query and test setup, not to every ORC workload. A useful benchmark should reproduce the workload of interest rather than promote a single published result into a universal ranking.

Rank #4
Sale
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
  • NEARLY 2X FASTER THAN OUR PREVIOUS GENERATION(8) – move 1,000 high-res photos in under 60 seconds(6) with up to 2000MB/s transfer speeds(2).
  • IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.
  • POCKET-SIZED – fits easily in pockets and small bags.
  • SPACE TO OWN YOUR AI CONTENT – speed and capacity to download your high-res clips and photo edits.
  • 256-BIT AES ENCRYPTION(4) – helps keep private files secure with password protection.

Build a benchmark that reflects real use

  1. Fix the schema and workload. Use the same data, column types, query mix, and output requirements for every format. Include ingest and writes if they matter; do not test only full-file reads when users run filtered queries.
  2. Measure both full reads and selective work. Include queries that read a subset of columns and filter rows. Columnar storage and predicate pushdown can avoid irrelevant data, but whether they do depends on reader implementation and layout.
  3. Record size and bytes read alongside time. Report file size, bytes scanned, elapsed time, and throughput. Compression and read performance change with data types, repetition, encodings, and codecs.
  4. Separate cold-cache and warm-cache runs. State cache conditions for every result. The 2024 comparative study used cold-cache results by default and warmed-cache results for selected experiments, illustrating why the condition matters.
  5. Include conversion and memory costs. Measure the time and memory to get data into the engine’s working representation, not just the format reader’s file-read time. Arrow IPC may reduce conversion when Arrow is already the target representation; Parquet may reduce storage.
  6. Test startup and streaming behavior. Record time to first useful batch as well as total completion time if consumers process incrementally. CSV and Arrow streams can be consumed as data arrives; Parquet and ORC need footer metadata before normal processing can begin.
  7. Hold layout choices accountable. Document row-group or stripe sizes, compression settings, partitioning, and file counts. In Arrow Dataset workflows, the documentation gives general guidance to avoid files below 20 MB or above 2 GB and layouts with more than 10,000 distinct partitions; these are not universal rules for every engine or filesystem.
  8. Publish the execution context. Include hardware, engine and library versions, schema, data types, query mix, cache state, compression, and layout so readers can judge whether the comparison applies to their own stack.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check reader and library support before choosing

Apache Arrow’s C++ Dataset API lists Parquet, Feather/Arrow IPC, CSV, and ORC among its supported formats. In that API, ORC can currently be read but not written; that restriction is specific to the documented C++ API, not a claim about every Arrow binding or library. The API also supports projection, predicate pushdown, and optional parallel reading. Confirm the features available in the exact engine and version used for the benchmark instead of assuming a format’s capabilities are exposed everywhere.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 2
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
Sandisk 1TB Portable SSD, Up to 800MB/s Read Speeds, Black (Old Model)
From Sandisk, a brand professional photographers trust to take on assignments.
$188.90
SaleBestseller No. 3
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$129.99
SaleBestseller No. 4
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.; POCKET-SIZED – fits easily in pockets and small bags.
$249.99
Bestseller No. 5
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$229.99
Best Value
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

A straightforward shortlist

  • Storage-conscious analytical scans: begin with Parquet and compare its read and write costs with the actual query mix.
  • Selective scans or Hadoop-centric execution: include ORC and verify that your reader uses its pruning features.
  • Arrow-native in-memory work or interchange: include Arrow IPC/Feather and measure memory, conversion, and file-size effects.
  • Text interoperability or sequential consumption: keep CSV as a meaningful baseline, especially where inspection or incremental text processing is part of the use case.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.