The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Apache Iceberg’s REST Catalog standardizes how clients talk to a catalog; it does not make two clients equally fast. Differences in client implementation, server capabilities, metadata and cache state, planning mode, engine settings, network round trips and data scanning can all change elapsed time. To find the cause, separate catalog setup and table loading from scan planning, execution and result delivery.
What does “one protocol” guarantee?
The REST Catalog protocol provides a common HTTP interface for catalog operations. Iceberg describes its interoperability goal this way: “a single client implementation works with any compliant server.” That is a statement about compatibility, not a performance guarantee. A compliant client and server can still take different paths or do different amounts of work. Apache Iceberg REST Catalog Protocol documentation
As an Amazon Associate I earn from qualifying purchases.
So “Why is Iceberg query planning slow?” and “Does Iceberg REST catalog improve query performance?” do not have one protocol-wide answer. The REST interface can change where work happens, but it does not by itself make planning or the eventual scan faster. Identify the slow phase before attributing a difference to the protocol.
Recommended Free Tools
Where the time goes
Catalog setup and network round trips
A REST client discovers server configuration during initialization with GET /v1/config. The response can provide defaults, enforce overrides and advertise optional endpoints. Clients may differ in which settings or optional features they use, and servers may not support every optional capability. When setup or catalog latency is in question, record the effective configuration, advertised endpoints and number of requests rather than treating the protocol as a single call. Apache Iceberg REST Catalog Protocol documentation
#1 Best Overall
Table loading, metadata and cache state
Loading a table ordinarily involves downloading its metadata. The REST protocol documents ETag-aware loading: a client can send If-None-Match and reuse cached table metadata when the server responds 304 Not Modified. It also documents lazy snapshot loading, which can avoid fetching full snapshot history when a client only needs branch and tag references. Consequently, cold and warm starts—and tables with different histories—may involve different amounts of metadata work. Apache Iceberg REST Catalog Protocol documentation
Scan planning: client-side or server-side
In the Java REST client, client-side scan planning is the documented default. The client reads table metadata and creates file scan tasks locally. Server-side planning is optional and requires the server to advertise support; the client sends the filter, snapshot and selected columns, and the server returns tasks. The asynchronous lifecycle can involve submitting a plan, polling with a plan ID and fetching task batches. Apache Iceberg REST Catalog Protocol documentation
Server-side planning may reduce metadata downloads by moving planning work to the catalog service, which may also use its own caches or indexes. It can also add server processing and polling or network wait. Neither mode is automatically faster: compare the work and elapsed time on both sides for the same workload.
Metadata pruning and file selection
Iceberg’s manifest list stores partition-value ranges for manifests; manifests contain data-file partition information and column statistics. Planning can use that information to prune manifests and exclude files that cannot match a predicate. Effective pruning can reduce planning work and later data I/O, but its benefit depends on the table metadata, predicate and layout. It is not a universal speed multiplier. Apache Iceberg 1.9.0 performance documentation
Engine planning, execution and result delivery
Having file tasks is not the same as finishing a query. The engine still optimizes the plan and reads data, and results still need to be delivered. Engine configuration can affect those phases: Trino’s Iceberg connector documentation, for example, covers cost-based optimization statistics, metadata caching and split sizing. The linked page is Trino’s versioned documentation for 483; check the version and settings actually deployed rather than assuming its defaults apply elsewhere. Trino 483 Iceberg connector documentation
How do I compare two Iceberg clients?
For a useful comparison, hold constant the server, table state, query and execution environment. Measure phases separately where instrumentation permits; the following breakdown is a practical measurement framework, not a claim that every client exposes these exact timers.
Rank #4
- Fix the comparison conditions. Use the same catalog server and configuration, table snapshot and metadata state, query text and parameters, storage and network region, client or engine resource limits, and concurrency.
- Record implementation and negotiation details. Capture client and server versions, advertised REST endpoints and effective configuration. Check feature support for those specific releases.
- Run cold- and warm-cache cases. Keep the two states distinct; metadata caching and ETag revalidation can change the amount of work.
- Measure the path. Where possible, record catalog request count and latency, metadata bytes fetched and loading or parsing time, scan-planning duration and task turnaround, engine planning, execution, data bytes and files scanned, and result-transfer time. Record server-side planning work as well when that mode is used.
- Repeat and report distributions. Include repeated trials and a distribution such as median and tail latency, not just a single wall-clock result.
These are methodological recommendations based on the documented request, metadata and planning lifecycle. They are not a published benchmark protocol or a claim that every implementation supplies every metric. If the total differs, the phase breakdown helps show whether the cause is setup, metadata, planning, execution, scanning or result delivery.
What existing benchmarks do—and do not—show
The CIDR 2023 paper Analyzing and Comparing Lakehouse Storage Systems reported that, in its own 3 TB TPC-DS experiment, query runtime was 1.4× faster on Delta than Hudi and 1.7× faster on Delta than Iceberg. The authors discuss factors including reading time, file sizes and counts, a custom Parquet reader, and differences in query plans. This is a result for particular formats and implementations in a particular Spark setup—not a comparison of two Iceberg clients using one REST server. CIDR 2023 paper
Best Value
The paper also describes metadata operations becoming a planning bottleneck for very small queries and notes that the Hudi system in that experiment cached query plans. That makes startup and metadata worth measuring separately; it does not establish that one Iceberg client plans faster. Apache Hudi’s 2026 project-authored article likewise emphasizes workload shape, configuration parity and tested versions, and treats older TPC-DS tests as historical evidence rather than a current general ranking. Apache Hudi project article, August 13, 2026
Those results cannot rank REST clients. A client-versus-client conclusion needs the same server and workload, documented versions and configuration, and measurements that distinguish planning from execution. Do not infer a universal winner from a benchmark that tests different formats, systems or workloads.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems




