Arrow data can stay in Arrow buffers when it moves between Python libraries inside one process, provided both sides support Arrow’s C Data or PyCapsule protocols. When the data comes from ClickHouse, ClickHouse Connect returns Arrow results directly, so you can avoid building Python rows. The documentation does not promise a copy-free path from the database server into your application, and moving Arrow data into ClickHouse is a separate question with thinner documentation.
What Arrow can keep in place
Apache Arrow is a columnar in-memory format and a set of interchange tools. In PyArrow, the core objects are typed arrays, record batches, tables, and buffers. A table is a set of columns, and each column is a chunked array: a sequence of arrays that share one type.
Arrow data is immutable. The Apache Arrow documentation for its data types and in-memory data model states: “Arrow data is immutable, so values can be selected but not assigned.” No individual author is named on that page. Immutability is what makes sharing safe. A slice can reference a range of an existing buffer instead of rewriting values, and a PyArrow buffer can wrap memory that already implements Python’s buffer protocol without allocating a second buffer. PyArrow documents the buffer-to-memoryview conversion as zero-copy.
Where copies appear
Buffer.to_pybytes()materializes a Python bytes object. The PyArrow documentation states that this copies the buffer.- Turning columns into Python lists, dictionaries, or row tuples creates a new Python object for every value, which takes the data out of Arrow memory entirely.
The in-process boundary: C Data Interface and PyCapsule
The Arrow C Data Interface shares Arrow structures between compatible implementations through pointers. The producer supplies a release callback, and the consumer calls it to signal that it no longer needs the memory. That callback is how lifetime is coordinated across implementations. The interface specification lists sharing data between independent runtimes or components in the same process as a goal. It lists inter-process sharing and persistence as non-goals.
#1 Best Overall
For Python libraries, the PyCapsule Interface exposes three methods on a producing object: __arrow_c_schema__, __arrow_c_array__, and __arrow_c_stream__. PyArrow constructors can consume these for schemas, arrays, tables, and streams. The PyArrow documentation says these conversions can be zero-copy when the participating structures and implementations support the interface. That is a conditional statement, not a guarantee for every conversion or dtype, so check the column types on both sides.
Getting Arrow results out of ClickHouse
ClickHouse Connect is the Python client covered by the current ClickHouse documentation. It has three Arrow-oriented result paths. The examples below assume a client created with clickhouse_connect.get_client(...).
Rank #2
query_arrow(): one bounded table
query_arrow() uses ClickHouse’s Arrow output format and returns a pyarrow.Table. Use it when the result is bounded and one table is the workflow you need.
query_arrow_stream(): batch by batch
query_arrow_stream() returns a stream context that yields PyArrow record batches. The ClickHouse documentation requires opening it in a with block, which ensures the stream is closed when you finish. Because you process one batch at a time, you do not need to hold the entire result in a single table.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
DataFrame output
Arrow-backed pandas output uses Arrow-backed dtypes and requires pandas 2.x. Polars output can be built from the Arrow table. ClickHouse documents both conversions as zero-copy “where possible.” The sources checked do not list the exceptions, so test the dtypes your own queries produce.
import clickhouse_connect
client = clickhouse_connect.get_client(host="localhost")
# Bounded result as one pyarrow.Table
table = client.query_arrow("SELECT number, toString(number) AS label FROM numbers(1000)")
print(table.schema)
# Streamed result as pyarrow.RecordBatch objects
total = 0
with client.query_arrow_stream("SELECT number FROM numbers(1000000)") as stream:
for batch in stream:
total += batch.num_rows
print(total)
Sending Arrow data into ClickHouse
“Transfer to ClickHouse” suggests inserts, and the insert direction is less documented. The sources checked point to a specialized insert_arrow method that accepts a PyArrow Table. Those references came from a translated reference mirror rather than an English primary documentation page. Confirm the method name, its arguments, and its copy behavior against the ClickHouse Connect release you install and the current official documentation. Do not assume the insert path avoids copies until you have verified it. The general client API documents ordinary query and insert methods and directs Arrow and DataFrame users to the specialized methods.
Rank #4
Where “zero-copy” stops being accurate
Three boundaries end the claim:
- The network transport. A ClickHouse query crosses a client–server boundary. Arrow output is the database’s result representation, but the sources reviewed do not promise that the server-to-client path avoids copies. A copy-free trip from a remote database into application objects is not something the documentation supports.
- Process and machine boundaries. The C Data Interface does not move data between processes or machines. When data must cross one, or be stored, use Arrow IPC. IPC relies on a serialized format rather than direct in-process buffer sharing.
- Python objects. Once values become bytes, lists, dictionaries, or rows, they have left Arrow memory.
A more accurate phrasing for a ClickHouse workflow is “Arrow results from ClickHouse, with zero-copy conversion to Arrow-backed DataFrames where the types allow it.”
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choosing an approach
Pick the method based on result size, whether you need streaming, whether the boundary is in-process or remote, whether downstream code accepts Arrow types, and how long the buffers must live.
Quick Recap
| Choice | Use it when | Copy considerations |
|---|---|---|
query_arrow() to a pyarrow.Table |
The result is bounded and a single table is the goal | Results arrive as Arrow output. The sources do not promise zero copies across the network path. |
query_arrow_stream() |
Results should be processed batch by batch | You avoid holding one complete result table. The sources give no memory-savings figures. |
| Arrow-backed pandas output | Existing analysis code expects a DataFrame | Zero-copy where possible. Requires pandas 2.x. Verify the dtypes. |
| Polars built from the Arrow table | Downstream code uses Polars | Zero-copy where possible, per ClickHouse Connect’s documentation. Verify the types in your workload. |
| PyCapsule or C Data handoff | Two compatible libraries share data in one process | Buffers can be shared without copying. Protocol support, type compatibility, and buffer lifetime all matter. |
| Arrow IPC | Data crosses processes or machines, or is persisted | Data is serialized. Direct buffer sharing does not apply. |
Implementation checklist
- Identify the boundary: same process, a network connection to ClickHouse, or a process, machine, or storage boundary.
- Pin versions in your requirements file. The ClickHouse Connect documentation is published from the moving
mainbranch of the ClickHouse docs repository, so method signatures and supported types can change. The PyArrow documentation available at the time of review listed version 25.0.1 as current. Pin pandas to 2.x if you use Arrow-backed pandas output. - Use
query_arrow()for bounded results. Usequery_arrow_stream()inside awithblock when you process batches. - Print
table.schemaafter retrieval, or inspect the dtypes of the converted DataFrame, and confirm the types match what your code expects. - Keep Arrow types downstream wherever consumers accept them. On hot paths, avoid
to_pybytes()and row-by-row Python conversion. - Keep Arrow objects alive for as long as any consumer references their buffers.
- Measure with your own data, hardware, and package versions. The sources checked contain no benchmarks, so no throughput, latency, or memory-savings figures are given here.
Scope and version notes
- The sources checked do not establish one tested combination of ClickHouse Connect, PyArrow, and pandas versions.
- This article does not report measured results for any of the transfer paths described.




