GH-49305: [Python] Expose RecordBatchFileReader.count_rows (#50646) ### Rationale for this change Resolves [49305](https://github.com/apache/arrow/issues/49305) `RecordBatchFileReader::CountRows` has existed in Arrow C++ (`cpp/src/arrow/ipc/reader.h`) but was never bound in Python, so the only way to get the total number of rows of an IPC file was: ```python num_rows = sum(reader.get_batch(i).num_rows for i in range(reader.num_record_batches)) ``` ### This findings are done in original opened Issue #49305 That deserializes every record batch just to read its length, which is wasteful and becomes expensive on remote filesystems. ### What changes are included in this PR? Adds `RecordBatchFileReader.count_rows()`: ```python with pa.ipc.open_file(source) as reader: reader.count_rows() ``` Three changes: * `python/pyarrow/includes/libarrow.pxd`: declare `CResult[int64_t] CountRows()` on `CRecordBatchFileReader`, which was the missing piece. * `python/pyarrow/ipc.pxi`: add `count_rows()` to `_RecordBatchFileReader`, released GIL around the call, with the same closed reader guard used by the existing `stats` property * `python/pyarrow/tests/test_ipc.py`: tests. On the naming: `count_rows()` follows the C++ method and is consistent with the existing `count_rows()` on `Dataset`, `Scanner` and `Fragment`. To be precise about the benefit, since the issue describes it as reading the count from the metadata: the C++ implementation still walks every block, but reads only each record batch's flatbuffer message header to pick up its length, and never touches the data buffers. So this is a reduction in bytes read rather than in the number of reads, and the gain shows up on remote filesystems and on files with large batches rather than in a local in memory benchmark. This is only added to the file reader. The stream reader has no footer and cannot count rows without consuming the stream. ### Are these changes tested? Yes, two tests in `python/pyarrow/tests/test_ipc.py`: * `test_file_count_rows`: count matches the sum of the written batch lengths, and counting does not consume the reader (count, `read_all()`, count again). * `test_file_count_rows_no_batches`: a file with a schema but no batches counts 0. Locally `test_ipc.py` passes (72 tests) and `test_feather.py` passes (83 passed, 8 skipped, 1 xfailed), the latter because the feather reader sits on the same file reader. The docstring example was run and produces the output shown. ### Are there any user-facing changes? Yes, a new public method `RecordBatchFileReader.count_rows()`. No existing behaviour changes. ### AI usage disclosure I used Claude Code to locate the unbound C++ method and the place where the declaration was missing, and to draft the binding, the docstring and the tests. I reviewed the result, rebuilt PyArrow locally, ran the test suites quoted above, and checked the C++ implementation of `CountRows` myself to confirm what it actually does before describing the benefit here. * GitHub Issue: #49305 Authored-by: Guja <127162872+GujaLomsadze@users.noreply.github.com> Signed-off-by: Antoine Pitrou <antoine@python.org>
Apache Arrow is a universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics. It contains a set of technologies that enable data systems to efficiently store, process, and move data.
Major components of the project include:
↗: Arrow-powered API, drivers, and libraries for access to databases and query engines↗↗↗↗↗↗↗The ↗ icon denotes that this component of the project is maintained in a separate repository.
Arrow is an Apache Software Foundation project. Learn more at arrow.apache.org.
The reference Arrow libraries contain many distinct software components:
The official Arrow libraries in this repository are in different stages of implementing the Arrow format and related features. See our current feature matrix on git main.
Please read our latest project contribution guide.
If you are using AI coding tools, please review our AI-generated code guidance.
Even if you do not plan to contribute to Apache Arrow itself or Arrow integrations in other projects, we'd be happy to have you involved:
We use runs-on for managing the project self-hosted runners. We use AWS for some of the required infrastructure for the project.