GH-49305: [Python] Expose RecordBatchFileReader.count_rows (#50646)

### Rationale for this change

Resolves [49305](https://github.com/apache/arrow/issues/49305)

`RecordBatchFileReader::CountRows` has existed in Arrow C++ (`cpp/src/arrow/ipc/reader.h`) but
was never bound in Python, so the only way to get the total number of rows of an IPC file was:

```python
num_rows = sum(reader.get_batch(i).num_rows for i in range(reader.num_record_batches))
```

### This findings are done in original opened Issue #49305

That deserializes every record batch just to read its length, which is wasteful and becomes
expensive on remote filesystems.

### What changes are included in this PR?

Adds `RecordBatchFileReader.count_rows()`:

```python
with pa.ipc.open_file(source) as reader:
    reader.count_rows()
```

Three changes:

* `python/pyarrow/includes/libarrow.pxd`: declare `CResult[int64_t] CountRows()` on
  `CRecordBatchFileReader`, which was the missing piece.
* `python/pyarrow/ipc.pxi`: add `count_rows()` to `_RecordBatchFileReader`, released GIL around
  the call, with the same closed reader guard used by the existing `stats` property
* `python/pyarrow/tests/test_ipc.py`: tests.

On the naming: `count_rows()` follows the C++ method and is consistent with the existing
`count_rows()` on `Dataset`, `Scanner` and `Fragment`.

To be precise about the benefit, since the issue describes it as reading the count from the
metadata: the C++ implementation still walks every block, but reads only each record batch's
flatbuffer message header to pick up its length, and never touches the data buffers. So this is
a reduction in bytes read rather than in the number of reads, and the gain shows up on remote
filesystems and on files with large batches rather than in a local in memory benchmark.

This is only added to the file reader. The stream reader has no footer and cannot count rows
without consuming the stream.

### Are these changes tested?

Yes, two tests in `python/pyarrow/tests/test_ipc.py`:

* `test_file_count_rows`: count matches the sum of the written batch lengths, and counting does
  not consume the reader (count, `read_all()`, count again).
* `test_file_count_rows_no_batches`: a file with a schema but no batches counts 0.

Locally `test_ipc.py` passes (72 tests) and `test_feather.py` passes (83 passed, 8 skipped,
1 xfailed), the latter because the feather reader sits on the same file reader. The docstring
example was run and produces the output shown.

### Are there any user-facing changes?

Yes, a new public method `RecordBatchFileReader.count_rows()`. No existing behaviour changes.

### AI usage disclosure

I used Claude Code to locate the unbound C++ method and the place where the declaration was
missing, and to draft the binding, the docstring and the tests. I reviewed the result, rebuilt
PyArrow locally, ran the test suites quoted above, and checked the C++ implementation of
`CountRows` myself to confirm what it actually does before describing the benefit here.
* GitHub Issue: #49305

Authored-by: Guja <127162872+GujaLomsadze@users.noreply.github.com>
Signed-off-by: Antoine Pitrou <antoine@python.org>
3 files changed
tree: 625ed6b2ba6b255d36e3e2cebb9e388afd678b8b
  1. .github/
  2. c_glib/
  3. ci/
  4. cpp/
  5. dev/
  6. docs/
  7. format/
  8. matlab/
  9. python/
  10. r/
  11. ruby/
  12. .asf.yaml
  13. .clang-format
  14. .clang-tidy
  15. .clang-tidy-ignore
  16. .dockerignore
  17. .editorconfig
  18. .env
  19. .gitattributes
  20. .gitignore
  21. .gitmodules
  22. .hadolint.yaml
  23. .pre-commit-config.yaml
  24. .rubocop.yml
  25. .shellcheckrc
  26. CHANGELOG.md
  27. cmake-format.py
  28. CODE_OF_CONDUCT.md
  29. compose.yaml
  30. CONTRIBUTING.md
  31. CPPLINT.cfg
  32. LICENSE.txt
  33. NOTICE.txt
  34. README.md
README.md

Apache Arrow

Fuzzing Status License BlueSky Follow

Powering In-Memory Analytics

Apache Arrow is a universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics. It contains a set of technologies that enable data systems to efficiently store, process, and move data.

Major components of the project include:

The ↗ icon denotes that this component of the project is maintained in a separate repository.

Arrow is an Apache Software Foundation project. Learn more at arrow.apache.org.

What's in the Arrow libraries?

The reference Arrow libraries contain many distinct software components:

  • Columnar vector and table-like containers (similar to data frames) supporting flat or nested types
  • Fast, language agnostic metadata messaging layer (using Google's FlatBuffers library)
  • Reference-counted off-heap buffer memory management, for zero-copy memory sharing and handling memory-mapped files
  • IO interfaces to local and remote filesystems
  • Self-describing binary wire formats (streaming and batch/file-like) for remote procedure calls (RPC) and interprocess communication (IPC)
  • Integration tests for verifying binary compatibility between the implementations (e.g. sending data from Java to C++)
  • Conversions to and from other in-memory data structures
  • Readers and writers for various widely-used file formats (such as Parquet, CSV)

Implementation status

The official Arrow libraries in this repository are in different stages of implementing the Arrow format and related features. See our current feature matrix on git main.

How to Contribute

Please read our latest project contribution guide.

If you are using AI coding tools, please review our AI-generated code guidance.

Getting involved

Even if you do not plan to contribute to Apache Arrow itself or Arrow integrations in other projects, we'd be happy to have you involved:

Continuous Integration Sponsors

We use runs-on for managing the project self-hosted runners. We use AWS for some of the required infrastructure for the project.