)]}'
{
  "commit": "a67cd19fff65b6c995be9a5eae56845157d95301",
  "tree": "60ea22b460b5702dfc2415a9234429290a287c04",
  "parents": [
    "ce4edd53203eb4bca96c10ebf3d2118299dad006"
  ],
  "author": {
    "name": "Jonas Dedden",
    "email": "university@jonas-dedden.de",
    "time": "Fri Dec 05 17:59:45 2025 +0100"
  },
  "committer": {
    "name": "GitHub",
    "email": "noreply@github.com",
    "time": "Fri Dec 05 11:59:45 2025 -0500"
  },
  "message": "Implement a `Vec\u003cRecordBatch\u003e` wrapper for `pyarrow.Table` convenience (#8790)\n\n# Rationale for this change\n\nWhen dealing with Parquet files that have an exceedingly large amount of\nBinary or UTF8 data in one row group, there can be issues when returning\na single RecordBatch because of index overflows\n(https://github.com/apache/arrow-rs/issues/7973).\n\nIn `pyarrow` this is usually solved by representing data as a\n`pyarrow.Table` object whose columns are `ChunkedArray`s, which\nbasically are just lists of Arrow Arrays, or alternatively, the\n`pyarrow.Table` is just a representation of a list of `RecordBatch`es.\n\nI\u0027d like to build a function in PyO3 that returns a `pyarrow.Table`,\nvery similar to [pyarrow\u0027s read_row_group\nmethod](https://arrow.apache.org/docs/python/generated/pyarrow.parquet.ParquetFile.html#pyarrow.parquet.ParquetFile.read_row_group).\nWith that, we could have feature parity with `pyarrow` in circumstances\nof potential index overflows without resorting to type changes (such as\nreading the data as `LargeString` or `StringView` columns).\nCurrently, AFAIS, there is no way in `arrow-pyarrow` to export a\n`pyarrow.Table` directly. Especially convenience methods from\n`Vec\u003cRecordBatch\u003e` seem to be missing. This PR tries to implement a\nconvenience wrapper that allows directly exporting `pyarrow.Table`.\n\n# What changes are included in this PR?\n\nA new struct `Table` in the crate `arrow-pyarrow` is added which can be\nconstructed from `Vec\u003cRecordBatch\u003e` or from `ArrowArrayStreamReader`.\nIt implements `FromPyArrow` and `IntoPyArrow`. \n\n`FromPyArrow` will support anything that either implements the\nArrowStreamReader protocol or is a RecordBatchReader, or has a\n`to_reader()` method which does that. `pyarrow.Table` does both of these\nthings.\n`IntoPyArrow` will result int a `pyarrow.Table` on the Python side,\nconstructed through `pyarrow.Table.from_batches(...)`.\n\n# Are these changes tested?\n\nYes, in `arrow-pyarrow-integration-tests`.\n\n# Are there any user-facing changes?\n\nA new `Table` convience wrapper is added!",
  "tree_diff": [
    {
      "type": "modify",
      "old_id": "7d5d63c1d50da6508c2876b9ea23e114f09f84df",
      "old_mode": 33188,
      "old_path": "arrow-pyarrow-integration-testing/src/lib.rs",
      "new_id": "a5690b3070409d83d9e3c5b106c3839785b4a6a3",
      "new_mode": 33188,
      "new_path": "arrow-pyarrow-integration-testing/src/lib.rs"
    },
    {
      "type": "modify",
      "old_id": "f5d53155fe78d0946f6046ff8f1c6f8aba6820ee",
      "old_mode": 33188,
      "old_path": "arrow-pyarrow-integration-testing/tests/test_sql.py",
      "new_id": "79220fb6a69f3540aae8be990bebd5592da84c80",
      "new_mode": 33188,
      "new_path": "arrow-pyarrow-integration-testing/tests/test_sql.py"
    },
    {
      "type": "modify",
      "old_id": "d4bbb201f027d019c8dbb239add7edbe9927493d",
      "old_mode": 33188,
      "old_path": "arrow-pyarrow/src/lib.rs",
      "new_id": "1f8941ef1cf53bdb0af049e4c8afb392ae29230b",
      "new_mode": 33188,
      "new_path": "arrow-pyarrow/src/lib.rs"
    }
  ]
}
