| % Generated by roxygen2: do not edit by hand |
| % Please edit documentation in R/dataset-format.R |
| \name{FragmentScanOptions} |
| \alias{FragmentScanOptions} |
| \alias{CsvFragmentScanOptions} |
| \alias{ParquetFragmentScanOptions} |
| \alias{JsonFragmentScanOptions} |
| \title{Format-specific scan options} |
| \description{ |
| A \code{FragmentScanOptions} holds options specific to a \code{FileFormat} and a scan |
| operation. |
| } |
| \section{Factory}{ |
| |
| \code{FragmentScanOptions$create()} takes the following arguments: |
| \itemize{ |
| \item \code{format}: A string identifier of the file format. Currently supported values: |
| \itemize{ |
| \item "parquet" |
| \item "csv"/"text", aliases for the same format. |
| } |
| \item \code{...}: Additional format-specific options |
| |
| \code{format = "parquet"}: |
| \itemize{ |
| \item \code{use_buffered_stream}: Read files through buffered input streams rather than |
| loading entire row groups at once. This may be enabled |
| to reduce memory overhead. Disabled by default. |
| \item \code{buffer_size}: Size of buffered stream, if enabled. Default is 8KB. |
| \item \code{pre_buffer}: Pre-buffer the raw Parquet data. This can improve performance |
| on high-latency filesystems. Disabled by default. |
| \item \code{thrift_string_size_limit}: Maximum string size allocated for decoding thrift |
| strings. May need to be increased in order to read |
| files with especially large headers. Default value |
| 100000000. |
| \item \code{thrift_container_size_limit}: Maximum size of thrift containers. May need to be |
| increased in order to read files with especially large |
| headers. Default value 1000000. |
| \code{format = "text"}: see \link{CsvConvertOptions}. Note that options can only be |
| specified with the Arrow C++ library naming. Also, "block_size" from |
| \link{CsvReadOptions} may be given. |
| } |
| } |
| |
| It returns the appropriate subclass of \code{FragmentScanOptions} |
| (e.g. \code{CsvFragmentScanOptions}). |
| } |
| |