fix: encoding of Variant object header field-id and offset sizes (#421)
### What
`VariantEncodingHelper` wrote and read the Variant object value header
with
`field_id_size_minus_one` and `field_offset_size_minus_one` in each
other's bit positions.
Per `apache/parquet-format` `VariantEncoding.md`, the object
`value_header` — the 6 bits above
the 2 basic-type bits — is laid out as:
```
5 4 3 2 1 0
+---+---+-------+-------+
value_header | R | | | |
+---+---+-------+-------+
^ ^ ^
| | +-- field_offset_size_minus_one
| +-- field_id_size_minus_one
+-- is_large
```
`MakeObjectHeader` and `ParseObjectHeader` had the two 2-bit fields
transposed, and the layout
comment above them documented the same transposition — so the block was
internally consistent
rather than wrong in one expression.
`is_large` was already correct. The array header and metadata header
helpers were checked and
match the spec. This affects the object header only.
### Impact
Reader and writer shared the inverted convention, so arrow-dotnet
round-tripped its own output
correctly. The bug was only observable across implementations, and only
when
`fieldIdSize != offsetSize` — when the two are equal, transposing them
is a no-op.
Those sizes are computed independently in `VariantValueWriter`
(`fieldIdSize` from the maximum
field ID, `offsetSize` from the encoded data length), so they diverge
routinely: for example an
object drawn from a >255-entry metadata dictionary (2-byte field IDs)
whose own field data is
under 256 bytes (1-byte offsets).
For `fieldIdSize=2, offsetSize=1, isLarge=false`, the spec-correct
header byte is `0x12`;
before this change we emitted `0x06`, and read `0x12` back as
`fieldIdSize=1, offsetSize=2`.
Such objects were silently misparsed in both directions — field IDs and
offsets read at the
wrong widths, surfacing as garbage field values or out-of-range offsets
rather than a clean
error.
### Changes
- `VariantEncodingHelper.MakeObjectHeader` / `ParseObjectHeader`: swap
the two shifts, and
correct the layout comment. The `out` parameters were already named
correctly, so neither
call site — `VariantValueWriter` or `VariantObjectReader` — needed
changes.
- `VariantEncodingHelperTests`: add `MakeObjectHeaderUsesSpecBitLayout`
and
Closes #420.An implementation of Arrow targeting .NET Standard.
See our current feature matrix for currently available features.
using System.Diagnostics; using System.IO; using System.Threading.Tasks; using Apache.Arrow; using Apache.Arrow.Ipc; public static async Task<RecordBatch> ReadArrowAsync(string filename) { using (var stream = File.OpenRead(filename)) using (var reader = new ArrowFileReader(stream)) { var recordBatch = await reader.ReadNextRecordBatchAsync(); Debug.WriteLine("Read record batch with {0} column(s)", recordBatch.ColumnCount); return recordBatch; } }
Apache.Arrow.Compression package. When reading compressed data, you must pass an Apache.Arrow.Compression.CompressionCodecFactory instance to the ArrowFileReader or ArrowStreamReader constructor, and when writing compressed data a CompressionCodecFactory must be set in the IpcOptions. Alternatively, a custom implementation of ICompressionCodecFactory can be used.Install the latest .NET Core SDK from https://dotnet.microsoft.com/download.
dotnet build
To build the NuGet package run the following command to build a debug flavor, preview package into the artifacts folder.
dotnet pack
When building the officially released version run: (see Note below about current git repository)
dotnet pack -c Release
Which will build the final/stable package.
NOTE: When building the officially released version, ensure that your git repository has the origin remote set to https://github.com/apache/arrow.git, which will ensure Source Link is set correctly. See https://github.com/dotnet/sourcelink/blob/main/docs/README.md for more information.
There are two output artifacts:
Apache.Arrow.<version>.nupkg - this contains the executable assembliesApache.Arrow.<version>.snupkg - this contains the debug symbols filesBoth of these artifacts can then be uploaded to https://www.nuget.org/packages/manage/upload.
Build from the Apache Arrow project root.
docker build -f csharp/build/docker/Dockerfile .
dotnet test
All build artifacts are placed in the artifacts folder in the project root.
This project follows the coding style specified in Coding Style.
See https://flatbuffers.dev/languages/c_sharp/ for how to get the flatc executable.
Run flatc --csharp on each .fbs file in the format folder. And replace the checked in .cs files under FlatBuf with the generated files.
Update the non-generated FlatBuffers .cs files with the files from the google/flatbuffers repo.