This benchmark compares Apache Fory C#, protobuf-net, and MessagePack-CSharp.
Serializer setup used in this benchmark:
fory: Fory.Builder().Compatible(true).Build()protobuf: protobuf-net runtime serializermsgpack: MessagePackSerializer.Typeless with TypelessContractlessStandardResolvercd benchmarks/csharp ./run.sh
This runs all benchmark cases and generates:
build/benchmark_results.jsonreport/README.mdreport/throughput.png./run.sh --help Options: --data <struct|sample|mediacontent|structlist|samplelist|mediacontentlist> --serializer <fory|protobuf|msgpack> --duration <seconds> --warmup <seconds> --external-equivalence --allocation-iterations <count> --baseline-worktree <dir>
Examples:
# Run only struct benchmarks ./run.sh --data struct # Run only Fory benchmarks ./run.sh --serializer fory # Use longer runs for stable numbers ./run.sh --duration 10 --warmup 2
The opt-in external-type mode compares the latest apache/main ordinary serializer, the current ordinary serializer, and the current external structural serializer for equivalent models from a Fory-free project:
./run.sh --external-equivalence --duration 10 --warmup 3
It covers direct class and struct roots, a direct holder field, list and map fields, and list and map roots. Each ordinary/external pair uses the same Fory configuration, numeric registration IDs, field schemas, values, and native carrier serializers.
Fory.CSharpBenchmark.csproj owns only the existing Fory, protobuf-net, and MessagePack benchmarks. Fory.ExternalTypeBenchmark.csproj is a parallel package that owns the external-type models and measurements. Building or running the ordinary benchmark therefore does not compile external-type models or their generated serializer instantiations.
The runner fetches apache/main and reuses the dedicated sibling worktree ../fory-benchmark-baseline by default. Use --baseline-worktree to select another dedicated path. If the worktree already exists, the runner refuses tracked changes before switching it to the fetched commit. It never resets or cleans the worktree, so untracked files are preserved and a conflicting file causes the switch to fail safely.
The runner builds the baseline and current external-type benchmark projects separately. For each selected lane, it launches two balanced adjacent triples in fixed order: apache-main ordinary external, followed by external ordinary apache-main. Each of the six workers runs in a fresh process and constructs, registers, validates, warms, times, and measures allocations for only its selected implementation. Each baseline/current and ordinary/external comparison is therefore adjacent in both orders. Keeping the implementations in separate processes prevents shared reference-generic Dynamic PGO profiles from biasing a later implementation while retaining normal .NET tiering.
Use --data to select one of these lanes:
class-rootstruct-rootholder-fieldlist-fieldlist-rootmap-fieldmap-rootEvery worker runs a fixed allocation pass with 100,000 operations by default. To select another fixed count, use:
./run.sh --external-equivalence --allocation-iterations 100000
Registration, serializer resolution, payload creation, delegates, metadata caches, and timed warmup for the selected implementation are completed before allocation counters start. Allocation results use GC.GetAllocatedBytesForCurrentThread() and report bytes per operation.
Results are written to build/external_equivalence_results.json. Per-process results, including a Base64 proof frame, are written under build/external_equivalence_sides/ with the lane, one-based sample index, and role in each filename. After every worker finishes, the merger rejects missing, duplicate, reordered, or reused samples, incompatible runtime metadata, serialized size or byte differences, and unstable per-role allocation samples. It fails if current ordinary latency is more than 1% above apache/main, current external latency is more than 1% above current ordinary, current external latency is more than 1% above apache/main, current ordinary allocates more than apache/main, or current external allocation differs from current ordinary. The complete JSON, including commit provenance and every violation, is written before a regression returns a nonzero exit code. This opt-in mode does not invoke the default Python report generator or update the published 36-case report.
The benchmark executable requires exactly one implementation for direct worker use:
dotnet run -c Release --no-build \ --project ./Fory.ExternalTypeBenchmark.csproj -- \ --external-implementation ordinary \ --data class-root \ --duration 10 \ --warmup 3 \ --allocation-iterations 100000 \ --output build/external_equivalence_sides/class-root-direct-ordinary.json
Use ordinary or external. run.sh owns the paired process ordering and normally supplies this selector.
Set FORY_BENCH_SCHEMA_MISMATCH=1 to run the Fory-only compatible-read schema-mismatch mode. This mode is off by default. When enabled, run with --serializer fory; protobuf-net and MessagePack benchmark modes fail with a configuration error. Fory serialization uses the normal v1 benchmark models, and Fory deserialization uses v2 models registered with the same Fory type IDs where one int32 field is widened to int64.
struct: 8-field integer objectsample: mixed primitive fields and arraysmediacontent: nested object with list fieldsstructlist: list of struct payloadssamplelist: list of sample payloadsmediacontentlist: list of media content payloadsEach case benchmarks:
Latest run (Darwin arm64, .NET 8.0.24, --duration 2 --warmup 0.5):
| Serializer | Mean ops/sec across all cases |
|---|---|
| fory | 2,032,057 |
| protobuf | 1,940,328 |
| msgpack | 1,901,489 |
Per-case winners vary by payload and operation. The full breakdown is generated at:
benchmarks/csharp/build/benchmark_results.jsonbenchmarks/csharp/report/README.mdbenchmarks/csharp/report/throughput.png