Go
Go’s serialization landscape mixes stdlib codecs (encoding/json, encoding/json/v2, encoding/gob), a competitive JSON performance tier (sonic, goccy, jsoniter, segmentio, ugorji), schemaless binary (MessagePack, CBOR, kelindar/binary, BSON), text documents (YAML, TOML), schema/IDL stacks (protobuf, Avro, Dagr), and columnar / fixed-wire rows (Arrow IPC, Parquet, SBE). The Go tree registers 30 serializers. encoding/json is the v1 API. encoding/json/v2 is the Go 1.27 package with stricter defaults.
Runtime
What it is
Go compiles to native machine code before the process starts. There is no Java-style virtual machine and no intermediate language such as .NET IL. The compiler still embeds a small runtime in every binary. That runtime includes a concurrent garbage collector, a scheduler for goroutines (Go’s lightweight threads), and the stacks those goroutines use. You do not install a separate “Go VM” in order to run the benchmark.
| This suite | |
|---|---|
| Language / module | Go 1.27 (go.mod toolchain go1.27.1) |
| Host bootstrap | Go 1.22 or newer. GOTOOLCHAIN=auto may download 1.27. |
| Prepare | ./scripts/install-host-requirements.sh go installs into ~/.local/go |
| Run | go/scripts/run-benchmarks.sh runs go build and then the binary |
| Memory | Concurrent garbage collector inside the Go runtime |
What this suite runs
The runner is a normal go build of the go/ tree with the compiler’s default optimizations. The install script only needs a bootstrap compiler. If that bootstrap is older than the version pinned in go.mod, the Go toolchain setting GOTOOLCHAIN=auto downloads the exact version the module asks for.
What changes the numbers
Go’s garbage collector is designed for short pauses, but allocation still matters. Rows that reuse an Encoder, an EncMode, or a buffer — sonic’s Pretouch, ugorji Handles — often pull ahead of encoding/json for that reason. SIMD libraries such as sonic also depend on the host CPU. encoding/gob and kelindar/binary are Go-only wire formats.
Suite-specific gotchas
protobuf, linkedin/goavro, arrow-ipc, parquet, parquet-uncompressed, and sbe have no native stream API on the bytes payload this suite times. Their stream rows are adapted: the timed path is still bytes, then a write or read of those bytes. The columnar rows are bytes-only measurements. There is no compliance decoder.
These times cannot be ranked against another language.
Where to go next
The steps to install the toolchain and run the benchmark are in go/README.md. The language overview is the Go documentation.
Benchmark runner
- Directory:
go/(repository root) - Output: monorepo
logs/go/YYYY-MM-DD-HHMMSS.csv(Language=go, times in nanoseconds) - Runner:
go/scripts/run-benchmarks.sh {smoke|all-single|full|research}orgo build && ./bin/serializer-benchmark-go <reps> - Registration:
go/serializers/registry.go
Serializers
| Serializer | Category | Package | Native path | Stream | Notes |
|---|---|---|---|---|---|
| dagr-packed | Schema | dagr + gen (go/gen/dagrv2) |
direct builder / lazy reader | adapted | Direct value structs built in Prepare (like protobuf's toProto); BuildAppend timed; lazy accessors → domain timed; N>1 suite frame |
| dagr-regular | Schema | dagr + gen (go/gen/dagrv2) |
arena serializer / lazy reader | adapted | Regular (vtable) nodes; generated arena built in Prepare; AppendTo<Graph> into one reused builder timed; lazy accessors → domain timed; N>1 suite frame |
| dagr-frozen | Schema | dagr + gen (go/gen/dagrv2) |
arena serializer / lazy reader | adapted | Frozen nodes; generated arena built in Prepare; AppendTo<Graph> into one reused builder timed; lazy accessors → domain timed; N>1 suite frame |
| dagr-frozen-packed | Schema | dagr + gen (go/gen/dagrv2) |
direct builder / lazy reader | adapted | Frozen+packed nodes; same call path as dagr-packed (BuildAppend timed) |
| encoding/gob | Native | stdlib | registered types | native | Buffer Reset between encodes |
| encoding/json | JSON | stdlib | struct tags | native | v1 API. Stream SetEscapeHTML(false) |
| encoding/json/v2 | JSON | stdlib | struct tags | native | v2 defaults. MarshalWrite / UnmarshalRead |
| fxamacker/cbor | CBOR | cbor/v2 | reused Enc/DecMode | native | Default EncOptions (not CoreDet) |
| goccy/go-json | JSON | goccy/go-json | drop-in API | native | Fast stdlib substitute |
| goccy/go-yaml | YAML | goccy/go-yaml | Marshal/Unmarshal | native | High-perf YAML |
| hamba/avro | Schema | hamba/avro/v2 | frozen API + schema cache | native | Stream NewEncoder/NewDecoder; schema parse once |
| ion-go | Binary | ion-go | MarshalBinary / Unmarshal |
native | Binary encoder stream; wire names are Go field names (no ion tags on the suite structs) |
| jsoniter | JSON | json-iterator/go | compatible config | native | Widely deployed |
| kelindar/binary | Binary | kelindar/binary | Encoder.Reset | native | Go-only compact packer |
| linkedin/goavro | Schema | goavro/v2 | BinaryFromNative maps | adapted | Bytes-only codec; OCF is a different format; map convert untimed |
| mongo-bson | Document | mongo-driver/bson | Encoder+JSON tags | native | Batch wrap {items}; length-prefixed stream read |
| pelletier/go-toml | TOML | go-toml/v2 | Marshal/Unmarshal | native | Batch wrapped {items} untimed |
| arrow-ipc | Columnar | arrow-go/v18 | Schema in prepare | adapted | IPC stream bytes, not the file format. Record batch built inside SerializeBytes. table_project reads the f_float_0 buffer. No compliance decoder |
| parquet | Columnar | arrow-go/v18 | Schema in prepare | adapted | pqarrow file. Snappy is set explicitly (arrow-go's writer default is uncompressed). table_project passes column 0. No compliance decoder |
| parquet-uncompressed | Columnar | arrow-go/v18 | Schema in prepare | adapted | Same writer with compression off. No compliance decoder |
| protobuf | Schema | protobuf + gen | Message in prepare | adapted | MarshalAppend; ToDomain untimed; no native stream API |
| sbe | Fixed wire | sbe-tool 1.40.2 | type id in Prepare | adapted | Flyweight filled inside SerializeBytes. table, table_project, signal. No nested_table. No compliance decoder |
| segmentio/encoding/json | JSON | segmentio/encoding | drop-in API | native | Production fork |
| shamaton/msgpack | MessagePack | msgpack/v3 | Marshal/Unmarshal | native | Stream MarshalWrite/UnmarshalRead |
| shamaton/msgpack (array) | MessagePack | msgpack/v3 | MarshalAsArray/UnmarshalAsArray | native | Struct-as-array (no field-name keys); stream MarshalWriteAsArray/UnmarshalReadAsArray |
| sonic | JSON | bytedance/sonic | ConfigDefault + Pretouch |
native | SIMD-oriented hot path |
| ugorji/cbor | CBOR | ugorji/go/codec | CborHandle + EncoderBytes | native | go-codec multi-format |
| ugorji/json | JSON | ugorji/go/codec | JsonHandle + EncoderBytes | native | go-codec multi-format |
| ugorji/msgpack | MessagePack | ugorji/go/codec | MsgpackHandle + EncoderBytes | native | go-codec multi-format |
| vmihailenco/msgpack | MessagePack | msgpack/v5 | reused Encoder | native | Encoder.Reset + buffer |
Specifics
Why each library exists, what problem it was written to solve, and how. Names link to the source repository (or the stdlib / in-tree path this suite times). A version after the name is the last measured SerializerVersion from this suite's latest bench.
dagr-packed
Dagr ("Data Graph") is a schema-driven binary format that can store shared nodes and cycles, built on an arena model. One Python DSL schema generates the code for every target language (dagr build), so there is no runtime library: the suite commits the generated code from schemas/v2/dagr/schema.py. The five suite types are emitted in all four node layouts, one row each. The graph data type is emitted only for regular and frozen, because a packed reference cannot store the person ring. On suite types, prepare builds the native value and the timed call writes every field (direct builder or arena serializer) and reads them back. This row uses the packed node layout (tagged, evolvable). It does not support the graph data type.
dagr-regular
Dagr ("Data Graph") is a schema-driven binary format that can store shared nodes and cycles, built on an arena model. One Python DSL schema generates the code for every target language (dagr build), so there is no runtime library: the suite commits the generated code from schemas/v2/dagr/schema.py. The five suite types are emitted in all four node layouts, one row each. The graph data type is emitted only for regular and frozen, because a packed reference cannot store the person ring. On suite types, prepare builds the native value and the timed call writes every field (direct builder or arena serializer) and reads them back. This row uses the regular node layout (vtable, evolvable). It also times the graph data type.
dagr-frozen
Dagr ("Data Graph") is a schema-driven binary format that can store shared nodes and cycles, built on an arena model. One Python DSL schema generates the code for every target language (dagr build), so there is no runtime library: the suite commits the generated code from schemas/v2/dagr/schema.py. The five suite types are emitted in all four node layouts, one row each. The graph data type is emitted only for regular and frozen, because a packed reference cannot store the person ring. On suite types, prepare builds the native value and the timed call writes every field (direct builder or arena serializer) and reads them back. This row uses the frozen node layout (positional, no evolution). It also times the graph data type.
dagr-frozen-packed
Dagr ("Data Graph") is a schema-driven binary format that can store shared nodes and cycles, built on an arena model. One Python DSL schema generates the code for every target language (dagr build), so there is no runtime library: the suite commits the generated code from schemas/v2/dagr/schema.py. The five suite types are emitted in all four node layouts, one row each. The graph data type is emitted only for regular and frozen, because a packed reference cannot store the person ring. On suite types, prepare builds the native value and the timed call writes every field (direct builder or arena serializer) and reads them back. This row uses the frozen+packed node layout (positional and inline, no evolution). It does not support the graph data type.
encoding/gob · go1.27.1
encoding/gob is Go's native binary stream for Go types. It was created so Go programs can RPC and persist values without an IDL. It is not a cross-language wire format.
encoding/json · go1.27.1
Go's encoding/json is the standard library JSON codec. It exists so every Go program can speak RFC 8259 with struct tags. This row is the baseline other Go JSON libraries try to beat. On Go 1.27 the package keeps v1 semantics. The stricter defaults are the separate encoding/json/v2 row.
encoding/json/v2 · go1.27.1
encoding/json/v2 is the Go 1.27 standard library JSON API with stricter defaults than encoding/json: invalid UTF-8 and duplicate object names are errors, and <, >, and & are not HTML-escaped. This row times Marshal / Unmarshal and MarshalWrite / UnmarshalRead.
fxamacker/cbor · 2.9.4
fxamacker/cbor is a widely used Go CBOR codec for RFC 8949. CBOR is the IETF binary JSON-like format. This library focuses on correctness options (including deterministic modes) and reusable Enc/DecMode values.
goccy/go-json · 0.10.6
goccy/go-json was written as a faster drop-in for encoding/json. The problem was stdlib JSON cost in high-QPS Go services. It keeps the same API and implements a faster encode/decode path.
goccy/go-yaml · 1.19.2
goccy/go-yaml is a high-performance YAML 1.2 library for Go. The problem was that go-yaml v2/v3 was often the slow path in config-heavy services. It aims at a faster Marshal/Unmarshal.
hamba/avro · 2.31.0
hamba/avro is a high-performance Avro library for Go. Avro was created for compact, schema-driven records. hamba focuses on a frozen API and schema cache so the timed path is encode/decode, not schema parse.
ion-go · 1.5.0
Amazon Ion was created at Amazon as a rich, self-describing superset of JSON (text and binary) for internal services. The problem was JSON's limited types. ion-go is the official Go implementation. This row times MarshalBinary / Unmarshal and the binary encoder stream. Suite structs carry json tags, not ion tags, so the wire uses the Go field names.
jsoniter · 1.1.12
json-iterator/go was created as a high-performance, stdlib-compatible JSON library for Go. The problem was the same stdlib bottleneck. It solves it with a compatible config and a faster implementation.
kelindar/binary · 1.0.19
kelindar/binary is a compact, Go-only packer. It was written for high-throughput in-process and Go-to-Go payloads where a public schema is not required. Encoder.Reset is the reuse path.
linkedin/goavro · 2.15.0
LinkedIn goavro is an Avro binary codec for Go, used in Kafka/Avro pipelines. Avro exists so producers and consumers share a schema. goavro speaks BinaryFromNative maps (OCF is a different format).
mongo-bson · 1.17.9
BSON (Binary JSON) was created for MongoDB so documents could be stored and traversed without a text parse. Official language drivers implement that spec. This row times that library's serialize/deserialize path.
pelletier/go-toml · 2.4.3
pelletier/go-toml is a TOML 1.0 parser/encoder for Go. TOML was created as an obvious config language. This library is a common Go implementation of that spec.
protobuf · 1.36.12
Protocol Buffers were created at Google so many languages could share a compact, evolving binary contract without hand-written parsers. The problem was ad-hoc binary formats and verbose XML. Protobuf solves it with an IDL, generated code, and a documented tag/length wire format.
segmentio/encoding/json · 0.5.4
segmentio/encoding/json is a production fork of a faster Go JSON stack, kept as a drop-in for encoding/json. The problem was stdlib JSON in large Go services. This row times that API.
shamaton/msgpack · 3.2.3
shamaton/msgpack is another Go MessagePack implementation, including a struct-as-array mode that omits field-name keys. The problem was MessagePack map overhead for known structs. Array mode solves that with positional fields.
shamaton/msgpack (array) · 3.2.3
shamaton/msgpack is another Go MessagePack implementation, including a struct-as-array mode that omits field-name keys. The problem was MessagePack map overhead for known structs. Array mode solves that with positional fields. This row uses struct-as-array (no field-name keys) and the matching stream helpers.
sonic · 1.15.4
Bytedance sonic was written to push Go JSON through JIT and SIMD on the hot path. The problem was that even fast reflection JSON was not enough at ByteDance scale. sonic solves it with ConfigDefault plus optional Pretouch.
ugorji/cbor · 1.3.2
ugorji/go (go-codec) was created as one Go library that speaks several formats (JSON, MessagePack, CBOR, Binc) through a shared handle/encoder model. The problem was maintaining a separate stack per format. This row times one handle of that multi-format codec. This row times the CborHandle.
ugorji/json · 1.3.2
ugorji/go (go-codec) was created as one Go library that speaks several formats (JSON, MessagePack, CBOR, Binc) through a shared handle/encoder model. The problem was maintaining a separate stack per format. This row times one handle of that multi-format codec. This row times the JsonHandle.
ugorji/msgpack · 1.3.2
ugorji/go (go-codec) was created as one Go library that speaks several formats (JSON, MessagePack, CBOR, Binc) through a shared handle/encoder model. The problem was maintaining a separate stack per format. This row times one handle of that multi-format codec. This row times the MsgpackHandle.
vmihailenco/msgpack · 5.4.1
vmihailenco/msgpack is a popular MessagePack library for Go. MessagePack exists as compact binary JSON. This implementation emphasizes a familiar Encoder/Decoder API with buffer reuse.
Call-path contract (same idea as Python/Rust)
prepare(fixture) # untimed: config, Pretouch, schema, proto convert
for rep:
serialize_bytes / stream # timed
deserialize_bytes / stream # timed (codec only)
ToDomain (if DomainConverter) # untimed (e.g. protobuf Message → model)
fidelity(expected, actual) # untimed
Caveats
- protobuf date fields may use millisecond timestamps; fidelity allows limited date-string drift where configured.
- encoding/gob and kelindar/binary are not cross-language wire formats.
- pelletier/go-toml wraps multi-instance cells as a TOML table with
items(TOML cannot use bare array roots). - Stream adapted for protobuf, linkedin/goavro, arrow-ipc, parquet, parquet-uncompressed, sbe, and the four dagr rows. OCF, gRPC, and the Arrow file format would change the wire format. The other registered Go codecs use native stream APIs. Columnar rows time the bytes API only. There is no compliance decoder.
- The four dagr rows have no Batch wrapper in their schema: multi-instance cells use the suite's cross-language frame (
u32 LE count+u32 LE len+ record, per instance — same as the Rust/C runners). Like protobuf (message built in Prepare), dagr-packed builds its direct value structs in Prepare. Unlike protobuf (ToDomainuntimed), every dagr row materializes the domain value from the lazy reader inside the timer, so their decode does strictly more work at the suite boundary. dagr-regular and dagr-frozen have no direct builder (it exists only for packed-rooted graphs): their native model is the generated arena, built in Prepare, and the timed encode is the generatedAppendTo<Graph>(b, dst, root)into one reuseddagr.Builder(it resets the builder, stores the arena, frames the record and appends it todst). - mongo-bson uses official Encoder/Decoder +
UseJSONStructTags(no JSON map bridge).
Also: go/README.md (call-path table). Serialization Categories.
Numbers
Measured numbers for this language live on the Dashboard (pre-filtered). Claim level is L1 (one machine, one session) — see Claims and replication.
Design choices
- Prepare outside the loop — configs, Pretouch, EncMode, Avro schema, protobuf messages, ugorji Handles, goavro maps.
- Optimal APIs — library-recommended encode/decode; no pretty-print; no JSON envelopes for binary codecs.
- Dual mode —
bytesandstreamwith honestStreamModemetadata (native vs adapted). - Shared domain types in
go/modelwith format struct tags for reflection codecs.