Serialization 401: Implementing Contemporary Serializers
This course teaches how Protocol Buffers works at the byte level. You will also learn how three language libraries turn those bytes into ordinary values and back again. A later set of articles then opens the timed call sites of the libraries that lead on this suite’s document fixture, one language at a time, and follows those calls into the library source.
| Jump | |
|---|---|
| Prereqs | 101 · 201 (schema / wire ideas) |
| Sibling | 301 production judgment — whether / which, not how bytes |
| Suite | Add a serializer · Dashboard |
Protocol Buffers is one standard in this suite, and a popular schema-driven binary format. You describe message shapes in a .proto file. Tools then generate code that can encode and decode those messages. This elective studies that one standard down to the bytes. The compliance catalog lists the other standards. The wire article comes first, then the paths taken by the Python, Rust, and C libraries. A small hands-on lab connects the theory to code you write yourself. After that, ten language articles compare the libraries that lead on this suite’s document fixture by reading their timed functions, not by repeating the 201 format essays.
The course is for people who want to implement, debug, or deeply integrate codecs. It is not only for people who choose a format from a list of options. You do not need to have written a serializer from scratch already. You do need intermediate reading comfort in at least one of the languages used in this suite. You also need a working memory of schema-dependent binary ideas from Serialization 201.
Who this is for
Take this elective if you need to implement, debug, or deeply integrate serializers. It comes after Serialization 201. It is deliberately more hands-on than the production-judgment track. It does not replace Serialization 301.
Course 301 is about multi-constraint product choices under real pressure. Course 401 is about wire rules and runtime paths. You will learn what each byte means. You will also learn how a library walks from a language value to those bytes and back.
In short, 401 teaches the how of encoding and decoding. 301 teaches the whether and which of production decisions.
Prerequisites
| Type | Requirement |
|---|---|
| Hard | 101 and 201 (schema identity, encode cost, evolution, dynamic vs IDL binary) |
| Soft | 301 is recommended (trust boundaries, multi-language systems, honest measurement) |
| Skills | Intermediate reading level in at least one of Python, Rust, or C |
Learning outcomes
By the end of this course you should be able to:
- Encode and decode core Protocol Buffers wire structures on paper and with tables. That includes tags (small keys that name which field is next), varints (variable-length integers), length-delimited fields, nested messages, simple repeated fields, and the rule for skipping unknown fields.
- Trace encode and decode paths in Python (
google.protobuf), Rust (prost), and C (protobuf-c). You should also explain buffer ownership: who allocates the bytes, who frees them, and how long a buffer must stay valid. - Contrast the classic C runtime (protobuf-c) with the embedded-oriented design of nanopb. Nanopb is a C library that prefers static size budgets over free-form heap trees.
- Build a mini subset encoder/decoder. Validate it against golden byte sequences and at least one official parser.
- State deliberate omissions honestly. This lab is a teaching subset. It is not a full production Protocol Buffers implementation.
- Read two library call paths in one language and explain, from the source, why one finishes more encode-and-decode cycles per second (or writes fewer bytes) on this suite’s document fixture.
How this course fits the program
| Course | Role |
|---|---|
| 101 | Foundations |
| 201 | Mechanisms |
| 301 | Production judgment (core advanced track) |
| 401 (this course) | Implementer elective — wire format, language paths, and a thin subset lab |
This course teaches wire encoding, runtime paths, and a thin subset lab. It is not a full reimplementation of Protocol Buffers. It is also not a multi-constraint product-choice guide. That role belongs to 301.
Modules
The table below lists every article in this course. For each article, it states what you should be able to do after reading it.
| Article | You should be able to… |
|---|---|
| Protobuf wire format step-by-step | Read and emit tags, varints, length-delimited (LEN) fields, nested messages, unpacked and packed repeated fields; skip unknowns safely |
| Lab: mini encoder/decoder | Build a MiniUser subset codec; pass goldens G1–G5, bounds tests, and an official-parser check |
| Python: google.protobuf path | Trace codegen → backend → SerializeToString / ParseFromString and who owns the bytes |
| Rust: prost path | Trace encoded_len / encode_raw / merge_field and monomorphized (per-type specialized) codegen |
| C: protobuf-c path | Trace descriptor-driven pack/unpack and heap free discipline |
| C: nanopb vs protobuf-c | Choose a heap-friendly C library versus a static-budget C library for a deployment |
| Same bytes, three runtimes | Design interop matrix tests; separate bit-identical encodings from logical fidelity |
Language comparisons in code
These articles do not repeat the 201 “text versus binary” essays. Each one opens two timed call sites in this repository, then follows those calls into the library. The fixture is document, one instance, unless the article says otherwise. Measured numbers live on the Dashboard. Both call sites in a pair must follow the same timing contract.
The Standard column is the contract each side implements. When both sides share a standard, the gap is the library. When the column names two standards, the gap mixes the encoding and the library. No public spec is the Dashboard bucket for a library with no public spec. A same-standard “Open this slice” link sets that Standard. A link that names two standards opens on All, so both columns stay on the chart.
The first ten pages take the speed or size leader in each language and ask why it leads. The later pages hold one variable still: same library and two encodings, same JSON and two libraries, same Protocol Buffers bytes and three JavaScript libraries, an in-place crate used as a classical decoder, a Dashboard row that times a different library than the row name, and the Kotlin pairs that keep JSON or keep a schema format while changing the library.
| Article | Standard | You should be able to… |
|---|---|---|
| Python: msgspec-msgpack vs orjson | MessagePack and JSON | Show why a positional MessagePack Struct decodes faster than Rust JSON over dictionaries |
| Python: msgspec JSON vs MessagePack | JSON and MessagePack | Hold the library still and isolate JSON tokens from MessagePack type codes |
| Python: orjson vs json | JSON | Show why the same 448-byte JSON can differ by a factor of five |
| Rust: Speedy vs Bincode | No public spec | Show why generated write_to plus fixed-width integers is faster than Serde plus variable-length integers |
| Rust: Speedy vs Postcard | No public spec | Show compactness as a width choice that can stay on Serde |
| Rust: rkyv vs Speedy | No public spec | Show that in-place access only helps if the timed path uses it |
| C: custom-binary vs ubj | No public spec and UBJSON | Show that ubj is the same packed record plus a 37-byte envelope and a second copy |
| C++: Bitsery vs YAS | No public spec | Show why one-byte lengths and a reused buffer beat eight-byte lengths and a 20 KiB stream |
| C++: the simdjson row | JSON | Read a Dashboard row whose encode is nlohmann dump and whose decode is simdjson parse plus a DOM walk |
| C#: BinaryPack vs Bond Fast | No public spec and Bond | Show positional IL stores versus a type-and-identifier prefix on every field |
| Go: kelindar/binary vs hamba/avro | No public spec and Avro | Show two cached positional plans, and why the Avro schema walk costs a little more |
| Java: Protostuff vs protobuf-java | No public spec and Protocol Buffers | Show why equal 155-byte messages still differ: POJO merge versus generated parseFrom |
| Kotlin: Protostuff vs protobuf | No public spec and Protocol Buffers | Show the same 155-byte pair on Kotlin, and why the Kotlin DSL row matches protobuf-java |
| Kotlin: kotlinx-json vs Moshi | JSON | Show why the same 440-byte JSON can differ between kotlinx.serialization and a generated Moshi adapter |
| Kotlin: FlatBuffers vs protobuf | FlatBuffers and Protocol Buffers | Show why vtable loads can beat a smaller Protocol Buffers stream on the Kotlin harness |
| PHP: JSON vs protobuf | JSON and Protocol Buffers | Show that PHP’s native JSON engine is faster than official userland protobuf, even though JSON is larger |
| Fortran: json-fortran vs rojff | JSON | Show two pure-Fortran JSON call sites, and why integers past int32 become JSON numbers |
| Zig: std.json vs serde.json | JSON | Show two comptime JSON call sites on the same suite document |
| Mojo: EmberJson vs mojo-avro | JSON and Avro | Show JSON text vs Avro binary on the same suite document |
| JavaScript: JSON vs google-protobuf | JSON and Protocol Buffers | Show that V8’s native JSON path is faster than a real JavaScript Protocol Buffers encode and decode, even though JSON is larger |
| JavaScript: three Protocol Buffers libraries | Protocol Buffers | Compare google-protobuf, protobufjs, and protobuf-es on the same 155 bytes |
| Swift: FlatBuffers vs SwiftProtobuf | FlatBuffers and Protocol Buffers | Show why vtable loads can beat a smaller Protocol Buffers stream |
Suggested path. This order matches the self-check below. The side navigation lists the same pages.
- Wire format
- Lab (start as soon as the wire article is readable)
- Python → Rust → C protobuf-c
- nanopb compare
- Cross-language fidelity
- One language-comparison article in a language you read fluently (table above)
The flagship schema in this benchmark suite is schemas/v2/protobuf/benchmark_v2.proto. Teaching pages intentionally use a much smaller message called MiniUser. MiniUser is not the suite schema. It exists so you can study hex dumps without drowning in fields.
Lab notebooks (Python / Colab)
| Notebook | Use with |
|---|---|
| Wire format playground | Wire format |
| MiniUser encoder lab | Lab article |
Index and install notes live in the notebooks README.
If you want the same golden hex sequences G1–G5 in another language, see the multi-language homework companions: companions/go · companions/rust.
Four Protocol Buffers libraries at a glance
In this section we compare how four libraries relate to the same binary layout.
The wire format is the layout of field numbers, wire types, and payloads on the byte stream. That layout is shared. Python, Rust, and C can all encode and decode the same Protocol Buffers bytes. What differs is codec engineering. Each library walks the schema in its own way. Each library allocates buffers differently. Each library reports errors differently.
Python (google.protobuf) |
Rust (prost) |
C (protobuf-c) |
C (nanopb) | |
|---|---|---|---|---|
| Schema at encode time | Runtime descriptors plus a backend (upb or pure Python) | Monomorphized per-type code (specialized at compile time for each message type) | Runtime descriptor tables | Field list plus static maximum sizes |
| Output ownership | New immutable bytes object |
Caller-owned Vec<u8> |
Caller-allocated buffer | Static buffer or stream budget |
| Decode ownership | Garbage-collected Message object | Owned Rust struct | Heap message that you must free with free_unpacked |
Preallocated static struct |
| Typical failure | Parse error raised as a Python exception | DecodeError |
NULL return, or a leak if you skip free |
Encode/decode fails when data exceeds a configured max |
Details appear in the language-path articles and in nanopb compare.
Honesty rules
The program-wide rules still apply. There are no universal winners. Implementation quality beats brand name. The Dashboard owns measured numbers.
401-specific honesty:
- Wire truth is shared; runtimes differ. Python, Rust, and C can all encode and decode the same binary layout. They can still own buffers differently.
- The subset lab labels its omissions. Packed repeated fields, zigzag signed integers, maps, oneofs, and full production hardening are out of scope on purpose.
- Suite benchmark runners illustrate integration. They are not the reference design for how you should structure production Protocol Buffers.
- Dashboard numbers are optional cost context. Speed tables are not the focus of the Protocol Buffers sequence. The language-comparison articles quote one L1 slice so you can attach a number to a line of code. They do not name a universal library.
- Hostile input is a 301 topic. For operational controls on untrusted payloads, see 301 untrusted input. Codec-side bounds still belong in every decoder. Examples include truncated varints and overlong lengths.
- Language tours are parallel, not ranked. This course does not crown “Rust wins.”
Assessment (self-check)
Use the following checklist to test yourself after you finish the modules.
- Complete the lab golden vectors G1–G5 (including the empty G2). Also complete unknown-field skip, bounds failures, and at least one official-parser cross-check.
- Explain pack/unpack ownership in one of Python, Rust, or C. Who allocates? Who frees? What must stay valid during the call?
- State when nanopb is preferable to protobuf-c, and when the reverse is true. See nanopb compare.
- Design a three-language encode/decode matrix test. Say when bit-identity (
memcmpof encodings) is required. Say when logical equality is enough. See cross-language fidelity. - Open one language-comparison article. Quote the two timed functions. State whether the speed gap is an encoding difference (the bytes on the wire), an implementation difference (how the library writes those bytes), or a runner difference (work moved into untimed
prepare). When both call sites share a standard, the gap is the library. When the Standard column names two contracts, the encoding is part of the gap.
Where to go next
- Serialization 201 if schema-dependent concepts feel rusty.
- Serialization 301 for multi-constraint product choices.
- Shared suite schema in this repository:
schemas/v2/protobuf/benchmark_v2.proto.