Benchmarks
This section explains how this project measures serializers, not the theory of formats themselves. On the site it lives under the Benchmarks tab. Live interactive results live under the Dashboard tab.
If Learn / Serialization 101 answers “what is serialization?”, Benchmarks answers “how do we time it fairly, and what do the numbers mean?”
You do not need to be a performance engineer to read these pages. They are written for introductory computer science students and for anyone who wants a clear picture of the suite before diving into tables and plots.
Learning goals
After reading this section you should be able to:
- Say what a benchmark is measuring here (serialize and deserialize time, size, and correctness).
- Find the right page for layout, modes, statistics, metrics, test data, or published numbers.
- Compare serializers within one language (and ideally one category) instead of treating absolute times as a global ranking.
- Open a language Overview (what libraries we measure, roster and caveats) under the Languages tab.
- Open the Dashboard for measured numbers and interactive comparisons.
For the ideas behind formats, start with Serialization 101 under Learn.
The big picture in one paragraph
Each programming language has a small program called a benchmark runner. The benchmark runner builds sample data, runs each serializer many times, checks that the data comes back correctly, and writes a spreadsheet-like file (CSV). A shared analysis program then cleans those runs, computes summaries, and produces the tables and charts you see on the site. Everyone shares the same data shapes and the same analysis rules so comparisons stay fair.
benchmark runner (C#, Python, Rust, …)
│ timed serialize + deserialize
▼
logs/<language>/*.csv
│
▼
analyze-benchmarks
│
├──► reports/stats_<lang>_latest.json (Dashboard sync input)
└──► reports/<docs_dir>/results.md (unpublished PR-diff report)
Contribute
| Task | Guide |
|---|---|
| Add one library to an existing language | Adding a serializer |
| Add a new language tree | Adding a language |
Shared terms (quick)
| Say this | Not this | Means |
|---|---|---|
| data type | “fixture” | One of message, document, telemetry, strings, event |
| benchmark runner | “harness” | Per-language program that times serializers and writes the CSV |
| batch size N | — | How many instances in one serialize call (1 or 100) |
| category / family | — | JSON, schemaless binary, schema-driven, language-native |
| I/O mode (bytes / stream) | — | How the library was called (buffer vs stream), not payload size |
| run mode (smoke / full / …) | — | How heavy the experiment is (repetitions / intent) |
Full mode guide: Modes. Data-type glossary: Test data — vocabulary.
How to read this section
Suggested order for a first visit
- Architecture — lab design and what is timed
- Timing honesty — what
preparemay do, and whatserializemust do - Modes + Test data — what the columns mean
- Methodology — warmup, filters, uncertainty
- Metrics when a column name is opaque
- Claims and replication before publishing a blog or paper claim
| Page | What you will learn | Start here if you want… |
|---|---|---|
| Architecture | How the repository is organized and how timing works | The measurement design |
| Timing honesty | What belongs in prepare vs the timed encode/decode |
“Is this row measuring a real encode?” |
| Modes | I/O modes (bytes/stream) and run modes (smoke…research) | Why the Dashboard Mode filter has two paths; which run preset to use |
| Categories | Four families of serializers (JSON, binary, schema-driven, native) | Fair “apples to apples” groups |
| Test data | The five sample data types and how sizes are chosen | What we serialize |
| Methodology | Warmup, outliers, confidence intervals, effect sizes, exploratory ranks | How CSVs become published numbers |
| Metrics | Names and meanings of every reported measurement | “What does this column mean?” |
| Claims and replication | L1 / L2 / L3: what you may claim from one run vs many | Honest generalization language |
| Adding a serializer | Drop-in checklist for one library | Author path |
| Adding a language | Checklist to plug in a new language benchmark runner | Extending the suite |
| Dashboard | Live measured numbers (Overview, Details, Compare) | The published L1 artifact |
Theory tracks: 101 · 201 · 301 · 401. Live charts: Dashboard.
Languages in this suite
Each language has a hand-written Overview (roster, caveats, how to read this runner). Measured numbers live on the Dashboard. Unpublished CLI reports land under reports/<docs_dir>/results.md for PR diffs — do not commit them to the site.
| Language | Serializers (registered) | Overview | Dashboard |
|---|---|---|---|
| C | 20 | Overview | Dashboard |
| C# | 38 | Overview | Dashboard |
| C++ | 27+ | Overview | Dashboard |
| Go | 19 | Overview | Dashboard |
| Java | 18 | Overview | Dashboard |
| JavaScript | 20 † | Overview | Dashboard |
| Kotlin | 26 | Overview | Dashboard |
| Mojo | 6 | Overview | Dashboard |
| PHP | 15 | Overview | Dashboard |
| Python | 16 | Overview | Dashboard |
| Rust | 16 | Overview | Dashboard |
| Swift | 14 | Overview | Dashboard |
| Zig | 17 | Overview | Dashboard |
† In JavaScript, simdjson is optional (it needs a native addon). If that addon fails to build, the rest of the run still continues without it.
Log language ids (the Language column in CSVs) use short names such as csharp, python, rust, c, javascript, go, java, kotlin, php, cpp, swift, zig, and mojo. Documentation folders sometimes differ (for example docs/c-sharp/ for C#).
To refresh published numbers after a local run, pack Dashboard data with python3 dashboard/scripts/sync-data.py. Regeneration and claim levels: Claims and replication.
One rule for fair comparison
Prefer:
Same language + same category + same data type + same I/O mode
Absolute times across languages (for example “Python vs C++”) mix runtimes, garbage collectors, and allocators. Those numbers can still be informative as a rough direction, but they are not a precise ranking. Details live under methodology limitations.