Skip to content

Using this suite without fooling yourself

Problem

Benchmark tables are easy to misuse. A single chart often becomes a policy decision. Someone says “library A is 3× faster than library B.” Nobody asks whether A and B implement the same job. Nobody checks the same payload shape. Nobody checks the same language. Nobody checks the same timing rules. Organizations then switch codecs. They observe little improvement. They conclude that “benchmarks lie.” The real issue was misaligned comparison.

This multi-language suite is built to support fair, local comparisons. It answers a question inside one language, one standard, one data set, and one data type. In this section you will learn how to read the numbers as a first-year student should. Read them as answers to carefully stated questions.


Short answer

Treat every published number as the answer to a narrow question. For a given language, standard, data set, and data type, how do registered serializers compare? Compare encode time, decode time, size, and related metrics. Do that after the analysis pipeline’s warmup and optional outlier rules.

The Dashboard filter bar is that question. Language, Standard, Data set, and Data type are the cell. Data set is Suite or Columnar. The Suite types are message, document, telemetry, strings, and event. Published averages such as all@all use only those five types. The Columnar types are table, table_project, nested_table, and signal. On table_project, serialize writes the full row and deserialize reads only f_float_0.

The Dashboard sets the data set from the standard. Where a language registers them on both sets, JSON, Avro, Protocol Buffers, and FlatBuffers still need a matching data type. Arrow IPC, Parquet, and ORC are Columnar only. SBE (Simple Binary Encoding) is Columnar only as a data set. It is schema-driven. Each record is one stride: the bytes from the start of that record to the start of the next. Variable-length fields change that distance. SBE is not a columnar file format like Parquet.

A family is a teaching cut. It groups formats that solve roughly the same product job: JSON text, schemaless binary, schema-driven, language-native, or columnar. See Serialization categories. A family is not a filter. A library is one implementation of a standard. The filter is the standard.

A cross-language or cross-standard chart needs the workload stated again before it becomes architecture policy. When the decision is about trust, evolution, or multi-hop design, suite timings are inputs. They are one part of the argument. See the rest of Serialization 301.


Constraints that matter

Constraint Why it breaks naive rankings
Language and runtime Different virtual machines, garbage collectors, and standard libraries; “Protobuf in Python” is not “Protobuf in C.”
Standard JSON and YAML are both readable text. They are different contracts. Protocol Buffers and Avro are both schema-driven. They are different contracts.
Data set and data type Suite message and Columnar table are different jobs, including inside one standard that has both.
Payload shape Dense structs versus deep graphs change cost centers (see Encode/decode cost).
Implementation Several libraries can share a format label and differ by an order of magnitude.
What is timed Benchmark runner paths measure serialize and deserialize of prepared fixtures—not network round-trip time, disk I/O, or your production validation layer.
Analysis policy Warmup exclusion and outlier filters change means; raw CSV is not the published table unless you re-run analysis with the same config.

Warmup means early iterations that may be slower while the runtime heats up. That includes just-in-time (JIT) compilation—the runtime optimizes frequently used code while the program runs—and caches. Analysis often drops those so the table reflects steady state. It does not reflect cold start.

This matters because a mixed chart can look like a tournament. The honest use is a same-language, same-standard shortlist.


Decision frame

Use this checklist before quoting a Dashboard number:

  1. Same language? If no, stop. Use the numbers only as rough orientation. Do not use them as a pick.
  2. Same standard? This is the Dashboard Standard control. Cross-standard only when the product question is which contract. Then speed is one axis among others. The family is the teaching cut that led you to the standard. See Serialization categories.
  3. Same data set? Suite or Columnar. The Dashboard sets this from the standard. When a standard has both, read the label and then check the data type.
  4. Same data type? message, document, table, and table_project are different jobs. table_project writes every column and reads only f_float_0.
  5. Which metric, and which sample policy? Mean encode, mean decode, size, operations per second, or tails. On Overview the toggle is Ops/Sec or Latency. On Compare the metric row is separate. The Samples control (for example IQR 1.5) decides which runs enter the chart. Pick the one your reliability target (service-level objective) cares about. State it with the number.
  6. Still missing? Compression on the wire, authentication, schema registry behavior, and multi-hop delivery of one message to many consumers (fan-out) sit outside the core tables. Design a separate experiment for those.

Bytes versus stream is a caveat on languages that still publish a stream row. Keep that row off the parent-row chart. It is a language Overview note, and it is a separate axis from the four filters above.

  Question: “Is X better than Y for us?”
        │
        ▼
  Fix language + standard + data set + data type + metric
        │
        ▼
  Read the Dashboard for that slice only
        │
        ▼
  Re-check product constraints (trust, evolution, operations)
        │
        ▼
  Decide — or design a measurement this suite cannot do

In other words: fix the experimental cell first. Then read numbers inside that cell. Then bring product constraints back into the decision.


Failure modes

Mistake What goes wrong
Global leaderboard thinking Picking the top row of a mixed chart as “the company format.”
Cross-language crowning Mandating a library because it looked fast in another runtime.
Format brand equals one speed “We moved to binary” without naming the implementation.
Ignoring warmup and cold path Production cold starts differ from steady-state tables (analysis drops RepetitionIndex == 0 by default).
Confusing size with latency The smallest payload on a local network may not win 99th-percentile latency (p99: 99% of requests are faster than this) if CPU or allocations dominate.
Treating fidelity notes as optional Some codecs are registered with documented shape limits; Overview caveats bound the claim.
Policy from means only Garbage-collection pauses and tail latency may not appear as a single mean encode time.

Real-world sketch

A team sees that a schema-driven library is fastest on the Rust Dashboard for a dense fixture. The team mandates it for a public multi-language HTTP API. Clients are browser and mobile. Operators need human-readable debug logs. The public contract is already JSON.

The suite result answered “fastest schema path in Rust for this fixture.” It did not answer “best public API contract.” A better use of the suite is to compare JSON implementations within each language that must speak the public contract. Measure schema-driven codecs only on internal hops that already accept an IDL.


In this suite

Resource Use it for
Serialization categories Families as orientation, then the named standard
Language Overview Registered names, categories, fidelity caveats (inventory source of truth)
Dashboard Published timings and sizes (filter by language)
Analysis methodology Warmup, outliers, grouping keys, units
Metrics catalog What each field means
Test Data Fixture meanings and size knobs
Benchmark architecture What the benchmark runner times
Dashboard (top-level Dashboard tab) Interactive slices of the same analysis story

Grouping key for fair peers (conceptually):
(Language, Standard, Data set, Data type) — then compare SerializerName rows inside that cell.

Illustrative only: prose in theory pages must not invent winners. When you need a number, open the Dashboard for the language you will actually run.


Experiments

Question: Am I about to quote a fair suite comparison for a real decision, or a misaligned chart?

Setup

  1. Write the decision question in one sentence. One example is “Which JSON library in Python for message-shaped RPC?”
  2. Open categories and the language Overview / Dashboard for the runtime you will ship.
  3. Note analysis configuration from methodology if you will re-derive tables. Include warmup and outlier policy.

Procedure

  1. Apply the checklist in Decision frame. Check language, then standard, then data set, then data type, then metric.
  2. Pull only the matching rows from the Dashboard. Leave cross-standard and cross-language ranks out of the policy claim.
  3. Record which metric column you will use. Examples include encode versus decode versus size versus operations per second.
  4. List product constraints the suite does not measure. Examples include trust, registry, and network RTT.
  5. Either decide from that slice or design an out-of-suite experiment for the missing constraints.

Decision rule

  • If any checklist box fails, do not use the number as architecture policy.
  • If the slice is fair but the product question is not about performance, treat suite data as supporting. Do not treat it as decisive.

Metrics

Metric / signal Role
Comparison validity (same language, standard, data set, data type) Primary gate. Binary pass/fail before any number.
Chosen reliability-target (SLO) metric (for example decode median or size) The one number allowed in the argument
total_median_ns / ser_median_ns / deser_median_ns Default speed ranks on the Dashboard
median_size_bytes Density and bandwidth axis
mean_fidelity Fixture round trip (write, then read). Not specification compliance. Non-faithful rows are out.
serializer_version Reproducibility of the claim
runs, warmup, outliers removed Trust in the statistic
Dashboard or CSV filter state Document what you hid

Conclusion style: “Under Python, standard JSON, Suite, message at n=1, A beats B on deserialize median; size is similar; fidelity is 1.0.”

Outside the claim: a chart that mixes standards, and a cross-language champion.


What this suite cannot tell you

  • End-to-end service latency. That includes queueing, network, TLS, and framework overhead.
  • Correctness of your schema evolution policy when services are updated gradually so old and new versions run at the same time. See 201 schema evolution and later 301 contract articles.
  • Security of deserializing untrusted bytes. That includes native formats and hostile inputs.
  • Whether gzip or zstd on the wire wins for your message size mix. Compression is orthogonal. See 201 compression vs format.
  • Cross-language byte identity for every registered pair. Benchmark runners are per-language unless you design a fidelity experiment.
  • Business constraints. Those include compliance, team skill, vendor lock-in, and existing public contracts.

Common mistakes

  • Screenshotting one latency distribution into an architecture decision record without stating language, standard, and data type.
  • Averaging ranks across languages “to be fair.”
  • Changing fixture generation parameters and comparing to old published snapshots without regenerating both sides.
  • Using language-native serializers’ speed as an argument for network interchange.

Key takeaways

  • Suite numbers answer narrow, local questions. They do not answer “what format should the industry use.”
  • Fix language, standard, data set, data type, and metric before comparing serializers.
  • Implementation quality and payload shape often dominate format brand.
  • Analysis policies are part of the claim. Know warmup and outlier rules.
  • Use Dashboard numbers as evidence inside a larger 301 decision about trust, contracts, and workload.
  • Explicitly list what you still must measure outside this benchmark runner.