Skip to content

Techniques

File layout

A file starts and ends with the four bytes PAR1. Between them are the row groups, then optional page indexes and Bloom filters, then the Thrift compact footer, then a 4-byte little-endian footer length, then PAR1 again.

The footer is a bare Thrift struct. Field ids are deltas packed into the high nibble of a header byte. This library writes that codec itself. It does not link a Thrift runtime.

The writer uses footer version 1. The reader accepts version 1 and version 2, which is what current PyArrow writes.

Levels

A required leaf under the root has maximum definition level 0 and maximum repetition level 0. No level bytes are stored. An optional leaf has definition level 1. A list uses the standard three-level group: an optional list group, a repeated element group, and the element. Empty lists and null lists differ by definition level, so both survive a round trip.

Encodings

Plain encoding stores values back to back. Booleans are packed from the least significant bit. Dictionary pages store plain values, and data pages store RLE indexes after a one-byte bit width. Delta encodings follow the Parquet block and miniblock layout, with 128 values per block and 4 miniblocks. The reader decodes ALP from its page header, offset array, and per-vector exceptions. The writer does not emit ALP pages.

Compression

Page compression wraps the bytes the encoding produced. Snappy, gzip, Zstd, and LZ4_RAW are implemented in Mojo. Gzip readers accept several members and check the CRC-32 and ISIZE. Zstd frames that use raw or RLE blocks are written; compressed Zstd blocks from other writers are decoded.

Statistics and indexes

When statistics are on, each column chunk records a null count and the minimum and maximum non-null values. Integer minimum and maximum are also written in the older min and max fields so older readers still see them. Floating columns record a NaN count.

Page checksums are CRC-32 of the stored page body, excluding the page header. A mismatch is a decode error. The reader accepts column indexes, offset indexes, and Bloom filters when a file contains them. The default writer does not emit those optional sections.

Limits

Limit Value
Schema depth 1024
Decompressed page or codec output 64 MiB, and the shared buffer cap is 256 MiB
Encryption Rejected
LZO and Brotli Rejected with a compression error