Files
ddidderr 039dec6c68 feat(rust): port block entropy statistics
Continue the incremental zstd_compress.c migration with its sequence
statistics and seqStore entropy-compression layer.  The following now
live in rust/src/zstd_compress_stats.rs:

- ZSTD_seqToCodes(), exported under its original name because the C
  dictionary builder (zdict.c) and decodecorpus link against it,
- ZSTD_buildSequencesStatistics() and its dummy variant, whose result
  struct no longer crosses the language boundary,
- ZSTD_entropyCompressSeqStore_internal(), _wExtLitBuffer(), and
  ZSTD_entropyCompressSeqStore(),
- ZSTD_buildBlockEntropyStats() with its literals/sequences helpers,
- ZSTD_copyBlockSequences() and the ZSTD_updateRep() rules it shares
  with the superblock writer.

Boundary: ZSTD_CCtx and ZSTD_CCtx_params stay private to C.  These
paths read exactly two parameter fields, so the C shims keep the
original static/exported function names and forward the strategy and
ZSTD_literalsCompressionIsDisabled() as int scalars alongside the
seqStore and entropy-table leaves.  The leaf layouts (SeqDef,
SeqStore_t, ZSTD_hufCTables_t, ZSTD_fseCTables_t, entropy metadata,
SeqCollector, ZSTD_Sequence) are pinned by compile-time asserts in
zstd_compress.c and by both-pointer-width layout tests in Rust.
ZSTD_buildSeqStore, block dispatch/splitting, and the block-size
estimation helpers remain C for a later slice.

The superblock module previously round-tripped through the C export of
ZSTD_buildBlockEntropyStats and mirrored the entropy leaf structs
privately.  It now calls the crate-internal builder directly, and the
shared struct definitions moved to zstd_compress_stats; consequently
ZSTD_rust_compressSuperBlock() takes the two parameter scalars instead
of an opaque ZSTD_CCtx_params pointer, extracted by its C shim.  Its
layout and repcode tests moved with the definitions.

One C helper family gets no shim: the static
ZSTD_entropyCompressSeqStore_wExtLitBuffer() had a single caller and
was folded into the Rust implementation.

Byte-identity was verified against the pre-change compressor: COPYING,
datagen -g5000000 -s7, and datagen -g300000 -s21 -P90, each at levels
1/3/9/19 and --fast=5, plus a superblock-heavy pass at level 19 with
--target-compressed-block-size=1024; all 18 frames are byte-identical,
covering the repeat-mode state machine, longOffsets, RLE/raw fallback,
and dstSize_tooSmall paths through the block splitter and superblock.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets -- -D
  warnings && cargo test --all-targets (131 tests: new reference
  vectors for seqToCodes, RLE table headers, empty-seqStore repeat
  copies, literal-stats type selection, and repcode resolution in
  copyBlockSequences)
- cargo build --release --no-default-features --features compression
  and --features decompression
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests (covers
  ZSTD_generateSequences and ZDICT training over the Rust seqToCodes)
- make -C tests test-rust-lib-smoke
- ./tests/zstreamtest -i1 and ./tests/zstreamtest --newapi -t1 -i1
- ./tests/decodecorpus -t -T1s (decodecorpus links Rust seqToCodes)
- /tmp/ref vs /tmp/got frame diff as described above: byte-identical
2026-07-11 23:11:01 +02:00
..

Rust rewrite

This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage.

Component map

The crate is organized from low-level representation helpers toward the public zstd ABI:

  • Common primitives
    • mem, bits, bitstream, and cpu implement byte-order, bitstream, and target-feature operations used by the codecs.
    • errors, debug, xxhash, and zstd_common provide common exported ABI functions and state.
    • common contains shared frame constants and internal data types.
  • Entropy coding
    • entropy_common reads FSE normalized counts and Huffman statistics.
    • fse_decompress builds FSE decoding tables and decodes FSE streams.
    • fse_compress normalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams.
    • huf_compress builds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads.
    • huf_decompress builds Huffman decoding tables and decodes X1 and X2 Huffman streams.
  • Compression primitives
    • hist counts byte frequencies for FSE and Huffman compression.
    • zstd_presplit chooses split points for full compression blocks.
    • zstd_compress_literals emits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state.
    • zstd_compress_stats converts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the public ZSTD_Sequence format. Its C shims extract the sequence store, the entropy-table leaves, and the two ZSTD_CCtx_params scalars these paths read.
    • zstd_compress_frame serializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-owned ZSTD_CCtx_params layout never crosses the language boundary.
    • zstd_compress_params owns the compression-level tables (formerly clevels.h), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the public ZSTD_* symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy.
    • zstd_fast and zstd_double_fast implement the single- and two-table fast block match finders, including attached and external dictionary paths.
    • zstd_lazy implements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants.
    • zstd_opt_tree maintains the binary-tree index used by optimal matching; the dynamic-programming optimal parser itself remains in C for now.
    • zstd_ldm implements long-distance-match parameter selection, table maintenance, sequence generation, and sequence consumption.
  • Runtime support
    • threading provides platform pthread wrappers required by zstd headers.
    • pool implements the bounded worker pool used by multithreaded compression.
  • Dictionary support
    • zstd_ddict owns, loads, copies, and references decode dictionaries.
  • Block decompression
    • zstd_decompress_block decodes literal and sequence sections, maintains FSE/Huffman repeat state, and executes compressed-block sequences.
    • zstd_decompress owns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains configuration-dependent context allocation plus legacy and trace leaves.
  • Command-line frontend
    • zstd_cli owns the Rust parser, safety policy, and dispatch. It is built by the separate cli/ static-library package only for program archives, so library builds do not acquire program-only dependencies. The C fileio backend still owns file opening, safe replacement, sparse writes, metadata, and streaming I/O.

The optimal block matcher, high-level frame compression, dictionary-building, legacy decoding callbacks, and the CLI file-I/O backend are still C. They must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result.

Compatibility boundary

The public ABI continues to come from the existing headers under lib/. Exported Rust functions therefore use C layout and calling conventions. A C source file whose implementation has moved to Rust remains in the original makefile source list as a small shim so header configuration and platform preprocessor behavior stay available during the transition.

The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation.

Validation

Run focused Rust checks from this directory:

cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo build --release

The program-only Rust archive has its own feature matrix and should be checked from rust/cli as well:

cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features compression --all-targets
cargo test --no-default-features --features decompression --all-targets

Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example:

make -C tests fuzzer
./tests/fuzzer -i1 --no-big-tests
make -C tests test-rust-lib-smoke

Broader tests/Makefile targets remain the authoritative integration gates as more of the library and CLI are rewritten.