Continue the incremental zstd_compress.c migration with its sequence statistics and seqStore entropy-compression layer. The following now live in rust/src/zstd_compress_stats.rs: - ZSTD_seqToCodes(), exported under its original name because the C dictionary builder (zdict.c) and decodecorpus link against it, - ZSTD_buildSequencesStatistics() and its dummy variant, whose result struct no longer crosses the language boundary, - ZSTD_entropyCompressSeqStore_internal(), _wExtLitBuffer(), and ZSTD_entropyCompressSeqStore(), - ZSTD_buildBlockEntropyStats() with its literals/sequences helpers, - ZSTD_copyBlockSequences() and the ZSTD_updateRep() rules it shares with the superblock writer. Boundary: ZSTD_CCtx and ZSTD_CCtx_params stay private to C. These paths read exactly two parameter fields, so the C shims keep the original static/exported function names and forward the strategy and ZSTD_literalsCompressionIsDisabled() as int scalars alongside the seqStore and entropy-table leaves. The leaf layouts (SeqDef, SeqStore_t, ZSTD_hufCTables_t, ZSTD_fseCTables_t, entropy metadata, SeqCollector, ZSTD_Sequence) are pinned by compile-time asserts in zstd_compress.c and by both-pointer-width layout tests in Rust. ZSTD_buildSeqStore, block dispatch/splitting, and the block-size estimation helpers remain C for a later slice. The superblock module previously round-tripped through the C export of ZSTD_buildBlockEntropyStats and mirrored the entropy leaf structs privately. It now calls the crate-internal builder directly, and the shared struct definitions moved to zstd_compress_stats; consequently ZSTD_rust_compressSuperBlock() takes the two parameter scalars instead of an opaque ZSTD_CCtx_params pointer, extracted by its C shim. Its layout and repcode tests moved with the definitions. One C helper family gets no shim: the static ZSTD_entropyCompressSeqStore_wExtLitBuffer() had a single caller and was folded into the Rust implementation. Byte-identity was verified against the pre-change compressor: COPYING, datagen -g5000000 -s7, and datagen -g300000 -s21 -P90, each at levels 1/3/9/19 and --fast=5, plus a superblock-heavy pass at level 19 with --target-compressed-block-size=1024; all 18 frames are byte-identical, covering the repeat-mode state machine, longOffsets, RLE/raw fallback, and dstSize_tooSmall paths through the block splitter and superblock. Test plan: - cd rust && cargo fmt --check && cargo clippy --all-targets -- -D warnings && cargo test --all-targets (131 tests: new reference vectors for seqToCodes, RLE table headers, empty-seqStore repeat copies, literal-stats type selection, and repcode resolution in copyBlockSequences) - cargo build --release --no-default-features --features compression and --features decompression - make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests (covers ZSTD_generateSequences and ZDICT training over the Rust seqToCodes) - make -C tests test-rust-lib-smoke - ./tests/zstreamtest -i1 and ./tests/zstreamtest --newapi -t1 -i1 - ./tests/decodecorpus -t -T1s (decodecorpus links Rust seqToCodes) - /tmp/ref vs /tmp/got frame diff as described above: byte-identical
Rust rewrite
This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage.
Component map
The crate is organized from low-level representation helpers toward the public zstd ABI:
- Common primitives
mem,bits,bitstream, andcpuimplement byte-order, bitstream, and target-feature operations used by the codecs.errors,debug,xxhash, andzstd_commonprovide common exported ABI functions and state.commoncontains shared frame constants and internal data types.
- Entropy coding
entropy_commonreads FSE normalized counts and Huffman statistics.fse_decompressbuilds FSE decoding tables and decodes FSE streams.fse_compressnormalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams.huf_compressbuilds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads.huf_decompressbuilds Huffman decoding tables and decodes X1 and X2 Huffman streams.
- Compression primitives
histcounts byte frequencies for FSE and Huffman compression.zstd_presplitchooses split points for full compression blocks.zstd_compress_literalsemits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state.zstd_compress_statsconverts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the publicZSTD_Sequenceformat. Its C shims extract the sequence store, the entropy-table leaves, and the twoZSTD_CCtx_paramsscalars these paths read.zstd_compress_frameserializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-ownedZSTD_CCtx_paramslayout never crosses the language boundary.zstd_compress_paramsowns the compression-level tables (formerlyclevels.h), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the publicZSTD_*symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy.zstd_fastandzstd_double_fastimplement the single- and two-table fast block match finders, including attached and external dictionary paths.zstd_lazyimplements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants.zstd_opt_treemaintains the binary-tree index used by optimal matching; the dynamic-programming optimal parser itself remains in C for now.zstd_ldmimplements long-distance-match parameter selection, table maintenance, sequence generation, and sequence consumption.
- Runtime support
threadingprovides platform pthread wrappers required by zstd headers.poolimplements the bounded worker pool used by multithreaded compression.
- Dictionary support
zstd_ddictowns, loads, copies, and references decode dictionaries.
- Block decompression
zstd_decompress_blockdecodes literal and sequence sections, maintains FSE/Huffman repeat state, and executes compressed-block sequences.zstd_decompressowns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains configuration-dependent context allocation plus legacy and trace leaves.
- Command-line frontend
zstd_cliowns the Rust parser, safety policy, and dispatch. It is built by the separatecli/static-library package only for program archives, so library builds do not acquire program-only dependencies. The Cfileiobackend still owns file opening, safe replacement, sparse writes, metadata, and streaming I/O.
The optimal block matcher, high-level frame compression, dictionary-building, legacy decoding callbacks, and the CLI file-I/O backend are still C. They must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result.
Compatibility boundary
The public ABI continues to come from the existing headers under lib/.
Exported Rust functions therefore use C layout and calling conventions. A C
source file whose implementation has moved to Rust remains in the original
makefile source list as a small shim so header configuration and platform
preprocessor behavior stay available during the transition.
The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation.
Validation
Run focused Rust checks from this directory:
cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo build --release
The program-only Rust archive has its own feature matrix and should be checked
from rust/cli as well:
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features compression --all-targets
cargo test --no-default-features --features decompression --all-targets
Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example:
make -C tests fuzzer
./tests/fuzzer -i1 --no-big-tests
make -C tests test-rust-lib-smoke
Broader tests/Makefile targets remain the authoritative integration gates as
more of the library and CLI are rewritten.