Files
zstd-rs/rust
ddidderr de6c837d7f fix(test): use direct zero-filled dictionary probe storage
Use vec![0; size] in the local-dictionary allocator probe so the focused
Rust tests satisfy the repository's strict clippy configuration. The probe
still owns the allocation until the test releases it and production code is
unchanged.

Test Plan:
- git diff --check
- ulimit -v 41943040; CARGO_BUILD_JOBS=1 cargo clippy --manifest-path rust/Cargo.toml --all-targets -- -A clippy::manual-bits -D warnings
2026-07-21 21:11:10 +02:00
..

Rust rewrite

This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage.

Component map

The crate is organized from low-level representation helpers toward the public zstd ABI:

  • Common primitives
    • mem, bits, bitstream, and cpu implement byte-order, bitstream, and target-feature operations used by the codecs.
    • errors, debug, xxhash, and zstd_common provide common exported ABI functions and state.
    • common contains shared frame constants and internal data types.
  • Entropy coding
    • entropy_common reads FSE normalized counts and Huffman statistics.
    • fse_decompress builds FSE decoding tables and decodes FSE streams.
    • fse_compress normalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams.
    • huf_compress builds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads.
    • huf_decompress builds Huffman decoding tables and decodes X1 and X2 Huffman streams.
  • Compression primitives
    • hist counts byte frequencies for FSE and Huffman compression.
    • zstd_presplit chooses split points for full compression blocks.
    • zstd_compress_literals emits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state.
    • zstd_compress_stats converts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the public ZSTD_Sequence format. Its C shims extract the sequence store, the entropy-table leaves, and the two ZSTD_CCtx_params scalars these paths read.
    • zstd_compress_block_split searches for profitable sequence-store partitions, while zstd_compress emits those partitions through the Rust single-block serializer. C retains split discovery's context setup and the outer block-dispatch decision. zstd_compress also owns ordinary sequence-block entropy emission, sequence collection, the legacy RLE compatibility gate, sequence-store construction, and its branch/fallback policy; C supplies only the private matchfinder, LDM block-preparation/consumption, external-sequence-producer, and state-preparation callbacks. Rust also owns the frame-chunk block loop, including block sizing, target/split/internal dispatch, output accounting, and frame-state updates; C supplies the private block-compression callbacks. Rust also owns the single-threaded buffered/stable stream state machine, including direct versus buffered output, pending-output draining, and frame reset policy. The external-sequence-and-literals block loop and public sequence conversion are Rust-owned as well; Rust also owns overflow-correction branch/order and the window rebase while C retains private workspace and index callbacks plus CCtx-facing adapters.
    • zstd_compress_frame serializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-owned ZSTD_CCtx_params layout never crosses the language boundary.
    • zstd_compress also owns the transparent single-threaded and multithreaded stream-initialization policy: dictionary selection, parameter resolution, initial buffer sizing, and the ordered setup decisions are projected into Rust while C retains the private contexts and mutation callbacks. The public sequence APIs likewise use Rust-owned validation, frame-header, checksum, and output-accounting orchestration around C-owned block state. The public end-of-frame path, the ZSTD_compress2_c fallback, and the ZSTD_compressStream2_c fallback also use Rust-owned orchestration boundaries while C retains the context reset, stream adapter, checksum/epilogue, and trace callbacks. The stream2 fallback keeps transparent initialization/defer, buffer validation, serial-versus-MT dispatch, and result accounting in Rust; C supplies only the private CCtx callbacks and projected state. Public advanced one-shot compression now uses the same Rust-owned begin-then-end ordering boundary while C retains the private begin/end callbacks. Public frame-progression dispatch is Rust-owned as well; C projects the scalar single-thread inputs and keeps the private MT context behind a callback. Compression-size estimation also uses Rust-owned row-matchfinder mode selection and compression-level budget progression around C-owned parameter construction and workspace-size callbacks.
    • zstd_compress_params owns the compression-level tables (formerly clevels.h), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the public ZSTD_* symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy.
    • zstd_fast and zstd_double_fast implement the single- and two-table fast block match finders, including attached and external dictionary paths.
    • zstd_lazy implements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants.
    • zstd_opt_tree maintains the binary-tree index used by optimal matching; zstd_opt owns the dynamic-programming price model, optimal parse, and sequence emission.
    • zstd_ldm implements long-distance-match parameter selection, table maintenance, sequence generation, sequence consumption, and sequence-store encoding.
  • Dictionary building
    • divsufsort constructs the suffix array that drives the legacy ZDICT trainer (ZDICT_trainFromBuffer_legacy); dict_builder_zdict, dict_builder_cover, and dict_builder_fastcover own the sample analysis, training, and dictionary assembly. The corresponding C translation units are declaration-only ABI shims.
  • Runtime support
    • threading provides platform pthread wrappers required by zstd headers.
    • pool implements the bounded worker pool used by multithreaded compression.
    • The zstdmt_compress integration keeps job descriptors and synchronization private to C while Rust owns input-retention scans, reusable input-range overlap decisions, outer scheduling and end-directive adjustments, job-creation decisions, compression-job stage sequencing and error flow, pending-output decisions through scalar job projections, and frame progression's job-ring scan, error normalization, and active-worker accounting. Rust also owns frame-block preparation ordering and the MT serial turn/skip policy, including LDM-before-checksum sequencing; C callbacks retain the private match-state window, workspace operations, synchronization, LDM state, and checksum state. Rust also owns serial-reset LDM normalization/table sizing/publication, final MT completion ordering, and normal job-table teardown; C retains the private job storage, context mutation, and synchronization callbacks.
  • Dictionary support
    • zstd_ddict owns, loads, copies, and references decode dictionaries.
  • Legacy decoding
    • legacy hosts one frozen module per historical format; legacy::zstd_v01 through legacy::zstd_v07 port the seven self-contained historical decoders. Their original C translation units remain declaration-only shims for the native build.
  • Block decompression
    • zstd_decompress_block decodes literal and sequence sections, maintains FSE/Huffman repeat state, executes compressed-block sequences, and selects the short or long sequence decoder from projected configuration and history state. C retains the decoder-context layout and configuration projection; the public and fullbench block-decoder wrappers are Rust-owned.
    • zstd_decompress owns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains the configuration-dependent context layout and platform details, plus legacy and trace leaves; Rust owns the default-window, no-forward-progress, and heap-mode policies, decoder storage allocation, custom memory dispatch, and the ZSTD_copyDCtx prefix copy up to the projected inBuff field.
  • Command-line frontend
    • zstd_cli owns the Rust parser, safety policy, and dispatch. It is built by the separate cli/ static-library package only for program archives, so library builds do not acquire program-only dependencies. The C fileio layer retains the format-specific codec callbacks, private asynchronous-pool adapters, metadata, zstd codec/error mapping, and the adaptive progression and diagnostic callbacks, while Rust owns the mixed-format probe/dispatch loop, zstd stream-compression I/O loop, scalar adaptive decisions, optional-format decompression loops, decompression result policy/final accounting, both shared- and separate-destination multi-file compression schedulers, and both shared- and separate-destination decompression schedulers. Rust also owns the per-file --list status/order policy, the scalar compression metadata-transfer and file-removal status policies; C retains the private stat probe, metadata operations, destination-name construction, and the private file/resource/format callbacks for these scheduler boundaries. Rust already owns the file preference policy, filename decisions, source/destination opening and destination retry ordering, dictionary buffers, asynchronous I/O pools, and pass-through copy leaf.
    • timefn provides the monotonic nanosecond clock behind UTIL_time_t, while benchfn owns the benchmark run/timing loop (BMK_benchFunction, BMK_benchTimedFn) and benchzstd owns benchmark orchestration and reporting. Both live in Rust; the C translation units are ABI shims. C test binaries (fullbench, fuzzer, zstreamtest, paramgrill, ...) link a helpers-only build of the archive, produced without the package's cli feature, because the parser layer requires the C fileio backend that tests do not compile.

Dictionary-ingestion dispatch, CCtx dictionary/prefix attachment dispatch, the CLI zstd compression stream loop, scalar adaptive decisions, and the shared- and separate-destination multi-file compression schedulers, shared- and separate-destination decompression scheduling, single-threaded stream initialization and the buffered/stable stream state machine, MT stream initialization, MT serial-reset LDM policy and publication, MT outer scheduling and flush policy, MT completion trace/reset ordering, MT compression-job stage sequencing and error flow, MT frame-progression job aggregation, frame-block preparation ordering, MT serial turn/skip and LDM/checksum sequencing, overflow-correction policy/order, public sequence-API orchestration, sequence-store and block policy, external-producer invocation and success-path validation, external-sequence-store reset, external-sequence/literals block loop, optional-format decompression loops, block decoder wrappers, decompression result policy, and public end-of-frame compression orchestration, public CCtx-copy pledge/frame-parameter policy, public CCtx allocation validation and allocate-before-init ordering, public static-CCtx workspace validation and initialization dispatch, and public CDict constructor parameter selection, default-level normalization, and workspace/object teardown ordering, public static-CDict workspace validation and initialization dispatch, public advanced-CDict one-shot validation and begin/end ordering, advanced-CDict parameter selection, dedicated-search fallback, row-matchfinder resolution, and create/init/failure ordering, public advanced-compression parameter validation/init policy, public usingDict parameter selection, dictionary-presence handling, and default-level normalization, and public usingCDict frame-policy construction and begin/end sequencing, legacy public CDict-begin frame-policy construction and unknown-source pledge, and the compressBegin_usingDict familys unknown-source parameter selection and default-level normalization, and public advanced-begin parameter validation and init-then-begin ordering now run in Rust. Compression-size estimation, including row-matchfinder selection and compression-level budget progression, and the CLI adaptive feedback state machine also run in Rust behind their private-state callback boundaries. CDict advanced private workspace construction, private static-CCtx and static-CDict workspace construction and dictionary-content allocation/loading, and advanced-CDict dictionary-content loading remain in C. Rust now owns advanced-CDict custom-memory validation, workspace-size query/allocation, allocation/create/init cleanup ordering, the CCtx workspace-size formula, the scalar CCtx-reset plan, post-match-state storage reservation order, and match-state reset policy/order for both CCtx and CDict; C retains private layout-size inputs, workspace resize/layout, private field publication, and allocator callbacks. Private CCtx reset/matchfinder/workspace operations and codec, progression-snapshot, and diagnostic callbacks remain in C. Rust owns adaptive refresh timing. CDict initialization ordering and scalar publication, shared compression-begin dictionary selection, CDict reset attach-versus-copy selection, and CDict-begin parameter selection, initialization ordering, and source-window policy now run in Rust behind private-state bridges; their C callbacks and final begin/reset/attach/copy operations remain. The remaining C paths must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result.

Legacy decoding

Each lib/legacy/zstd_v0N.c file is a frozen snapshot of the entropy coders and frame logic of one historical release. The Rust ports in src/legacy/ keep that property: every version owns its own frozen FSE/Huff0 and frame logic, ported line by line, and must never reuse the modern entropy modules or share code with other legacy versions. Outputs and error codes must be byte-identical to the original C files. Their only shared dependency is the errors module, matching the C files' error_private.h include.

Cargo features legacy-v01 .. legacy-v07 gate the per-version modules and are never default features. All seven modules are now available; the build systems derive the enabled feature list from the C configuration:

  • lib/Makefile and programs/Makefile map ZSTD_LEGACY_SUPPORT=N to the features for versions >= N (0 disables legacy), matching the ZSTD_LEGACY_FILES selection in lib/libzstd.mk.
  • tests/Makefile always enables all seven features because the test objects compile every lib/legacy/*.c file regardless of dispatch level.
  • build/meson maps legacy_level like the makefiles; build/cmake enables all seven whenever ZSTD_LEGACY_SUPPORT is on because it always compiles all seven C files.

Every build system also encodes the legacy selection in the Rust target directory name (for example c1-d1-default-legacy5), for the same reason the HUF mode is encoded there: a cached archive built for one configuration must never be linked into a build that expects another.

Each port adds src/legacy/zstd_v0N.rs, registers it in src/legacy/mod.rs behind its feature, and reduces lib/legacy/zstd_v0N.c to a declaration-only shim. For v0.1 the streaming ZSTDv01_Dctx state lives entirely in Rust: C code only ever holds an opaque pointer, so the C-side struct definition is gone.

Compatibility boundary

The public ABI continues to come from the existing headers under lib/. Exported Rust functions therefore use C layout and calling conventions. A C source file whose implementation has moved to Rust remains in the original makefile source list as a small shim so header configuration and platform preprocessor behavior stay available during the transition.

The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression/dictionary-builder modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation.

Validation

Run focused Rust checks from this directory:

cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo build --release

The program-only Rust archive has its own feature matrix and should be checked from rust/cli as well:

cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features cli,compression --all-targets
cargo test --no-default-features --features cli,decompression --all-targets
cargo test --no-default-features --all-targets

Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example:

make -C tests fuzzer
./tests/fuzzer -i1 --no-big-tests
make -C tests test-rust-lib-smoke

Broader tests/Makefile targets remain the authoritative integration gates as more of the library and CLI are rewritten.