# Rust rewrite This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage. ## Component map The crate is organized from low-level representation helpers toward the public zstd ABI: - Common primitives - `mem`, `bits`, `bitstream`, and `cpu` implement byte-order, bitstream, and target-feature operations used by the codecs. - `errors`, `debug`, `xxhash`, and `zstd_common` provide common exported ABI functions and state. - `common` contains shared frame constants and internal data types. - Entropy coding - `entropy_common` reads FSE normalized counts and Huffman statistics. - `fse_decompress` builds FSE decoding tables and decodes FSE streams. - `fse_compress` normalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams. - `huf_compress` builds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads. - `huf_decompress` builds Huffman decoding tables and decodes X1 and X2 Huffman streams. - Compression primitives - `hist` counts byte frequencies for FSE and Huffman compression. - `zstd_presplit` chooses split points for full compression blocks. - `zstd_compress_literals` emits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state. - `zstd_compress_stats` converts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the public `ZSTD_Sequence` format. Its C shims extract the sequence store, the entropy-table leaves, and the two `ZSTD_CCtx_params` scalars these paths read. - `zstd_compress_block_split` searches for profitable sequence-store partitions, while `zstd_compress` emits those partitions through the Rust single-block serializer. C retains split discovery's context setup and the outer block-dispatch decision. `zstd_compress` also owns ordinary sequence-block entropy emission, sequence collection, the legacy RLE compatibility gate, and sequence-store construction; C supplies only the private matchfinder, LDM, external-sequence-producer, and state-preparation callbacks. Rust also owns the frame-chunk block loop, including block sizing, target/split/internal dispatch, output accounting, and frame-state updates; C supplies the private block-compression callbacks. Rust also owns the single-threaded buffered/stable stream state machine, including direct versus buffered output, pending-output draining, and frame reset policy. The external-sequence-and-literals block loop and public sequence conversion are Rust-owned as well; C retains only CCtx-facing adapters. - `zstd_compress_frame` serializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-owned `ZSTD_CCtx_params` layout never crosses the language boundary. - `zstd_compress` also owns the transparent single-threaded and multithreaded stream-initialization policy: dictionary selection, parameter resolution, initial buffer sizing, and the ordered setup decisions are projected into Rust while C retains the private contexts and mutation callbacks. The public sequence APIs likewise use Rust-owned validation, frame-header, checksum, and output-accounting orchestration around C-owned block state. The `ZSTD_compress2_c` fallback also uses a Rust-owned stable-buffer orchestration boundary while C retains the context reset and stream adapter. - `zstd_compress_params` owns the compression-level tables (formerly `clevels.h`), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the public `ZSTD_*` symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy. - `zstd_fast` and `zstd_double_fast` implement the single- and two-table fast block match finders, including attached and external dictionary paths. - `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants. - `zstd_opt_tree` maintains the binary-tree index used by optimal matching; `zstd_opt` owns the dynamic-programming price model, optimal parse, and sequence emission. - `zstd_ldm` implements long-distance-match parameter selection, table maintenance, sequence generation, and sequence consumption. - Dictionary building - `divsufsort` constructs the suffix array that drives the legacy `ZDICT` trainer (`ZDICT_trainFromBuffer_legacy`); `dict_builder_zdict`, `dict_builder_cover`, and `dict_builder_fastcover` own the sample analysis, training, and dictionary assembly. The corresponding C translation units are declaration-only ABI shims. - Runtime support - `threading` provides platform pthread wrappers required by zstd headers. - `pool` implements the bounded worker pool used by multithreaded compression. - The `zstdmt_compress` integration keeps job descriptors and synchronization private to C while Rust owns input-retention scans, reusable input-range overlap decisions, outer scheduling and end-directive adjustments, job-creation decisions, compression-job stage sequencing and error flow, and pending-output decisions through scalar job projections. - Dictionary support - `zstd_ddict` owns, loads, copies, and references decode dictionaries. - Legacy decoding - `legacy` hosts one frozen module per historical format; `legacy::zstd_v01` through `legacy::zstd_v07` port the seven self-contained historical decoders. Their original C translation units remain declaration-only shims for the native build. - Block decompression - `zstd_decompress_block` decodes literal and sequence sections, maintains FSE/Huffman repeat state, executes compressed-block sequences, and selects the short or long sequence decoder from projected configuration and history state. C retains the decoder-context layout and configuration projection; the public and fullbench block-decoder wrappers are Rust-owned. - `zstd_decompress` owns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains the configuration-dependent context layout and platform details, plus legacy and trace leaves; Rust owns decoder storage allocation and custom memory dispatch. - Command-line frontend - `zstd_cli` owns the Rust parser, safety policy, and dispatch. It is built by the separate `cli/` static-library package only for program archives, so library builds do not acquire program-only dependencies. The C `fileio` layer retains the format-specific codec callbacks, private asynchronous-pool adapters, metadata, zstd codec/error mapping, and adaptive-policy integration, while Rust owns the mixed-format probe/dispatch loop, zstd stream-compression I/O loop, scalar adaptive decisions, optional-format decompression loops, and decompression result policy/final accounting. Rust already owns the file preference policy, filename decisions, source/destination opening, dictionary buffers, asynchronous I/O pools, and pass-through copy leaf. - `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`, while `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`, `BMK_benchTimedFn`) and `benchzstd` owns benchmark orchestration and reporting. Both live in Rust; the C translation units are ABI shims. C test binaries (fullbench, fuzzer, zstreamtest, paramgrill, ...) link a helpers-only build of the archive, produced without the package's `cli` feature, because the parser layer requires the C `fileio` backend that tests do not compile. Dictionary-ingestion dispatch, CCtx dictionary/prefix attachment dispatch, the CLI zstd compression stream loop, scalar adaptive decisions, and the shared-destination multi-file compression scheduler, single-threaded stream initialization and the buffered/stable stream state machine, MT stream initialization, MT outer scheduling and flush policy, MT compression-job stage sequencing and error flow, public sequence-API orchestration, sequence-store and block policy, external-sequence/literals block loop, optional-format decompression loops, block decoder wrappers, and decompression result policy now run in Rust. CDict lifecycle, the private dictionary-content loader, reset policy, private CCtx/matchfinder/workspace operations, and codec/adaptive-policy callbacks remain in C. The remaining C paths must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result. ## Legacy decoding Each `lib/legacy/zstd_v0N.c` file is a frozen snapshot of the entropy coders and frame logic of one historical release. The Rust ports in `src/legacy/` keep that property: every version owns its own frozen FSE/Huff0 and frame logic, ported line by line, and must never reuse the modern entropy modules or share code with other legacy versions. Outputs and error codes must be byte-identical to the original C files. Their only shared dependency is the `errors` module, matching the C files' `error_private.h` include. Cargo features `legacy-v01` .. `legacy-v07` gate the per-version modules and are never default features. All seven modules are now available; the build systems derive the enabled feature list from the C configuration: - `lib/Makefile` and `programs/Makefile` map `ZSTD_LEGACY_SUPPORT=N` to the features for versions >= N (0 disables legacy), matching the `ZSTD_LEGACY_FILES` selection in `lib/libzstd.mk`. - `tests/Makefile` always enables all seven features because the test objects compile every `lib/legacy/*.c` file regardless of dispatch level. - `build/meson` maps `legacy_level` like the makefiles; `build/cmake` enables all seven whenever `ZSTD_LEGACY_SUPPORT` is on because it always compiles all seven C files. Every build system also encodes the legacy selection in the Rust target directory name (for example `c1-d1-default-legacy5`), for the same reason the HUF mode is encoded there: a cached archive built for one configuration must never be linked into a build that expects another. Each port adds `src/legacy/zstd_v0N.rs`, registers it in `src/legacy/mod.rs` behind its feature, and reduces `lib/legacy/zstd_v0N.c` to a declaration-only shim. For v0.1 the streaming `ZSTDv01_Dctx` state lives entirely in Rust: C code only ever holds an opaque pointer, so the C-side struct definition is gone. ## Compatibility boundary The public ABI continues to come from the existing headers under `lib/`. Exported Rust functions therefore use C layout and calling conventions. A C source file whose implementation has moved to Rust remains in the original makefile source list as a small shim so header configuration and platform preprocessor behavior stay available during the transition. The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression/dictionary-builder modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation. ## Validation Run focused Rust checks from this directory: ```sh cargo fmt --check cargo clippy --all-targets -- -D warnings cargo test --all-targets cargo build --release ``` The program-only Rust archive has its own feature matrix and should be checked from `rust/cli` as well: ```sh cargo clippy --all-targets -- -D warnings cargo test --all-targets cargo test --no-default-features --features cli,compression --all-targets cargo test --no-default-features --features cli,decompression --all-targets cargo test --no-default-features --all-targets ``` Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example: ```sh make -C tests fuzzer ./tests/fuzzer -i1 --no-big-tests make -C tests test-rust-lib-smoke ``` Broader `tests/Makefile` targets remain the authoritative integration gates as more of the library and CLI are rewritten.