# Rust rewrite This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage. ## Component map The crate is organized from low-level representation helpers toward the public zstd ABI: - Common primitives - `mem`, `bits`, `bitstream`, and `cpu` implement byte-order, bitstream, and target-feature operations used by the codecs. - `errors`, `debug`, `xxhash`, and `zstd_common` provide common exported ABI functions and state. - `common` contains shared frame constants and internal data types. - Entropy coding - `entropy_common` reads FSE normalized counts and Huffman statistics. - `fse_decompress` builds FSE decoding tables and decodes FSE streams. - `fse_compress` normalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams. - `huf_compress` builds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads. - `huf_decompress` builds Huffman decoding tables and decodes X1 and X2 Huffman streams. - Compression primitives - `hist` counts byte frequencies for FSE and Huffman compression. - `zstd_presplit` chooses split points for full compression blocks. - `zstd_compress_literals` emits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state. - `zstd_compress_stats` converts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the public `ZSTD_Sequence` format. Its C shims extract the sequence store, the entropy-table leaves, and the two `ZSTD_CCtx_params` scalars these paths read. - `zstd_compress_frame` serializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-owned `ZSTD_CCtx_params` layout never crosses the language boundary. - `zstd_compress_params` owns the compression-level tables (formerly `clevels.h`), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the public `ZSTD_*` symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy. - `zstd_fast` and `zstd_double_fast` implement the single- and two-table fast block match finders, including attached and external dictionary paths. - `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants. - `zstd_opt_tree` maintains the binary-tree index used by optimal matching; the dynamic-programming optimal parser itself remains in C for now. - `zstd_ldm` implements long-distance-match parameter selection, table maintenance, sequence generation, and sequence consumption. - Dictionary building - `divsufsort` constructs the suffix array that drives the legacy `ZDICT` trainer (`ZDICT_trainFromBuffer_legacy`). The sample analysis and dictionary assembly in `zdict.c`, `cover.c`, and `fastcover.c` remain C. - Runtime support - `threading` provides platform pthread wrappers required by zstd headers. - `pool` implements the bounded worker pool used by multithreaded compression. - Dictionary support - `zstd_ddict` owns, loads, copies, and references decode dictionaries. - Legacy decoding - `legacy` hosts one frozen module per historical format; `legacy::zstd_v01` ports the self-contained v0.1 decoder. Versions v0.2 through v0.7 are still C. - Block decompression - `zstd_decompress_block` decodes literal and sequence sections, maintains FSE/Huffman repeat state, and executes compressed-block sequences. - `zstd_decompress` owns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains configuration-dependent context allocation plus legacy and trace leaves. - Command-line frontend - `zstd_cli` owns the Rust parser, safety policy, and dispatch. It is built by the separate `cli/` static-library package only for program archives, so library builds do not acquire program-only dependencies. The C `fileio` backend still owns file opening, safe replacement, sparse writes, metadata, and streaming I/O. - `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`, and `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`, `BMK_benchTimedFn`) used by the CLI benchmark mode and by C test tools. Both live in the `cli/` package, but C test binaries (fullbench, fuzzer, zstreamtest, paramgrill, ...) link a helpers-only build of that archive, produced without the package's `cli` feature, because the parser layer requires the C `fileio` backend that tests do not compile. Benchmark orchestration and reporting (`benchzstd.c`) remain C, reached from the Rust parser through the `ZSTD_NOBENCH`-gated bridge in `zstdcli.c`. The optimal block matcher, high-level frame compression, dictionary-building except suffix-array construction, the legacy v0.2-v0.7 decoders, benchmark orchestration (`benchzstd`), and the CLI file-I/O backend are still C. They must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result. ## Legacy decoding Each `lib/legacy/zstd_v0N.c` file is a frozen snapshot of the entropy coders and frame logic of one historical release. The Rust ports in `src/legacy/` keep that property: every version owns its own frozen FSE/Huff0 and frame logic, ported line by line, and must never reuse the modern entropy modules or share code with other legacy versions. Outputs and error codes must be byte-identical to the original C files. Their only shared dependency is the `errors` module, matching the C files' `error_private.h` include. Cargo features `legacy-v01` .. `legacy-v07` gate the per-version modules and are never default features. The build systems derive the feature list from the C configuration: - `lib/Makefile` and `programs/Makefile` map `ZSTD_LEGACY_SUPPORT=N` to the features for versions >= N (0 disables legacy), matching the `ZSTD_LEGACY_FILES` selection in `lib/libzstd.mk`. - `tests/Makefile` always enables all seven features because the test objects compile every `lib/legacy/*.c` file regardless of dispatch level. - `build/meson` maps `legacy_level` like the makefiles; `build/cmake` enables all seven whenever `ZSTD_LEGACY_SUPPORT` is on because it always compiles all seven C files. Every build system also encodes the legacy selection in the Rust target directory name (for example `c1-d1-default-legacy5`), for the same reason the HUF mode is encoded there: a cached archive built for one configuration must never be linked into a build that expects another. A feature whose version has not been ported yet gates nothing; the original C file still provides that decoder, so mixed C/Rust legacy levels link cleanly. Porting a version means adding `src/legacy/zstd_v0N.rs`, registering it in `src/legacy/mod.rs` behind its feature, and reducing `lib/legacy/zstd_v0N.c` to a declaration-only shim. For v0.1 the streaming `ZSTDv01_Dctx` state lives entirely in Rust: C code only ever holds an opaque pointer, so the C-side struct definition is gone. ## Compatibility boundary The public ABI continues to come from the existing headers under `lib/`. Exported Rust functions therefore use C layout and calling conventions. A C source file whose implementation has moved to Rust remains in the original makefile source list as a small shim so header configuration and platform preprocessor behavior stay available during the transition. The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression/dictionary-builder modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation. ## Validation Run focused Rust checks from this directory: ```sh cargo fmt --check cargo clippy --all-targets -- -D warnings cargo test --all-targets cargo build --release ``` The program-only Rust archive has its own feature matrix and should be checked from `rust/cli` as well: ```sh cargo clippy --all-targets -- -D warnings cargo test --all-targets cargo test --no-default-features --features cli,compression --all-targets cargo test --no-default-features --features cli,decompression --all-targets cargo test --no-default-features --all-targets ``` Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example: ```sh make -C tests fuzzer ./tests/fuzzer -i1 --no-big-tests make -C tests test-rust-lib-smoke ``` Broader `tests/Makefile` targets remain the authoritative integration gates as more of the library and CLI are rewritten.