Files
zstd-rs/rust
ddidderr 57a2f3bef4 feat(cli): port the lorem generator to Rust
Move LOREM_genBlock and LOREM_genBuffer into the Rust CLI helper archive,
preserving the frozen word distribution, seeded sentence and paragraph layout,
tail filling, and C-compatible buffer ABI. Use per-call state so generation is
reentrant without changing the generated bytes.

Focused tests cover C reference vectors, distribution construction,
determinism, zero-sized null buffers, tiny tail-filled buffers, and continued
nonzero-fill generation.

Test Plan:
- cargo test --manifest-path rust/cli/Cargo.toml --no-default-features lorem::tests
- cargo clippy --manifest-path rust/cli/Cargo.toml --all-targets -- -D warnings

Refs: programs/lorem.c compatibility port
2026-07-12 09:16:59 +02:00
..

Rust rewrite

This directory contains the in-progress Rust replacement for the zstd library and command-line program. During the migration, the crate is built as a static library and linked into the original C test programs. Production C translation units become declaration-only shims as their implementations move to Rust; the original C tests remain unchanged and provide compatibility coverage.

Component map

The crate is organized from low-level representation helpers toward the public zstd ABI:

  • Common primitives
    • mem, bits, bitstream, and cpu implement byte-order, bitstream, and target-feature operations used by the codecs.
    • errors, debug, xxhash, and zstd_common provide common exported ABI functions and state.
    • common contains shared frame constants and internal data types.
  • Entropy coding
    • entropy_common reads FSE normalized counts and Huffman statistics.
    • fse_decompress builds FSE decoding tables and decodes FSE streams.
    • fse_compress normalizes counts, writes FSE headers, builds compression tables, and encodes FSE streams.
    • huf_compress builds Huffman compression tables, writes table headers, and encodes one- and four-stream Huffman payloads.
    • huf_decompress builds Huffman decoding tables and decodes X1 and X2 Huffman streams.
  • Compression primitives
    • hist counts byte frequencies for FSE and Huffman compression.
    • zstd_presplit chooses split points for full compression blocks.
    • zstd_compress_literals emits raw, RLE, and Huffman literal sections while preserving the compressor's Huffman-table repeat state.
    • zstd_compress_stats converts stored sequences into symbol codes, selects each block's symbol encoding types, compresses a seqStore's literals and sequences into a compressed-block body, builds the block entropy statistics shared with the superblock writer and the block splitter, and exports collected sequences in the public ZSTD_Sequence format. Its C shims extract the sequence store, the entropy-table leaves, and the two ZSTD_CCtx_params scalars these paths read.
    • zstd_compress_frame serializes frame headers, skippable frames, and the last empty block; it takes scalar frame parameters so the C-owned ZSTD_CCtx_params layout never crosses the language boundary.
    • zstd_compress_params owns the compression-level tables (formerly clevels.h), parameter bounds, clamping, validation, table selection, source/dictionary adjustment, and match-state/CDict size estimation. The C integration layer keeps the public ZSTD_* symbols and feeds the leaves configuration-owned scalars: the excluded-block-compressor strategy cascade, struct sizes, and sanitizer redzone policy.
    • zstd_fast and zstd_double_fast implement the single- and two-table fast block match finders, including attached and external dictionary paths.
    • zstd_lazy implements greedy, lazy, lazy2, and binary-tree matching, including row-based and dictionary search variants.
    • zstd_opt_tree maintains the binary-tree index used by optimal matching; the dynamic-programming optimal parser itself remains in C for now.
    • zstd_ldm implements long-distance-match parameter selection, table maintenance, sequence generation, and sequence consumption.
  • Dictionary building
    • divsufsort constructs the suffix array that drives the legacy ZDICT trainer (ZDICT_trainFromBuffer_legacy). The sample analysis and dictionary assembly in zdict.c, cover.c, and fastcover.c remain C.
  • Runtime support
    • threading provides platform pthread wrappers required by zstd headers.
    • pool implements the bounded worker pool used by multithreaded compression.
  • Dictionary support
    • zstd_ddict owns, loads, copies, and references decode dictionaries.
  • Legacy decoding
    • legacy hosts one frozen module per historical format; legacy::zstd_v01 ports the self-contained v0.1 decoder. Versions v0.2 through v0.7 are still C.
  • Block decompression
    • zstd_decompress_block decodes literal and sequence sections, maintains FSE/Huffman repeat state, and executes compressed-block sequences.
    • zstd_decompress owns the public decompression context, one-shot, dictionary, parameter, and streaming state machines. Its C shim retains configuration-dependent context allocation plus legacy and trace leaves.
  • Command-line frontend
    • zstd_cli owns the Rust parser, safety policy, and dispatch. It is built by the separate cli/ static-library package only for program archives, so library builds do not acquire program-only dependencies. The C fileio backend still owns file opening, safe replacement, sparse writes, metadata, and streaming I/O.
    • timefn provides the monotonic nanosecond clock behind UTIL_time_t, and benchfn owns the benchmark run/timing loop (BMK_benchFunction, BMK_benchTimedFn) used by the CLI benchmark mode and by C test tools. Both live in the cli/ package, but C test binaries (fullbench, fuzzer, zstreamtest, paramgrill, ...) link a helpers-only build of that archive, produced without the package's cli feature, because the parser layer requires the C fileio backend that tests do not compile. Benchmark orchestration and reporting (benchzstd.c) remain C, reached from the Rust parser through the ZSTD_NOBENCH-gated bridge in zstdcli.c.

The optimal block matcher, high-level frame compression, dictionary-building except suffix-array construction, the legacy v0.2-v0.7 decoders, benchmark orchestration (benchzstd), and the CLI file-I/O backend are still C. They must move before the rewrite is complete. Keeping that boundary explicit prevents a passing hybrid build from being mistaken for the final all-Rust result.

Legacy decoding

Each lib/legacy/zstd_v0N.c file is a frozen snapshot of the entropy coders and frame logic of one historical release. The Rust ports in src/legacy/ keep that property: every version owns its own frozen FSE/Huff0 and frame logic, ported line by line, and must never reuse the modern entropy modules or share code with other legacy versions. Outputs and error codes must be byte-identical to the original C files. Their only shared dependency is the errors module, matching the C files' error_private.h include.

Cargo features legacy-v01 .. legacy-v07 gate the per-version modules and are never default features. The build systems derive the feature list from the C configuration:

  • lib/Makefile and programs/Makefile map ZSTD_LEGACY_SUPPORT=N to the features for versions >= N (0 disables legacy), matching the ZSTD_LEGACY_FILES selection in lib/libzstd.mk.
  • tests/Makefile always enables all seven features because the test objects compile every lib/legacy/*.c file regardless of dispatch level.
  • build/meson maps legacy_level like the makefiles; build/cmake enables all seven whenever ZSTD_LEGACY_SUPPORT is on because it always compiles all seven C files.

Every build system also encodes the legacy selection in the Rust target directory name (for example c1-d1-default-legacy5), for the same reason the HUF mode is encoded there: a cached archive built for one configuration must never be linked into a build that expects another.

A feature whose version has not been ported yet gates nothing; the original C file still provides that decoder, so mixed C/Rust legacy levels link cleanly. Porting a version means adding src/legacy/zstd_v0N.rs, registering it in src/legacy/mod.rs behind its feature, and reducing lib/legacy/zstd_v0N.c to a declaration-only shim. For v0.1 the streaming ZSTDv01_Dctx state lives entirely in Rust: C code only ever holds an opaque pointer, so the C-side struct definition is gone.

Compatibility boundary

The public ABI continues to come from the existing headers under lib/. Exported Rust functions therefore use C layout and calling conventions. A C source file whose implementation has moved to Rust remains in the original makefile source list as a small shim so header configuration and platform preprocessor behavior stay available during the transition.

The library, test, and program makefiles select an archive directory for the active C configuration: enabled compression/decompression/dictionary-builder modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The native static archive flattens Rust object members rather than nesting a Rust archive, while the native shared library retains all migrated Rust exports. When the HUF mode changes, the test and program paths also rebuild cached C outputs before linking. This prevents original C tests from using a stale or configuration-incompatible implementation.

Validation

Run focused Rust checks from this directory:

cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo build --release

The program-only Rust archive has its own feature matrix and should be checked from rust/cli as well:

cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features cli,compression --all-targets
cargo test --no-default-features --features cli,decompression --all-targets
cargo test --no-default-features --all-targets

Then run original compatibility tests from the repository root, starting with the narrow target for the component being migrated. For example:

make -C tests fuzzer
./tests/fuzzer -i1 --no-big-tests
make -C tests test-rust-lib-smoke

Broader tests/Makefile targets remain the authoritative integration gates as more of the library and CLI are rewritten.