Files
zstd-rs/rust/README.md
T
ddidderr 3b2e8155cd docs(rust): document compress2 and multi-file boundaries
Record the Rust-owned stable-buffer fallback orchestration and shared-output
multi-file scheduler in the migration boundary documentation while retaining
the C-owned private context, resource, and callback responsibilities.

Test Plan:
- git diff --cached --check
- documentation-only boundary update; code validation is covered by
  17cfedd56 and 035001b76

Commit is intentionally unsigned because GPG pinentry hangs in this
non-interactive environment.
2026-07-19 10:32:02 +02:00

235 lines
13 KiB
Markdown

# Rust rewrite
This directory contains the in-progress Rust replacement for the zstd library
and command-line program. During the migration, the crate is built as a static
library and linked into the original C test programs. Production C translation
units become declaration-only shims as their implementations move to Rust; the
original C tests remain unchanged and provide compatibility coverage.
## Component map
The crate is organized from low-level representation helpers toward the public
zstd ABI:
- Common primitives
- `mem`, `bits`, `bitstream`, and `cpu` implement byte-order, bitstream, and
target-feature operations used by the codecs.
- `errors`, `debug`, `xxhash`, and `zstd_common` provide common exported ABI
functions and state.
- `common` contains shared frame constants and internal data types.
- Entropy coding
- `entropy_common` reads FSE normalized counts and Huffman statistics.
- `fse_decompress` builds FSE decoding tables and decodes FSE streams.
- `fse_compress` normalizes counts, writes FSE headers, builds compression
tables, and encodes FSE streams.
- `huf_compress` builds Huffman compression tables, writes table headers,
and encodes one- and four-stream Huffman payloads.
- `huf_decompress` builds Huffman decoding tables and decodes X1 and X2
Huffman streams.
- Compression primitives
- `hist` counts byte frequencies for FSE and Huffman compression.
- `zstd_presplit` chooses split points for full compression blocks.
- `zstd_compress_literals` emits raw, RLE, and Huffman literal sections
while preserving the compressor's Huffman-table repeat state.
- `zstd_compress_stats` converts stored sequences into symbol codes,
selects each block's symbol encoding types, compresses a seqStore's
literals and sequences into a compressed-block body, builds the block
entropy statistics shared with the superblock writer and the block
splitter, and exports collected sequences in the public `ZSTD_Sequence`
format. Its C shims extract the sequence store, the entropy-table
leaves, and the two `ZSTD_CCtx_params` scalars these paths read.
- `zstd_compress_block_split` searches for profitable sequence-store
partitions, while `zstd_compress` emits those partitions through the
Rust single-block serializer. C retains split discovery's context setup
and the outer block-dispatch decision. `zstd_compress` also owns ordinary
sequence-block entropy emission, sequence collection, the legacy RLE
compatibility gate, and sequence-store construction; C supplies only the
private matchfinder, LDM, external-sequence-producer, and state-preparation
callbacks. Rust also owns the frame-chunk block loop, including block
sizing, target/split/internal dispatch, output accounting, and frame-state
updates; C supplies the private block-compression callbacks. Rust also
owns the single-threaded buffered/stable stream state machine, including
direct versus buffered output, pending-output draining, and frame reset
policy. The external-sequence-and-literals block loop and public sequence
conversion are Rust-owned as well; C retains only CCtx-facing adapters.
- `zstd_compress_frame` serializes frame headers, skippable frames, and the
last empty block; it takes scalar frame parameters so the C-owned
`ZSTD_CCtx_params` layout never crosses the language boundary.
- `zstd_compress` also owns the transparent single-threaded and multithreaded
stream-initialization policy: dictionary selection, parameter resolution,
initial buffer sizing, and the ordered setup decisions are projected into
Rust while C retains the private contexts and mutation callbacks.
The public sequence APIs likewise use Rust-owned validation, frame-header,
checksum, and output-accounting orchestration around C-owned block state.
The `ZSTD_compress2_c` fallback also uses a Rust-owned stable-buffer
orchestration boundary while C retains the context reset and stream adapter.
- `zstd_compress_params` owns the compression-level tables (formerly
`clevels.h`), parameter bounds, clamping, validation, table selection,
source/dictionary adjustment, and match-state/CDict size estimation.
The C integration layer keeps the public `ZSTD_*` symbols and feeds the
leaves configuration-owned scalars: the excluded-block-compressor
strategy cascade, struct sizes, and sanitizer redzone policy.
- `zstd_fast` and `zstd_double_fast` implement the single- and two-table
fast block match finders, including attached and external dictionary paths.
- `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching,
including row-based and dictionary search variants.
- `zstd_opt_tree` maintains the binary-tree index used by optimal matching;
`zstd_opt` owns the dynamic-programming price model, optimal parse, and
sequence emission.
- `zstd_ldm` implements long-distance-match parameter selection, table
maintenance, sequence generation, and sequence consumption.
- Dictionary building
- `divsufsort` constructs the suffix array that drives the legacy `ZDICT`
trainer (`ZDICT_trainFromBuffer_legacy`); `dict_builder_zdict`,
`dict_builder_cover`, and `dict_builder_fastcover` own the sample analysis,
training, and dictionary assembly. The corresponding C translation units
are declaration-only ABI shims.
- Runtime support
- `threading` provides platform pthread wrappers required by zstd headers.
- `pool` implements the bounded worker pool used by multithreaded compression.
- The `zstdmt_compress` integration keeps job descriptors and synchronization
private to C while Rust owns input-retention scans, reusable input-range
overlap decisions, outer scheduling and end-directive adjustments,
job-creation decisions, compression-job stage sequencing and error flow,
and pending-output decisions through scalar job projections.
- Dictionary support
- `zstd_ddict` owns, loads, copies, and references decode dictionaries.
- Legacy decoding
- `legacy` hosts one frozen module per historical format; `legacy::zstd_v01`
through `legacy::zstd_v07` port the seven self-contained historical
decoders. Their original C translation units remain declaration-only
shims for the native build.
- Block decompression
- `zstd_decompress_block` decodes literal and sequence sections, maintains
FSE/Huffman repeat state, executes compressed-block sequences, and selects
the short or long sequence decoder from projected configuration and history
state. C retains the decoder-context layout and configuration projection;
the public and fullbench block-decoder wrappers are Rust-owned.
- `zstd_decompress` owns the public decompression context, one-shot,
dictionary, parameter, and streaming state machines. Its C shim retains
the configuration-dependent context layout and platform details, plus
legacy and trace leaves; Rust owns decoder storage allocation and custom
memory dispatch.
- Command-line frontend
- `zstd_cli` owns the Rust parser, safety policy, and dispatch. It is built
by the separate `cli/` static-library package only for program archives,
so library builds do not acquire program-only dependencies. The C
`fileio` layer retains the format-specific codec callbacks, private
asynchronous-pool adapters, metadata, zstd codec/error mapping, and
adaptive-policy integration, while Rust owns the mixed-format probe/dispatch
loop, zstd stream-compression I/O loop, scalar adaptive decisions,
optional-format decompression loops, and decompression result policy/final
accounting.
Rust already owns the file preference policy, filename decisions,
source/destination opening, dictionary buffers, asynchronous I/O pools,
and pass-through copy leaf.
- `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`,
while `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`,
`BMK_benchTimedFn`) and `benchzstd` owns benchmark orchestration and
reporting. Both live in Rust; the C translation units are ABI shims.
C test binaries (fullbench, fuzzer, zstreamtest, paramgrill, ...) link a
helpers-only build of the archive, produced without the package's `cli`
feature, because the parser layer requires the C `fileio` backend that
tests do not compile.
Dictionary-ingestion dispatch, CCtx dictionary/prefix attachment dispatch, the
CLI zstd compression stream loop, scalar adaptive decisions, and the
shared-destination multi-file compression scheduler,
single-threaded stream initialization and the buffered/stable stream state
machine, MT stream initialization, MT outer scheduling and flush policy, MT
compression-job stage sequencing and error flow, public sequence-API
orchestration, sequence-store and block policy,
external-sequence/literals block loop, optional-format decompression loops,
block decoder wrappers, and decompression result policy now run in Rust. CDict
lifecycle, the private dictionary-content loader, reset policy, private
CCtx/matchfinder/workspace operations, and codec/adaptive-policy callbacks
remain in C. The remaining C paths must move before the rewrite is complete.
Keeping that boundary explicit prevents a passing hybrid build from being
mistaken for the final all-Rust result.
## Legacy decoding
Each `lib/legacy/zstd_v0N.c` file is a frozen snapshot of the entropy coders
and frame logic of one historical release. The Rust ports in `src/legacy/`
keep that property: every version owns its own frozen FSE/Huff0 and frame
logic, ported line by line, and must never reuse the modern entropy modules
or share code with other legacy versions. Outputs and error codes must be
byte-identical to the original C files. Their only shared dependency is the
`errors` module, matching the C files' `error_private.h` include.
Cargo features `legacy-v01` .. `legacy-v07` gate the per-version modules and
are never default features. All seven modules are now available; the build
systems derive the enabled feature list from the C configuration:
- `lib/Makefile` and `programs/Makefile` map `ZSTD_LEGACY_SUPPORT=N` to the
features for versions >= N (0 disables legacy), matching the
`ZSTD_LEGACY_FILES` selection in `lib/libzstd.mk`.
- `tests/Makefile` always enables all seven features because the test
objects compile every `lib/legacy/*.c` file regardless of dispatch level.
- `build/meson` maps `legacy_level` like the makefiles; `build/cmake`
enables all seven whenever `ZSTD_LEGACY_SUPPORT` is on because it always
compiles all seven C files.
Every build system also encodes the legacy selection in the Rust target
directory name (for example `c1-d1-default-legacy5`), for the same reason the
HUF mode is encoded there: a cached archive built for one configuration must
never be linked into a build that expects another.
Each port adds `src/legacy/zstd_v0N.rs`, registers it in `src/legacy/mod.rs`
behind its feature, and reduces `lib/legacy/zstd_v0N.c` to a
declaration-only shim. For v0.1 the streaming `ZSTDv01_Dctx` state lives
entirely in Rust: C code only ever holds an opaque pointer, so the C-side
struct definition is gone.
## Compatibility boundary
The public ABI continues to come from the existing headers under `lib/`.
Exported Rust functions therefore use C layout and calling conventions. A C
source file whose implementation has moved to Rust remains in the original
makefile source list as a small shim so header configuration and platform
preprocessor behavior stay available during the transition.
The library, test, and program makefiles select an archive directory for the
active C configuration: enabled compression/decompression/dictionary-builder
modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit
C binaries. The
native static archive flattens Rust object members rather than nesting a Rust
archive, while the native shared library retains all migrated Rust exports.
When the HUF mode changes, the test and program paths also rebuild cached C
outputs before linking. This prevents original C tests from using a stale or
configuration-incompatible implementation.
## Validation
Run focused Rust checks from this directory:
```sh
cargo fmt --check
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo build --release
```
The program-only Rust archive has its own feature matrix and should be checked
from `rust/cli` as well:
```sh
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features cli,compression --all-targets
cargo test --no-default-features --features cli,decompression --all-targets
cargo test --no-default-features --all-targets
```
Then run original compatibility tests from the repository root, starting with
the narrow target for the component being migrated. For example:
```sh
make -C tests fuzzer
./tests/fuzzer -i1 --no-big-tests
make -C tests test-rust-lib-smoke
```
Broader `tests/Makefile` targets remain the authoritative integration gates as
more of the library and CLI are rewritten.