Continue the incremental zstd_compress.c migration with its sequence statistics and seqStore entropy-compression layer. The following now live in rust/src/zstd_compress_stats.rs: - ZSTD_seqToCodes(), exported under its original name because the C dictionary builder (zdict.c) and decodecorpus link against it, - ZSTD_buildSequencesStatistics() and its dummy variant, whose result struct no longer crosses the language boundary, - ZSTD_entropyCompressSeqStore_internal(), _wExtLitBuffer(), and ZSTD_entropyCompressSeqStore(), - ZSTD_buildBlockEntropyStats() with its literals/sequences helpers, - ZSTD_copyBlockSequences() and the ZSTD_updateRep() rules it shares with the superblock writer. Boundary: ZSTD_CCtx and ZSTD_CCtx_params stay private to C. These paths read exactly two parameter fields, so the C shims keep the original static/exported function names and forward the strategy and ZSTD_literalsCompressionIsDisabled() as int scalars alongside the seqStore and entropy-table leaves. The leaf layouts (SeqDef, SeqStore_t, ZSTD_hufCTables_t, ZSTD_fseCTables_t, entropy metadata, SeqCollector, ZSTD_Sequence) are pinned by compile-time asserts in zstd_compress.c and by both-pointer-width layout tests in Rust. ZSTD_buildSeqStore, block dispatch/splitting, and the block-size estimation helpers remain C for a later slice. The superblock module previously round-tripped through the C export of ZSTD_buildBlockEntropyStats and mirrored the entropy leaf structs privately. It now calls the crate-internal builder directly, and the shared struct definitions moved to zstd_compress_stats; consequently ZSTD_rust_compressSuperBlock() takes the two parameter scalars instead of an opaque ZSTD_CCtx_params pointer, extracted by its C shim. Its layout and repcode tests moved with the definitions. One C helper family gets no shim: the static ZSTD_entropyCompressSeqStore_wExtLitBuffer() had a single caller and was folded into the Rust implementation. Byte-identity was verified against the pre-change compressor: COPYING, datagen -g5000000 -s7, and datagen -g300000 -s21 -P90, each at levels 1/3/9/19 and --fast=5, plus a superblock-heavy pass at level 19 with --target-compressed-block-size=1024; all 18 frames are byte-identical, covering the repeat-mode state machine, longOffsets, RLE/raw fallback, and dstSize_tooSmall paths through the block splitter and superblock. Test plan: - cd rust && cargo fmt --check && cargo clippy --all-targets -- -D warnings && cargo test --all-targets (131 tests: new reference vectors for seqToCodes, RLE table headers, empty-seqStore repeat copies, literal-stats type selection, and repcode resolution in copyBlockSequences) - cargo build --release --no-default-features --features compression and --features decompression - make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests (covers ZSTD_generateSequences and ZDICT training over the Rust seqToCodes) - make -C tests test-rust-lib-smoke - ./tests/zstreamtest -i1 and ./tests/zstreamtest --newapi -t1 -i1 - ./tests/decodecorpus -t -T1s (decodecorpus links Rust seqToCodes) - /tmp/ref vs /tmp/got frame diff as described above: byte-identical
187 lines
9.6 KiB
Markdown
187 lines
9.6 KiB
Markdown
# Rust rewrite
|
|
|
|
This directory contains the in-progress Rust replacement for the zstd library
|
|
and command-line program. During the migration, the crate is built as a static
|
|
library and linked into the original C test programs. Production C translation
|
|
units become declaration-only shims as their implementations move to Rust; the
|
|
original C tests remain unchanged and provide compatibility coverage.
|
|
|
|
## Component map
|
|
|
|
The crate is organized from low-level representation helpers toward the public
|
|
zstd ABI:
|
|
|
|
- Common primitives
|
|
- `mem`, `bits`, `bitstream`, and `cpu` implement byte-order, bitstream, and
|
|
target-feature operations used by the codecs.
|
|
- `errors`, `debug`, `xxhash`, and `zstd_common` provide common exported ABI
|
|
functions and state.
|
|
- `common` contains shared frame constants and internal data types.
|
|
- Entropy coding
|
|
- `entropy_common` reads FSE normalized counts and Huffman statistics.
|
|
- `fse_decompress` builds FSE decoding tables and decodes FSE streams.
|
|
- `fse_compress` normalizes counts, writes FSE headers, builds compression
|
|
tables, and encodes FSE streams.
|
|
- `huf_compress` builds Huffman compression tables, writes table headers,
|
|
and encodes one- and four-stream Huffman payloads.
|
|
- `huf_decompress` builds Huffman decoding tables and decodes X1 and X2
|
|
Huffman streams.
|
|
- Compression primitives
|
|
- `hist` counts byte frequencies for FSE and Huffman compression.
|
|
- `zstd_presplit` chooses split points for full compression blocks.
|
|
- `zstd_compress_literals` emits raw, RLE, and Huffman literal sections
|
|
while preserving the compressor's Huffman-table repeat state.
|
|
- `zstd_compress_stats` converts stored sequences into symbol codes,
|
|
selects each block's symbol encoding types, compresses a seqStore's
|
|
literals and sequences into a compressed-block body, builds the block
|
|
entropy statistics shared with the superblock writer and the block
|
|
splitter, and exports collected sequences in the public `ZSTD_Sequence`
|
|
format. Its C shims extract the sequence store, the entropy-table
|
|
leaves, and the two `ZSTD_CCtx_params` scalars these paths read.
|
|
- `zstd_compress_frame` serializes frame headers, skippable frames, and the
|
|
last empty block; it takes scalar frame parameters so the C-owned
|
|
`ZSTD_CCtx_params` layout never crosses the language boundary.
|
|
- `zstd_compress_params` owns the compression-level tables (formerly
|
|
`clevels.h`), parameter bounds, clamping, validation, table selection,
|
|
source/dictionary adjustment, and match-state/CDict size estimation.
|
|
The C integration layer keeps the public `ZSTD_*` symbols and feeds the
|
|
leaves configuration-owned scalars: the excluded-block-compressor
|
|
strategy cascade, struct sizes, and sanitizer redzone policy.
|
|
- `zstd_fast` and `zstd_double_fast` implement the single- and two-table
|
|
fast block match finders, including attached and external dictionary paths.
|
|
- `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching,
|
|
including row-based and dictionary search variants.
|
|
- `zstd_opt_tree` maintains the binary-tree index used by optimal matching;
|
|
the dynamic-programming optimal parser itself remains in C for now.
|
|
- `zstd_ldm` implements long-distance-match parameter selection, table
|
|
maintenance, sequence generation, and sequence consumption.
|
|
- Dictionary building
|
|
- `divsufsort` constructs the suffix array that drives the legacy `ZDICT`
|
|
trainer (`ZDICT_trainFromBuffer_legacy`). The sample analysis and
|
|
dictionary assembly in `zdict.c`, `cover.c`, and `fastcover.c` remain C.
|
|
- Runtime support
|
|
- `threading` provides platform pthread wrappers required by zstd headers.
|
|
- `pool` implements the bounded worker pool used by multithreaded compression.
|
|
- Dictionary support
|
|
- `zstd_ddict` owns, loads, copies, and references decode dictionaries.
|
|
- Legacy decoding
|
|
- `legacy` hosts one frozen module per historical format; `legacy::zstd_v01`
|
|
ports the self-contained v0.1 decoder. Versions v0.2 through v0.7 are
|
|
still C.
|
|
- Block decompression
|
|
- `zstd_decompress_block` decodes literal and sequence sections, maintains
|
|
FSE/Huffman repeat state, and executes compressed-block sequences.
|
|
- `zstd_decompress` owns the public decompression context, one-shot,
|
|
dictionary, parameter, and streaming state machines. Its C shim retains
|
|
configuration-dependent context allocation plus legacy and trace leaves.
|
|
- Command-line frontend
|
|
- `zstd_cli` owns the Rust parser, safety policy, and dispatch. It is built
|
|
by the separate `cli/` static-library package only for program archives,
|
|
so library builds do not acquire program-only dependencies. The C
|
|
`fileio` backend still owns file opening, safe replacement, sparse writes,
|
|
metadata, and streaming I/O.
|
|
- `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`,
|
|
and `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`,
|
|
`BMK_benchTimedFn`) used by the CLI benchmark mode and by C test tools.
|
|
Both live in the `cli/` package, but C test binaries (fullbench, fuzzer,
|
|
zstreamtest, paramgrill, ...) link a helpers-only build of that archive,
|
|
produced without the package's `cli` feature, because the parser layer
|
|
requires the C `fileio` backend that tests do not compile. Benchmark
|
|
orchestration and reporting (`benchzstd.c`) remain C, reached from the
|
|
Rust parser through the `ZSTD_NOBENCH`-gated bridge in `zstdcli.c`.
|
|
|
|
The optimal block matcher, high-level frame compression, dictionary-building
|
|
except suffix-array construction, the legacy v0.2-v0.7 decoders, benchmark
|
|
orchestration (`benchzstd`), and the CLI file-I/O backend are still C. They
|
|
must move before the rewrite is complete. Keeping that boundary explicit
|
|
prevents a passing hybrid build from being mistaken for the final all-Rust
|
|
result.
|
|
|
|
## Legacy decoding
|
|
|
|
Each `lib/legacy/zstd_v0N.c` file is a frozen snapshot of the entropy coders
|
|
and frame logic of one historical release. The Rust ports in `src/legacy/`
|
|
keep that property: every version owns its own frozen FSE/Huff0 and frame
|
|
logic, ported line by line, and must never reuse the modern entropy modules
|
|
or share code with other legacy versions. Outputs and error codes must be
|
|
byte-identical to the original C files. Their only shared dependency is the
|
|
`errors` module, matching the C files' `error_private.h` include.
|
|
|
|
Cargo features `legacy-v01` .. `legacy-v07` gate the per-version modules and
|
|
are never default features. The build systems derive the feature list from
|
|
the C configuration:
|
|
|
|
- `lib/Makefile` and `programs/Makefile` map `ZSTD_LEGACY_SUPPORT=N` to the
|
|
features for versions >= N (0 disables legacy), matching the
|
|
`ZSTD_LEGACY_FILES` selection in `lib/libzstd.mk`.
|
|
- `tests/Makefile` always enables all seven features because the test
|
|
objects compile every `lib/legacy/*.c` file regardless of dispatch level.
|
|
- `build/meson` maps `legacy_level` like the makefiles; `build/cmake`
|
|
enables all seven whenever `ZSTD_LEGACY_SUPPORT` is on because it always
|
|
compiles all seven C files.
|
|
|
|
Every build system also encodes the legacy selection in the Rust target
|
|
directory name (for example `c1-d1-default-legacy5`), for the same reason the
|
|
HUF mode is encoded there: a cached archive built for one configuration must
|
|
never be linked into a build that expects another.
|
|
|
|
A feature whose version has not been ported yet gates nothing; the original
|
|
C file still provides that decoder, so mixed C/Rust legacy levels link
|
|
cleanly. Porting a version means adding `src/legacy/zstd_v0N.rs`, registering
|
|
it in `src/legacy/mod.rs` behind its feature, and reducing
|
|
`lib/legacy/zstd_v0N.c` to a declaration-only shim. For v0.1 the streaming
|
|
`ZSTDv01_Dctx` state lives entirely in Rust: C code only ever holds an opaque
|
|
pointer, so the C-side struct definition is gone.
|
|
|
|
## Compatibility boundary
|
|
|
|
The public ABI continues to come from the existing headers under `lib/`.
|
|
Exported Rust functions therefore use C layout and calling conventions. A C
|
|
source file whose implementation has moved to Rust remains in the original
|
|
makefile source list as a small shim so header configuration and platform
|
|
preprocessor behavior stay available during the transition.
|
|
|
|
The library, test, and program makefiles select an archive directory for the
|
|
active C configuration: enabled compression/decompression/dictionary-builder
|
|
modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit
|
|
C binaries. The
|
|
native static archive flattens Rust object members rather than nesting a Rust
|
|
archive, while the native shared library retains all migrated Rust exports.
|
|
When the HUF mode changes, the test and program paths also rebuild cached C
|
|
outputs before linking. This prevents original C tests from using a stale or
|
|
configuration-incompatible implementation.
|
|
|
|
## Validation
|
|
|
|
Run focused Rust checks from this directory:
|
|
|
|
```sh
|
|
cargo fmt --check
|
|
cargo clippy --all-targets -- -D warnings
|
|
cargo test --all-targets
|
|
cargo build --release
|
|
```
|
|
|
|
The program-only Rust archive has its own feature matrix and should be checked
|
|
from `rust/cli` as well:
|
|
|
|
```sh
|
|
cargo clippy --all-targets -- -D warnings
|
|
cargo test --all-targets
|
|
cargo test --no-default-features --features cli,compression --all-targets
|
|
cargo test --no-default-features --features cli,decompression --all-targets
|
|
cargo test --no-default-features --all-targets
|
|
```
|
|
|
|
Then run original compatibility tests from the repository root, starting with
|
|
the narrow target for the component being migrated. For example:
|
|
|
|
```sh
|
|
make -C tests fuzzer
|
|
./tests/fuzzer -i1 --no-big-tests
|
|
make -C tests test-rust-lib-smoke
|
|
```
|
|
|
|
Broader `tests/Makefile` targets remain the authoritative integration gates as
|
|
more of the library and CLI are rewritten.
|