Compare commits

...
Author SHA1 Message Date
ddidderr 322f99dfd0 feat(rust): port legacy v0.1 decoder
Port lib/legacy/zstd_v01.c (the frozen zstd v0.1 decoder) to
rust/src/legacy/zstd_v01.rs as the first legacy-format port on the new
scaffolding, and reduce the C file to a declaration-only shim that
keeps its header includes for configuration and platform preprocessor
behavior.

Frozen-decoder policy: zstd_v01.c embeds its own v0.1-era FSE and
Huff0 snapshot, distinct from every other release. The Rust port is a
line-by-line translation with the same table layouts (FSE_DTable as a
u32 header word plus packed newState/symbol/nbBits entries, the Huff0
u16 DTable with byte/nbBits pairs), the same arithmetic including
wrap-around and pointer-comparison quirks (e.g. the offset-vs-base
address check in ZSTD_execSequence), the same internal FSE error space
(size_t)-1..-7, and the same public ZSTD error codes. It reuses no
modern Rust entropy module; its only crate dependency is `errors`,
matching the C file's error_private.h include. The 32-bit-only reload
points are kept as compile-time conditions on usize::BITS.

Symbol takeover boundary: all nine ZSTDv01_* entry points from
zstd_v01.h now come from Rust as context-free #[no_mangle] extern "C"
functions (isError, decompress, decompressDCtx,
findFrameSizeInfoLegacy, createDCtx, freeDCtx, resetDCtx,
nextSrcSizeToDecompress, decompressContinue). zstd_legacy.h only uses
the first four for v0.1; streaming for v0.1-v0.3 intentionally returns
version_unsupported there, unchanged. The ZSTDv01_Dctx struct
definition moves entirely into Rust: C code only ever holds an opaque
pointer (zstd_v01.h forward-declares the type), and the context is
malloc/free-allocated exactly like the C version so create/free may
pair across the language boundary.

Byte-identity verification against the pristine pre-migration C build
(f8745da6, pure C, ZSTD_LEGACY_SUPPORT=1):

- Real v0.1 frames were generated by building the v0.1.0 git tag and
  compressing text, random, and 426 KB multi-block inputs. A one-shot
  ZSTD_decompress harness linked once against the pristine C libzstd.a
  and once against the Rust-backed libzstd.a produced bit-identical
  outputs for all frames.
- A direct ZSTDv01_* probe (one-shot decode, dst-too-small, truncated
  input, bad magic, findFrameSizeInfoLegacy, and the streaming
  continue loop) printed identical results, including exact error
  codes (-70 dstSize_tooSmall, -72 srcSize_wrong, -10 prefix_unknown)
  and identical dBound values.
- zstd -l -v on v0.1 files matches the pristine binary; CLI streaming
  decode of v0.1 fails with the same "Version not supported" in both,
  by design of zstd_legacy.h.

Unit tests embed three v0.1.0-generated fixtures (entropy-coded,
raw-block, and four-block frames) plus the truncation, bad-magic,
small-destination, and streaming-API cases, all asserting the exact C
error codes above. Note that `make -C tests test-legacy` only covers
v0.4+ frames, so the embedded fixtures and the harness comparison are
the actual v0.1 coverage.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets
  --features legacy-v01 -- -D warnings && cargo test --all-targets
  --features legacy-v01 (127 tests, 9 for v0.1)
- cargo clippy/test --no-default-features --features
  decompression,legacy-v01 (module builds standalone)
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests, also with
  ZSTD_LEGACY_SUPPORT=1 (mixed Rust v0.1 + C v0.2-0.7 link)
- make -C tests test-rust-lib-smoke && make -C tests test-legacy
- make -C programs zstd (default and ZSTD_LEGACY_SUPPORT=1); nm shows
  the nine ZSTDv01_* symbols provided by Rust at level 1
- make -C lib libzstd.a ZSTD_LEGACY_SUPPORT=0 (no legacy symbols) and
  meson -Dlegacy_level=1 shared library exporting all nine
2026-07-11 23:08:19 +02:00
ddidderr bdd35c838d build(rust): add legacy feature scaffolding
The legacy decoders (lib/legacy/zstd_v01.c .. zstd_v07.c) are next in
the Rust migration. Each of those files is a frozen snapshot of the
FSE/Huff0 entropy coders and frame logic of one historical release, so
their ports must not reuse the modern Rust entropy modules and must not
share code with each other: outputs and error codes have to stay
byte-identical to the frozen C forever. This commit installs the
build-system scaffolding so seven per-version ports can land
independently, each adding only its own module file plus a one-line
registration in rust/src/legacy/mod.rs.

Cargo grows features legacy-v01 .. legacy-v07. They are never default
features: the C build defaults differ per build system, so each build
system passes the list explicitly, derived from its own legacy
configuration:

- lib/Makefile and programs/Makefile map ZSTD_LEGACY_SUPPORT=N to the
  features for versions N..7 (0 disables legacy), mirroring the
  ZSTD_LEGACY_FILES selection in lib/libzstd.mk.
- tests/Makefile always enables all seven features because its
  ZSTDLEGACY_FILES wildcard compiles every lib/legacy/*.c regardless of
  the dispatch level.
- build/meson maps legacy_level exactly like the makefiles; build/cmake
  enables all seven whenever ZSTD_LEGACY_SUPPORT is ON because it
  always compiles all seven C files (ZSTD_LEGACY_LEVEL only selects the
  C dispatch).

Every build system also encodes the legacy selection in the Rust target
directory name (e.g. c1-d1-default-legacy5), for the same reason the
HUF mode is encoded there: a cached archive built for one configuration
must never be linked into a build expecting another. In tests/Makefile
the legacy level additionally flows into the existing HUF C-mode stamp,
so the flat C test objects (which bake -DZSTD_LEGACY_SUPPORT into the
dispatch) are rebuilt whenever the level changes. In programs/Makefile
the compress-only, decompress-only, and CLI archives keep
level-independent directories (RUST_HUF_MODE) because they are only
linked into ZSTD_LEGACY_SUPPORT=0 program variants and carry no legacy
features.

A feature whose version has not been ported yet gates nothing: the
module registration in rust/src/legacy/mod.rs is added by each port,
so enabling e.g. legacy-v05 today simply leaves that decoder in C.
This is what makes mixed C/Rust legacy levels link cleanly while the
seven ports land in any order.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets
  -- -D warnings && cargo test --all-targets
- cargo clippy with --no-default-features --features
  decompression,legacy-v01 and with all seven legacy features
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests
- make -C tests test-rust-lib-smoke; make -C tests test-legacy
- make -C lib libzstd.a with ZSTD_LEGACY_SUPPORT=0, 1 and default (5)
- cmake configure and meson setup (including -Dlegacy_level=1) emit the
  expected --features lists and legacy-suffixed target directories
2026-07-11 23:07:55 +02:00
ddidderr fef5f4478a feat(rust): port benchmark loop and CLI bench mode
Move the implementation of programs/benchfn.c into rust/src/benchfn.rs and
wire benchmark mode (-b/-e/-i) into the Rust CLI frontend, which previously
rejected those options as not yet implemented. `zstd -b1 -i0 FILE` and range
runs like `zstd -b5e6 -i0 FILE` work again, including the synthetic-sample
benchmark when no file is given.

benchfn.rs is a faithful port of the run/timing state machine:
BMK_benchFunction keeps the exact loop accounting (first-loop blockResults
and errorFn checks, dstSize summed on the first loop only, 0xE5 warm-up of
result buffers, nbLoops minimum of 1) and BMK_benchTimedFn keeps the same
convergence behavior (x10 workload growth for short runs, budget-based
nbLoops estimation, runs below half the run budget re-tried rather than
reported, best qualifying run returned). Arithmetic that C leaves to
unsigned wrap-around uses wrapping operations so debug builds cannot panic
where release C would wrap.

ABI notes: BMK_runTime_t and BMK_runOutcome_t are returned by value across
the C boundary and BMK_benchParams_t is passed by value, so all three are
repr(C) mirrors of benchfn.h; their field offsets are pinned by const
asserts in Rust and matching C static asserts in the benchfn.c shim, which
is now declaration-only. BMK_timedFnState_t stays opaque, fits the 64-byte
BMK_timedFnState_shell (compile-time checked), and is malloc/free-managed
so creation and destruction remain interchangeable with C callers.

The CLI parses -b (bench mode), -e (range end, digits attach directly,
defaulting to 0 like readU32FromChar) and -i (duration in seconds), then
dispatches through a new ZSTD_rust_cli_bench bridge in the zstdcli.c shim.
The bridge exists because benchmark availability is a C preprocessor
property (ZSTD_NOBENCH): orchestration and reporting stay in C benchzstd.c,
stripped variants (zstd-small, zstd-compress, zstd-decompress) compile the
stub branch and report "benchmark mode is not available in this build", and
the Rust side never references benchmark symbols directly. Level clamping
against ZSTD_maxCLevel() happens in the bridge, where the symbol is
guaranteed to exist whenever benchmarking is compiled in. -T selects the
worker count, defaulting to single-threaded like the C bench path; -S
(separate files) and --priority=rt remain unimplemented.

Makefile updates only extend the Rust source prerequisite lists with
benchfn.rs; the helpers-archive plumbing from the timefn commit already
links fullbench(-lib/-dll/32) and paramgrill, the benchfn consumers among
the C tests. Original C test sources are untouched.

Known pre-existing issues, unchanged by this commit: tests/fullbench-lib
fails to link at the base commit too (libzstd.a precedes fullbench.c in its
link line), and the cli-tests basic/help.sh, compression/levels.sh,
compression/golden.sh, and decompression/pass-through.sh scripts fail
identically with a base-commit binary because the Rust CLI frontend is
still a partial reimplementation.

Test Plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets -- -D warnings
  && cargo test --all-targets && cargo build --release
- cd rust/cli && cargo fmt --check && cargo clippy --all-targets -- -D
  warnings && cargo test --all-targets; repeat tests with
  --no-default-features plus features compression / decompression / (none)
- make -C programs zstd; ./programs/zstd -b1 -i0 lib/common/xxhash.c;
  ./programs/zstd -b5e6 -i0 programs/fileio.c; ./programs/zstd -b1 -i0
  (synthetic); echo roundtrip via zstd | zstd -d
- make -C programs zstd-small zstd-compress zstd-decompress zstd-nolegacy
  zstd-dictBuilder; zstd-small -b reports benchmark unavailable; compress/
  decompress roundtrip across the split binaries
- make -C tests fullbench fuzzer zstreamtest paramgrill decodecorpus
  poolTests fullbench32 fuzzer32; ./tests/fullbench -i0 (exercises Rust
  BMK_benchTimedFn from C); ./tests/fullbench32 -i0; ./tests/fuzzer -i1
  --no-big-tests; ./tests/poolTests; make -C tests test-rust-lib-smoke
- cli-tests subset: basic/version.sh, compression/basic.sh,
  compression/multiple-files.sh pass; failing scripts match the base commit

Refs: rust/README.md
2026-07-11 14:26:21 +02:00
ddidderr 54f5c29742 feat(rust): port program timing helpers
Move the implementation of programs/timefn.c into rust/src/timefn.rs. The
file provides the monotonic nanosecond clock (UTIL_getTime, span helpers,
UTIL_waitForNextTick, UTIL_support_MT_measurements) used by the CLI and by
several C test tools. timefn.c remains as a declaration-only shim so the
original source lists and header configuration keep working, and it pins the
ABI with static asserts: UTIL_time_t is returned by value and must stay a
plain 64-bit counter, which the Rust #[repr(C)] mirror also asserts.

Platform selection mirrors the C preprocessor structure: Windows uses
QueryPerformanceCounter, Apple targets use mach_absolute_time, and other
POSIX systems use libc clock_gettime(CLOCK_MONOTONIC). Only the unix path is
exercised by this environment; the Windows and Apple paths are written from
the C source and compile-checked logically but are untested here. The C90
clock() fallback is unreachable on Rust-supported targets, so multi-threaded
measurement support is always reported.

The symbols live in the program-only zstd-cli-rs package, keeping them out
of library builds. Linking that archive into C test binaries surfaced a
structural problem: rustc's local ThinLTO promotes internal symbols across
codegen units, so extracting the timefn object could drag in the zstd_cli
parser object, whose FIO_* externs test binaries cannot satisfy. The parser
is therefore gated behind a new additive `cli` cargo feature (default on).
Program archives build with cli,compression,decompression as before, while
tests/Makefile links a helpers-only archive (rust/target/cli-helpers) built
with --no-default-features, which contains no fileio references at all.

tests/Makefile gains build rules for the helpers archive and adds it as a
prerequisite of every binary that compiles the timefn shim: fullbench(32),
fullbench-lib, fullbench-dll, fuzzer(32), zstreamtest(32/asan/tsan/ubsan),
paramgrill, decodecorpus, and poolTests. Prerequisite order places the
archive after all C objects in `$^` link lines; the known-broken -dll
recipes filter to %.c, so they name the archive explicitly. Original C test
sources are untouched; only link inputs changed.

zstd-cli-rs now depends on libc (already used by the core crate) for
clock_gettime and the Mach timebase bindings.

Test Plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets -- -D warnings
  && cargo test --all-targets && cargo build --release
- cd rust/cli && cargo fmt --check && cargo clippy --all-targets -- -D
  warnings && cargo test --all-targets; repeat tests with
  --no-default-features plus features compression / decompression / (none)
- make -C programs zstd; roundtrip echo hello | zstd | zstd -d
- make -C tests fullbench fuzzer zstreamtest paramgrill decodecorpus
  poolTests; ./tests/fullbench -i0; ./tests/fuzzer -i1 --no-big-tests;
  ./tests/poolTests; make -C tests test-rust-lib-smoke
- verified with nm that the helpers archive member defining UTIL_getTime has
  no FIO_*/ZSTD_* undefined references

Refs: rust/README.md
2026-07-11 14:25:32 +02:00
ddidderr caf12dda22 feat(rust): port divsufsort
Move the dictionary builder's suffix-array construction from
lib/dictBuilder/divsufsort.c to rust/src/divsufsort.rs, the first
dictBuilder module to migrate.  It rides on the dict-builder cargo
feature dimension introduced by the previous commit.

divsufsort() is a self-contained algorithm (two-stage sort of type-B*
substrings via sssort, rank refinement via trsort, then induced sorting
of the full array), so its context-free signature allows a direct symbol
takeover: the Rust #[no_mangle] export provides the existing `divsufsort`
symbol and the C file becomes a declaration-only shim that just keeps the
header's prototypes in the build.  Only divsufsort() moved; divbwt() has
no callers anywhere in zstd, so it is now declaration-only, keeping the
Rust export surface minimal.  The unused openMP parameter is retained for
signature compatibility (zstd never defines LIBBSC_OPENMP).

The port is a mechanical translation of the exact configuration zstd
compiles: ALPHABET_SIZE=256, SS_INSERTIONSORT_THRESHOLD=8,
SS_BLOCKSIZE=1024, SS_MISORT_STACKSIZE=16, SS_SMERGE_STACKSIZE=32,
TR_STACKSIZE=64.  Every C `int*` cursor into the SA buffer becomes an
`isize` index into a single `&mut [i32]` slice, preserving the pointer
arithmetic (including transient one-before-the-range cursors and the
bitwise-complement rank marking) while staying bounds-checked; all value
arithmetic keeps C int semantics.  The C -1/-2 error results are
preserved, with Vec::try_reserve_exact standing in for the bucket-array
malloc failure path.  Behavior is bit-identical by construction and by
measurement (see test plan); runtime on an 11 MB training buffer is
within ~5% of the C build end-to-end.

Users see no behavioral change: dictionaries trained through
ZDICT_trainFromBuffer_legacy() are byte-identical to the C build.  The
only external difference is that the never-called `divbwt` symbol is no
longer defined in the library.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets
  -- -D warnings && cargo test --all-targets && cargo build --release
  (125 tests pass; new unit tests cover empty/one/two-byte inputs,
  all-equal bytes, an exact hand-computed "abracadabra" SA, and
  fixed-seed LCG buffers at 256-, 4-, and 2-symbol alphabets verified
  against a naive reference sort plus permutation/sorted invariants)
- Feature matrix: cargo build --release --no-default-features
  --features compression,decompression (and decompression-only,
  compression-only, compression,dict-builder); `divsufsort` is exported
  only when dict-builder is enabled
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests (includes
  ZDICT training tests): pass
- make -C tests test-rust-lib-smoke: pass
- make -C tests test-invalidDictionaries: pass
- make -C programs zstd zstd-dictBuilder zstd-small zstd-compress
  zstd-decompress: build; compress/decompress round-trip verified
- Byte-identity vs pristine C build (commit 959e4852): a harness calling
  ZDICT_trainFromBuffer_legacy() (the only zstd path reaching
  divsufsort) and divsufsort() directly, linked against both libzstd.a
  builds, produces byte-identical dictionaries (80,288 B and full
  112,640 B capacity) and byte-identical suffix arrays on a 1 MB source
  set and an 11 MB binary/repetitive set; a differential driver over 148
  random and structured buffers (sizes 3..6000, alphabets 1..256,
  Fibonacci word, sawtooth, 6 KB near-constant) shows zero mismatches.
  The CLI --train path could not be exercised because the Rust CLI
  frontend rejects --train in both the pristine and ported builds (a
  pre-existing migration gap unrelated to this change).
2026-07-11 14:23:39 +02:00
ddidderr 91c80fc8e7 build(rust): add dict-builder cargo feature dimension
The dictionary-builder sources (lib/dictBuilder) are about to start moving
to Rust, beginning with divsufsort. The Rust crate previously only modeled
the compression/decompression module split plus the forced-HUF decoder
modes, so no build could express "this C configuration includes (or
excludes) dictBuilder" to Cargo. Without that, a Rust archive could carry
dictBuilder modules into a build whose C side disabled them, or worse,
omit a migrated implementation from a build whose C shims require it.

Add a `dict-builder` cargo feature and thread it through every build that
consumes the Rust static archive, mirroring exactly how each build system
already gates the dictBuilder C sources:

- rust/Cargo.toml: new `dict-builder` feature, included in the default
  set because the C library builds dictBuilder by default
  (ZSTD_LIB_DICTBUILDER ?= 1). The feature is empty until the first
  dictBuilder module lands.
- lib/Makefile: RUST_CARGO_FEATURES gains dict-builder when
  ZSTD_LIB_DICTBUILDER is enabled, following the existing
  ZSTD_LIB_COMPRESSION/ZSTD_LIB_DECOMPRESSION pattern. The archive
  directory naming grows a matching `b<0|1>` dimension
  (c1-d1-b1-default etc.) so differently configured archives never
  collide; the repeated config prefix is factored into
  RUST_MODULE_CONFIG.
- programs/Makefile: the full-featured archives now request
  compression,decompression,dict-builder (equal to the default set, so
  the target directory stays shared with tests). The partial-library
  variants gain the `b0` name dimension, and zstd-dictBuilder gets its
  own lib-c1-d0-b1 archive because it compiles the dictBuilder C sources
  without decompression; it previously shared the compression-only
  archive, which will lack the migrated dictBuilder symbols.
- tests/Makefile: no flag change needed since tests use the crate default
  feature set; a comment now records that dict-builder arrives that way.
- build/cmake/lib/CMakeLists.txt: ZSTD_BUILD_DICTBUILDER now adds the
  dict-builder feature and a `b<0|1>` component in the Rust build-config
  directory name, in lockstep with the DictBuilderSources gating.
- build/meson/lib/meson.build: meson compiles the dictBuilder sources
  unconditionally, so the feature list and config name gain dict-builder
  unconditionally (c1-d1-b1-<huf-mode>).

The `dict-builder` feature deliberately does not imply `compression`.
lib/Makefile forces ZSTD_LIB_DICTBUILDER=0 when compression is disabled,
but CMake does not couple the two options, so encoding the C-side
constraint in Cargo would make the Rust archive diverge from the C source
list in that (already unsupported) CMake configuration.

Test plan:
- cd rust && cargo build --release
- cargo build --release --no-default-features \
    --features compression,decompression
- cargo build --release --no-default-features \
    --features compression,dict-builder
- Full validation (fuzzer, smoke tests, dictionary byte-identity) runs
  with the follow-up commit that ports divsufsort onto this scaffolding.
2026-07-11 14:23:39 +02:00
23 changed files with 6408 additions and 4434 deletions
+1
View File
@@ -34,6 +34,7 @@ install/
# Build artefacts
/rust/target/
/rust/cli/target/
contrib/linux-kernel/linux/
projects/
bin/
+18 -1
View File
@@ -142,6 +142,7 @@ endif()
set(_zstd_rust_features)
set(_zstd_rust_compression 0)
set(_zstd_rust_decompression 0)
set(_zstd_rust_dictbuilder 0)
if(ZSTD_BUILD_COMPRESSION)
list(APPEND _zstd_rust_features compression)
set(_zstd_rust_compression 1)
@@ -150,6 +151,10 @@ if(ZSTD_BUILD_DECOMPRESSION)
list(APPEND _zstd_rust_features decompression)
set(_zstd_rust_decompression 1)
endif()
if(ZSTD_BUILD_DICTBUILDER)
list(APPEND _zstd_rust_features dict-builder)
set(_zstd_rust_dictbuilder 1)
endif()
set(_zstd_rust_huf_mode default)
if(_zstd_huf_force_x1)
@@ -164,6 +169,18 @@ elseif(_zstd_huf_force_x2)
endif()
endif()
# This CMake build compiles every lib/legacy/zstd_v0N.c whenever legacy
# support is enabled (ZSTD_LEGACY_LEVEL only selects the C dispatch), so the
# Rust archive enables every per-version legacy feature to match. The build
# configuration encodes the switch so archives never mix.
set(_zstd_rust_legacy 0)
if(ZSTD_LEGACY_SUPPORT)
set(_zstd_rust_legacy 1)
list(APPEND _zstd_rust_features
legacy-v01 legacy-v02 legacy-v03 legacy-v04
legacy-v05 legacy-v06 legacy-v07)
endif()
set(_zstd_rust_target "${ZSTD_RUST_TARGET}")
if(NOT _zstd_rust_target AND CMAKE_SYSTEM_NAME STREQUAL "Linux"
AND CMAKE_SIZEOF_VOID_P EQUAL 4)
@@ -190,7 +207,7 @@ if(_zstd_rust_features)
endif()
set(_zstd_rust_build_config
"c${_zstd_rust_compression}-d${_zstd_rust_decompression}-${_zstd_rust_huf_mode}")
"c${_zstd_rust_compression}-d${_zstd_rust_decompression}-b${_zstd_rust_dictbuilder}-${_zstd_rust_huf_mode}-legacy${_zstd_rust_legacy}")
set(ZSTD_RUST_MANIFEST "${ZSTD_SOURCE_DIR}/rust/Cargo.toml")
set(ZSTD_RUST_TARGET_DIR
"${CMAKE_CURRENT_BINARY_DIR}/rust-target/${_zstd_rust_build_config}")
+14 -2
View File
@@ -73,7 +73,9 @@ if rust_huf_force_x1 and rust_huf_force_x2
error('HUF_FORCE_DECOMPRESS_X1 and HUF_FORCE_DECOMPRESS_X2 are mutually exclusive')
endif
rust_features = ['compression', 'decompression']
# Meson always compiles the dictBuilder sources above, so the Rust archive
# must always carry the matching dict-builder module set.
rust_features = ['compression', 'decompression', 'dict-builder']
rust_huf_mode = 'default'
rust_huf_c_args = []
if rust_huf_force_x1
@@ -86,6 +88,16 @@ elif rust_huf_force_x2
rust_huf_c_args += '-DHUF_FORCE_DECOMPRESS_X2'
endif
# Mirror the legacy source selection below: legacy_level N compiles
# lib/legacy/zstd_v0N.c .. zstd_v07.c, so the Rust archive enables the
# matching per-version features. The build configuration encodes the level
# so archives built for different legacy levels never mix.
foreach i : [1, 2, 3, 4, 5, 6, 7]
if legacy_level != 0 and legacy_level <= i
rust_features += 'legacy-v0@0@'.format(i)
endif
endforeach
rust_target = get_option('zstd_rust_target')
if rust_target == ''
if host_machine_os == os_linux and \
@@ -96,7 +108,7 @@ if rust_target == ''
endif
endif
rust_build_config = 'c1-d1-' + rust_huf_mode
rust_build_config = 'c1-d1-b1-' + rust_huf_mode + '-legacy@0@'.format(legacy_level)
rust_target_dir = join_paths(meson.current_build_dir(), 'rust-target', rust_build_config)
is_msvc = cc_id == compiler_msvc or cc_id == 'clang-cl'
rust_staticlib_name = is_msvc ? 'zstd_rs.lib' : 'libzstd_rs.a'
+19 -3
View File
@@ -84,22 +84,38 @@ endif
ifneq ($(ZSTD_LIB_DECOMPRESSION),0)
RUST_CARGO_FEATURES += decompression
endif
ifneq ($(ZSTD_LIB_DICTBUILDER),0)
RUST_CARGO_FEATURES += dict-builder
endif
RUST_MODULE_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-b$(ZSTD_LIB_DICTBUILDER)
RUST_HUF_FEATURE :=
RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-default
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-default
ifneq ($(RUST_HUF_FORCE_X1),0)
RUST_HUF_FEATURE := huf-force-decompress-x1
RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-huf-force-decompress-x1
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-huf-force-decompress-x1
endif
ifneq ($(RUST_HUF_FORCE_X2),0)
RUST_HUF_FEATURE := huf-force-decompress-x2
RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-huf-force-decompress-x2
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-huf-force-decompress-x2
endif
ifneq ($(RUST_HUF_FEATURE),)
ifneq ($(ZSTD_LIB_DECOMPRESSION),0)
RUST_CARGO_FEATURES += $(RUST_HUF_FEATURE)
endif
endif
# Legacy decoders follow the ZSTD_LEGACY_FILES selection exactly: level N
# enables versions v0.N .. v0.7 (0 disables legacy). The build directory
# also encodes the level, so archives for different legacy levels never mix.
RUST_LEGACY_FEATURES :=
ifneq ($(ZSTD_LEGACY_SUPPORT), 0)
ifeq ($(shell test $(ZSTD_LEGACY_SUPPORT) -lt 8; echo $$?), 0)
RUST_LEGACY_FEATURES := $(addprefix legacy-v0,$(wordlist $(ZSTD_LEGACY_SUPPORT),7,1 2 3 4 5 6 7))
endif
endif
RUST_CARGO_FEATURES += $(RUST_LEGACY_FEATURES)
RUST_BUILD_CONFIG := $(RUST_BUILD_CONFIG)-legacy$(ZSTD_LEGACY_SUPPORT)
RUST_CARGO_FEATURES := $(subst $(space),$(comma),$(strip $(RUST_CARGO_FEATURES)))
RUST_TARGET ?=
+5 -1886
View File
@@ -24,1890 +24,9 @@
* OTHER DEALINGS IN THE SOFTWARE.
*/
/*- Compiler specifics -*/
#ifdef __clang__
#pragma clang diagnostic ignored "-Wshorten-64-to-32"
#endif
#if defined(_MSC_VER)
# pragma warning(disable : 4244)
# pragma warning(disable : 4127) /* C4127 : Condition expression is constant */
#endif
/*- Dependencies -*/
#include <assert.h>
#include <stdio.h>
#include <stdlib.h>
/* divsufsort() is implemented in rust/src/divsufsort.rs, which provides the
* symbol directly. This translation unit keeps the header's prototypes in
* the build so the dictionary builder continues to compile against the
* original interface. divbwt() has no callers in zstd and is declaration-
* only; it moves to Rust if a user ever appears. */
#include "divsufsort.h"
/*- Constants -*/
#if defined(INLINE)
# undef INLINE
#endif
#if !defined(INLINE)
# define INLINE __inline
#endif
#if defined(ALPHABET_SIZE) && (ALPHABET_SIZE < 1)
# undef ALPHABET_SIZE
#endif
#if !defined(ALPHABET_SIZE)
# define ALPHABET_SIZE (256)
#endif
#define BUCKET_A_SIZE (ALPHABET_SIZE)
#define BUCKET_B_SIZE (ALPHABET_SIZE * ALPHABET_SIZE)
#if defined(SS_INSERTIONSORT_THRESHOLD)
# if SS_INSERTIONSORT_THRESHOLD < 1
# undef SS_INSERTIONSORT_THRESHOLD
# define SS_INSERTIONSORT_THRESHOLD (1)
# endif
#else
# define SS_INSERTIONSORT_THRESHOLD (8)
#endif
#if defined(SS_BLOCKSIZE)
# if SS_BLOCKSIZE < 0
# undef SS_BLOCKSIZE
# define SS_BLOCKSIZE (0)
# elif 32768 <= SS_BLOCKSIZE
# undef SS_BLOCKSIZE
# define SS_BLOCKSIZE (32767)
# endif
#else
# define SS_BLOCKSIZE (1024)
#endif
/* minstacksize = log(SS_BLOCKSIZE) / log(3) * 2 */
#if SS_BLOCKSIZE == 0
# define SS_MISORT_STACKSIZE (96)
#elif SS_BLOCKSIZE <= 4096
# define SS_MISORT_STACKSIZE (16)
#else
# define SS_MISORT_STACKSIZE (24)
#endif
#define SS_SMERGE_STACKSIZE (32)
#define TR_INSERTIONSORT_THRESHOLD (8)
#define TR_STACKSIZE (64)
/*- Macros -*/
#ifndef SWAP
# define SWAP(_a, _b) do { t = (_a); (_a) = (_b); (_b) = t; } while(0)
#endif /* SWAP */
#ifndef MIN
# define MIN(_a, _b) (((_a) < (_b)) ? (_a) : (_b))
#endif /* MIN */
#ifndef MAX
# define MAX(_a, _b) (((_a) > (_b)) ? (_a) : (_b))
#endif /* MAX */
#define STACK_PUSH(_a, _b, _c, _d)\
do {\
assert(ssize < STACK_SIZE);\
stack[ssize].a = (_a), stack[ssize].b = (_b),\
stack[ssize].c = (_c), stack[ssize++].d = (_d);\
} while(0)
#define STACK_PUSH5(_a, _b, _c, _d, _e)\
do {\
assert(ssize < STACK_SIZE);\
stack[ssize].a = (_a), stack[ssize].b = (_b),\
stack[ssize].c = (_c), stack[ssize].d = (_d), stack[ssize++].e = (_e);\
} while(0)
#define STACK_POP(_a, _b, _c, _d)\
do {\
assert(0 <= ssize);\
if(ssize == 0) { return; }\
(_a) = stack[--ssize].a, (_b) = stack[ssize].b,\
(_c) = stack[ssize].c, (_d) = stack[ssize].d;\
} while(0)
#define STACK_POP5(_a, _b, _c, _d, _e)\
do {\
assert(0 <= ssize);\
if(ssize == 0) { return; }\
(_a) = stack[--ssize].a, (_b) = stack[ssize].b,\
(_c) = stack[ssize].c, (_d) = stack[ssize].d, (_e) = stack[ssize].e;\
} while(0)
#define BUCKET_A(_c0) bucket_A[(_c0)]
#if ALPHABET_SIZE == 256
#define BUCKET_B(_c0, _c1) (bucket_B[((_c1) << 8) | (_c0)])
#define BUCKET_BSTAR(_c0, _c1) (bucket_B[((_c0) << 8) | (_c1)])
#else
#define BUCKET_B(_c0, _c1) (bucket_B[(_c1) * ALPHABET_SIZE + (_c0)])
#define BUCKET_BSTAR(_c0, _c1) (bucket_B[(_c0) * ALPHABET_SIZE + (_c1)])
#endif
/*- Private Functions -*/
static const int lg_table[256]= {
-1,0,1,1,2,2,2,2,3,3,3,3,3,3,3,3,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,
5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7
};
#if (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE)
static INLINE
int
ss_ilg(int n) {
#if SS_BLOCKSIZE == 0
return (n & 0xffff0000) ?
((n & 0xff000000) ?
24 + lg_table[(n >> 24) & 0xff] :
16 + lg_table[(n >> 16) & 0xff]) :
((n & 0x0000ff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff]);
#elif SS_BLOCKSIZE < 256
return lg_table[n];
#else
return (n & 0xff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff];
#endif
}
#endif /* (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE) */
#if SS_BLOCKSIZE != 0
static const int sqq_table[256] = {
0, 16, 22, 27, 32, 35, 39, 42, 45, 48, 50, 53, 55, 57, 59, 61,
64, 65, 67, 69, 71, 73, 75, 76, 78, 80, 81, 83, 84, 86, 87, 89,
90, 91, 93, 94, 96, 97, 98, 99, 101, 102, 103, 104, 106, 107, 108, 109,
110, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126,
128, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 144, 145, 146, 147, 148, 149, 150, 150, 151, 152, 153, 154, 155, 155,
156, 157, 158, 159, 160, 160, 161, 162, 163, 163, 164, 165, 166, 167, 167, 168,
169, 170, 170, 171, 172, 173, 173, 174, 175, 176, 176, 177, 178, 178, 179, 180,
181, 181, 182, 183, 183, 184, 185, 185, 186, 187, 187, 188, 189, 189, 190, 191,
192, 192, 193, 193, 194, 195, 195, 196, 197, 197, 198, 199, 199, 200, 201, 201,
202, 203, 203, 204, 204, 205, 206, 206, 207, 208, 208, 209, 209, 210, 211, 211,
212, 212, 213, 214, 214, 215, 215, 216, 217, 217, 218, 218, 219, 219, 220, 221,
221, 222, 222, 223, 224, 224, 225, 225, 226, 226, 227, 227, 228, 229, 229, 230,
230, 231, 231, 232, 232, 233, 234, 234, 235, 235, 236, 236, 237, 237, 238, 238,
239, 240, 240, 241, 241, 242, 242, 243, 243, 244, 244, 245, 245, 246, 246, 247,
247, 248, 248, 249, 249, 250, 250, 251, 251, 252, 252, 253, 253, 254, 254, 255
};
static INLINE
int
ss_isqrt(int x) {
int y, e;
if(x >= (SS_BLOCKSIZE * SS_BLOCKSIZE)) { return SS_BLOCKSIZE; }
e = (x & 0xffff0000) ?
((x & 0xff000000) ?
24 + lg_table[(x >> 24) & 0xff] :
16 + lg_table[(x >> 16) & 0xff]) :
((x & 0x0000ff00) ?
8 + lg_table[(x >> 8) & 0xff] :
0 + lg_table[(x >> 0) & 0xff]);
if(e >= 16) {
y = sqq_table[x >> ((e - 6) - (e & 1))] << ((e >> 1) - 7);
if(e >= 24) { y = (y + 1 + x / y) >> 1; }
y = (y + 1 + x / y) >> 1;
} else if(e >= 8) {
y = (sqq_table[x >> ((e - 6) - (e & 1))] >> (7 - (e >> 1))) + 1;
} else {
return sqq_table[x] >> 4;
}
return (x < (y * y)) ? y - 1 : y;
}
#endif /* SS_BLOCKSIZE != 0 */
/*---------------------------------------------------------------------------*/
/* Compares two suffixes. */
static INLINE
int
ss_compare(const unsigned char *T,
const int *p1, const int *p2,
int depth) {
const unsigned char *U1, *U2, *U1n, *U2n;
for(U1 = T + depth + *p1,
U2 = T + depth + *p2,
U1n = T + *(p1 + 1) + 2,
U2n = T + *(p2 + 1) + 2;
(U1 < U1n) && (U2 < U2n) && (*U1 == *U2);
++U1, ++U2) {
}
return U1 < U1n ?
(U2 < U2n ? *U1 - *U2 : 1) :
(U2 < U2n ? -1 : 0);
}
/*---------------------------------------------------------------------------*/
#if (SS_BLOCKSIZE != 1) && (SS_INSERTIONSORT_THRESHOLD != 1)
/* Insertionsort for small size groups */
static
void
ss_insertionsort(const unsigned char *T, const int *PA,
int *first, int *last, int depth) {
int *i, *j;
int t;
int r;
for(i = last - 2; first <= i; --i) {
for(t = *i, j = i + 1; 0 < (r = ss_compare(T, PA + t, PA + *j, depth));) {
do { *(j - 1) = *j; } while((++j < last) && (*j < 0));
if(last <= j) { break; }
}
if(r == 0) { *j = ~*j; }
*(j - 1) = t;
}
}
#endif /* (SS_BLOCKSIZE != 1) && (SS_INSERTIONSORT_THRESHOLD != 1) */
/*---------------------------------------------------------------------------*/
#if (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE)
static INLINE
void
ss_fixdown(const unsigned char *Td, const int *PA,
int *SA, int i, int size) {
int j, k;
int v;
int c, d, e;
for(v = SA[i], c = Td[PA[v]]; (j = 2 * i + 1) < size; SA[i] = SA[k], i = k) {
d = Td[PA[SA[k = j++]]];
if(d < (e = Td[PA[SA[j]]])) { k = j; d = e; }
if(d <= c) { break; }
}
SA[i] = v;
}
/* Simple top-down heapsort. */
static
void
ss_heapsort(const unsigned char *Td, const int *PA, int *SA, int size) {
int i, m;
int t;
m = size;
if((size % 2) == 0) {
m--;
if(Td[PA[SA[m / 2]]] < Td[PA[SA[m]]]) { SWAP(SA[m], SA[m / 2]); }
}
for(i = m / 2 - 1; 0 <= i; --i) { ss_fixdown(Td, PA, SA, i, m); }
if((size % 2) == 0) { SWAP(SA[0], SA[m]); ss_fixdown(Td, PA, SA, 0, m); }
for(i = m - 1; 0 < i; --i) {
t = SA[0], SA[0] = SA[i];
ss_fixdown(Td, PA, SA, 0, i);
SA[i] = t;
}
}
/*---------------------------------------------------------------------------*/
/* Returns the median of three elements. */
static INLINE
int *
ss_median3(const unsigned char *Td, const int *PA,
int *v1, int *v2, int *v3) {
int *t;
if(Td[PA[*v1]] > Td[PA[*v2]]) { SWAP(v1, v2); }
if(Td[PA[*v2]] > Td[PA[*v3]]) {
if(Td[PA[*v1]] > Td[PA[*v3]]) { return v1; }
else { return v3; }
}
return v2;
}
/* Returns the median of five elements. */
static INLINE
int *
ss_median5(const unsigned char *Td, const int *PA,
int *v1, int *v2, int *v3, int *v4, int *v5) {
int *t;
if(Td[PA[*v2]] > Td[PA[*v3]]) { SWAP(v2, v3); }
if(Td[PA[*v4]] > Td[PA[*v5]]) { SWAP(v4, v5); }
if(Td[PA[*v2]] > Td[PA[*v4]]) { SWAP(v2, v4); SWAP(v3, v5); }
if(Td[PA[*v1]] > Td[PA[*v3]]) { SWAP(v1, v3); }
if(Td[PA[*v1]] > Td[PA[*v4]]) { SWAP(v1, v4); SWAP(v3, v5); }
if(Td[PA[*v3]] > Td[PA[*v4]]) { return v4; }
return v3;
}
/* Returns the pivot element. */
static INLINE
int *
ss_pivot(const unsigned char *Td, const int *PA, int *first, int *last) {
int *middle;
int t;
t = last - first;
middle = first + t / 2;
if(t <= 512) {
if(t <= 32) {
return ss_median3(Td, PA, first, middle, last - 1);
} else {
t >>= 2;
return ss_median5(Td, PA, first, first + t, middle, last - 1 - t, last - 1);
}
}
t >>= 3;
first = ss_median3(Td, PA, first, first + t, first + (t << 1));
middle = ss_median3(Td, PA, middle - t, middle, middle + t);
last = ss_median3(Td, PA, last - 1 - (t << 1), last - 1 - t, last - 1);
return ss_median3(Td, PA, first, middle, last);
}
/*---------------------------------------------------------------------------*/
/* Binary partition for substrings. */
static INLINE
int *
ss_partition(const int *PA,
int *first, int *last, int depth) {
int *a, *b;
int t;
for(a = first - 1, b = last;;) {
for(; (++a < b) && ((PA[*a] + depth) >= (PA[*a + 1] + 1));) { *a = ~*a; }
for(; (a < --b) && ((PA[*b] + depth) < (PA[*b + 1] + 1));) { }
if(b <= a) { break; }
t = ~*b;
*b = *a;
*a = t;
}
if(first < a) { *first = ~*first; }
return a;
}
/* Multikey introsort for medium size groups. */
static
void
ss_mintrosort(const unsigned char *T, const int *PA,
int *first, int *last,
int depth) {
#define STACK_SIZE SS_MISORT_STACKSIZE
struct { int *a, *b, c; int d; } stack[STACK_SIZE];
const unsigned char *Td;
int *a, *b, *c, *d, *e, *f;
int s, t;
int ssize;
int limit;
int v, x = 0;
for(ssize = 0, limit = ss_ilg(last - first);;) {
if((last - first) <= SS_INSERTIONSORT_THRESHOLD) {
#if 1 < SS_INSERTIONSORT_THRESHOLD
if(1 < (last - first)) { ss_insertionsort(T, PA, first, last, depth); }
#endif
STACK_POP(first, last, depth, limit);
continue;
}
Td = T + depth;
if(limit-- == 0) { ss_heapsort(Td, PA, first, last - first); }
if(limit < 0) {
for(a = first + 1, v = Td[PA[*first]]; a < last; ++a) {
if((x = Td[PA[*a]]) != v) {
if(1 < (a - first)) { break; }
v = x;
first = a;
}
}
if(Td[PA[*first] - 1] < v) {
first = ss_partition(PA, first, a, depth);
}
if((a - first) <= (last - a)) {
if(1 < (a - first)) {
STACK_PUSH(a, last, depth, -1);
last = a, depth += 1, limit = ss_ilg(a - first);
} else {
first = a, limit = -1;
}
} else {
if(1 < (last - a)) {
STACK_PUSH(first, a, depth + 1, ss_ilg(a - first));
first = a, limit = -1;
} else {
last = a, depth += 1, limit = ss_ilg(a - first);
}
}
continue;
}
/* choose pivot */
a = ss_pivot(Td, PA, first, last);
v = Td[PA[*a]];
SWAP(*first, *a);
/* partition */
for(b = first; (++b < last) && ((x = Td[PA[*b]]) == v);) { }
if(((a = b) < last) && (x < v)) {
for(; (++b < last) && ((x = Td[PA[*b]]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
}
for(c = last; (b < --c) && ((x = Td[PA[*c]]) == v);) { }
if((b < (d = c)) && (x > v)) {
for(; (b < --c) && ((x = Td[PA[*c]]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
for(; b < c;) {
SWAP(*b, *c);
for(; (++b < c) && ((x = Td[PA[*b]]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
for(; (b < --c) && ((x = Td[PA[*c]]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
if(a <= d) {
c = b - 1;
if((s = a - first) > (t = b - a)) { s = t; }
for(e = first, f = b - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
if((s = d - c) > (t = last - d - 1)) { s = t; }
for(e = b, f = last - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
a = first + (b - a), c = last - (d - c);
b = (v <= Td[PA[*a] - 1]) ? a : ss_partition(PA, a, c, depth);
if((a - first) <= (last - c)) {
if((last - c) <= (c - b)) {
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
STACK_PUSH(c, last, depth, limit);
last = a;
} else if((a - first) <= (c - b)) {
STACK_PUSH(c, last, depth, limit);
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
last = a;
} else {
STACK_PUSH(c, last, depth, limit);
STACK_PUSH(first, a, depth, limit);
first = b, last = c, depth += 1, limit = ss_ilg(c - b);
}
} else {
if((a - first) <= (c - b)) {
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
STACK_PUSH(first, a, depth, limit);
first = c;
} else if((last - c) <= (c - b)) {
STACK_PUSH(first, a, depth, limit);
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
first = c;
} else {
STACK_PUSH(first, a, depth, limit);
STACK_PUSH(c, last, depth, limit);
first = b, last = c, depth += 1, limit = ss_ilg(c - b);
}
}
} else {
limit += 1;
if(Td[PA[*first] - 1] < v) {
first = ss_partition(PA, first, last, depth);
limit = ss_ilg(last - first);
}
depth += 1;
}
}
#undef STACK_SIZE
}
#endif /* (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE) */
/*---------------------------------------------------------------------------*/
#if SS_BLOCKSIZE != 0
static INLINE
void
ss_blockswap(int *a, int *b, int n) {
int t;
for(; 0 < n; --n, ++a, ++b) {
t = *a, *a = *b, *b = t;
}
}
static INLINE
void
ss_rotate(int *first, int *middle, int *last) {
int *a, *b, t;
int l, r;
l = middle - first, r = last - middle;
for(; (0 < l) && (0 < r);) {
if(l == r) { ss_blockswap(first, middle, l); break; }
if(l < r) {
a = last - 1, b = middle - 1;
t = *a;
do {
*a-- = *b, *b-- = *a;
if(b < first) {
*a = t;
last = a;
if((r -= l + 1) <= l) { break; }
a -= 1, b = middle - 1;
t = *a;
}
} while(1);
} else {
a = first, b = middle;
t = *a;
do {
*a++ = *b, *b++ = *a;
if(last <= b) {
*a = t;
first = a + 1;
if((l -= r + 1) <= r) { break; }
a += 1, b = middle;
t = *a;
}
} while(1);
}
}
}
/*---------------------------------------------------------------------------*/
static
void
ss_inplacemerge(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int depth) {
const int *p;
int *a, *b;
int len, half;
int q, r;
int x;
for(;;) {
if(*(last - 1) < 0) { x = 1; p = PA + ~*(last - 1); }
else { x = 0; p = PA + *(last - 1); }
for(a = first, len = middle - first, half = len >> 1, r = -1;
0 < len;
len = half, half >>= 1) {
b = a + half;
q = ss_compare(T, PA + ((0 <= *b) ? *b : ~*b), p, depth);
if(q < 0) {
a = b + 1;
half -= (len & 1) ^ 1;
} else {
r = q;
}
}
if(a < middle) {
if(r == 0) { *a = ~*a; }
ss_rotate(a, middle, last);
last -= middle - a;
middle = a;
if(first == middle) { break; }
}
--last;
if(x != 0) { while(*--last < 0) { } }
if(middle == last) { break; }
}
}
/*---------------------------------------------------------------------------*/
/* Merge-forward with internal buffer. */
static
void
ss_mergeforward(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int depth) {
int *a, *b, *c, *bufend;
int t;
int r;
bufend = buf + (middle - first) - 1;
ss_blockswap(buf, first, middle - first);
for(t = *(a = first), b = buf, c = middle;;) {
r = ss_compare(T, PA + *b, PA + *c, depth);
if(r < 0) {
do {
*a++ = *b;
if(bufend <= b) { *bufend = t; return; }
*b++ = *a;
} while(*b < 0);
} else if(r > 0) {
do {
*a++ = *c, *c++ = *a;
if(last <= c) {
while(b < bufend) { *a++ = *b, *b++ = *a; }
*a = *b, *b = t;
return;
}
} while(*c < 0);
} else {
*c = ~*c;
do {
*a++ = *b;
if(bufend <= b) { *bufend = t; return; }
*b++ = *a;
} while(*b < 0);
do {
*a++ = *c, *c++ = *a;
if(last <= c) {
while(b < bufend) { *a++ = *b, *b++ = *a; }
*a = *b, *b = t;
return;
}
} while(*c < 0);
}
}
}
/* Merge-backward with internal buffer. */
static
void
ss_mergebackward(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int depth) {
const int *p1, *p2;
int *a, *b, *c, *bufend;
int t;
int r;
int x;
bufend = buf + (last - middle) - 1;
ss_blockswap(buf, middle, last - middle);
x = 0;
if(*bufend < 0) { p1 = PA + ~*bufend; x |= 1; }
else { p1 = PA + *bufend; }
if(*(middle - 1) < 0) { p2 = PA + ~*(middle - 1); x |= 2; }
else { p2 = PA + *(middle - 1); }
for(t = *(a = last - 1), b = bufend, c = middle - 1;;) {
r = ss_compare(T, p1, p2, depth);
if(0 < r) {
if(x & 1) { do { *a-- = *b, *b-- = *a; } while(*b < 0); x ^= 1; }
*a-- = *b;
if(b <= buf) { *buf = t; break; }
*b-- = *a;
if(*b < 0) { p1 = PA + ~*b; x |= 1; }
else { p1 = PA + *b; }
} else if(r < 0) {
if(x & 2) { do { *a-- = *c, *c-- = *a; } while(*c < 0); x ^= 2; }
*a-- = *c, *c-- = *a;
if(c < first) {
while(buf < b) { *a-- = *b, *b-- = *a; }
*a = *b, *b = t;
break;
}
if(*c < 0) { p2 = PA + ~*c; x |= 2; }
else { p2 = PA + *c; }
} else {
if(x & 1) { do { *a-- = *b, *b-- = *a; } while(*b < 0); x ^= 1; }
*a-- = ~*b;
if(b <= buf) { *buf = t; break; }
*b-- = *a;
if(x & 2) { do { *a-- = *c, *c-- = *a; } while(*c < 0); x ^= 2; }
*a-- = *c, *c-- = *a;
if(c < first) {
while(buf < b) { *a-- = *b, *b-- = *a; }
*a = *b, *b = t;
break;
}
if(*b < 0) { p1 = PA + ~*b; x |= 1; }
else { p1 = PA + *b; }
if(*c < 0) { p2 = PA + ~*c; x |= 2; }
else { p2 = PA + *c; }
}
}
}
/* D&C based merge. */
static
void
ss_swapmerge(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int bufsize, int depth) {
#define STACK_SIZE SS_SMERGE_STACKSIZE
#define GETIDX(a) ((0 <= (a)) ? (a) : (~(a)))
#define MERGE_CHECK(a, b, c)\
do {\
if(((c) & 1) ||\
(((c) & 2) && (ss_compare(T, PA + GETIDX(*((a) - 1)), PA + *(a), depth) == 0))) {\
*(a) = ~*(a);\
}\
if(((c) & 4) && ((ss_compare(T, PA + GETIDX(*((b) - 1)), PA + *(b), depth) == 0))) {\
*(b) = ~*(b);\
}\
} while(0)
struct { int *a, *b, *c; int d; } stack[STACK_SIZE];
int *l, *r, *lm, *rm;
int m, len, half;
int ssize;
int check, next;
for(check = 0, ssize = 0;;) {
if((last - middle) <= bufsize) {
if((first < middle) && (middle < last)) {
ss_mergebackward(T, PA, first, middle, last, buf, depth);
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
continue;
}
if((middle - first) <= bufsize) {
if(first < middle) {
ss_mergeforward(T, PA, first, middle, last, buf, depth);
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
continue;
}
for(m = 0, len = MIN(middle - first, last - middle), half = len >> 1;
0 < len;
len = half, half >>= 1) {
if(ss_compare(T, PA + GETIDX(*(middle + m + half)),
PA + GETIDX(*(middle - m - half - 1)), depth) < 0) {
m += half + 1;
half -= (len & 1) ^ 1;
}
}
if(0 < m) {
lm = middle - m, rm = middle + m;
ss_blockswap(lm, middle, m);
l = r = middle, next = 0;
if(rm < last) {
if(*rm < 0) {
*rm = ~*rm;
if(first < lm) { for(; *--l < 0;) { } next |= 4; }
next |= 1;
} else if(first < lm) {
for(; *r < 0; ++r) { }
next |= 2;
}
}
if((l - first) <= (last - r)) {
STACK_PUSH(r, rm, last, (next & 3) | (check & 4));
middle = lm, last = l, check = (check & 3) | (next & 4);
} else {
if((next & 2) && (r == middle)) { next ^= 6; }
STACK_PUSH(first, lm, l, (check & 3) | (next & 4));
first = r, middle = rm, check = (next & 3) | (check & 4);
}
} else {
if(ss_compare(T, PA + GETIDX(*(middle - 1)), PA + *middle, depth) == 0) {
*middle = ~*middle;
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
}
}
#undef STACK_SIZE
}
#endif /* SS_BLOCKSIZE != 0 */
/*---------------------------------------------------------------------------*/
/* Substring sort */
static
void
sssort(const unsigned char *T, const int *PA,
int *first, int *last,
int *buf, int bufsize,
int depth, int n, int lastsuffix) {
int *a;
#if SS_BLOCKSIZE != 0
int *b, *middle, *curbuf;
int j, k, curbufsize, limit;
#endif
int i;
if(lastsuffix != 0) { ++first; }
#if SS_BLOCKSIZE == 0
ss_mintrosort(T, PA, first, last, depth);
#else
if((bufsize < SS_BLOCKSIZE) &&
(bufsize < (last - first)) &&
(bufsize < (limit = ss_isqrt(last - first)))) {
if(SS_BLOCKSIZE < limit) { limit = SS_BLOCKSIZE; }
buf = middle = last - limit, bufsize = limit;
} else {
middle = last, limit = 0;
}
for(a = first, i = 0; SS_BLOCKSIZE < (middle - a); a += SS_BLOCKSIZE, ++i) {
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, a, a + SS_BLOCKSIZE, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, a, a + SS_BLOCKSIZE, depth);
#endif
curbufsize = last - (a + SS_BLOCKSIZE);
curbuf = a + SS_BLOCKSIZE;
if(curbufsize <= bufsize) { curbufsize = bufsize, curbuf = buf; }
for(b = a, k = SS_BLOCKSIZE, j = i; j & 1; b -= k, k <<= 1, j >>= 1) {
ss_swapmerge(T, PA, b - k, b, b + k, curbuf, curbufsize, depth);
}
}
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, a, middle, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, a, middle, depth);
#endif
for(k = SS_BLOCKSIZE; i != 0; k <<= 1, i >>= 1) {
if(i & 1) {
ss_swapmerge(T, PA, a - k, a, middle, buf, bufsize, depth);
a -= k;
}
}
if(limit != 0) {
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, middle, last, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, middle, last, depth);
#endif
ss_inplacemerge(T, PA, first, middle, last, depth);
}
#endif
if(lastsuffix != 0) {
/* Insert last type B* suffix. */
int PAi[2]; PAi[0] = PA[*(first - 1)], PAi[1] = n - 2;
for(a = first, i = *(first - 1);
(a < last) && ((*a < 0) || (0 < ss_compare(T, &(PAi[0]), PA + *a, depth)));
++a) {
*(a - 1) = *a;
}
*(a - 1) = i;
}
}
/*---------------------------------------------------------------------------*/
static INLINE
int
tr_ilg(int n) {
return (n & 0xffff0000) ?
((n & 0xff000000) ?
24 + lg_table[(n >> 24) & 0xff] :
16 + lg_table[(n >> 16) & 0xff]) :
((n & 0x0000ff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff]);
}
/*---------------------------------------------------------------------------*/
/* Simple insertionsort for small size groups. */
static
void
tr_insertionsort(const int *ISAd, int *first, int *last) {
int *a, *b;
int t, r;
for(a = first + 1; a < last; ++a) {
for(t = *a, b = a - 1; 0 > (r = ISAd[t] - ISAd[*b]);) {
do { *(b + 1) = *b; } while((first <= --b) && (*b < 0));
if(b < first) { break; }
}
if(r == 0) { *b = ~*b; }
*(b + 1) = t;
}
}
/*---------------------------------------------------------------------------*/
static INLINE
void
tr_fixdown(const int *ISAd, int *SA, int i, int size) {
int j, k;
int v;
int c, d, e;
for(v = SA[i], c = ISAd[v]; (j = 2 * i + 1) < size; SA[i] = SA[k], i = k) {
d = ISAd[SA[k = j++]];
if(d < (e = ISAd[SA[j]])) { k = j; d = e; }
if(d <= c) { break; }
}
SA[i] = v;
}
/* Simple top-down heapsort. */
static
void
tr_heapsort(const int *ISAd, int *SA, int size) {
int i, m;
int t;
m = size;
if((size % 2) == 0) {
m--;
if(ISAd[SA[m / 2]] < ISAd[SA[m]]) { SWAP(SA[m], SA[m / 2]); }
}
for(i = m / 2 - 1; 0 <= i; --i) { tr_fixdown(ISAd, SA, i, m); }
if((size % 2) == 0) { SWAP(SA[0], SA[m]); tr_fixdown(ISAd, SA, 0, m); }
for(i = m - 1; 0 < i; --i) {
t = SA[0], SA[0] = SA[i];
tr_fixdown(ISAd, SA, 0, i);
SA[i] = t;
}
}
/*---------------------------------------------------------------------------*/
/* Returns the median of three elements. */
static INLINE
int *
tr_median3(const int *ISAd, int *v1, int *v2, int *v3) {
int *t;
if(ISAd[*v1] > ISAd[*v2]) { SWAP(v1, v2); }
if(ISAd[*v2] > ISAd[*v3]) {
if(ISAd[*v1] > ISAd[*v3]) { return v1; }
else { return v3; }
}
return v2;
}
/* Returns the median of five elements. */
static INLINE
int *
tr_median5(const int *ISAd,
int *v1, int *v2, int *v3, int *v4, int *v5) {
int *t;
if(ISAd[*v2] > ISAd[*v3]) { SWAP(v2, v3); }
if(ISAd[*v4] > ISAd[*v5]) { SWAP(v4, v5); }
if(ISAd[*v2] > ISAd[*v4]) { SWAP(v2, v4); SWAP(v3, v5); }
if(ISAd[*v1] > ISAd[*v3]) { SWAP(v1, v3); }
if(ISAd[*v1] > ISAd[*v4]) { SWAP(v1, v4); SWAP(v3, v5); }
if(ISAd[*v3] > ISAd[*v4]) { return v4; }
return v3;
}
/* Returns the pivot element. */
static INLINE
int *
tr_pivot(const int *ISAd, int *first, int *last) {
int *middle;
int t;
t = last - first;
middle = first + t / 2;
if(t <= 512) {
if(t <= 32) {
return tr_median3(ISAd, first, middle, last - 1);
} else {
t >>= 2;
return tr_median5(ISAd, first, first + t, middle, last - 1 - t, last - 1);
}
}
t >>= 3;
first = tr_median3(ISAd, first, first + t, first + (t << 1));
middle = tr_median3(ISAd, middle - t, middle, middle + t);
last = tr_median3(ISAd, last - 1 - (t << 1), last - 1 - t, last - 1);
return tr_median3(ISAd, first, middle, last);
}
/*---------------------------------------------------------------------------*/
typedef struct _trbudget_t trbudget_t;
struct _trbudget_t {
int chance;
int remain;
int incval;
int count;
};
static INLINE
void
trbudget_init(trbudget_t *budget, int chance, int incval) {
budget->chance = chance;
budget->remain = budget->incval = incval;
}
static INLINE
int
trbudget_check(trbudget_t *budget, int size) {
if(size <= budget->remain) { budget->remain -= size; return 1; }
if(budget->chance == 0) { budget->count += size; return 0; }
budget->remain += budget->incval - size;
budget->chance -= 1;
return 1;
}
/*---------------------------------------------------------------------------*/
static INLINE
void
tr_partition(const int *ISAd,
int *first, int *middle, int *last,
int **pa, int **pb, int v) {
int *a, *b, *c, *d, *e, *f;
int t, s;
int x = 0;
for(b = middle - 1; (++b < last) && ((x = ISAd[*b]) == v);) { }
if(((a = b) < last) && (x < v)) {
for(; (++b < last) && ((x = ISAd[*b]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
}
for(c = last; (b < --c) && ((x = ISAd[*c]) == v);) { }
if((b < (d = c)) && (x > v)) {
for(; (b < --c) && ((x = ISAd[*c]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
for(; b < c;) {
SWAP(*b, *c);
for(; (++b < c) && ((x = ISAd[*b]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
for(; (b < --c) && ((x = ISAd[*c]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
if(a <= d) {
c = b - 1;
if((s = a - first) > (t = b - a)) { s = t; }
for(e = first, f = b - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
if((s = d - c) > (t = last - d - 1)) { s = t; }
for(e = b, f = last - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
first += (b - a), last -= (d - c);
}
*pa = first, *pb = last;
}
static
void
tr_copy(int *ISA, const int *SA,
int *first, int *a, int *b, int *last,
int depth) {
/* sort suffixes of middle partition
by using sorted order of suffixes of left and right partition. */
int *c, *d, *e;
int s, v;
v = b - SA - 1;
for(c = first, d = a - 1; c <= d; ++c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*++d = s;
ISA[s] = d - SA;
}
}
for(c = last - 1, e = d + 1, d = b; e < d; --c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*--d = s;
ISA[s] = d - SA;
}
}
}
static
void
tr_partialcopy(int *ISA, const int *SA,
int *first, int *a, int *b, int *last,
int depth) {
int *c, *d, *e;
int s, v;
int rank, lastrank, newrank = -1;
v = b - SA - 1;
lastrank = -1;
for(c = first, d = a - 1; c <= d; ++c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*++d = s;
rank = ISA[s + depth];
if(lastrank != rank) { lastrank = rank; newrank = d - SA; }
ISA[s] = newrank;
}
}
lastrank = -1;
for(e = d; first <= e; --e) {
rank = ISA[*e];
if(lastrank != rank) { lastrank = rank; newrank = e - SA; }
if(newrank != rank) { ISA[*e] = newrank; }
}
lastrank = -1;
for(c = last - 1, e = d + 1, d = b; e < d; --c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*--d = s;
rank = ISA[s + depth];
if(lastrank != rank) { lastrank = rank; newrank = d - SA; }
ISA[s] = newrank;
}
}
}
static
void
tr_introsort(int *ISA, const int *ISAd,
int *SA, int *first, int *last,
trbudget_t *budget) {
#define STACK_SIZE TR_STACKSIZE
struct { const int *a; int *b, *c; int d, e; }stack[STACK_SIZE];
int *a, *b, *c;
int t;
int v, x = 0;
int incr = ISAd - ISA;
int limit, next;
int ssize, trlink = -1;
for(ssize = 0, limit = tr_ilg(last - first);;) {
if(limit < 0) {
if(limit == -1) {
/* tandem repeat partition */
tr_partition(ISAd - incr, first, first, last, &a, &b, last - SA - 1);
/* update ranks */
if(a < last) {
for(c = first, v = a - SA - 1; c < a; ++c) { ISA[*c] = v; }
}
if(b < last) {
for(c = a, v = b - SA - 1; c < b; ++c) { ISA[*c] = v; }
}
/* push */
if(1 < (b - a)) {
STACK_PUSH5(NULL, a, b, 0, 0);
STACK_PUSH5(ISAd - incr, first, last, -2, trlink);
trlink = ssize - 2;
}
if((a - first) <= (last - b)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, tr_ilg(last - b), trlink);
last = a, limit = tr_ilg(a - first);
} else if(1 < (last - b)) {
first = b, limit = tr_ilg(last - b);
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
} else {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, tr_ilg(a - first), trlink);
first = b, limit = tr_ilg(last - b);
} else if(1 < (a - first)) {
last = a, limit = tr_ilg(a - first);
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
} else if(limit == -2) {
/* tandem repeat copy */
a = stack[--ssize].b, b = stack[ssize].c;
if(stack[ssize].d == 0) {
tr_copy(ISA, SA, first, a, b, last, ISAd - ISA);
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
tr_partialcopy(ISA, SA, first, a, b, last, ISAd - ISA);
}
STACK_POP5(ISAd, first, last, limit, trlink);
} else {
/* sorted partition */
if(0 <= *first) {
a = first;
do { ISA[*a] = a - SA; } while((++a < last) && (0 <= *a));
first = a;
}
if(first < last) {
a = first; do { *a = ~*a; } while(*++a < 0);
next = (ISA[*a] != ISAd[*a]) ? tr_ilg(a - first + 1) : -1;
if(++a < last) { for(b = first, v = a - SA - 1; b < a; ++b) { ISA[*b] = v; } }
/* push */
if(trbudget_check(budget, a - first)) {
if((a - first) <= (last - a)) {
STACK_PUSH5(ISAd, a, last, -3, trlink);
ISAd += incr, last = a, limit = next;
} else {
if(1 < (last - a)) {
STACK_PUSH5(ISAd + incr, first, a, next, trlink);
first = a, limit = -3;
} else {
ISAd += incr, last = a, limit = next;
}
}
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
if(1 < (last - a)) {
first = a, limit = -3;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
continue;
}
if((last - first) <= TR_INSERTIONSORT_THRESHOLD) {
tr_insertionsort(ISAd, first, last);
limit = -3;
continue;
}
if(limit-- == 0) {
tr_heapsort(ISAd, first, last - first);
for(a = last - 1; first < a; a = b) {
for(x = ISAd[*a], b = a - 1; (first <= b) && (ISAd[*b] == x); --b) { *b = ~*b; }
}
limit = -3;
continue;
}
/* choose pivot */
a = tr_pivot(ISAd, first, last);
SWAP(*first, *a);
v = ISAd[*first];
/* partition */
tr_partition(ISAd, first, first + 1, last, &a, &b, v);
if((last - first) != (b - a)) {
next = (ISA[*a] != v) ? tr_ilg(b - a) : -1;
/* update ranks */
for(c = first, v = a - SA - 1; c < a; ++c) { ISA[*c] = v; }
if(b < last) { for(c = a, v = b - SA - 1; c < b; ++c) { ISA[*c] = v; } }
/* push */
if((1 < (b - a)) && (trbudget_check(budget, b - a))) {
if((a - first) <= (last - b)) {
if((last - b) <= (b - a)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
STACK_PUSH5(ISAd, b, last, limit, trlink);
last = a;
} else if(1 < (last - b)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
first = b;
} else {
ISAd += incr, first = a, last = b, limit = next;
}
} else if((a - first) <= (b - a)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, limit, trlink);
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
last = a;
} else {
STACK_PUSH5(ISAd, b, last, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
STACK_PUSH5(ISAd, b, last, limit, trlink);
STACK_PUSH5(ISAd, first, a, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
if((a - first) <= (b - a)) {
if(1 < (last - b)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
STACK_PUSH5(ISAd, first, a, limit, trlink);
first = b;
} else if(1 < (a - first)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
last = a;
} else {
ISAd += incr, first = a, last = b, limit = next;
}
} else if((last - b) <= (b - a)) {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, limit, trlink);
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
first = b;
} else {
STACK_PUSH5(ISAd, first, a, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
STACK_PUSH5(ISAd, first, a, limit, trlink);
STACK_PUSH5(ISAd, b, last, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
}
} else {
if((1 < (b - a)) && (0 <= trlink)) { stack[trlink].d = -1; }
if((a - first) <= (last - b)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, limit, trlink);
last = a;
} else if(1 < (last - b)) {
first = b;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
} else {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, limit, trlink);
first = b;
} else if(1 < (a - first)) {
last = a;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
}
} else {
if(trbudget_check(budget, last - first)) {
limit = tr_ilg(last - first), ISAd += incr;
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
}
#undef STACK_SIZE
}
/*---------------------------------------------------------------------------*/
/* Tandem repeat sort */
static
void
trsort(int *ISA, int *SA, int n, int depth) {
int *ISAd;
int *first, *last;
trbudget_t budget;
int t, skip, unsorted;
trbudget_init(&budget, tr_ilg(n) * 2 / 3, n);
/* trbudget_init(&budget, tr_ilg(n) * 3 / 4, n); */
for(ISAd = ISA + depth; -n < *SA; ISAd += ISAd - ISA) {
first = SA;
skip = 0;
unsorted = 0;
do {
if((t = *first) < 0) { first -= t; skip += t; }
else {
if(skip != 0) { *(first + skip) = skip; skip = 0; }
last = SA + ISA[t] + 1;
if(1 < (last - first)) {
budget.count = 0;
tr_introsort(ISA, ISAd, SA, first, last, &budget);
if(budget.count != 0) { unsorted += budget.count; }
else { skip = first - last; }
} else if((last - first) == 1) {
skip = -1;
}
first = last;
}
} while(first < (SA + n));
if(skip != 0) { *(first + skip) = skip; }
if(unsorted == 0) { break; }
}
}
/*---------------------------------------------------------------------------*/
/* Sorts suffixes of type B*. */
static
int
sort_typeBstar(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int openMP) {
int *PAb, *ISAb, *buf;
#ifdef LIBBSC_OPENMP
int *curbuf;
int l;
#endif
int i, j, k, t, m, bufsize;
int c0, c1;
#ifdef LIBBSC_OPENMP
int d0, d1;
#endif
(void)openMP;
/* Initialize bucket arrays. */
for(i = 0; i < BUCKET_A_SIZE; ++i) { bucket_A[i] = 0; }
for(i = 0; i < BUCKET_B_SIZE; ++i) { bucket_B[i] = 0; }
/* Count the number of occurrences of the first one or two characters of each
type A, B and B* suffix. Moreover, store the beginning position of all
type B* suffixes into the array SA. */
for(i = n - 1, m = n, c0 = T[n - 1]; 0 <= i;) {
/* type A suffix. */
do { ++BUCKET_A(c1 = c0); } while((0 <= --i) && ((c0 = T[i]) >= c1));
if(0 <= i) {
/* type B* suffix. */
++BUCKET_BSTAR(c0, c1);
SA[--m] = i;
/* type B suffix. */
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) <= c1); --i, c1 = c0) {
++BUCKET_B(c0, c1);
}
}
}
m = n - m;
/*
note:
A type B* suffix is lexicographically smaller than a type B suffix that
begins with the same first two characters.
*/
/* Calculate the index of start/end point of each bucket. */
for(c0 = 0, i = 0, j = 0; c0 < ALPHABET_SIZE; ++c0) {
t = i + BUCKET_A(c0);
BUCKET_A(c0) = i + j; /* start point */
i = t + BUCKET_B(c0, c0);
for(c1 = c0 + 1; c1 < ALPHABET_SIZE; ++c1) {
j += BUCKET_BSTAR(c0, c1);
BUCKET_BSTAR(c0, c1) = j; /* end point */
i += BUCKET_B(c0, c1);
}
}
if(0 < m) {
/* Sort the type B* suffixes by their first two characters. */
PAb = SA + n - m; ISAb = SA + m;
for(i = m - 2; 0 <= i; --i) {
t = PAb[i], c0 = T[t], c1 = T[t + 1];
SA[--BUCKET_BSTAR(c0, c1)] = i;
}
t = PAb[m - 1], c0 = T[t], c1 = T[t + 1];
SA[--BUCKET_BSTAR(c0, c1)] = m - 1;
/* Sort the type B* substrings using sssort. */
#ifdef LIBBSC_OPENMP
if (openMP)
{
buf = SA + m;
c0 = ALPHABET_SIZE - 2, c1 = ALPHABET_SIZE - 1, j = m;
#pragma omp parallel default(shared) private(bufsize, curbuf, k, l, d0, d1)
{
bufsize = (n - (2 * m)) / omp_get_num_threads();
curbuf = buf + omp_get_thread_num() * bufsize;
k = 0;
for(;;) {
#pragma omp critical(sssort_lock)
{
if(0 < (l = j)) {
d0 = c0, d1 = c1;
do {
k = BUCKET_BSTAR(d0, d1);
if(--d1 <= d0) {
d1 = ALPHABET_SIZE - 1;
if(--d0 < 0) { break; }
}
} while(((l - k) <= 1) && (0 < (l = k)));
c0 = d0, c1 = d1, j = k;
}
}
if(l == 0) { break; }
sssort(T, PAb, SA + k, SA + l,
curbuf, bufsize, 2, n, *(SA + k) == (m - 1));
}
}
}
else
{
buf = SA + m, bufsize = n - (2 * m);
for(c0 = ALPHABET_SIZE - 2, j = m; 0 < j; --c0) {
for(c1 = ALPHABET_SIZE - 1; c0 < c1; j = i, --c1) {
i = BUCKET_BSTAR(c0, c1);
if(1 < (j - i)) {
sssort(T, PAb, SA + i, SA + j,
buf, bufsize, 2, n, *(SA + i) == (m - 1));
}
}
}
}
#else
buf = SA + m, bufsize = n - (2 * m);
for(c0 = ALPHABET_SIZE - 2, j = m; 0 < j; --c0) {
for(c1 = ALPHABET_SIZE - 1; c0 < c1; j = i, --c1) {
i = BUCKET_BSTAR(c0, c1);
if(1 < (j - i)) {
sssort(T, PAb, SA + i, SA + j,
buf, bufsize, 2, n, *(SA + i) == (m - 1));
}
}
}
#endif
/* Compute ranks of type B* substrings. */
for(i = m - 1; 0 <= i; --i) {
if(0 <= SA[i]) {
j = i;
do { ISAb[SA[i]] = i; } while((0 <= --i) && (0 <= SA[i]));
SA[i + 1] = i - j;
if(i <= 0) { break; }
}
j = i;
do { ISAb[SA[i] = ~SA[i]] = j; } while(SA[--i] < 0);
ISAb[SA[i]] = j;
}
/* Construct the inverse suffix array of type B* suffixes using trsort. */
trsort(ISAb, SA, m, 1);
/* Set the sorted order of type B* suffixes. */
for(i = n - 1, j = m, c0 = T[n - 1]; 0 <= i;) {
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) >= c1); --i, c1 = c0) { }
if(0 <= i) {
t = i;
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) <= c1); --i, c1 = c0) { }
SA[ISAb[--j]] = ((t == 0) || (1 < (t - i))) ? t : ~t;
}
}
/* Calculate the index of start/end point of each bucket. */
BUCKET_B(ALPHABET_SIZE - 1, ALPHABET_SIZE - 1) = n; /* end point */
for(c0 = ALPHABET_SIZE - 2, k = m - 1; 0 <= c0; --c0) {
i = BUCKET_A(c0 + 1) - 1;
for(c1 = ALPHABET_SIZE - 1; c0 < c1; --c1) {
t = i - BUCKET_B(c0, c1);
BUCKET_B(c0, c1) = i; /* end point */
/* Move all type B* suffixes to the correct position. */
for(i = t, j = BUCKET_BSTAR(c0, c1);
j <= k;
--i, --k) { SA[i] = SA[k]; }
}
BUCKET_BSTAR(c0, c0 + 1) = i - BUCKET_B(c0, c0) + 1; /* start point */
BUCKET_B(c0, c0) = i; /* end point */
}
}
return m;
}
/* Constructs the suffix array by using the sorted order of type B* suffixes. */
static
void
construct_SA(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m) {
int *i, *j, *k;
int s;
int c0, c1, c2;
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
*j = ~s;
c0 = T[--s];
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else {
assert(((s == 0) && (T[s] == c1)) || (s < 0));
*j = ~s;
}
}
}
}
/* Construct the suffix array by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
*k++ = (T[n - 2] < c2) ? ~(n - 1) : (n - 1);
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
c0 = T[--s];
if((s == 0) || (T[s - 1] < c0)) { s = ~s; }
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
*k++ = s;
} else {
assert(s < 0);
*i = ~s;
}
}
}
/* Constructs the burrows-wheeler transformed string directly
by using the sorted order of type B* suffixes. */
static
int
construct_BWT(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m) {
int *i, *j, *k, *orig;
int s;
int c0, c1, c2;
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
c0 = T[--s];
*j = ~((int)c0);
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else if(s != 0) {
*j = ~s;
#ifndef NDEBUG
} else {
assert(T[s] == c1);
#endif
}
}
}
}
/* Construct the BWTed string by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
*k++ = (T[n - 2] < c2) ? ~((int)T[n - 2]) : (n - 1);
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n, orig = SA; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
c0 = T[--s];
*i = c0;
if((0 < s) && (T[s - 1] < c0)) { s = ~((int)T[s - 1]); }
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
*k++ = s;
} else if(s != 0) {
*i = ~s;
} else {
orig = i;
}
}
return orig - SA;
}
/* Constructs the burrows-wheeler transformed string directly
by using the sorted order of type B* suffixes. */
static
int
construct_BWT_indexes(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m,
unsigned char * num_indexes, int * indexes) {
int *i, *j, *k, *orig;
int s;
int c0, c1, c2;
int mod = n / 8;
{
mod |= mod >> 1; mod |= mod >> 2;
mod |= mod >> 4; mod |= mod >> 8;
mod |= mod >> 16; mod >>= 1;
*num_indexes = (unsigned char)((n - 1) / (mod + 1));
}
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = j - SA;
c0 = T[--s];
*j = ~((int)c0);
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else if(s != 0) {
*j = ~s;
#ifndef NDEBUG
} else {
assert(T[s] == c1);
#endif
}
}
}
}
/* Construct the BWTed string by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
if (T[n - 2] < c2) {
if (((n - 1) & mod) == 0) indexes[(n - 1) / (mod + 1) - 1] = k - SA;
*k++ = ~((int)T[n - 2]);
}
else {
*k++ = n - 1;
}
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n, orig = SA; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = i - SA;
c0 = T[--s];
*i = c0;
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
if((0 < s) && (T[s - 1] < c0)) {
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = k - SA;
*k++ = ~((int)T[s - 1]);
} else
*k++ = s;
} else if(s != 0) {
*i = ~s;
} else {
orig = i;
}
}
return orig - SA;
}
/*---------------------------------------------------------------------------*/
/*- Function -*/
int
divsufsort(const unsigned char *T, int *SA, int n, int openMP) {
int *bucket_A, *bucket_B;
int m;
int err = 0;
/* Check arguments. */
if((T == NULL) || (SA == NULL) || (n < 0)) { return -1; }
else if(n == 0) { return 0; }
else if(n == 1) { SA[0] = 0; return 0; }
else if(n == 2) { m = (T[0] < T[1]); SA[m ^ 1] = 0, SA[m] = 1; return 0; }
bucket_A = (int *)malloc(BUCKET_A_SIZE * sizeof(int));
bucket_B = (int *)malloc(BUCKET_B_SIZE * sizeof(int));
/* Suffixsort. */
if((bucket_A != NULL) && (bucket_B != NULL)) {
m = sort_typeBstar(T, SA, bucket_A, bucket_B, n, openMP);
construct_SA(T, SA, bucket_A, bucket_B, n, m);
} else {
err = -2;
}
free(bucket_B);
free(bucket_A);
return err;
}
int
divbwt(const unsigned char *T, unsigned char *U, int *A, int n, unsigned char * num_indexes, int * indexes, int openMP) {
int *B;
int *bucket_A, *bucket_B;
int m, pidx, i;
/* Check arguments. */
if((T == NULL) || (U == NULL) || (n < 0)) { return -1; }
else if(n <= 1) { if(n == 1) { U[0] = T[0]; } return n; }
if((B = A) == NULL) { B = (int *)malloc((size_t)(n + 1) * sizeof(int)); }
bucket_A = (int *)malloc(BUCKET_A_SIZE * sizeof(int));
bucket_B = (int *)malloc(BUCKET_B_SIZE * sizeof(int));
/* Burrows-Wheeler Transform. */
if((B != NULL) && (bucket_A != NULL) && (bucket_B != NULL)) {
m = sort_typeBstar(T, B, bucket_A, bucket_B, n, openMP);
if (num_indexes == NULL || indexes == NULL) {
pidx = construct_BWT(T, B, bucket_A, bucket_B, n, m);
} else {
pidx = construct_BWT_indexes(T, B, bucket_A, bucket_B, n, m, num_indexes, indexes);
}
/* Copy to output string. */
U[0] = T[n - 1];
for(i = 0; i < pidx; ++i) { U[i + 1] = (unsigned char)B[i]; }
for(i += 1; i < n; ++i) { U[i] = (unsigned char)B[i]; }
pidx += 1;
} else {
pidx = -2;
}
free(bucket_B);
free(bucket_A);
if(A == NULL) { free(B); }
return pidx;
}
+4 -2108
View File
@@ -17,2111 +17,7 @@
#include "../common/compiler.h"
#include "../common/error_private.h"
/******************************************
* Static allocation
******************************************/
/* You can statically allocate FSE CTable/DTable as a table of unsigned using below macro */
#define FSE_DTABLE_SIZE_U32(maxTableLog) (1 + (1<<maxTableLog))
/* You can statically allocate Huff0 DTable as a table of unsigned short using below macro */
#define HUF_DTABLE_SIZE_U16(maxTableLog) (1 + (1<<maxTableLog))
#define HUF_CREATE_STATIC_DTABLE(DTable, maxTableLog) \
unsigned short DTable[HUF_DTABLE_SIZE_U16(maxTableLog)] = { maxTableLog }
/******************************************
* Error Management
******************************************/
#define FSE_LIST_ERRORS(ITEM) \
ITEM(FSE_OK_NoError) ITEM(FSE_ERROR_GENERIC) \
ITEM(FSE_ERROR_tableLog_tooLarge) ITEM(FSE_ERROR_maxSymbolValue_tooLarge) ITEM(FSE_ERROR_maxSymbolValue_tooSmall) \
ITEM(FSE_ERROR_dstSize_tooSmall) ITEM(FSE_ERROR_srcSize_wrong)\
ITEM(FSE_ERROR_corruptionDetected) \
ITEM(FSE_ERROR_maxCode)
#define FSE_GENERATE_ENUM(ENUM) ENUM,
typedef enum { FSE_LIST_ERRORS(FSE_GENERATE_ENUM) } FSE_errorCodes; /* enum is exposed, to detect & handle specific errors; compare function result to -enum value */
/******************************************
* FSE symbol compression API
******************************************/
/*
This API consists of small unitary functions, which highly benefit from being inlined.
You will want to enable link-time-optimization to ensure these functions are properly inlined in your binary.
Visual seems to do it automatically.
For gcc or clang, you'll need to add -flto flag at compilation and linking stages.
If none of these solutions is applicable, include "fse.c" directly.
*/
typedef unsigned FSE_CTable; /* don't allocate that. It's just a way to be more restrictive than void* */
typedef unsigned FSE_DTable; /* don't allocate that. It's just a way to be more restrictive than void* */
typedef struct
{
size_t bitContainer;
int bitPos;
char* startPtr;
char* ptr;
char* endPtr;
} FSE_CStream_t;
typedef struct
{
ptrdiff_t value;
const void* stateTable;
const void* symbolTT;
unsigned stateLog;
} FSE_CState_t;
typedef struct
{
size_t bitContainer;
unsigned bitsConsumed;
const char* ptr;
const char* start;
} FSE_DStream_t;
typedef struct
{
size_t state;
const void* table; /* precise table may vary, depending on U16 */
} FSE_DState_t;
typedef enum { FSE_DStream_unfinished = 0,
FSE_DStream_endOfBuffer = 1,
FSE_DStream_completed = 2,
FSE_DStream_tooFar = 3 } FSE_DStream_status; /* result of FSE_reloadDStream() */
/* 1,2,4,8 would be better for bitmap combinations, but slows down performance a bit ... ?! */
/****************************************************************
* Tuning parameters
****************************************************************/
/* MEMORY_USAGE :
* Memory usage formula : N->2^N Bytes (examples : 10 -> 1KB; 12 -> 4KB ; 16 -> 64KB; 20 -> 1MB; etc.)
* Increasing memory usage improves compression ratio
* Reduced memory usage can improve speed, due to cache effect
* Recommended max value is 14, for 16KB, which nicely fits into Intel x86 L1 cache */
#define FSE_MAX_MEMORY_USAGE 14
#define FSE_DEFAULT_MEMORY_USAGE 13
/* FSE_MAX_SYMBOL_VALUE :
* Maximum symbol value authorized.
* Required for proper stack allocation */
#define FSE_MAX_SYMBOL_VALUE 255
/****************************************************************
* template functions type & suffix
****************************************************************/
#define FSE_FUNCTION_TYPE BYTE
#define FSE_FUNCTION_EXTENSION
/****************************************************************
* Byte symbol type
****************************************************************/
typedef struct
{
unsigned short newState;
unsigned char symbol;
unsigned char nbBits;
} FSE_decode_t; /* size == U32 */
/****************************************************************
* Compiler specifics
****************************************************************/
#ifdef _MSC_VER /* Visual Studio */
# define FORCE_INLINE static __forceinline
# include <intrin.h> /* For Visual 2005 */
# pragma warning(disable : 4127) /* disable: C4127: conditional expression is constant */
# pragma warning(disable : 4214) /* disable: C4214: non-int bitfields */
#else
# define GCC_VERSION (__GNUC__ * 100 + __GNUC_MINOR__)
# if defined (__cplusplus) || defined (__STDC_VERSION__) && __STDC_VERSION__ >= 199901L /* C99 */
# ifdef __GNUC__
# define FORCE_INLINE static inline __attribute__((always_inline))
# else
# define FORCE_INLINE static inline
# endif
# else
# define FORCE_INLINE static
# endif /* __STDC_VERSION__ */
#endif
/****************************************************************
* Includes
****************************************************************/
#include <stdlib.h> /* malloc, free, qsort */
#include <string.h> /* memcpy, memset */
#include <stdio.h> /* printf (debug) */
#ifndef MEM_ACCESS_MODULE
#define MEM_ACCESS_MODULE
/****************************************************************
* Basic Types
*****************************************************************/
#if defined (__STDC_VERSION__) && __STDC_VERSION__ >= 199901L /* C99 */
# include <stdint.h>
typedef uint8_t BYTE;
typedef uint16_t U16;
typedef int16_t S16;
typedef uint32_t U32;
typedef int32_t S32;
typedef uint64_t U64;
typedef int64_t S64;
#else
typedef unsigned char BYTE;
typedef unsigned short U16;
typedef signed short S16;
typedef unsigned int U32;
typedef signed int S32;
typedef unsigned long long U64;
typedef signed long long S64;
#endif
#endif /* MEM_ACCESS_MODULE */
/****************************************************************
* Memory I/O
*****************************************************************/
static unsigned FSE_32bits(void)
{
return sizeof(void*)==4;
}
static unsigned FSE_isLittleEndian(void)
{
const union { U32 i; BYTE c[4]; } one = { 1 }; /* don't use static : performance detrimental */
return one.c[0];
}
static U16 FSE_read16(const void* memPtr)
{
U16 val; memcpy(&val, memPtr, sizeof(val)); return val;
}
static U32 FSE_read32(const void* memPtr)
{
U32 val; memcpy(&val, memPtr, sizeof(val)); return val;
}
static U64 FSE_read64(const void* memPtr)
{
U64 val; memcpy(&val, memPtr, sizeof(val)); return val;
}
static U16 FSE_readLE16(const void* memPtr)
{
if (FSE_isLittleEndian())
return FSE_read16(memPtr);
else
{
const BYTE* p = (const BYTE*)memPtr;
return (U16)(p[0] + (p[1]<<8));
}
}
static U32 FSE_readLE32(const void* memPtr)
{
if (FSE_isLittleEndian())
return FSE_read32(memPtr);
else
{
const BYTE* p = (const BYTE*)memPtr;
return (U32)((U32)p[0] + ((U32)p[1]<<8) + ((U32)p[2]<<16) + ((U32)p[3]<<24));
}
}
static U64 FSE_readLE64(const void* memPtr)
{
if (FSE_isLittleEndian())
return FSE_read64(memPtr);
else
{
const BYTE* p = (const BYTE*)memPtr;
return (U64)((U64)p[0] + ((U64)p[1]<<8) + ((U64)p[2]<<16) + ((U64)p[3]<<24)
+ ((U64)p[4]<<32) + ((U64)p[5]<<40) + ((U64)p[6]<<48) + ((U64)p[7]<<56));
}
}
static size_t FSE_readLEST(const void* memPtr)
{
if (FSE_32bits())
return (size_t)FSE_readLE32(memPtr);
else
return (size_t)FSE_readLE64(memPtr);
}
/****************************************************************
* Constants
*****************************************************************/
#define FSE_MAX_TABLELOG (FSE_MAX_MEMORY_USAGE-2)
#define FSE_MAX_TABLESIZE (1U<<FSE_MAX_TABLELOG)
#define FSE_MAXTABLESIZE_MASK (FSE_MAX_TABLESIZE-1)
#define FSE_DEFAULT_TABLELOG (FSE_DEFAULT_MEMORY_USAGE-2)
#define FSE_MIN_TABLELOG 5
#define FSE_TABLELOG_ABSOLUTE_MAX 15
#if FSE_MAX_TABLELOG > FSE_TABLELOG_ABSOLUTE_MAX
#error "FSE_MAX_TABLELOG > FSE_TABLELOG_ABSOLUTE_MAX is not supported"
#endif
/****************************************************************
* Error Management
****************************************************************/
#define FSE_STATIC_ASSERT(c) { enum { FSE_static_assert = 1/(int)(!!(c)) }; } /* use only *after* variable declarations */
/****************************************************************
* Complex types
****************************************************************/
typedef struct
{
int deltaFindState;
U32 deltaNbBits;
} FSE_symbolCompressionTransform; /* total 8 bytes */
typedef U32 DTable_max_t[FSE_DTABLE_SIZE_U32(FSE_MAX_TABLELOG)];
/****************************************************************
* Internal functions
****************************************************************/
FORCE_INLINE unsigned FSE_highbit32 (U32 val)
{
# if defined(_MSC_VER) /* Visual */
unsigned long r;
return _BitScanReverse(&r, val) ? (unsigned)r : 0;
# elif defined(__GNUC__) && (GCC_VERSION >= 304) /* GCC Intrinsic */
return __builtin_clz (val) ^ 31;
# else /* Software version */
static const unsigned DeBruijnClz[32] = { 0, 9, 1, 10, 13, 21, 2, 29, 11, 14, 16, 18, 22, 25, 3, 30, 8, 12, 20, 28, 15, 17, 24, 7, 19, 27, 23, 6, 26, 5, 4, 31 };
U32 v = val;
unsigned r;
v |= v >> 1;
v |= v >> 2;
v |= v >> 4;
v |= v >> 8;
v |= v >> 16;
r = DeBruijnClz[ (U32) (v * 0x07C4ACDDU) >> 27];
return r;
# endif
}
/****************************************************************
* Templates
****************************************************************/
/*
designed to be included
for type-specific functions (template emulation in C)
Objective is to write these functions only once, for improved maintenance
*/
/* safety checks */
#ifndef FSE_FUNCTION_EXTENSION
# error "FSE_FUNCTION_EXTENSION must be defined"
#endif
#ifndef FSE_FUNCTION_TYPE
# error "FSE_FUNCTION_TYPE must be defined"
#endif
/* Function names */
#define FSE_CAT(X,Y) X##Y
#define FSE_FUNCTION_NAME(X,Y) FSE_CAT(X,Y)
#define FSE_TYPE_NAME(X,Y) FSE_CAT(X,Y)
static U32 FSE_tableStep(U32 tableSize) { return (tableSize>>1) + (tableSize>>3) + 3; }
#define FSE_DECODE_TYPE FSE_decode_t
typedef struct {
U16 tableLog;
U16 fastMode;
} FSE_DTableHeader; /* sizeof U32 */
static size_t FSE_buildDTable
(FSE_DTable* dt, const short* normalizedCounter, unsigned maxSymbolValue, unsigned tableLog)
{
void* ptr = dt;
FSE_DTableHeader* const DTableH = (FSE_DTableHeader*)ptr;
FSE_DECODE_TYPE* const tableDecode = (FSE_DECODE_TYPE*)(ptr) + 1; /* because dt is unsigned, 32-bits aligned on 32-bits */
const U32 tableSize = 1 << tableLog;
const U32 tableMask = tableSize-1;
const U32 step = FSE_tableStep(tableSize);
U16 symbolNext[FSE_MAX_SYMBOL_VALUE+1];
U32 position = 0;
U32 highThreshold = tableSize-1;
const S16 largeLimit= (S16)(1 << (tableLog-1));
U32 noLarge = 1;
U32 s;
/* Sanity Checks */
if (maxSymbolValue > FSE_MAX_SYMBOL_VALUE) return (size_t)-FSE_ERROR_maxSymbolValue_tooLarge;
if (tableLog > FSE_MAX_TABLELOG) return (size_t)-FSE_ERROR_tableLog_tooLarge;
/* Init, lay down lowprob symbols */
DTableH[0].tableLog = (U16)tableLog;
for (s=0; s<=maxSymbolValue; s++)
{
if (normalizedCounter[s]==-1)
{
tableDecode[highThreshold--].symbol = (FSE_FUNCTION_TYPE)s;
symbolNext[s] = 1;
}
else
{
if (normalizedCounter[s] >= largeLimit) noLarge=0;
symbolNext[s] = normalizedCounter[s];
}
}
/* Spread symbols */
for (s=0; s<=maxSymbolValue; s++)
{
int i;
for (i=0; i<normalizedCounter[s]; i++)
{
tableDecode[position].symbol = (FSE_FUNCTION_TYPE)s;
position = (position + step) & tableMask;
while (position > highThreshold) position = (position + step) & tableMask; /* lowprob area */
}
}
if (position!=0) return (size_t)-FSE_ERROR_GENERIC; /* position must reach all cells once, otherwise normalizedCounter is incorrect */
/* Build Decoding table */
{
U32 i;
for (i=0; i<tableSize; i++)
{
FSE_FUNCTION_TYPE symbol = (FSE_FUNCTION_TYPE)(tableDecode[i].symbol);
U16 nextState = symbolNext[symbol]++;
tableDecode[i].nbBits = (BYTE) (tableLog - FSE_highbit32 ((U32)nextState) );
tableDecode[i].newState = (U16) ( (nextState << tableDecode[i].nbBits) - tableSize);
}
}
DTableH->fastMode = (U16)noLarge;
return 0;
}
/******************************************
* FSE byte symbol
******************************************/
#ifndef FSE_COMMONDEFS_ONLY
static unsigned FSE_isError(size_t code) { return (code > (size_t)(-FSE_ERROR_maxCode)); }
static short FSE_abs(short a)
{
return a<0? -a : a;
}
/****************************************************************
* Header bitstream management
****************************************************************/
static size_t FSE_readNCount (short* normalizedCounter, unsigned* maxSVPtr, unsigned* tableLogPtr,
const void* headerBuffer, size_t hbSize)
{
const BYTE* const istart = (const BYTE*) headerBuffer;
const BYTE* const iend = istart + hbSize;
const BYTE* ip = istart;
int nbBits;
int remaining;
int threshold;
U32 bitStream;
int bitCount;
unsigned charnum = 0;
int previous0 = 0;
if (hbSize < 4) return (size_t)-FSE_ERROR_srcSize_wrong;
bitStream = FSE_readLE32(ip);
nbBits = (bitStream & 0xF) + FSE_MIN_TABLELOG; /* extract tableLog */
if (nbBits > FSE_TABLELOG_ABSOLUTE_MAX) return (size_t)-FSE_ERROR_tableLog_tooLarge;
bitStream >>= 4;
bitCount = 4;
*tableLogPtr = nbBits;
remaining = (1<<nbBits)+1;
threshold = 1<<nbBits;
nbBits++;
while ((remaining>1) && (charnum<=*maxSVPtr))
{
if (previous0)
{
unsigned n0 = charnum;
while ((bitStream & 0xFFFF) == 0xFFFF)
{
n0+=24;
if (ip < iend-5)
{
ip+=2;
bitStream = FSE_readLE32(ip) >> bitCount;
}
else
{
bitStream >>= 16;
bitCount+=16;
}
}
while ((bitStream & 3) == 3)
{
n0+=3;
bitStream>>=2;
bitCount+=2;
}
n0 += bitStream & 3;
bitCount += 2;
if (n0 > *maxSVPtr) return (size_t)-FSE_ERROR_maxSymbolValue_tooSmall;
while (charnum < n0) normalizedCounter[charnum++] = 0;
if ((ip <= iend-7) || (ip + (bitCount>>3) <= iend-4))
{
ip += bitCount>>3;
bitCount &= 7;
bitStream = FSE_readLE32(ip) >> bitCount;
}
else
bitStream >>= 2;
}
{
const short max = (short)((2*threshold-1)-remaining);
short count;
if ((bitStream & (threshold-1)) < (U32)max)
{
count = (short)(bitStream & (threshold-1));
bitCount += nbBits-1;
}
else
{
count = (short)(bitStream & (2*threshold-1));
if (count >= threshold) count -= max;
bitCount += nbBits;
}
count--; /* extra accuracy */
remaining -= FSE_abs(count);
normalizedCounter[charnum++] = count;
previous0 = !count;
while (remaining < threshold)
{
nbBits--;
threshold >>= 1;
}
{
if ((ip <= iend-7) || (ip + (bitCount>>3) <= iend-4))
{
ip += bitCount>>3;
bitCount &= 7;
}
else
{
bitCount -= (int)(8 * (iend - 4 - ip));
ip = iend - 4;
}
bitStream = FSE_readLE32(ip) >> (bitCount & 31);
}
}
}
if (remaining != 1) return (size_t)-FSE_ERROR_GENERIC;
*maxSVPtr = charnum-1;
ip += (bitCount+7)>>3;
if ((size_t)(ip-istart) > hbSize) return (size_t)-FSE_ERROR_srcSize_wrong;
return ip-istart;
}
/*********************************************************
* Decompression (Byte symbols)
*********************************************************/
static size_t FSE_buildDTable_rle (FSE_DTable* dt, BYTE symbolValue)
{
void* ptr = dt;
FSE_DTableHeader* const DTableH = (FSE_DTableHeader*)ptr;
FSE_decode_t* const cell = (FSE_decode_t*)(ptr) + 1; /* because dt is unsigned */
DTableH->tableLog = 0;
DTableH->fastMode = 0;
cell->newState = 0;
cell->symbol = symbolValue;
cell->nbBits = 0;
return 0;
}
static size_t FSE_buildDTable_raw (FSE_DTable* dt, unsigned nbBits)
{
void* ptr = dt;
FSE_DTableHeader* const DTableH = (FSE_DTableHeader*)ptr;
FSE_decode_t* const dinfo = (FSE_decode_t*)(ptr) + 1; /* because dt is unsigned */
const unsigned tableSize = 1 << nbBits;
const unsigned tableMask = tableSize - 1;
const unsigned maxSymbolValue = tableMask;
unsigned s;
/* Sanity checks */
if (nbBits < 1) return (size_t)-FSE_ERROR_GENERIC; /* min size */
/* Build Decoding Table */
DTableH->tableLog = (U16)nbBits;
DTableH->fastMode = 1;
for (s=0; s<=maxSymbolValue; s++)
{
dinfo[s].newState = 0;
dinfo[s].symbol = (BYTE)s;
dinfo[s].nbBits = (BYTE)nbBits;
}
return 0;
}
/* FSE_initDStream
* Initialize a FSE_DStream_t.
* srcBuffer must point at the beginning of an FSE block.
* The function result is the size of the FSE_block (== srcSize).
* If srcSize is too small, the function will return an errorCode;
*/
static size_t FSE_initDStream(FSE_DStream_t* bitD, const void* srcBuffer, size_t srcSize)
{
if (srcSize < 1) return (size_t)-FSE_ERROR_srcSize_wrong;
if (srcSize >= sizeof(size_t))
{
U32 contain32;
bitD->start = (const char*)srcBuffer;
bitD->ptr = (const char*)srcBuffer + srcSize - sizeof(size_t);
bitD->bitContainer = FSE_readLEST(bitD->ptr);
contain32 = ((const BYTE*)srcBuffer)[srcSize-1];
if (contain32 == 0) return (size_t)-FSE_ERROR_GENERIC; /* stop bit not present */
bitD->bitsConsumed = 8 - FSE_highbit32(contain32);
}
else
{
U32 contain32;
bitD->start = (const char*)srcBuffer;
bitD->ptr = bitD->start;
bitD->bitContainer = *(const BYTE*)(bitD->start);
switch(srcSize)
{
case 7: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[6]) << (sizeof(size_t)*8 - 16);
/* fallthrough */
case 6: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[5]) << (sizeof(size_t)*8 - 24);
/* fallthrough */
case 5: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[4]) << (sizeof(size_t)*8 - 32);
/* fallthrough */
case 4: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[3]) << 24;
/* fallthrough */
case 3: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[2]) << 16;
/* fallthrough */
case 2: bitD->bitContainer += (size_t)(((const BYTE*)(bitD->start))[1]) << 8;
/* fallthrough */
default:;
}
contain32 = ((const BYTE*)srcBuffer)[srcSize-1];
if (contain32 == 0) return (size_t)-FSE_ERROR_GENERIC; /* stop bit not present */
bitD->bitsConsumed = 8 - FSE_highbit32(contain32);
bitD->bitsConsumed += (U32)(sizeof(size_t) - srcSize)*8;
}
return srcSize;
}
/*!FSE_lookBits
* Provides next n bits from the bitContainer.
* bitContainer is not modified (bits are still present for next read/look)
* On 32-bits, maxNbBits==25
* On 64-bits, maxNbBits==57
* return : value extracted.
*/
static size_t FSE_lookBits(FSE_DStream_t* bitD, U32 nbBits)
{
const U32 bitMask = sizeof(bitD->bitContainer)*8 - 1;
return ((bitD->bitContainer << (bitD->bitsConsumed & bitMask)) >> 1) >> ((bitMask-nbBits) & bitMask);
}
static size_t FSE_lookBitsFast(FSE_DStream_t* bitD, U32 nbBits) /* only if nbBits >= 1 !! */
{
const U32 bitMask = sizeof(bitD->bitContainer)*8 - 1;
return (bitD->bitContainer << (bitD->bitsConsumed & bitMask)) >> (((bitMask+1)-nbBits) & bitMask);
}
static void FSE_skipBits(FSE_DStream_t* bitD, U32 nbBits)
{
bitD->bitsConsumed += nbBits;
}
/*!FSE_readBits
* Read next n bits from the bitContainer.
* On 32-bits, don't read more than maxNbBits==25
* On 64-bits, don't read more than maxNbBits==57
* Use the fast variant *only* if n >= 1.
* return : value extracted.
*/
static size_t FSE_readBits(FSE_DStream_t* bitD, U32 nbBits)
{
size_t value = FSE_lookBits(bitD, nbBits);
FSE_skipBits(bitD, nbBits);
return value;
}
static size_t FSE_readBitsFast(FSE_DStream_t* bitD, U32 nbBits) /* only if nbBits >= 1 !! */
{
size_t value = FSE_lookBitsFast(bitD, nbBits);
FSE_skipBits(bitD, nbBits);
return value;
}
static unsigned FSE_reloadDStream(FSE_DStream_t* bitD)
{
if (bitD->bitsConsumed > (sizeof(bitD->bitContainer)*8)) /* should never happen */
return FSE_DStream_tooFar;
if (bitD->ptr >= bitD->start + sizeof(bitD->bitContainer))
{
bitD->ptr -= bitD->bitsConsumed >> 3;
bitD->bitsConsumed &= 7;
bitD->bitContainer = FSE_readLEST(bitD->ptr);
return FSE_DStream_unfinished;
}
if (bitD->ptr == bitD->start)
{
if (bitD->bitsConsumed < sizeof(bitD->bitContainer)*8) return FSE_DStream_endOfBuffer;
return FSE_DStream_completed;
}
{
U32 nbBytes = bitD->bitsConsumed >> 3;
U32 result = FSE_DStream_unfinished;
if (bitD->ptr - nbBytes < bitD->start)
{
nbBytes = (U32)(bitD->ptr - bitD->start); /* ptr > start */
result = FSE_DStream_endOfBuffer;
}
bitD->ptr -= nbBytes;
bitD->bitsConsumed -= nbBytes*8;
bitD->bitContainer = FSE_readLEST(bitD->ptr); /* reminder : srcSize > sizeof(bitD) */
return result;
}
}
static void FSE_initDState(FSE_DState_t* DStatePtr, FSE_DStream_t* bitD, const FSE_DTable* dt)
{
const void* ptr = dt;
const FSE_DTableHeader* const DTableH = (const FSE_DTableHeader*)ptr;
DStatePtr->state = FSE_readBits(bitD, DTableH->tableLog);
FSE_reloadDStream(bitD);
DStatePtr->table = dt + 1;
}
static BYTE FSE_decodeSymbol(FSE_DState_t* DStatePtr, FSE_DStream_t* bitD)
{
const FSE_decode_t DInfo = ((const FSE_decode_t*)(DStatePtr->table))[DStatePtr->state];
const U32 nbBits = DInfo.nbBits;
BYTE symbol = DInfo.symbol;
size_t lowBits = FSE_readBits(bitD, nbBits);
DStatePtr->state = DInfo.newState + lowBits;
return symbol;
}
static BYTE FSE_decodeSymbolFast(FSE_DState_t* DStatePtr, FSE_DStream_t* bitD)
{
const FSE_decode_t DInfo = ((const FSE_decode_t*)(DStatePtr->table))[DStatePtr->state];
const U32 nbBits = DInfo.nbBits;
BYTE symbol = DInfo.symbol;
size_t lowBits = FSE_readBitsFast(bitD, nbBits);
DStatePtr->state = DInfo.newState + lowBits;
return symbol;
}
/* FSE_endOfDStream
Tells if bitD has reached end of bitStream or not */
static unsigned FSE_endOfDStream(const FSE_DStream_t* bitD)
{
return ((bitD->ptr == bitD->start) && (bitD->bitsConsumed == sizeof(bitD->bitContainer)*8));
}
static unsigned FSE_endOfDState(const FSE_DState_t* DStatePtr)
{
return DStatePtr->state == 0;
}
FORCE_INLINE size_t FSE_decompress_usingDTable_generic(
void* dst, size_t maxDstSize,
const void* cSrc, size_t cSrcSize,
const FSE_DTable* dt, const unsigned fast)
{
BYTE* const ostart = (BYTE*) dst;
BYTE* op = ostart;
BYTE* const omax = op + maxDstSize;
BYTE* const olimit = omax-3;
FSE_DStream_t bitD;
FSE_DState_t state1;
FSE_DState_t state2;
size_t errorCode;
/* Init */
errorCode = FSE_initDStream(&bitD, cSrc, cSrcSize); /* replaced last arg by maxCompressed Size */
if (FSE_isError(errorCode)) return errorCode;
FSE_initDState(&state1, &bitD, dt);
FSE_initDState(&state2, &bitD, dt);
#define FSE_GETSYMBOL(statePtr) fast ? FSE_decodeSymbolFast(statePtr, &bitD) : FSE_decodeSymbol(statePtr, &bitD)
/* 4 symbols per loop */
for ( ; (FSE_reloadDStream(&bitD)==FSE_DStream_unfinished) && (op<olimit) ; op+=4)
{
op[0] = FSE_GETSYMBOL(&state1);
if (FSE_MAX_TABLELOG*2+7 > sizeof(bitD.bitContainer)*8) /* This test must be static */
FSE_reloadDStream(&bitD);
op[1] = FSE_GETSYMBOL(&state2);
if (FSE_MAX_TABLELOG*4+7 > sizeof(bitD.bitContainer)*8) /* This test must be static */
{ if (FSE_reloadDStream(&bitD) > FSE_DStream_unfinished) { op+=2; break; } }
op[2] = FSE_GETSYMBOL(&state1);
if (FSE_MAX_TABLELOG*2+7 > sizeof(bitD.bitContainer)*8) /* This test must be static */
FSE_reloadDStream(&bitD);
op[3] = FSE_GETSYMBOL(&state2);
}
/* tail */
/* note : FSE_reloadDStream(&bitD) >= FSE_DStream_partiallyFilled; Ends at exactly FSE_DStream_completed */
while (1)
{
if ( (FSE_reloadDStream(&bitD)>FSE_DStream_completed) || (op==omax) || (FSE_endOfDStream(&bitD) && (fast || FSE_endOfDState(&state1))) )
break;
*op++ = FSE_GETSYMBOL(&state1);
if ( (FSE_reloadDStream(&bitD)>FSE_DStream_completed) || (op==omax) || (FSE_endOfDStream(&bitD) && (fast || FSE_endOfDState(&state2))) )
break;
*op++ = FSE_GETSYMBOL(&state2);
}
/* end ? */
if (FSE_endOfDStream(&bitD) && FSE_endOfDState(&state1) && FSE_endOfDState(&state2))
return op-ostart;
if (op==omax) return (size_t)-FSE_ERROR_dstSize_tooSmall; /* dst buffer is full, but cSrc unfinished */
return (size_t)-FSE_ERROR_corruptionDetected;
}
static size_t FSE_decompress_usingDTable(void* dst, size_t originalSize,
const void* cSrc, size_t cSrcSize,
const FSE_DTable* dt)
{
FSE_DTableHeader DTableH;
memcpy(&DTableH, dt, sizeof(DTableH)); /* memcpy() into local variable, to avoid strict aliasing warning */
/* select fast mode (static) */
if (DTableH.fastMode) return FSE_decompress_usingDTable_generic(dst, originalSize, cSrc, cSrcSize, dt, 1);
return FSE_decompress_usingDTable_generic(dst, originalSize, cSrc, cSrcSize, dt, 0);
}
static size_t FSE_decompress(void* dst, size_t maxDstSize, const void* cSrc, size_t cSrcSize)
{
const BYTE* const istart = (const BYTE*)cSrc;
const BYTE* ip = istart;
short counting[FSE_MAX_SYMBOL_VALUE+1];
DTable_max_t dt; /* Static analyzer seems unable to understand this table will be properly initialized later */
unsigned tableLog;
unsigned maxSymbolValue = FSE_MAX_SYMBOL_VALUE;
size_t errorCode;
if (cSrcSize<2) return (size_t)-FSE_ERROR_srcSize_wrong; /* too small input size */
/* normal FSE decoding mode */
errorCode = FSE_readNCount (counting, &maxSymbolValue, &tableLog, istart, cSrcSize);
if (FSE_isError(errorCode)) return errorCode;
if (errorCode >= cSrcSize) return (size_t)-FSE_ERROR_srcSize_wrong; /* too small input size */
ip += errorCode;
cSrcSize -= errorCode;
errorCode = FSE_buildDTable (dt, counting, maxSymbolValue, tableLog);
if (FSE_isError(errorCode)) return errorCode;
/* always return, even if it is an error code */
return FSE_decompress_usingDTable (dst, maxDstSize, ip, cSrcSize, dt);
}
/* *******************************************************
* Huff0 : Huffman block compression
*********************************************************/
#define HUF_MAX_SYMBOL_VALUE 255
#define HUF_DEFAULT_TABLELOG 12 /* used by default, when not specified */
#define HUF_MAX_TABLELOG 12 /* max possible tableLog; for allocation purpose; can be modified */
#define HUF_ABSOLUTEMAX_TABLELOG 16 /* absolute limit of HUF_MAX_TABLELOG. Beyond that value, code does not work */
#if (HUF_MAX_TABLELOG > HUF_ABSOLUTEMAX_TABLELOG)
# error "HUF_MAX_TABLELOG is too large !"
#endif
typedef struct HUF_CElt_s {
U16 val;
BYTE nbBits;
} HUF_CElt ;
typedef struct nodeElt_s {
U32 count;
U16 parent;
BYTE byte;
BYTE nbBits;
} nodeElt;
/* *******************************************************
* Huff0 : Huffman block decompression
*********************************************************/
typedef struct {
BYTE byte;
BYTE nbBits;
} HUF_DElt;
static size_t HUF_readDTable (U16* DTable, const void* src, size_t srcSize)
{
BYTE huffWeight[HUF_MAX_SYMBOL_VALUE + 1];
U32 rankVal[HUF_ABSOLUTEMAX_TABLELOG + 1]; /* large enough for values from 0 to 16 */
U32 weightTotal;
U32 maxBits;
const BYTE* ip = (const BYTE*) src;
size_t iSize;
size_t oSize;
U32 n;
U32 nextRankStart;
void* ptr = DTable+1;
HUF_DElt* const dt = (HUF_DElt*)ptr;
if (!srcSize) return (size_t)-FSE_ERROR_srcSize_wrong;
iSize = ip[0];
FSE_STATIC_ASSERT(sizeof(HUF_DElt) == sizeof(U16)); /* if compilation fails here, assertion is false */
//memset(huffWeight, 0, sizeof(huffWeight)); /* should not be necessary, but some analyzer complain ... */
if (iSize >= 128) /* special header */
{
if (iSize >= (242)) /* RLE */
{
static int l[14] = { 1, 2, 3, 4, 7, 8, 15, 16, 31, 32, 63, 64, 127, 128 };
oSize = l[iSize-242];
memset(huffWeight, 1, sizeof(huffWeight));
iSize = 0;
}
else /* Incompressible */
{
oSize = iSize - 127;
iSize = ((oSize+1)/2);
if (iSize+1 > srcSize) return (size_t)-FSE_ERROR_srcSize_wrong;
ip += 1;
for (n=0; n<oSize; n+=2)
{
huffWeight[n] = ip[n/2] >> 4;
huffWeight[n+1] = ip[n/2] & 15;
}
}
}
else /* header compressed with FSE (normal case) */
{
if (iSize+1 > srcSize) return (size_t)-FSE_ERROR_srcSize_wrong;
oSize = FSE_decompress(huffWeight, HUF_MAX_SYMBOL_VALUE, ip+1, iSize); /* max 255 values decoded, last one is implied */
if (FSE_isError(oSize)) return oSize;
}
/* collect weight stats */
memset(rankVal, 0, sizeof(rankVal));
weightTotal = 0;
for (n=0; n<oSize; n++)
{
if (huffWeight[n] >= HUF_ABSOLUTEMAX_TABLELOG) return (size_t)-FSE_ERROR_corruptionDetected;
rankVal[huffWeight[n]]++;
weightTotal += (1 << huffWeight[n]) >> 1;
}
if (weightTotal == 0) return (size_t)-FSE_ERROR_corruptionDetected;
/* get last non-null symbol weight (implied, total must be 2^n) */
maxBits = FSE_highbit32(weightTotal) + 1;
if (maxBits > DTable[0]) return (size_t)-FSE_ERROR_tableLog_tooLarge; /* DTable is too small */
DTable[0] = (U16)maxBits;
{
U32 total = 1 << maxBits;
U32 rest = total - weightTotal;
U32 verif = 1 << FSE_highbit32(rest);
U32 lastWeight = FSE_highbit32(rest) + 1;
if (verif != rest) return (size_t)-FSE_ERROR_corruptionDetected; /* last value must be a clean power of 2 */
huffWeight[oSize] = (BYTE)lastWeight;
rankVal[lastWeight]++;
}
/* check tree construction validity */
if ((rankVal[1] < 2) || (rankVal[1] & 1)) return (size_t)-FSE_ERROR_corruptionDetected; /* by construction : at least 2 elts of rank 1, must be even */
/* Prepare ranks */
nextRankStart = 0;
for (n=1; n<=maxBits; n++)
{
U32 current = nextRankStart;
nextRankStart += (rankVal[n] << (n-1));
rankVal[n] = current;
}
/* fill DTable */
for (n=0; n<=oSize; n++)
{
const U32 w = huffWeight[n];
const U32 length = (1 << w) >> 1;
U32 i;
HUF_DElt D;
D.byte = (BYTE)n; D.nbBits = (BYTE)(maxBits + 1 - w);
for (i = rankVal[w]; i < rankVal[w] + length; i++)
dt[i] = D;
rankVal[w] += length;
}
return iSize+1;
}
static BYTE HUF_decodeSymbol(FSE_DStream_t* Dstream, const HUF_DElt* dt, const U32 dtLog)
{
const size_t val = FSE_lookBitsFast(Dstream, dtLog); /* note : dtLog >= 1 */
const BYTE c = dt[val].byte;
FSE_skipBits(Dstream, dt[val].nbBits);
return c;
}
static size_t HUF_decompress_usingDTable( /* -3% slower when non static */
void* dst, size_t maxDstSize,
const void* cSrc, size_t cSrcSize,
const U16* DTable)
{
if (cSrcSize < 6) return (size_t)-FSE_ERROR_srcSize_wrong;
{
BYTE* const ostart = (BYTE*) dst;
BYTE* op = ostart;
BYTE* const omax = op + maxDstSize;
BYTE* const olimit = maxDstSize < 15 ? op : omax-15;
const void* ptr = DTable;
const HUF_DElt* const dt = (const HUF_DElt*)(ptr)+1;
const U32 dtLog = DTable[0];
size_t errorCode;
U32 reloadStatus;
/* Init */
const U16* jumpTable = (const U16*)cSrc;
const size_t length1 = FSE_readLE16(jumpTable);
const size_t length2 = FSE_readLE16(jumpTable+1);
const size_t length3 = FSE_readLE16(jumpTable+2);
const size_t length4 = cSrcSize - 6 - length1 - length2 - length3; /* check coherency !! */
const char* const start1 = (const char*)(cSrc) + 6;
const char* const start2 = start1 + length1;
const char* const start3 = start2 + length2;
const char* const start4 = start3 + length3;
FSE_DStream_t bitD1, bitD2, bitD3, bitD4;
if (length1+length2+length3+6 >= cSrcSize) return (size_t)-FSE_ERROR_srcSize_wrong;
errorCode = FSE_initDStream(&bitD1, start1, length1);
if (FSE_isError(errorCode)) return errorCode;
errorCode = FSE_initDStream(&bitD2, start2, length2);
if (FSE_isError(errorCode)) return errorCode;
errorCode = FSE_initDStream(&bitD3, start3, length3);
if (FSE_isError(errorCode)) return errorCode;
errorCode = FSE_initDStream(&bitD4, start4, length4);
if (FSE_isError(errorCode)) return errorCode;
reloadStatus=FSE_reloadDStream(&bitD2);
/* 16 symbols per loop */
for ( ; (reloadStatus<FSE_DStream_completed) && (op<olimit); /* D2-3-4 are supposed to be synchronized and finish together */
op+=16, reloadStatus = FSE_reloadDStream(&bitD2) | FSE_reloadDStream(&bitD3) | FSE_reloadDStream(&bitD4), FSE_reloadDStream(&bitD1))
{
#define HUF_DECODE_SYMBOL_0(n, Dstream) \
op[n] = HUF_decodeSymbol(&Dstream, dt, dtLog);
#define HUF_DECODE_SYMBOL_1(n, Dstream) \
op[n] = HUF_decodeSymbol(&Dstream, dt, dtLog); \
if (FSE_32bits() && (HUF_MAX_TABLELOG>12)) FSE_reloadDStream(&Dstream)
#define HUF_DECODE_SYMBOL_2(n, Dstream) \
op[n] = HUF_decodeSymbol(&Dstream, dt, dtLog); \
if (FSE_32bits()) FSE_reloadDStream(&Dstream)
HUF_DECODE_SYMBOL_1( 0, bitD1);
HUF_DECODE_SYMBOL_1( 1, bitD2);
HUF_DECODE_SYMBOL_1( 2, bitD3);
HUF_DECODE_SYMBOL_1( 3, bitD4);
HUF_DECODE_SYMBOL_2( 4, bitD1);
HUF_DECODE_SYMBOL_2( 5, bitD2);
HUF_DECODE_SYMBOL_2( 6, bitD3);
HUF_DECODE_SYMBOL_2( 7, bitD4);
HUF_DECODE_SYMBOL_1( 8, bitD1);
HUF_DECODE_SYMBOL_1( 9, bitD2);
HUF_DECODE_SYMBOL_1(10, bitD3);
HUF_DECODE_SYMBOL_1(11, bitD4);
HUF_DECODE_SYMBOL_0(12, bitD1);
HUF_DECODE_SYMBOL_0(13, bitD2);
HUF_DECODE_SYMBOL_0(14, bitD3);
HUF_DECODE_SYMBOL_0(15, bitD4);
}
if (reloadStatus!=FSE_DStream_completed) /* not complete : some bitStream might be FSE_DStream_unfinished */
return (size_t)-FSE_ERROR_corruptionDetected;
/* tail */
{
/* bitTail = bitD1; */ /* *much* slower : -20% !??! */
FSE_DStream_t bitTail;
bitTail.ptr = bitD1.ptr;
bitTail.bitsConsumed = bitD1.bitsConsumed;
bitTail.bitContainer = bitD1.bitContainer; /* required in case of FSE_DStream_endOfBuffer */
bitTail.start = start1;
for ( ; (FSE_reloadDStream(&bitTail) < FSE_DStream_completed) && (op<omax) ; op++)
{
HUF_DECODE_SYMBOL_0(0, bitTail);
}
if (FSE_endOfDStream(&bitTail))
return op-ostart;
}
if (op==omax) return (size_t)-FSE_ERROR_dstSize_tooSmall; /* dst buffer is full, but cSrc unfinished */
return (size_t)-FSE_ERROR_corruptionDetected;
}
}
static size_t HUF_decompress (void* dst, size_t maxDstSize, const void* cSrc, size_t cSrcSize)
{
HUF_CREATE_STATIC_DTABLE(DTable, HUF_MAX_TABLELOG);
const BYTE* ip = (const BYTE*) cSrc;
size_t errorCode;
errorCode = HUF_readDTable (DTable, cSrc, cSrcSize);
if (FSE_isError(errorCode)) return errorCode;
if (errorCode >= cSrcSize) return (size_t)-FSE_ERROR_srcSize_wrong;
ip += errorCode;
cSrcSize -= errorCode;
return HUF_decompress_usingDTable (dst, maxDstSize, ip, cSrcSize, DTable);
}
#endif /* FSE_COMMONDEFS_ONLY */
/*
zstd - standard compression library
Copyright (C) 2014-2015, Yann Collet.
BSD 2-Clause License (https://opensource.org/licenses/bsd-license.php)
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are
met:
* Redistributions of source code must retain the above copyright
notice, this list of conditions and the following disclaimer.
* Redistributions in binary form must reproduce the above
copyright notice, this list of conditions and the following disclaimer
in the documentation and/or other materials provided with the
distribution.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS
"AS IS" AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT
LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR
A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT
OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL,
SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT
LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE,
DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
You can contact the author at :
- zstd source repository : https://github.com/Cyan4973/zstd
- ztsd public forum : https://groups.google.com/forum/#!forum/lz4c
*/
/****************************************************************
* Tuning parameters
*****************************************************************/
/* MEMORY_USAGE :
* Memory usage formula : N->2^N Bytes (examples : 10 -> 1KB; 12 -> 4KB ; 16 -> 64KB; 20 -> 1MB; etc.)
* Increasing memory usage improves compression ratio
* Reduced memory usage can improve speed, due to cache effect */
#define ZSTD_MEMORY_USAGE 17
/**************************************
CPU Feature Detection
**************************************/
/*
* Automated efficient unaligned memory access detection
* Based on known hardware architectures
* This list will be updated thanks to feedbacks
*/
#if defined(CPU_HAS_EFFICIENT_UNALIGNED_MEMORY_ACCESS) \
|| defined(__ARM_FEATURE_UNALIGNED) \
|| defined(__i386__) || defined(__x86_64__) \
|| defined(_M_IX86) || defined(_M_X64) \
|| defined(__ARM_ARCH_7__) || defined(__ARM_ARCH_8__) \
|| (defined(_M_ARM) && (_M_ARM >= 7))
# define ZSTD_UNALIGNED_ACCESS 1
#else
# define ZSTD_UNALIGNED_ACCESS 0
#endif
/********************************************************
* Includes
*********************************************************/
#include <stdlib.h> /* calloc */
#include <string.h> /* memcpy, memmove */
#include <stdio.h> /* debug : printf */
/********************************************************
* Compiler specifics
*********************************************************/
#ifdef __AVX2__
# include <immintrin.h> /* AVX2 intrinsics */
#endif
#ifdef _MSC_VER /* Visual Studio */
# include <intrin.h> /* For Visual 2005 */
# pragma warning(disable : 4127) /* disable: C4127: conditional expression is constant */
# pragma warning(disable : 4324) /* disable: C4324: padded structure */
#endif
#ifndef MEM_ACCESS_MODULE
#define MEM_ACCESS_MODULE
/********************************************************
* Basic Types
*********************************************************/
#if defined (__STDC_VERSION__) && __STDC_VERSION__ >= 199901L /* C99 */
# if defined(_AIX)
# include <inttypes.h>
# else
# include <stdint.h> /* intptr_t */
# endif
typedef uint8_t BYTE;
typedef uint16_t U16;
typedef int16_t S16;
typedef uint32_t U32;
typedef int32_t S32;
typedef uint64_t U64;
#else
typedef unsigned char BYTE;
typedef unsigned short U16;
typedef signed short S16;
typedef unsigned int U32;
typedef signed int S32;
typedef unsigned long long U64;
#endif
#endif /* MEM_ACCESS_MODULE */
/********************************************************
* Constants
*********************************************************/
static const U32 ZSTD_magicNumber = 0xFD2FB51E; /* 3rd version : seqNb header */
#define HASH_LOG (ZSTD_MEMORY_USAGE - 2)
#define HASH_TABLESIZE (1 << HASH_LOG)
#define HASH_MASK (HASH_TABLESIZE - 1)
#define KNUTH 2654435761
#define BIT7 128
#define BIT6 64
#define BIT5 32
#define BIT4 16
#define KB *(1 <<10)
#define MB *(1 <<20)
#define GB *(1U<<30)
#define BLOCKSIZE (128 KB) /* define, for static allocation */
#define WORKPLACESIZE (BLOCKSIZE*3)
#define MINMATCH 4
#define MLbits 7
#define LLbits 6
#define Offbits 5
#define MaxML ((1<<MLbits )-1)
#define MaxLL ((1<<LLbits )-1)
#define MaxOff ((1<<Offbits)-1)
#define LitFSELog 11
#define MLFSELog 10
#define LLFSELog 10
#define OffFSELog 9
#define MAX(a,b) ((a)<(b)?(b):(a))
#define MaxSeq MAX(MaxLL, MaxML)
#define LITERAL_NOENTROPY 63
#define COMMAND_NOENTROPY 7 /* to remove */
#define ZSTD_CONTENTSIZE_ERROR (0ULL - 2)
static const size_t ZSTD_blockHeaderSize = 3;
static const size_t ZSTD_frameHeaderSize = 4;
/********************************************************
* Memory operations
*********************************************************/
static unsigned ZSTD_32bits(void) { return sizeof(void*)==4; }
static unsigned ZSTD_isLittleEndian(void)
{
const union { U32 i; BYTE c[4]; } one = { 1 }; /* don't use static : performance detrimental */
return one.c[0];
}
static U16 ZSTD_read16(const void* p) { U16 r; memcpy(&r, p, sizeof(r)); return r; }
static void ZSTD_copy4(void* dst, const void* src) { memcpy(dst, src, 4); }
static void ZSTD_copy8(void* dst, const void* src) { memcpy(dst, src, 8); }
#define COPY8(d,s) { ZSTD_copy8(d,s); d+=8; s+=8; }
static void ZSTD_wildcopy(void* dst, const void* src, ptrdiff_t length)
{
const BYTE* ip = (const BYTE*)src;
BYTE* op = (BYTE*)dst;
BYTE* const oend = op + length;
while (op < oend) COPY8(op, ip);
}
static U16 ZSTD_readLE16(const void* memPtr)
{
if (ZSTD_isLittleEndian()) return ZSTD_read16(memPtr);
else
{
const BYTE* p = (const BYTE*)memPtr;
return (U16)((U16)p[0] + ((U16)p[1]<<8));
}
}
static U32 ZSTD_readLE24(const void* memPtr)
{
return ZSTD_readLE16(memPtr) + (((const BYTE*)memPtr)[2] << 16);
}
static U32 ZSTD_readBE32(const void* memPtr)
{
const BYTE* p = (const BYTE*)memPtr;
return (U32)(((U32)p[0]<<24) + ((U32)p[1]<<16) + ((U32)p[2]<<8) + ((U32)p[3]<<0));
}
/**************************************
* Local structures
***************************************/
typedef struct ZSTD_Cctx_s ZSTD_Cctx;
typedef enum { bt_compressed, bt_raw, bt_rle, bt_end } blockType_t;
typedef struct
{
blockType_t blockType;
U32 origSize;
} blockProperties_t;
typedef struct {
void* buffer;
U32* offsetStart;
U32* offset;
BYTE* offCodeStart;
BYTE* offCode;
BYTE* litStart;
BYTE* lit;
BYTE* litLengthStart;
BYTE* litLength;
BYTE* matchLengthStart;
BYTE* matchLength;
BYTE* dumpsStart;
BYTE* dumps;
} SeqStore_t;
typedef struct ZSTD_Cctx_s
{
const BYTE* base;
U32 current;
U32 nextUpdate;
SeqStore_t seqStore;
#ifdef __AVX2__
__m256i hashTable[HASH_TABLESIZE>>3];
#else
U32 hashTable[HASH_TABLESIZE];
#endif
BYTE buffer[WORKPLACESIZE];
} cctxi_t;
/**************************************
* Error Management
**************************************/
/* published entry point */
unsigned ZSTDv01_isError(size_t code) { return ERR_isError(code); }
/**************************************
* Tool functions
**************************************/
#define ZSTD_VERSION_MAJOR 0 /* for breaking interface changes */
#define ZSTD_VERSION_MINOR 1 /* for new (non-breaking) interface capabilities */
#define ZSTD_VERSION_RELEASE 3 /* for tweaks, bug-fixes, or development */
#define ZSTD_VERSION_NUMBER (ZSTD_VERSION_MAJOR *100*100 + ZSTD_VERSION_MINOR *100 + ZSTD_VERSION_RELEASE)
/**************************************************************
* Decompression code
**************************************************************/
static size_t ZSTDv01_getcBlockSize(const void* src, size_t srcSize, blockProperties_t* bpPtr)
{
const BYTE* const in = (const BYTE* const)src;
BYTE headerFlags;
U32 cSize;
if (srcSize < 3) return ERROR(srcSize_wrong);
headerFlags = *in;
cSize = in[2] + (in[1]<<8) + ((in[0] & 7)<<16);
bpPtr->blockType = (blockType_t)(headerFlags >> 6);
bpPtr->origSize = (bpPtr->blockType == bt_rle) ? cSize : 0;
if (bpPtr->blockType == bt_end) return 0;
if (bpPtr->blockType == bt_rle) return 1;
return cSize;
}
static size_t ZSTD_copyUncompressedBlock(void* dst, size_t maxDstSize, const void* src, size_t srcSize)
{
if (srcSize > maxDstSize) return ERROR(dstSize_tooSmall);
if (srcSize > 0) {
memcpy(dst, src, srcSize);
}
return srcSize;
}
static size_t ZSTD_decompressLiterals(void* ctx,
void* dst, size_t maxDstSize,
const void* src, size_t srcSize)
{
BYTE* op = (BYTE*)dst;
BYTE* const oend = op + maxDstSize;
const BYTE* ip = (const BYTE*)src;
size_t errorCode;
size_t litSize;
/* check : minimum 2, for litSize, +1, for content */
if (srcSize <= 3) return ERROR(corruption_detected);
litSize = ip[1] + (ip[0]<<8);
litSize += ((ip[-3] >> 3) & 7) << 16; /* mmmmh.... */
op = oend - litSize;
(void)ctx;
if (litSize > maxDstSize) return ERROR(dstSize_tooSmall);
errorCode = HUF_decompress(op, litSize, ip+2, srcSize-2);
if (FSE_isError(errorCode)) return ERROR(GENERIC);
return litSize;
}
static size_t ZSTDv01_decodeLiteralsBlock(void* ctx,
void* dst, size_t maxDstSize,
const BYTE** litStart, size_t* litSize,
const void* src, size_t srcSize)
{
const BYTE* const istart = (const BYTE* const)src;
const BYTE* ip = istart;
BYTE* const ostart = (BYTE* const)dst;
BYTE* const oend = ostart + maxDstSize;
blockProperties_t litbp;
size_t litcSize = ZSTDv01_getcBlockSize(src, srcSize, &litbp);
if (ZSTDv01_isError(litcSize)) return litcSize;
if (litcSize > srcSize - ZSTD_blockHeaderSize) return ERROR(srcSize_wrong);
ip += ZSTD_blockHeaderSize;
switch(litbp.blockType)
{
case bt_raw:
*litStart = ip;
ip += litcSize;
*litSize = litcSize;
break;
case bt_rle:
{
size_t rleSize = litbp.origSize;
if (rleSize>maxDstSize) return ERROR(dstSize_tooSmall);
if (!srcSize) return ERROR(srcSize_wrong);
if (rleSize > 0) {
memset(oend - rleSize, *ip, rleSize);
}
*litStart = oend - rleSize;
*litSize = rleSize;
ip++;
break;
}
case bt_compressed:
{
size_t decodedLitSize = ZSTD_decompressLiterals(ctx, dst, maxDstSize, ip, litcSize);
if (ZSTDv01_isError(decodedLitSize)) return decodedLitSize;
*litStart = oend - decodedLitSize;
*litSize = decodedLitSize;
ip += litcSize;
break;
}
case bt_end:
default:
return ERROR(GENERIC);
}
return ip-istart;
}
static size_t ZSTDv01_decodeSeqHeaders(int* nbSeq, const BYTE** dumpsPtr, size_t* dumpsLengthPtr,
FSE_DTable* DTableLL, FSE_DTable* DTableML, FSE_DTable* DTableOffb,
const void* src, size_t srcSize)
{
const BYTE* const istart = (const BYTE* const)src;
const BYTE* ip = istart;
const BYTE* const iend = istart + srcSize;
U32 LLtype, Offtype, MLtype;
U32 LLlog, Offlog, MLlog;
size_t dumpsLength;
/* check */
if (srcSize < 5) return ERROR(srcSize_wrong);
/* SeqHead */
*nbSeq = ZSTD_readLE16(ip); ip+=2;
LLtype = *ip >> 6;
Offtype = (*ip >> 4) & 3;
MLtype = (*ip >> 2) & 3;
if (*ip & 2)
{
dumpsLength = ip[2];
dumpsLength += ip[1] << 8;
ip += 3;
}
else
{
dumpsLength = ip[1];
dumpsLength += (ip[0] & 1) << 8;
ip += 2;
}
*dumpsPtr = ip;
ip += dumpsLength;
*dumpsLengthPtr = dumpsLength;
/* check */
if (ip > iend-3) return ERROR(srcSize_wrong); /* min : all 3 are "raw", hence no header, but at least xxLog bits per type */
/* sequences */
{
S16 norm[MaxML+1]; /* assumption : MaxML >= MaxLL and MaxOff */
size_t headerSize;
/* Build DTables */
switch(LLtype)
{
case bt_rle :
LLlog = 0;
FSE_buildDTable_rle(DTableLL, *ip++); break;
case bt_raw :
LLlog = LLbits;
FSE_buildDTable_raw(DTableLL, LLbits); break;
default :
{ U32 max = MaxLL;
headerSize = FSE_readNCount(norm, &max, &LLlog, ip, iend-ip);
if (FSE_isError(headerSize)) return ERROR(GENERIC);
if (LLlog > LLFSELog) return ERROR(corruption_detected);
ip += headerSize;
FSE_buildDTable(DTableLL, norm, max, LLlog);
} }
switch(Offtype)
{
case bt_rle :
Offlog = 0;
if (ip > iend-2) return ERROR(srcSize_wrong); /* min : "raw", hence no header, but at least xxLog bits */
FSE_buildDTable_rle(DTableOffb, *ip++); break;
case bt_raw :
Offlog = Offbits;
FSE_buildDTable_raw(DTableOffb, Offbits); break;
default :
{ U32 max = MaxOff;
headerSize = FSE_readNCount(norm, &max, &Offlog, ip, iend-ip);
if (FSE_isError(headerSize)) return ERROR(GENERIC);
if (Offlog > OffFSELog) return ERROR(corruption_detected);
ip += headerSize;
FSE_buildDTable(DTableOffb, norm, max, Offlog);
} }
switch(MLtype)
{
case bt_rle :
MLlog = 0;
if (ip > iend-2) return ERROR(srcSize_wrong); /* min : "raw", hence no header, but at least xxLog bits */
FSE_buildDTable_rle(DTableML, *ip++); break;
case bt_raw :
MLlog = MLbits;
FSE_buildDTable_raw(DTableML, MLbits); break;
default :
{ U32 max = MaxML;
headerSize = FSE_readNCount(norm, &max, &MLlog, ip, iend-ip);
if (FSE_isError(headerSize)) return ERROR(GENERIC);
if (MLlog > MLFSELog) return ERROR(corruption_detected);
ip += headerSize;
FSE_buildDTable(DTableML, norm, max, MLlog);
} } }
return ip-istart;
}
typedef struct {
size_t litLength;
size_t offset;
size_t matchLength;
} seq_t;
typedef struct {
FSE_DStream_t DStream;
FSE_DState_t stateLL;
FSE_DState_t stateOffb;
FSE_DState_t stateML;
size_t prevOffset;
const BYTE* dumps;
const BYTE* dumpsEnd;
} seqState_t;
static void ZSTD_decodeSequence(seq_t* seq, seqState_t* seqState)
{
size_t litLength;
size_t prevOffset;
size_t offset;
size_t matchLength;
const BYTE* dumps = seqState->dumps;
const BYTE* const de = seqState->dumpsEnd;
/* Literal length */
litLength = FSE_decodeSymbol(&(seqState->stateLL), &(seqState->DStream));
prevOffset = litLength ? seq->offset : seqState->prevOffset;
seqState->prevOffset = seq->offset;
if (litLength == MaxLL)
{
const U32 add = dumps<de ? *dumps++ : 0;
if (add < 255) litLength += add;
else
{
if (dumps<=(de-3))
{
litLength = ZSTD_readLE24(dumps);
dumps += 3;
}
}
}
/* Offset */
{
U32 offsetCode, nbBits;
offsetCode = FSE_decodeSymbol(&(seqState->stateOffb), &(seqState->DStream));
if (ZSTD_32bits()) FSE_reloadDStream(&(seqState->DStream));
nbBits = offsetCode - 1;
if (offsetCode==0) nbBits = 0; /* cmove */
offset = ((size_t)1 << (nbBits & ((sizeof(offset)*8)-1))) + FSE_readBits(&(seqState->DStream), nbBits);
if (ZSTD_32bits()) FSE_reloadDStream(&(seqState->DStream));
if (offsetCode==0) offset = prevOffset;
}
/* MatchLength */
matchLength = FSE_decodeSymbol(&(seqState->stateML), &(seqState->DStream));
if (matchLength == MaxML)
{
const U32 add = dumps<de ? *dumps++ : 0;
if (add < 255) matchLength += add;
else
{
if (dumps<=(de-3))
{
matchLength = ZSTD_readLE24(dumps);
dumps += 3;
}
}
}
matchLength += MINMATCH;
/* save result */
seq->litLength = litLength;
seq->offset = offset;
seq->matchLength = matchLength;
seqState->dumps = dumps;
}
static size_t ZSTD_execSequence(BYTE* op,
seq_t sequence,
const BYTE** litPtr, const BYTE* const litLimit,
BYTE* const base, BYTE* const oend)
{
static const int dec32table[] = {0, 1, 2, 1, 4, 4, 4, 4}; /* added */
static const int dec64table[] = {8, 8, 8, 7, 8, 9,10,11}; /* subtracted */
const BYTE* const ostart = op;
BYTE* const oLitEnd = op + sequence.litLength;
const size_t litLength = sequence.litLength;
BYTE* const endMatch = op + litLength + sequence.matchLength; /* risk : address space overflow (32-bits) */
const BYTE* const litEnd = *litPtr + litLength;
/* checks */
size_t const seqLength = sequence.litLength + sequence.matchLength;
if (seqLength > (size_t)(oend - op)) return ERROR(dstSize_tooSmall);
if (sequence.litLength > (size_t)(litLimit - *litPtr)) return ERROR(corruption_detected);
/* Now we know there are no overflow in literal nor match lengths, can use pointer checks */
if (sequence.offset > (U32)(oLitEnd - base)) return ERROR(corruption_detected);
if (endMatch > oend) return ERROR(dstSize_tooSmall); /* overwrite beyond dst buffer */
if (litEnd > litLimit) return ERROR(corruption_detected); /* overRead beyond lit buffer */
if (sequence.matchLength > (size_t)(*litPtr-op)) return ERROR(dstSize_tooSmall); /* overwrite literal segment */
/* copy Literals */
ZSTD_memmove(op, *litPtr, sequence.litLength); /* note : v0.1 seems to allow scenarios where output or input are close to end of buffer */
op += litLength;
*litPtr = litEnd; /* update for next sequence */
/* check : last match must be at a minimum distance of 8 from end of dest buffer */
if (oend-op < 8) return ERROR(dstSize_tooSmall);
/* copy Match */
{
const U32 overlapRisk = (((size_t)(litEnd - endMatch)) < 12);
const BYTE* match = op - sequence.offset; /* possible underflow at op - offset ? */
size_t qutt = 12;
U64 saved[2];
/* check */
if (match < base) return ERROR(corruption_detected);
if (sequence.offset > (size_t)base) return ERROR(corruption_detected);
/* save beginning of literal sequence, in case of write overlap */
if (overlapRisk)
{
if ((endMatch + qutt) > oend) qutt = oend-endMatch;
memcpy(saved, endMatch, qutt);
}
if (sequence.offset < 8)
{
const int dec64 = dec64table[sequence.offset];
op[0] = match[0];
op[1] = match[1];
op[2] = match[2];
op[3] = match[3];
match += dec32table[sequence.offset];
ZSTD_copy4(op+4, match);
match -= dec64;
} else { ZSTD_copy8(op, match); }
op += 8; match += 8;
if (endMatch > oend-(16-MINMATCH))
{
if (op < oend-8)
{
ZSTD_wildcopy(op, match, (oend-8) - op);
match += (oend-8) - op;
op = oend-8;
}
while (op<endMatch) *op++ = *match++;
}
else
ZSTD_wildcopy(op, match, (ptrdiff_t)sequence.matchLength-8); /* works even if matchLength < 8 */
/* restore, in case of overlap */
if (overlapRisk) memcpy(endMatch, saved, qutt);
}
return endMatch-ostart;
}
typedef struct ZSTDv01_Dctx_s
{
U32 LLTable[FSE_DTABLE_SIZE_U32(LLFSELog)];
U32 OffTable[FSE_DTABLE_SIZE_U32(OffFSELog)];
U32 MLTable[FSE_DTABLE_SIZE_U32(MLFSELog)];
void* previousDstEnd;
void* base;
size_t expected;
blockType_t bType;
U32 phase;
} dctx_t;
static size_t ZSTD_decompressSequences(
void* ctx,
void* dst, size_t maxDstSize,
const void* seqStart, size_t seqSize,
const BYTE* litStart, size_t litSize)
{
dctx_t* dctx = (dctx_t*)ctx;
const BYTE* ip = (const BYTE*)seqStart;
const BYTE* const iend = ip + seqSize;
BYTE* const ostart = (BYTE* const)dst;
BYTE* op = ostart;
BYTE* const oend = ostart + maxDstSize;
size_t errorCode, dumpsLength;
const BYTE* litPtr = litStart;
const BYTE* const litEnd = litStart + litSize;
int nbSeq;
const BYTE* dumps;
U32* DTableLL = dctx->LLTable;
U32* DTableML = dctx->MLTable;
U32* DTableOffb = dctx->OffTable;
BYTE* const base = (BYTE*) (dctx->base);
/* Build Decoding Tables */
errorCode = ZSTDv01_decodeSeqHeaders(&nbSeq, &dumps, &dumpsLength,
DTableLL, DTableML, DTableOffb,
ip, iend-ip);
if (ZSTDv01_isError(errorCode)) return errorCode;
ip += errorCode;
/* Regen sequences */
{
seq_t sequence;
seqState_t seqState;
memset(&sequence, 0, sizeof(sequence));
seqState.dumps = dumps;
seqState.dumpsEnd = dumps + dumpsLength;
seqState.prevOffset = 1;
errorCode = FSE_initDStream(&(seqState.DStream), ip, iend-ip);
if (FSE_isError(errorCode)) return ERROR(corruption_detected);
FSE_initDState(&(seqState.stateLL), &(seqState.DStream), DTableLL);
FSE_initDState(&(seqState.stateOffb), &(seqState.DStream), DTableOffb);
FSE_initDState(&(seqState.stateML), &(seqState.DStream), DTableML);
for ( ; (FSE_reloadDStream(&(seqState.DStream)) <= FSE_DStream_completed) && (nbSeq>0) ; )
{
size_t oneSeqSize;
nbSeq--;
ZSTD_decodeSequence(&sequence, &seqState);
oneSeqSize = ZSTD_execSequence(op, sequence, &litPtr, litEnd, base, oend);
if (ZSTDv01_isError(oneSeqSize)) return oneSeqSize;
op += oneSeqSize;
}
/* check if reached exact end */
if ( !FSE_endOfDStream(&(seqState.DStream)) ) return ERROR(corruption_detected); /* requested too much : data is corrupted */
if (nbSeq<0) return ERROR(corruption_detected); /* requested too many sequences : data is corrupted */
/* last literal segment */
{
size_t lastLLSize = litEnd - litPtr;
if (op+lastLLSize > oend) return ERROR(dstSize_tooSmall);
if (lastLLSize > 0) {
if (op != litPtr) memmove(op, litPtr, lastLLSize);
op += lastLLSize;
}
}
}
return op-ostart;
}
static size_t ZSTD_decompressBlock(
void* ctx,
void* dst, size_t maxDstSize,
const void* src, size_t srcSize)
{
/* blockType == blockCompressed, srcSize is trusted */
const BYTE* ip = (const BYTE*)src;
const BYTE* litPtr = NULL;
size_t litSize = 0;
size_t errorCode;
/* Decode literals sub-block */
errorCode = ZSTDv01_decodeLiteralsBlock(ctx, dst, maxDstSize, &litPtr, &litSize, src, srcSize);
if (ZSTDv01_isError(errorCode)) return errorCode;
ip += errorCode;
srcSize -= errorCode;
return ZSTD_decompressSequences(ctx, dst, maxDstSize, ip, srcSize, litPtr, litSize);
}
size_t ZSTDv01_decompressDCtx(void* ctx, void* dst, size_t maxDstSize, const void* src, size_t srcSize)
{
const BYTE* ip = (const BYTE*)src;
const BYTE* iend = ip + srcSize;
BYTE* const ostart = (BYTE* const)dst;
BYTE* op = ostart;
BYTE* const oend = ostart + maxDstSize;
size_t remainingSize = srcSize;
U32 magicNumber;
size_t errorCode=0;
blockProperties_t blockProperties;
/* Frame Header */
if (srcSize < ZSTD_frameHeaderSize+ZSTD_blockHeaderSize) return ERROR(srcSize_wrong);
magicNumber = ZSTD_readBE32(src);
if (magicNumber != ZSTD_magicNumber) return ERROR(prefix_unknown);
ip += ZSTD_frameHeaderSize; remainingSize -= ZSTD_frameHeaderSize;
/* Loop on each block */
while (1)
{
size_t blockSize = ZSTDv01_getcBlockSize(ip, iend-ip, &blockProperties);
if (ZSTDv01_isError(blockSize)) return blockSize;
ip += ZSTD_blockHeaderSize;
remainingSize -= ZSTD_blockHeaderSize;
if (blockSize > remainingSize) return ERROR(srcSize_wrong);
switch(blockProperties.blockType)
{
case bt_compressed:
errorCode = ZSTD_decompressBlock(ctx, op, oend-op, ip, blockSize);
break;
case bt_raw :
errorCode = ZSTD_copyUncompressedBlock(op, oend-op, ip, blockSize);
break;
case bt_rle :
return ERROR(GENERIC); /* not yet supported */
break;
case bt_end :
/* end of frame */
if (remainingSize) return ERROR(srcSize_wrong);
break;
default:
return ERROR(GENERIC);
}
if (blockSize == 0) break; /* bt_end */
if (ZSTDv01_isError(errorCode)) return errorCode;
op += errorCode;
ip += blockSize;
remainingSize -= blockSize;
}
return op-ostart;
}
size_t ZSTDv01_decompress(void* dst, size_t maxDstSize, const void* src, size_t srcSize)
{
dctx_t ctx;
ctx.base = dst;
return ZSTDv01_decompressDCtx(&ctx, dst, maxDstSize, src, srcSize);
}
/* ZSTD_errorFrameSizeInfoLegacy() :
assumes `cSize` and `dBound` are _not_ NULL */
static void ZSTD_errorFrameSizeInfoLegacy(size_t* cSize, unsigned long long* dBound, size_t ret)
{
*cSize = ret;
*dBound = ZSTD_CONTENTSIZE_ERROR;
}
void ZSTDv01_findFrameSizeInfoLegacy(const void *src, size_t srcSize, size_t* cSize, unsigned long long* dBound)
{
const BYTE* ip = (const BYTE*)src;
size_t remainingSize = srcSize;
size_t nbBlocks = 0;
U32 magicNumber;
blockProperties_t blockProperties;
/* Frame Header */
if (srcSize < ZSTD_frameHeaderSize+ZSTD_blockHeaderSize) {
ZSTD_errorFrameSizeInfoLegacy(cSize, dBound, ERROR(srcSize_wrong));
return;
}
magicNumber = ZSTD_readBE32(src);
if (magicNumber != ZSTD_magicNumber) {
ZSTD_errorFrameSizeInfoLegacy(cSize, dBound, ERROR(prefix_unknown));
return;
}
ip += ZSTD_frameHeaderSize; remainingSize -= ZSTD_frameHeaderSize;
/* Loop on each block */
while (1)
{
size_t blockSize = ZSTDv01_getcBlockSize(ip, remainingSize, &blockProperties);
if (ZSTDv01_isError(blockSize)) {
ZSTD_errorFrameSizeInfoLegacy(cSize, dBound, blockSize);
return;
}
ip += ZSTD_blockHeaderSize;
remainingSize -= ZSTD_blockHeaderSize;
if (blockSize > remainingSize) {
ZSTD_errorFrameSizeInfoLegacy(cSize, dBound, ERROR(srcSize_wrong));
return;
}
if (blockSize == 0) break; /* bt_end */
ip += blockSize;
remainingSize -= blockSize;
nbBlocks++;
}
*cSize = ip - (const BYTE*)src;
*dBound = nbBlocks * BLOCKSIZE;
}
/*******************************
* Streaming Decompression API
*******************************/
size_t ZSTDv01_resetDCtx(ZSTDv01_Dctx* dctx)
{
dctx->expected = ZSTD_frameHeaderSize;
dctx->phase = 0;
dctx->previousDstEnd = NULL;
dctx->base = NULL;
return 0;
}
ZSTDv01_Dctx* ZSTDv01_createDCtx(void)
{
ZSTDv01_Dctx* dctx = (ZSTDv01_Dctx*)malloc(sizeof(ZSTDv01_Dctx));
if (dctx==NULL) return NULL;
ZSTDv01_resetDCtx(dctx);
return dctx;
}
size_t ZSTDv01_freeDCtx(ZSTDv01_Dctx* dctx)
{
free(dctx);
return 0;
}
size_t ZSTDv01_nextSrcSizeToDecompress(ZSTDv01_Dctx* dctx)
{
return ((dctx_t*)dctx)->expected;
}
size_t ZSTDv01_decompressContinue(ZSTDv01_Dctx* dctx, void* dst, size_t maxDstSize, const void* src, size_t srcSize)
{
dctx_t* ctx = (dctx_t*)dctx;
/* Sanity check */
if (srcSize != ctx->expected) return ERROR(srcSize_wrong);
if (dst != ctx->previousDstEnd) /* not contiguous */
ctx->base = dst;
/* Decompress : frame header */
if (ctx->phase == 0)
{
/* Check frame magic header */
U32 magicNumber = ZSTD_readBE32(src);
if (magicNumber != ZSTD_magicNumber) return ERROR(prefix_unknown);
ctx->phase = 1;
ctx->expected = ZSTD_blockHeaderSize;
return 0;
}
/* Decompress : block header */
if (ctx->phase == 1)
{
blockProperties_t bp;
size_t blockSize = ZSTDv01_getcBlockSize(src, ZSTD_blockHeaderSize, &bp);
if (ZSTDv01_isError(blockSize)) return blockSize;
if (bp.blockType == bt_end)
{
ctx->expected = 0;
ctx->phase = 0;
}
else
{
ctx->expected = blockSize;
ctx->bType = bp.blockType;
ctx->phase = 2;
}
return 0;
}
/* Decompress : block content */
{
size_t rSize;
switch(ctx->bType)
{
case bt_compressed:
rSize = ZSTD_decompressBlock(ctx, dst, maxDstSize, src, srcSize);
break;
case bt_raw :
rSize = ZSTD_copyUncompressedBlock(dst, maxDstSize, src, srcSize);
break;
case bt_rle :
return ERROR(GENERIC); /* not yet handled */
break;
case bt_end : /* should never happen (filtered at phase 1) */
rSize = 0;
break;
default:
return ERROR(GENERIC);
}
ctx->phase = 1;
ctx->expected = ZSTD_blockHeaderSize;
if (ZSTDv01_isError(rSize)) return rSize;
ctx->previousDstEnd = (void*)( ((char*)dst) + rSize);
return rSize;
}
}
/* Implementation moved to Rust (rust/src/legacy/zstd_v01.rs).
* The frozen v0.1 decoder, including its embedded FSE/Huff0 snapshot and the
* ZSTDv01_Dctx streaming state, now lives entirely in Rust; C code only ever
* holds an opaque ZSTDv01_Dctx pointer. */
+46 -14
View File
@@ -32,7 +32,8 @@ RUST_CLI_MANIFEST := $(RUST_CLI_DIR)/Cargo.toml
RUST_SOURCES := $(RUST_MANIFEST) $(RUST_DIR)/Cargo.lock \
$(shell find $(RUST_DIR)/src -type f -name '*.rs' -print)
RUST_CLI_SOURCES := $(RUST_CLI_MANIFEST) $(RUST_CLI_DIR)/Cargo.lock \
$(RUST_CLI_DIR)/src/lib.rs $(RUST_DIR)/src/zstd_cli.rs
$(RUST_CLI_DIR)/src/lib.rs $(RUST_DIR)/src/zstd_cli.rs \
$(RUST_DIR)/src/timefn.rs $(RUST_DIR)/src/benchfn.rs
# Keep Rust's HUF implementation in lockstep with libzstd.mk's C selection.
# Forced modes may arrive as libzstd.mk variables or as direct -D flags in
@@ -55,26 +56,46 @@ endif
endif
RUST_HUF_FEATURE :=
RUST_BUILD_CONFIG := default
RUST_HUF_MODE := default
ifneq ($(RUST_HUF_FORCE_X1),0)
RUST_HUF_FEATURE := huf-force-decompress-x1
RUST_BUILD_CONFIG := huf-force-decompress-x1
RUST_HUF_MODE := huf-force-decompress-x1
endif
ifneq ($(RUST_HUF_FORCE_X2),0)
RUST_HUF_FEATURE := huf-force-decompress-x2
RUST_BUILD_CONFIG := huf-force-decompress-x2
RUST_HUF_MODE := huf-force-decompress-x2
endif
# The full-featured zstd program mirrors libzstd.mk's legacy file selection:
# ZSTD_LEGACY_SUPPORT=N compiles v0.N .. v0.7, so the Rust archive enables the
# matching per-version features. The archive directory encodes the level so
# builds for different legacy levels never share cached Rust outputs. The
# compress-only, decompress-only, and CLI archives further below are used
# solely by ZSTD_LEGACY_SUPPORT=0 program variants and stay legacy-free.
empty :=
space := $(empty) $(empty)
comma := ,
RUST_LEGACY_FEATURES :=
ifneq ($(ZSTD_LEGACY_SUPPORT), 0)
ifeq ($(shell test $(ZSTD_LEGACY_SUPPORT) -lt 8; echo $$?), 0)
RUST_LEGACY_FEATURES := $(addprefix legacy-v0,$(wordlist $(ZSTD_LEGACY_SUPPORT),7,1 2 3 4 5 6 7))
endif
endif
RUST_BUILD_CONFIG := $(RUST_HUF_MODE)-legacy$(ZSTD_LEGACY_SUPPORT)
RUST_TARGET_DIR := $(RUST_DIR)/target/$(RUST_BUILD_CONFIG)
RUST_STATICLIB := $(RUST_TARGET_DIR)/release/libzstd_rs.a
RUST_TARGET_32 ?= i686-unknown-linux-gnu
RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a
RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_TARGET_DIR) --no-default-features
RUST_CARGO_FLAGS += --features compression,decompression
RUST_CARGO_FLAGS += --features compression,decompression,dict-builder
ifneq ($(RUST_HUF_FEATURE),)
RUST_CARGO_FLAGS += --features $(RUST_HUF_FEATURE)
endif
ifneq ($(RUST_LEGACY_FEATURES),)
RUST_CARGO_FLAGS += --features $(subst $(space),$(comma),$(strip $(RUST_LEGACY_FEATURES)))
endif
$(RUST_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_CARGO_FLAGS)
@@ -82,13 +103,13 @@ $(RUST_STATICLIB): $(RUST_SOURCES)
$(RUST_STATICLIB_32): $(RUST_SOURCES)
$(CARGO) build $(RUST_CARGO_FLAGS) --target $(RUST_TARGET_32)
RUST_CLI_BUILD_CONFIG := cli-c1-d1-$(RUST_BUILD_CONFIG)
RUST_CLI_BUILD_CONFIG := cli-c1-d1-$(RUST_HUF_MODE)
RUST_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_CLI_BUILD_CONFIG)
RUST_CLI_STATICLIB := $(RUST_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_CLI_STATICLIB_32 := $(RUST_CLI_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_cli_rs.a
RUST_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_CLI_TARGET_DIR) \
--no-default-features --features compression,decompression
--no-default-features --features cli,compression,decompression
$(RUST_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_CLI_CARGO_FLAGS)
@@ -96,7 +117,7 @@ $(RUST_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(RUST_CLI_STATICLIB_32): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_CLI_CARGO_FLAGS) --target $(RUST_TARGET_32)
RUST_DECOMPRESS_BUILD_CONFIG := lib-c0-d1-$(RUST_BUILD_CONFIG)
RUST_DECOMPRESS_BUILD_CONFIG := lib-c0-d1-b0-$(RUST_HUF_MODE)
RUST_DECOMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DECOMPRESS_BUILD_CONFIG)
RUST_DECOMPRESS_STATICLIB := $(RUST_DECOMPRESS_TARGET_DIR)/release/libzstd_rs.a
RUST_DECOMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
@@ -109,17 +130,17 @@ endif
$(RUST_DECOMPRESS_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_DECOMPRESS_CARGO_FLAGS)
RUST_DECOMPRESS_CLI_BUILD_CONFIG := cli-c0-d1-$(RUST_BUILD_CONFIG)
RUST_DECOMPRESS_CLI_BUILD_CONFIG := cli-c0-d1-$(RUST_HUF_MODE)
RUST_DECOMPRESS_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DECOMPRESS_CLI_BUILD_CONFIG)
RUST_DECOMPRESS_CLI_STATICLIB := $(RUST_DECOMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_DECOMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_DECOMPRESS_CLI_TARGET_DIR) \
--no-default-features --features decompression
--no-default-features --features cli,decompression
$(RUST_DECOMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_DECOMPRESS_CLI_CARGO_FLAGS)
RUST_COMPRESS_BUILD_CONFIG := lib-c1-d0-$(RUST_BUILD_CONFIG)
RUST_COMPRESS_BUILD_CONFIG := lib-c1-d0-b0-$(RUST_HUF_MODE)
RUST_COMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_BUILD_CONFIG)
RUST_COMPRESS_STATICLIB := $(RUST_COMPRESS_TARGET_DIR)/release/libzstd_rs.a
RUST_COMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
@@ -129,12 +150,23 @@ RUST_COMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
$(RUST_COMPRESS_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_COMPRESS_CARGO_FLAGS)
RUST_COMPRESS_CLI_BUILD_CONFIG := cli-c1-d0-$(RUST_BUILD_CONFIG)
RUST_DICTBUILDER_BUILD_CONFIG := lib-c1-d0-b1-$(RUST_HUF_MODE)
RUST_DICTBUILDER_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DICTBUILDER_BUILD_CONFIG)
RUST_DICTBUILDER_STATICLIB := $(RUST_DICTBUILDER_TARGET_DIR)/release/libzstd_rs.a
RUST_DICTBUILDER_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_DICTBUILDER_TARGET_DIR) \
--no-default-features \
--features compression,dict-builder
$(RUST_DICTBUILDER_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_DICTBUILDER_CARGO_FLAGS)
RUST_COMPRESS_CLI_BUILD_CONFIG := cli-c1-d0-$(RUST_HUF_MODE)
RUST_COMPRESS_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_CLI_BUILD_CONFIG)
RUST_COMPRESS_CLI_STATICLIB := $(RUST_COMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_COMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_COMPRESS_CLI_TARGET_DIR) \
--no-default-features --features compression
--no-default-features --features cli,compression
$(RUST_COMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_COMPRESS_CLI_CARGO_FLAGS)
@@ -412,7 +444,7 @@ zstd-compress: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) zstdcli.c util.c ti
## zstd-dictBuilder: executable supporting dictionary creation and compression (only)
CLEAN += zstd-dictBuilder
zstd-dictBuilder: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) $(ZDICT_SRC) zstdcli.c util.c timefn.c fileio.c fileio_asyncio.c dibio.c $(RUST_COMPRESS_STATICLIB) $(RUST_COMPRESS_CLI_STATICLIB)
zstd-dictBuilder: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) $(ZDICT_SRC) zstdcli.c util.c timefn.c fileio.c fileio_asyncio.c dibio.c $(RUST_DICTBUILDER_STATICLIB) $(RUST_COMPRESS_CLI_STATICLIB)
$(CC) $(FLAGS) -DZSTD_NOBENCH -DZSTD_NODECOMPRESS -DZSTD_NOTRACE $^ -o $@$(EXT)
RUST_DIRECT_LINK_TARGETS := zstd32 zstd-nolegacy zstd-small zstd-frugal \
+17 -243
View File
@@ -8,249 +8,23 @@
* You may select, at your option, one of the above-listed licenses.
*/
/* The implementation lives in rust/src/benchfn.rs, built into the Rust CLI
* static archive. This translation unit stays in the original source lists so
* build configuration keeps working while the implementation is in Rust. */
/* *************************************
* Includes
***************************************/
#include <stdlib.h> /* malloc, free */
#include <string.h> /* memset */
#include <assert.h> /* assert */
#include "timefn.h" /* UTIL_time_t, UTIL_getTime */
#include <stddef.h> /* size_t, offsetof */
#include "benchfn.h"
/* *************************************
* Constants
***************************************/
#define TIMELOOP_MICROSEC SEC_TO_MICRO /* 1 second */
#define TIMELOOP_NANOSEC (1*1000000000ULL) /* 1 second */
#define KB *(1 <<10)
#define MB *(1 <<20)
#define GB *(1U<<30)
/* *************************************
* Debug errors
***************************************/
#if defined(DEBUG) && (DEBUG >= 1)
# include <stdio.h> /* fprintf */
# define DISPLAY(...) fprintf(stderr, __VA_ARGS__)
# define DEBUGOUTPUT(...) { if (DEBUG) DISPLAY(__VA_ARGS__); }
#else
# define DEBUGOUTPUT(...)
#endif
/* error without displaying */
#define RETURN_QUIET_ERROR(retValue, ...) { \
DEBUGOUTPUT("%s: %i: \n", __FILE__, __LINE__); \
DEBUGOUTPUT("Error : "); \
DEBUGOUTPUT(__VA_ARGS__); \
DEBUGOUTPUT(" \n"); \
return retValue; \
}
/* Abort execution if a condition is not met */
#define CONTROL(c) { if (!(c)) { DEBUGOUTPUT("error: %s \n", #c); abort(); } }
/* *************************************
* Benchmarking an arbitrary function
***************************************/
int BMK_isSuccessful_runOutcome(BMK_runOutcome_t outcome)
{
return outcome.error_tag_never_ever_use_directly == 0;
}
/* warning : this function will stop program execution if outcome is invalid !
* check outcome validity first, using BMK_isValid_runResult() */
BMK_runTime_t BMK_extract_runTime(BMK_runOutcome_t outcome)
{
CONTROL(outcome.error_tag_never_ever_use_directly == 0);
return outcome.internal_never_ever_use_directly;
}
size_t BMK_extract_errorResult(BMK_runOutcome_t outcome)
{
CONTROL(outcome.error_tag_never_ever_use_directly != 0);
return outcome.error_result_never_ever_use_directly;
}
static BMK_runOutcome_t BMK_runOutcome_error(size_t errorResult)
{
BMK_runOutcome_t b;
memset(&b, 0, sizeof(b));
b.error_tag_never_ever_use_directly = 1;
b.error_result_never_ever_use_directly = errorResult;
return b;
}
static BMK_runOutcome_t BMK_setValid_runTime(BMK_runTime_t runTime)
{
BMK_runOutcome_t outcome;
outcome.error_tag_never_ever_use_directly = 0;
outcome.internal_never_ever_use_directly = runTime;
return outcome;
}
/* initFn will be measured once, benchFn will be measured `nbLoops` times */
/* initFn is optional, provide NULL if none */
/* benchFn must return a size_t value that errorFn can interpret */
/* takes # of blocks and list of size & stuff for each. */
/* can report result of benchFn for each block into blockResult. */
/* blockResult is optional, provide NULL if this information is not required */
/* note : time per loop can be reported as zero if run time < timer resolution */
BMK_runOutcome_t BMK_benchFunction(BMK_benchParams_t p,
unsigned nbLoops)
{
nbLoops += !nbLoops; /* minimum nbLoops is 1 */
/* init */
{ size_t i;
for(i = 0; i < p.blockCount; i++) {
memset(p.dstBuffers[i], 0xE5, p.dstCapacities[i]); /* warm up and erase result buffer */
} }
/* benchmark */
{ size_t dstSize = 0;
UTIL_time_t const clockStart = UTIL_getTime();
unsigned loopNb, blockNb;
if (p.initFn != NULL) p.initFn(p.initPayload);
for (loopNb = 0; loopNb < nbLoops; loopNb++) {
for (blockNb = 0; blockNb < p.blockCount; blockNb++) {
size_t const res = p.benchFn(p.srcBuffers[blockNb], p.srcSizes[blockNb],
p.dstBuffers[blockNb], p.dstCapacities[blockNb],
p.benchPayload);
if (loopNb == 0) {
if (p.blockResults != NULL) p.blockResults[blockNb] = res;
if ((p.errorFn != NULL) && (p.errorFn(res))) {
RETURN_QUIET_ERROR(BMK_runOutcome_error(res),
"Function benchmark failed on block %u (of size %u) with error %i",
blockNb, (unsigned)p.srcSizes[blockNb], (int)res);
}
dstSize += res;
} }
} /* for (loopNb = 0; loopNb < nbLoops; loopNb++) */
{ PTime const totalTime = UTIL_clockSpanNano(clockStart);
BMK_runTime_t rt;
rt.nanoSecPerRun = (double)totalTime / nbLoops;
rt.sumOfReturn = dstSize;
return BMK_setValid_runTime(rt);
} }
}
/* ==== Benchmarking any function, providing intermediate results ==== */
struct BMK_timedFnState_s {
PTime timeSpent_ns;
PTime timeBudget_ns;
PTime runBudget_ns;
BMK_runTime_t fastestRun;
unsigned nbLoops;
UTIL_time_t coolTime;
}; /* typedef'd to BMK_timedFnState_t within bench.h */
BMK_timedFnState_t* BMK_createTimedFnState(unsigned total_ms, unsigned run_ms)
{
BMK_timedFnState_t* const r = (BMK_timedFnState_t*)malloc(sizeof(*r));
if (r == NULL) return NULL; /* malloc() error */
BMK_resetTimedFnState(r, total_ms, run_ms);
return r;
}
void BMK_freeTimedFnState(BMK_timedFnState_t* state) { free(state); }
BMK_timedFnState_t*
BMK_initStatic_timedFnState(void* buffer, size_t size, unsigned total_ms, unsigned run_ms)
{
typedef char check_size[ 2 * (sizeof(BMK_timedFnState_shell) >= sizeof(struct BMK_timedFnState_s)) - 1]; /* static assert : a compilation failure indicates that BMK_timedFnState_shell is not large enough */
typedef struct { check_size c; BMK_timedFnState_t tfs; } tfs_align; /* force tfs to be aligned at its next best position */
size_t const tfs_alignment = offsetof(tfs_align, tfs); /* provides the minimal alignment restriction for BMK_timedFnState_t */
BMK_timedFnState_t* const r = (BMK_timedFnState_t*)buffer;
if (buffer == NULL) return NULL;
if (size < sizeof(struct BMK_timedFnState_s)) return NULL;
if ((size_t)buffer % tfs_alignment) return NULL; /* buffer must be properly aligned */
BMK_resetTimedFnState(r, total_ms, run_ms);
return r;
}
void BMK_resetTimedFnState(BMK_timedFnState_t* timedFnState, unsigned total_ms, unsigned run_ms)
{
if (!total_ms) total_ms = 1 ;
if (!run_ms) run_ms = 1;
if (run_ms > total_ms) run_ms = total_ms;
timedFnState->timeSpent_ns = 0;
timedFnState->timeBudget_ns = (PTime)total_ms * TIMELOOP_NANOSEC / 1000;
timedFnState->runBudget_ns = (PTime)run_ms * TIMELOOP_NANOSEC / 1000;
timedFnState->fastestRun.nanoSecPerRun = (double)TIMELOOP_NANOSEC * 2000000000; /* hopefully large enough : must be larger than any potential measurement */
timedFnState->fastestRun.sumOfReturn = (size_t)(-1LL);
timedFnState->nbLoops = 1;
timedFnState->coolTime = UTIL_getTime();
}
/* Tells if nb of seconds set in timedFnState for all runs is spent.
* note : this function will return 1 if BMK_benchFunctionTimed() has actually errored. */
int BMK_isCompleted_TimedFn(const BMK_timedFnState_t* timedFnState)
{
return (timedFnState->timeSpent_ns >= timedFnState->timeBudget_ns);
}
#undef MIN
#define MIN(a,b) ( (a) < (b) ? (a) : (b) )
#define MINUSABLETIME (TIMELOOP_NANOSEC / 2) /* 0.5 seconds */
BMK_runOutcome_t BMK_benchTimedFn(BMK_timedFnState_t* cont,
BMK_benchParams_t p)
{
PTime const runBudget_ns = cont->runBudget_ns;
PTime const runTimeMin_ns = runBudget_ns / 2;
int completed = 0;
BMK_runTime_t bestRunTime = cont->fastestRun;
while (!completed) {
BMK_runOutcome_t const runResult = BMK_benchFunction(p, cont->nbLoops);
if(!BMK_isSuccessful_runOutcome(runResult)) { /* error : move out */
return runResult;
}
{ BMK_runTime_t const newRunTime = BMK_extract_runTime(runResult);
double const loopDuration_ns = newRunTime.nanoSecPerRun * cont->nbLoops;
cont->timeSpent_ns += (unsigned long long)loopDuration_ns;
/* estimate nbLoops for next run to last approximately 1 second */
if (loopDuration_ns > ((double)runBudget_ns / 50)) {
double const fastestRun_ns = MIN(bestRunTime.nanoSecPerRun, newRunTime.nanoSecPerRun);
cont->nbLoops = (unsigned)((double)runBudget_ns / fastestRun_ns) + 1;
} else {
/* previous run was too short : blindly increase workload by x multiplier */
const unsigned multiplier = 10;
assert(cont->nbLoops < ((unsigned)-1) / multiplier); /* avoid overflow */
cont->nbLoops *= multiplier;
}
if(loopDuration_ns < (double)runTimeMin_ns) {
/* don't report results for which benchmark run time was too small : increased risks of rounding errors */
assert(completed == 0);
continue;
} else {
if(newRunTime.nanoSecPerRun < bestRunTime.nanoSecPerRun) {
bestRunTime = newRunTime;
}
completed = 1;
}
}
} /* while (!completed) */
return BMK_setValid_runTime(bestRunTime);
}
/* BMK_runTime_t and BMK_runOutcome_t are returned by value across the C/Rust
* boundary, and BMK_benchParams_t is passed by value. The Rust #[repr(C)]
* definitions mirror the offsets pinned here. */
typedef char BMK_staticAssert_runTimeSumOffset[
(offsetof(BMK_runTime_t, sumOfReturn) == sizeof(double)) ? 1 : -1];
typedef char BMK_staticAssert_outcomeResultOffset[
(offsetof(BMK_runOutcome_t, error_result_never_ever_use_directly)
== sizeof(BMK_runTime_t)) ? 1 : -1];
typedef char BMK_staticAssert_outcomeTagOffset[
(offsetof(BMK_runOutcome_t, error_tag_never_ever_use_directly)
== sizeof(BMK_runTime_t) + sizeof(size_t)) ? 1 : -1];
typedef char BMK_staticAssert_shellAlignment[
(sizeof(BMK_timedFnState_shell) == BMK_TIMEDFNSTATE_SIZE) ? 1 : -1];
+8 -155
View File
@@ -8,161 +8,14 @@
* You may select, at your option, one of the above-listed licenses.
*/
/* === Dependencies === */
/* The implementation lives in rust/src/timefn.rs, built into the Rust CLI
* static archive. This translation unit stays in the original source lists so
* build configuration keeps working while the implementation is in Rust. */
#include "timefn.h"
#include "platform.h" /* set _POSIX_C_SOURCE */
#include <time.h> /* CLOCK_MONOTONIC, TIME_UTC */
/*-****************************************
* Time functions
******************************************/
#if defined(_WIN32) /* Windows */
#include <windows.h> /* LARGE_INTEGER */
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
static LARGE_INTEGER ticksPerSecond;
static int init = 0;
if (!init) {
if (!QueryPerformanceFrequency(&ticksPerSecond)) {
perror("timefn::QueryPerformanceFrequency");
abort();
}
init = 1;
}
{ UTIL_time_t r;
LARGE_INTEGER x;
QueryPerformanceCounter(&x);
r.t = (PTime)(x.QuadPart * 1000000000ULL / ticksPerSecond.QuadPart);
return r;
}
}
#elif defined(__APPLE__) && defined(__MACH__)
#include <mach/mach_time.h> /* mach_timebase_info_data_t, mach_timebase_info, mach_absolute_time */
UTIL_time_t UTIL_getTime(void)
{
static mach_timebase_info_data_t rate;
static int init = 0;
if (!init) {
mach_timebase_info(&rate);
init = 1;
}
{ UTIL_time_t r;
r.t = mach_absolute_time() * (PTime)rate.numer / (PTime)rate.denom;
return r;
}
}
/* POSIX.1-2001 (optional) */
#elif defined(CLOCK_MONOTONIC)
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
/* time must be initialized, othersize it may fail msan test.
* No good reason, likely a limitation of timespec_get() for some target */
struct timespec time = { 0, 0 };
if (clock_gettime(CLOCK_MONOTONIC, &time) != 0) {
perror("timefn::clock_gettime(CLOCK_MONOTONIC)");
abort();
}
{ UTIL_time_t r;
r.t = (PTime)time.tv_sec * 1000000000ULL + (PTime)time.tv_nsec;
return r;
}
}
/* C11 requires support of timespec_get().
* However, FreeBSD 11 claims C11 compliance while lacking timespec_get().
* Double confirm timespec_get() support by checking the definition of TIME_UTC.
* However, some versions of Android manage to simultaneously define TIME_UTC
* and lack timespec_get() support... */
#elif (defined (__STDC_VERSION__) && (__STDC_VERSION__ >= 201112L) /* C11 */) \
&& defined(TIME_UTC) && !defined(__ANDROID__)
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
/* time must be initialized, othersize it may fail msan test.
* No good reason, likely a limitation of timespec_get() for some target */
struct timespec time = { 0, 0 };
if (timespec_get(&time, TIME_UTC) != TIME_UTC) {
perror("timefn::timespec_get(TIME_UTC)");
abort();
}
{ UTIL_time_t r;
r.t = (PTime)time.tv_sec * 1000000000ULL + (PTime)time.tv_nsec;
return r;
}
}
#else /* relies on standard C90 (note : clock_t produces wrong measurements for multi-threaded workloads) */
UTIL_time_t UTIL_getTime(void)
{
UTIL_time_t r;
r.t = (PTime)clock() * 1000000000ULL / CLOCKS_PER_SEC;
return r;
}
#define TIME_MT_MEASUREMENTS_NOT_SUPPORTED
#endif
/* ==== Common functions, valid for all time API ==== */
PTime UTIL_getSpanTimeNano(UTIL_time_t clockStart, UTIL_time_t clockEnd)
{
return clockEnd.t - clockStart.t;
}
PTime UTIL_getSpanTimeMicro(UTIL_time_t begin, UTIL_time_t end)
{
return UTIL_getSpanTimeNano(begin, end) / 1000ULL;
}
PTime UTIL_clockSpanMicro(UTIL_time_t clockStart )
{
UTIL_time_t const clockEnd = UTIL_getTime();
return UTIL_getSpanTimeMicro(clockStart, clockEnd);
}
PTime UTIL_clockSpanNano(UTIL_time_t clockStart )
{
UTIL_time_t const clockEnd = UTIL_getTime();
return UTIL_getSpanTimeNano(clockStart, clockEnd);
}
void UTIL_waitForNextTick(void)
{
UTIL_time_t const clockStart = UTIL_getTime();
UTIL_time_t clockEnd;
do {
clockEnd = UTIL_getTime();
} while (UTIL_getSpanTimeNano(clockStart, clockEnd) == 0);
}
int UTIL_support_MT_measurements(void)
{
# if defined(TIME_MT_MEASUREMENTS_NOT_SUPPORTED)
return 0;
# else
return 1;
# endif
}
/* The Rust port mirrors this exact ABI: UTIL_time_t is returned by value and
* must remain a plain 64-bit nanosecond counter. */
typedef char UTIL_staticAssert_ptimeIs64Bit[(sizeof(PTime) == 8) ? 1 : -1];
typedef char UTIL_staticAssert_timeIsPlainCounter[
(sizeof(UTIL_time_t) == sizeof(PTime)) ? 1 : -1];
+51
View File
@@ -10,16 +10,67 @@
/* The CLI parser and control flow live in rust/src/zstd_cli.rs. Keep this
* translation unit as the stable C entry point used by program launchers. */
#include <stddef.h> /* size_t */
#define ZSTD_STATIC_LINKING_ONLY /* ZSTD_compressionParameters */
#include "../lib/zstd.h"
#ifndef ZSTD_NOBENCH
# include "benchzstd.h" /* BMK_benchFilesAdvanced, BMK_syntheticTest */
#endif
int ZSTD_rust_cli_main(int argCount, const char* const argv[]);
const char* ZSTD_rust_cli_expected_version(void);
int ZSTD_rust_cli_bench(const char* const* fileNames, unsigned nbFiles,
const char* dictFileName,
int startCLevel, int endCLevel,
const ZSTD_compressionParameters* compressionParams,
int displayLevel, unsigned nbSeconds,
size_t blockSize, int nbWorkers);
const char* ZSTD_rust_cli_expected_version(void)
{
return ZSTD_VERSION_STRING;
}
/* Benchmark bridge for the Rust CLI. Whether benchmarking exists is a C
* preprocessor property (ZSTD_NOBENCH), so the decision stays in this shim:
* the Rust frontend calls in unconditionally, and stripped program variants
* never reference benchmark symbols.
* @return the benchmark result code (>= 0), or -1 when unavailable. */
int ZSTD_rust_cli_bench(const char* const* fileNames, unsigned nbFiles,
const char* dictFileName,
int startCLevel, int endCLevel,
const ZSTD_compressionParameters* compressionParams,
int displayLevel, unsigned nbSeconds,
size_t blockSize, int nbWorkers)
{
#ifndef ZSTD_NOBENCH
BMK_advancedParams_t advancedParams = BMK_initAdvancedParams();
int startLevel = startCLevel;
int endLevel = endCLevel;
advancedParams.nbSeconds = nbSeconds;
advancedParams.blockSize = blockSize;
advancedParams.nbWorkers = nbWorkers;
if (startLevel > ZSTD_maxCLevel()) startLevel = ZSTD_maxCLevel();
if (endLevel > ZSTD_maxCLevel()) endLevel = ZSTD_maxCLevel();
if (endLevel < startLevel) endLevel = startLevel;
if (nbFiles == 0) {
/* No input file: benchmark a synthetic sample (lorem generator). */
return BMK_syntheticTest(-1.0, startLevel, endLevel,
compressionParams, displayLevel,
&advancedParams);
}
return BMK_benchFilesAdvanced(fileNames, nbFiles, dictFileName,
startLevel, endLevel,
compressionParams, displayLevel,
&advancedParams);
#else
(void)fileNames; (void)nbFiles; (void)dictFileName;
(void)startCLevel; (void)endCLevel; (void)compressionParams;
(void)displayLevel; (void)nbSeconds; (void)blockSize; (void)nbWorkers;
return -1;
#endif
}
int main(int argCount, const char* argv[])
{
return ZSTD_rust_cli_main(argCount, argv);
+14 -1
View File
@@ -7,11 +7,24 @@ edition = "2021"
crate-type = ["staticlib"]
[features]
default = ["compression", "decompression"]
default = ["compression", "decompression", "dict-builder"]
compression = []
decompression = []
dict-builder = []
huf-force-decompress-x1 = []
huf-force-decompress-x2 = []
# Legacy-format decoders (zstd v0.1 .. v0.7). Never default features: the
# build systems map ZSTD_LEGACY_SUPPORT=N to the features for versions >= N,
# exactly mirroring which lib/legacy/zstd_v0N.c files the C build compiles.
# A feature whose version is not yet ported to Rust gates nothing; the
# original C file still provides that decoder.
legacy-v01 = []
legacy-v02 = []
legacy-v03 = []
legacy-v04 = []
legacy-v05 = []
legacy-v06 = []
legacy-v07 = []
[dependencies]
libc = "0.2"
+65 -8
View File
@@ -48,11 +48,19 @@ zstd ABI:
the dynamic-programming optimal parser itself remains in C for now.
- `zstd_ldm` implements long-distance-match parameter selection, table
maintenance, sequence generation, and sequence consumption.
- Dictionary building
- `divsufsort` constructs the suffix array that drives the legacy `ZDICT`
trainer (`ZDICT_trainFromBuffer_legacy`). The sample analysis and
dictionary assembly in `zdict.c`, `cover.c`, and `fastcover.c` remain C.
- Runtime support
- `threading` provides platform pthread wrappers required by zstd headers.
- `pool` implements the bounded worker pool used by multithreaded compression.
- Dictionary support
- `zstd_ddict` owns, loads, copies, and references decode dictionaries.
- Legacy decoding
- `legacy` hosts one frozen module per historical format; `legacy::zstd_v01`
ports the self-contained v0.1 decoder. Versions v0.2 through v0.7 are
still C.
- Block decompression
- `zstd_decompress_block` decodes literal and sequence sections, maintains
FSE/Huffman repeat state, and executes compressed-block sequences.
@@ -65,11 +73,58 @@ zstd ABI:
so library builds do not acquire program-only dependencies. The C
`fileio` backend still owns file opening, safe replacement, sparse writes,
metadata, and streaming I/O.
- `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`,
and `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`,
`BMK_benchTimedFn`) used by the CLI benchmark mode and by C test tools.
Both live in the `cli/` package, but C test binaries (fullbench, fuzzer,
zstreamtest, paramgrill, ...) link a helpers-only build of that archive,
produced without the package's `cli` feature, because the parser layer
requires the C `fileio` backend that tests do not compile. Benchmark
orchestration and reporting (`benchzstd.c`) remain C, reached from the
Rust parser through the `ZSTD_NOBENCH`-gated bridge in `zstdcli.c`.
The optimal block matcher, high-level frame compression, dictionary-building,
legacy decoding callbacks, and the CLI file-I/O backend are still C. They must
move before the rewrite is complete. Keeping that boundary explicit prevents a
passing hybrid build from being mistaken for the final all-Rust result.
The optimal block matcher, high-level frame compression, dictionary-building
except suffix-array construction, the legacy v0.2-v0.7 decoders, benchmark
orchestration (`benchzstd`), and the CLI file-I/O backend are still C. They
must move before the rewrite is complete. Keeping that boundary explicit
prevents a passing hybrid build from being mistaken for the final all-Rust
result.
## Legacy decoding
Each `lib/legacy/zstd_v0N.c` file is a frozen snapshot of the entropy coders
and frame logic of one historical release. The Rust ports in `src/legacy/`
keep that property: every version owns its own frozen FSE/Huff0 and frame
logic, ported line by line, and must never reuse the modern entropy modules
or share code with other legacy versions. Outputs and error codes must be
byte-identical to the original C files. Their only shared dependency is the
`errors` module, matching the C files' `error_private.h` include.
Cargo features `legacy-v01` .. `legacy-v07` gate the per-version modules and
are never default features. The build systems derive the feature list from
the C configuration:
- `lib/Makefile` and `programs/Makefile` map `ZSTD_LEGACY_SUPPORT=N` to the
features for versions >= N (0 disables legacy), matching the
`ZSTD_LEGACY_FILES` selection in `lib/libzstd.mk`.
- `tests/Makefile` always enables all seven features because the test
objects compile every `lib/legacy/*.c` file regardless of dispatch level.
- `build/meson` maps `legacy_level` like the makefiles; `build/cmake`
enables all seven whenever `ZSTD_LEGACY_SUPPORT` is on because it always
compiles all seven C files.
Every build system also encodes the legacy selection in the Rust target
directory name (for example `c1-d1-default-legacy5`), for the same reason the
HUF mode is encoded there: a cached archive built for one configuration must
never be linked into a build that expects another.
A feature whose version has not been ported yet gates nothing; the original
C file still provides that decoder, so mixed C/Rust legacy levels link
cleanly. Porting a version means adding `src/legacy/zstd_v0N.rs`, registering
it in `src/legacy/mod.rs` behind its feature, and reducing
`lib/legacy/zstd_v0N.c` to a declaration-only shim. For v0.1 the streaming
`ZSTDv01_Dctx` state lives entirely in Rust: C code only ever holds an opaque
pointer, so the C-side struct definition is gone.
## Compatibility boundary
@@ -80,8 +135,9 @@ makefile source list as a small shim so header configuration and platform
preprocessor behavior stay available during the transition.
The library, test, and program makefiles select an archive directory for the
active C configuration: enabled compression/decompression modules, default or
forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The
active C configuration: enabled compression/decompression/dictionary-builder
modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit
C binaries. The
native static archive flattens Rust object members rather than nesting a Rust
archive, while the native shared library retains all migrated Rust exports.
When the HUF mode changes, the test and program paths also rebuild cached C
@@ -105,8 +161,9 @@ from `rust/cli` as well:
```sh
cargo clippy --all-targets -- -D warnings
cargo test --all-targets
cargo test --no-default-features --features compression --all-targets
cargo test --no-default-features --features decompression --all-targets
cargo test --no-default-features --features cli,compression --all-targets
cargo test --no-default-features --features cli,decompression --all-targets
cargo test --no-default-features --all-targets
```
Then run original compatibility tests from the repository root, starting with
+9
View File
@@ -2,6 +2,15 @@
# It is not intended for manual editing.
version = 4
[[package]]
name = "libc"
version = "0.2.186"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "68ab91017fe16c622486840e4c83c9a37afeff978bd239b5293d61ece587de66"
[[package]]
name = "zstd-cli-rs"
version = "0.1.0"
dependencies = [
"libc",
]
+8 -1
View File
@@ -7,6 +7,13 @@ edition = "2021"
crate-type = ["staticlib"]
[features]
default = ["compression", "decompression"]
default = ["cli", "compression", "decompression"]
# The command-line parser and dispatch layer, which requires the C fileio
# backend at link time. Program archives enable it; C test binaries link a
# helpers-only archive (timefn) built without it.
cli = []
compression = []
decompression = []
[dependencies]
libc = "0.2"
+5
View File
@@ -1,2 +1,7 @@
#[path = "../../src/benchfn.rs"]
mod benchfn;
#[path = "../../src/timefn.rs"]
mod timefn;
#[cfg(feature = "cli")]
#[path = "../../src/zstd_cli.rs"]
mod zstd_cli;
+575
View File
@@ -0,0 +1,575 @@
#![allow(non_camel_case_types)]
#![allow(non_snake_case)]
#![allow(clippy::missing_safety_doc)]
//! Benchmark loop for arbitrary functions over a set of blocks.
//!
//! Port of `programs/benchfn.c`. `BMK_benchFunction` measures one batch of
//! runs; `BMK_benchTimedFn` repeats batches, growing the loop count until a
//! run lasts long enough to be reported reliably against `run_ms`, within a
//! `total_ms` budget tracked by `BMK_timedFnState_t`.
//!
//! ABI notes: `BMK_runOutcome_t` and `BMK_runTime_t` are returned by value
//! across the C boundary and `BMK_benchParams_t` is passed by value, so all
//! three are `repr(C)` mirrors of the `benchfn.h` layout, pinned by asserts
//! here and in the C shim. `BMK_timedFnState_t` is opaque to C, but
//! `BMK_initStatic_timedFnState` guarantees it fits the 64-byte
//! `BMK_timedFnState_shell`, and `BMK_createTimedFnState` uses `malloc` so
//! creation and destruction stay interchangeable with C callers.
use std::os::raw::{c_int, c_uint, c_void};
use std::ptr;
use crate::timefn::{PTime, UTIL_clockSpanNano, UTIL_getTime, UTIL_time_t};
const TIMELOOP_NANOSEC: PTime = 1_000_000_000;
/// Valid benchmark result (`BMK_runTime_t` in benchfn.h).
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct BMK_runTime_t {
/// Time per iteration, over all blocks.
pub nanoSecPerRun: f64,
/// Sum of the benchmarked function's return values, first loop only.
pub sumOfReturn: usize,
}
/// Outcome variant of a benchmark run (`BMK_runOutcome_t` in benchfn.h):
/// either a valid `BMK_runTime_t` or an error result. C callers treat it as
/// opaque and use the accessor functions below.
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct BMK_runOutcome_t {
pub internal_never_ever_use_directly: BMK_runTime_t,
pub error_result_never_ever_use_directly: usize,
pub error_tag_never_ever_use_directly: c_int,
}
// These mirror the static asserts in the programs/benchfn.c shim: the structs
// cross the ABI by value, so field offsets must match the C header exactly.
const _: () =
assert!(std::mem::offset_of!(BMK_runTime_t, sumOfReturn) == std::mem::size_of::<f64>());
const _: () = assert!(
std::mem::offset_of!(BMK_runOutcome_t, error_result_never_ever_use_directly)
== std::mem::size_of::<BMK_runTime_t>()
);
const _: () = assert!(
std::mem::offset_of!(BMK_runOutcome_t, error_tag_never_ever_use_directly)
== std::mem::size_of::<BMK_runTime_t>() + std::mem::size_of::<usize>()
);
/// `size_t (*BMK_benchFn_t)(const void*, size_t, void*, size_t, void*)`
pub type BMK_benchFn_t = Option<
unsafe extern "C" fn(
src: *const c_void,
srcSize: usize,
dst: *mut c_void,
dstCapacity: usize,
customPayload: *mut c_void,
) -> usize,
>;
/// `size_t (*BMK_initFn_t)(void*)`
pub type BMK_initFn_t = Option<unsafe extern "C" fn(initPayload: *mut c_void) -> usize>;
/// `unsigned (*BMK_errorFn_t)(size_t)`
pub type BMK_errorFn_t = Option<unsafe extern "C" fn(result: usize) -> c_uint>;
/// Parameters of `BMK_benchFunction`, passed by value (`BMK_benchParams_t`).
#[repr(C)]
#[derive(Clone, Copy)]
pub struct BMK_benchParams_t {
pub benchFn: BMK_benchFn_t,
pub benchPayload: *mut c_void,
pub initFn: BMK_initFn_t,
pub initPayload: *mut c_void,
pub errorFn: BMK_errorFn_t,
pub blockCount: usize,
pub srcBuffers: *const *const c_void,
pub srcSizes: *const usize,
pub dstBuffers: *const *mut c_void,
pub dstCapacities: *const usize,
pub blockResults: *mut usize,
}
/// Aborts, like benchfn.c's `CONTROL`, when an accessor is used on the wrong
/// outcome variant.
fn control(condition: bool) {
if !condition {
std::process::abort();
}
}
fn error_outcome(errorResult: usize) -> BMK_runOutcome_t {
BMK_runOutcome_t {
internal_never_ever_use_directly: BMK_runTime_t {
nanoSecPerRun: 0.0,
sumOfReturn: 0,
},
error_result_never_ever_use_directly: errorResult,
error_tag_never_ever_use_directly: 1,
}
}
fn valid_outcome(runTime: BMK_runTime_t) -> BMK_runOutcome_t {
BMK_runOutcome_t {
internal_never_ever_use_directly: runTime,
error_result_never_ever_use_directly: 0,
error_tag_never_ever_use_directly: 0,
}
}
/// Tells if the outcome carries a valid measurement.
#[no_mangle]
pub extern "C" fn BMK_isSuccessful_runOutcome(outcome: BMK_runOutcome_t) -> c_int {
c_int::from(outcome.error_tag_never_ever_use_directly == 0)
}
/// Extracts the measurement; aborts if the outcome is an error, so validity
/// must be checked first with `BMK_isSuccessful_runOutcome`.
#[no_mangle]
pub extern "C" fn BMK_extract_runTime(outcome: BMK_runOutcome_t) -> BMK_runTime_t {
control(outcome.error_tag_never_ever_use_directly == 0);
outcome.internal_never_ever_use_directly
}
/// Extracts the faulty `benchFn` return value; aborts if the outcome is
/// valid, so failure must be checked first.
#[no_mangle]
pub extern "C" fn BMK_extract_errorResult(outcome: BMK_runOutcome_t) -> usize {
control(outcome.error_tag_never_ever_use_directly != 0);
outcome.error_result_never_ever_use_directly
}
/// Runs `initFn` once, then `benchFn` `nbLoops` times over every block, and
/// reports the mean time per loop. On the first loop, per-block results are
/// stored into `blockResults` (when provided) and checked with `errorFn`
/// (when provided); the first failing block aborts the measurement and
/// produces an error outcome carrying the faulty return value.
#[no_mangle]
pub unsafe extern "C" fn BMK_benchFunction(
p: BMK_benchParams_t,
mut nbLoops: c_uint,
) -> BMK_runOutcome_t {
// Minimum nbLoops is 1.
nbLoops += c_uint::from(nbLoops == 0);
// Warm up and erase the result buffers.
for blockNb in 0..p.blockCount {
unsafe {
let dst = *p.dstBuffers.add(blockNb);
ptr::write_bytes(dst.cast::<u8>(), 0xE5, *p.dstCapacities.add(blockNb));
}
}
let benchFn = p.benchFn.expect("benchFn is mandatory");
let mut dstSize = 0usize;
let clockStart = UTIL_getTime();
if let Some(initFn) = p.initFn {
unsafe { initFn(p.initPayload) };
}
for loopNb in 0..nbLoops {
for blockNb in 0..p.blockCount {
let res = unsafe {
benchFn(
*p.srcBuffers.add(blockNb),
*p.srcSizes.add(blockNb),
*p.dstBuffers.add(blockNb),
*p.dstCapacities.add(blockNb),
p.benchPayload,
)
};
if loopNb == 0 {
if !p.blockResults.is_null() {
unsafe { *p.blockResults.add(blockNb) = res };
}
if let Some(errorFn) = p.errorFn {
if unsafe { errorFn(res) } != 0 {
return error_outcome(res);
}
}
dstSize = dstSize.wrapping_add(res);
}
}
}
let totalTime = UTIL_clockSpanNano(clockStart);
valid_outcome(BMK_runTime_t {
nanoSecPerRun: totalTime as f64 / f64::from(nbLoops),
sumOfReturn: dstSize,
})
}
/// Benchmark session state (`struct BMK_timedFnState_s`), opaque to C.
#[repr(C)]
pub struct BMK_timedFnState_t {
timeSpent_ns: PTime,
timeBudget_ns: PTime,
runBudget_ns: PTime,
fastestRun: BMK_runTime_t,
nbLoops: c_uint,
coolTime: UTIL_time_t,
}
/// `BMK_TIMEDFNSTATE_SIZE` in benchfn.h: capacity of the caller-provided
/// `BMK_timedFnState_shell`, which the state must always fit.
const BMK_TIMEDFNSTATE_SIZE: usize = 64;
const _: () = assert!(std::mem::size_of::<BMK_timedFnState_t>() <= BMK_TIMEDFNSTATE_SIZE);
// The shell aligns via a `long long` member; the state must not need more.
const _: () = assert!(std::mem::align_of::<BMK_timedFnState_t>() <= std::mem::align_of::<u64>());
/// Allocates and initializes a benchmark session lasting a minimum of
/// `total_ms`, paced at intervals of approximately `run_ms`. Uses `malloc`
/// so ownership stays interchangeable with the original C implementation.
#[no_mangle]
pub extern "C" fn BMK_createTimedFnState(
total_ms: c_uint,
run_ms: c_uint,
) -> *mut BMK_timedFnState_t {
let state = unsafe { libc::malloc(std::mem::size_of::<BMK_timedFnState_t>()) }
.cast::<BMK_timedFnState_t>();
if state.is_null() {
return ptr::null_mut();
}
unsafe { BMK_resetTimedFnState(state, total_ms, run_ms) };
state
}
/// Releases a state obtained from `BMK_createTimedFnState`.
#[no_mangle]
pub unsafe extern "C" fn BMK_freeTimedFnState(state: *mut BMK_timedFnState_t) {
unsafe { libc::free(state.cast()) };
}
/// Places the session state into a caller-provided buffer, typically a
/// `BMK_timedFnState_shell`. Returns NULL when the buffer is missing, too
/// small, or misaligned.
#[no_mangle]
pub unsafe extern "C" fn BMK_initStatic_timedFnState(
buffer: *mut c_void,
size: usize,
total_ms: c_uint,
run_ms: c_uint,
) -> *mut BMK_timedFnState_t {
if buffer.is_null() {
return ptr::null_mut();
}
if size < std::mem::size_of::<BMK_timedFnState_t>() {
return ptr::null_mut();
}
if !(buffer as usize).is_multiple_of(std::mem::align_of::<BMK_timedFnState_t>()) {
return ptr::null_mut();
}
let state = buffer.cast::<BMK_timedFnState_t>();
unsafe { BMK_resetTimedFnState(state, total_ms, run_ms) };
state
}
/// Re-arms a session for a new benchmark of `total_ms`, paced at `run_ms`.
#[no_mangle]
pub unsafe extern "C" fn BMK_resetTimedFnState(
timedFnState: *mut BMK_timedFnState_t,
total_ms: c_uint,
run_ms: c_uint,
) {
let total_ms = if total_ms == 0 { 1 } else { total_ms };
let mut run_ms = if run_ms == 0 { 1 } else { run_ms };
if run_ms > total_ms {
run_ms = total_ms;
}
let state = BMK_timedFnState_t {
timeSpent_ns: 0,
timeBudget_ns: PTime::from(total_ms) * TIMELOOP_NANOSEC / 1000,
runBudget_ns: PTime::from(run_ms) * TIMELOOP_NANOSEC / 1000,
fastestRun: BMK_runTime_t {
// Must be larger than any potential measurement.
nanoSecPerRun: TIMELOOP_NANOSEC as f64 * 2_000_000_000.0,
sumOfReturn: usize::MAX,
},
nbLoops: 1,
coolTime: UTIL_getTime(),
};
unsafe { timedFnState.write(state) };
}
/// Tells if the total time budget of the session is spent. Also reports 1
/// after `BMK_benchTimedFn` returned an error.
#[no_mangle]
pub unsafe extern "C" fn BMK_isCompleted_TimedFn(timedFnState: *const BMK_timedFnState_t) -> c_int {
let state = unsafe { &*timedFnState };
c_int::from(state.timeSpent_ns >= state.timeBudget_ns)
}
/// Runs one measurement supposed to last about `run_ms`, automatically
/// scaling `nbLoops`. Runs shorter than half the run budget are re-tried
/// with a larger workload instead of being reported, limiting rounding-error
/// risks; the best (fastest) qualifying run is returned.
#[no_mangle]
pub unsafe extern "C" fn BMK_benchTimedFn(
cont: *mut BMK_timedFnState_t,
p: BMK_benchParams_t,
) -> BMK_runOutcome_t {
let cont = unsafe { &mut *cont };
let runBudget_ns = cont.runBudget_ns;
let runTimeMin_ns = runBudget_ns / 2;
let mut bestRunTime = cont.fastestRun;
loop {
let runResult = unsafe { BMK_benchFunction(p, cont.nbLoops) };
if BMK_isSuccessful_runOutcome(runResult) == 0 {
// Error: move out.
return runResult;
}
let newRunTime = BMK_extract_runTime(runResult);
let loopDuration_ns = newRunTime.nanoSecPerRun * f64::from(cont.nbLoops);
cont.timeSpent_ns = cont.timeSpent_ns.wrapping_add(loopDuration_ns as PTime);
// Estimate nbLoops for the next run to last approximately run_ms.
if loopDuration_ns > runBudget_ns as f64 / 50.0 {
let fastestRun_ns = bestRunTime.nanoSecPerRun.min(newRunTime.nanoSecPerRun);
cont.nbLoops = ((runBudget_ns as f64 / fastestRun_ns) as c_uint).wrapping_add(1);
} else {
// Previous run was too short: blindly increase workload by a
// x10 multiplier.
const MULTIPLIER: c_uint = 10;
debug_assert!(cont.nbLoops < c_uint::MAX / MULTIPLIER); // avoid overflow
cont.nbLoops = cont.nbLoops.wrapping_mul(MULTIPLIER);
}
if loopDuration_ns < runTimeMin_ns as f64 {
// Don't report results when the run time was too small, which
// increases the risk of rounding errors.
continue;
}
if newRunTime.nanoSecPerRun < bestRunTime.nanoSecPerRun {
bestRunTime = newRunTime;
}
return valid_outcome(bestRunTime);
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Test payload observed through `benchPayload`/`initPayload` pointers.
#[derive(Default)]
struct CallLog {
bench_calls: usize,
init_calls: usize,
}
/// Counts invocations and reports `srcSize`, like a size-preserving codec.
unsafe extern "C" fn counting_bench_fn(
_src: *const c_void,
srcSize: usize,
_dst: *mut c_void,
_dstCapacity: usize,
payload: *mut c_void,
) -> usize {
let log = unsafe { &mut *payload.cast::<CallLog>() };
log.bench_calls += 1;
srcSize
}
unsafe extern "C" fn counting_init_fn(payload: *mut c_void) -> usize {
let log = unsafe { &mut *payload.cast::<CallLog>() };
log.init_calls += 1;
0
}
/// Flags results of 5 bytes and above as errors.
unsafe extern "C" fn error_on_5(result: usize) -> c_uint {
c_uint::from(result >= 5)
}
struct Fixture {
srcs: Vec<Vec<u8>>,
dsts: Vec<Vec<u8>>,
src_ptrs: Vec<*const c_void>,
src_sizes: Vec<usize>,
dst_ptrs: Vec<*mut c_void>,
dst_capacities: Vec<usize>,
block_results: Vec<usize>,
log: CallLog,
}
impl Fixture {
fn new(block_sizes: &[usize]) -> Box<Self> {
let srcs: Vec<Vec<u8>> = block_sizes.iter().map(|size| vec![0u8; *size]).collect();
let mut dsts: Vec<Vec<u8>> = block_sizes.iter().map(|size| vec![0u8; *size]).collect();
let src_ptrs = srcs.iter().map(|src| src.as_ptr().cast()).collect();
let src_sizes = srcs.iter().map(Vec::len).collect();
let dst_ptrs = dsts.iter_mut().map(|dst| dst.as_mut_ptr().cast()).collect();
let dst_capacities = dsts.iter().map(Vec::len).collect();
let block_results = vec![0usize; block_sizes.len()];
Box::new(Self {
srcs,
dsts,
src_ptrs,
src_sizes,
dst_ptrs,
dst_capacities,
block_results,
log: CallLog::default(),
})
}
fn params(&mut self, errorFn: BMK_errorFn_t) -> BMK_benchParams_t {
let payload: *mut CallLog = &mut self.log;
BMK_benchParams_t {
benchFn: Some(counting_bench_fn),
benchPayload: payload.cast(),
initFn: Some(counting_init_fn),
initPayload: payload.cast(),
errorFn,
blockCount: self.srcs.len(),
srcBuffers: self.src_ptrs.as_ptr(),
srcSizes: self.src_sizes.as_ptr(),
dstBuffers: self.dst_ptrs.as_ptr(),
dstCapacities: self.dst_capacities.as_ptr(),
blockResults: self.block_results.as_mut_ptr(),
}
}
}
#[test]
fn bench_function_accounts_loops_blocks_and_first_loop_results() {
let mut fixture = Fixture::new(&[3, 8]);
let params = fixture.params(None);
let outcome = unsafe { BMK_benchFunction(params, 4) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
let run_time = BMK_extract_runTime(outcome);
// benchFn ran nbLoops times over each block; initFn ran once.
assert_eq!(fixture.log.bench_calls, 4 * 2);
assert_eq!(fixture.log.init_calls, 1);
// sumOfReturn and blockResults reflect the first loop only.
assert_eq!(run_time.sumOfReturn, 3 + 8);
assert_eq!(fixture.block_results, vec![3, 8]);
assert!(run_time.nanoSecPerRun >= 0.0);
}
#[test]
fn bench_function_treats_zero_loops_as_one_and_warms_up_buffers() {
let mut fixture = Fixture::new(&[4]);
let params = fixture.params(None);
let outcome = unsafe { BMK_benchFunction(params, 0) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
assert_eq!(fixture.log.bench_calls, 1);
// The result buffer was erased with the 0xE5 warm-up pattern.
assert_eq!(fixture.dsts[0], vec![0xE5; 4]);
}
#[test]
fn bench_function_reports_the_first_failing_block() {
let mut fixture = Fixture::new(&[3, 5, 7]);
let params = fixture.params(Some(error_on_5));
let outcome = unsafe { BMK_benchFunction(params, 10) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 0);
assert_eq!(BMK_extract_errorResult(outcome), 5);
// Execution stopped at the failing block, before the third one.
assert_eq!(fixture.log.bench_calls, 2);
// blockResults were recorded up to and including the failure.
assert_eq!(fixture.block_results[..2], [3, 5]);
}
#[test]
fn reset_clamps_budgets_and_rearms_the_loop_counter() {
let state = BMK_createTimedFnState(0, 7);
assert!(!state.is_null());
{
let state = unsafe { &*state };
// total_ms 0 becomes 1ms, and run_ms is clamped to total_ms.
assert_eq!(state.timeBudget_ns, 1_000_000);
assert_eq!(state.runBudget_ns, 1_000_000);
assert_eq!(state.nbLoops, 1);
assert_eq!(state.timeSpent_ns, 0);
assert_eq!(state.fastestRun.sumOfReturn, usize::MAX);
}
assert_eq!(unsafe { BMK_isCompleted_TimedFn(state) }, 0);
unsafe { BMK_resetTimedFnState(state, 2_000, 500) };
{
let state = unsafe { &*state };
assert_eq!(state.timeBudget_ns, 2_000_000_000);
assert_eq!(state.runBudget_ns, 500_000_000);
}
unsafe { BMK_freeTimedFnState(state) };
}
#[test]
fn static_state_initialization_validates_its_buffer() {
let mut shell = [0u64; BMK_TIMEDFNSTATE_SIZE / 8];
let buffer: *mut c_void = shell.as_mut_ptr().cast();
// A properly sized and aligned buffer is accepted.
let state = unsafe { BMK_initStatic_timedFnState(buffer, 64, 1_000, 100) };
assert!(!state.is_null());
assert_eq!(unsafe { BMK_isCompleted_TimedFn(state) }, 0);
// NULL, undersized, and misaligned buffers are rejected.
let too_small = std::mem::size_of::<BMK_timedFnState_t>() - 1;
unsafe {
assert!(BMK_initStatic_timedFnState(ptr::null_mut(), 64, 1, 1).is_null());
assert!(BMK_initStatic_timedFnState(buffer, too_small, 1, 1).is_null());
assert!(
BMK_initStatic_timedFnState(buffer.cast::<u8>().add(1).cast(), 63, 1, 1).is_null()
);
}
}
#[test]
fn timed_runs_grow_the_workload_and_spend_the_budget() {
let mut fixture = Fixture::new(&[16]);
let params = fixture.params(None);
let state = BMK_createTimedFnState(4, 2);
assert!(!state.is_null());
let mut rounds = 0usize;
while unsafe { BMK_isCompleted_TimedFn(state) } == 0 {
let outcome = unsafe { BMK_benchTimedFn(state, params) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
let run_time = BMK_extract_runTime(outcome);
assert_eq!(run_time.sumOfReturn, 16);
rounds += 1;
assert!(rounds < 1_000, "the time budget must eventually be spent");
}
{
let state = unsafe { &*state };
// A reported run had to last at least runBudget/2, which is only
// reachable for this trivial function with a grown loop counter.
assert!(state.nbLoops > 1);
assert!(state.timeSpent_ns >= state.timeBudget_ns);
}
// Every reported outcome came from a run of >= runBudget/2, and the
// budget accounting matches BMK_isCompleted_TimedFn.
assert!(rounds >= 1);
assert!(fixture.log.bench_calls >= rounds);
unsafe { BMK_freeTimedFnState(state) };
}
#[test]
fn timed_runs_propagate_errors_without_aborting() {
let mut fixture = Fixture::new(&[9]);
let params = fixture.params(Some(error_on_5));
let state = BMK_createTimedFnState(1_000, 100);
let outcome = unsafe { BMK_benchTimedFn(state, params) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 0);
assert_eq!(BMK_extract_errorResult(outcome), 9);
unsafe { BMK_freeTimedFnState(state) };
}
}
+2756
View File
@@ -0,0 +1,2756 @@
#![allow(clippy::missing_safety_doc)]
#![allow(clippy::too_many_arguments)]
//! Suffix-array construction for the dictionary builder.
//!
//! Port of `lib/dictBuilder/divsufsort.c` (libdivsufsort-lite, Copyright (c)
//! 2003-2008 Yuta Mori, MIT license) in the exact configuration zstd compiles
//! it with: `ALPHABET_SIZE = 256`, `SS_INSERTIONSORT_THRESHOLD = 8`,
//! `SS_BLOCKSIZE = 1024`, and no OpenMP. Only `divsufsort()` is exported;
//! `divbwt()` has no callers anywhere in zstd and was not ported.
//!
//! The C implementation walks raw `int*` cursors through the caller's SA
//! buffer, including transient one-before-the-range positions. Every such
//! cursor is translated to an `isize` index into one `&mut [i32]` slice
//! covering the whole buffer, so all arithmetic — including the
//! bitwise-complement rank marking and the C `int` value semantics — matches
//! the original exactly while staying bounds-checked.
use std::os::raw::c_int;
use std::slice;
const BUCKET_A_SIZE: usize = 256; /* ALPHABET_SIZE */
const BUCKET_B_SIZE: usize = 256 * 256; /* ALPHABET_SIZE * ALPHABET_SIZE */
const ALPHABET_SIZE: i32 = 256;
const SS_INSERTIONSORT_THRESHOLD: isize = 8;
const SS_BLOCKSIZE: isize = 1024;
/* minstacksize = log(SS_BLOCKSIZE) / log(3) * 2 */
const SS_MISORT_STACKSIZE: usize = 16;
const SS_SMERGE_STACKSIZE: usize = 32;
const TR_INSERTIONSORT_THRESHOLD: isize = 8;
const TR_STACKSIZE: usize = 64;
#[rustfmt::skip]
static LG_TABLE: [i32; 256] = [
-1,0,1,1,2,2,2,2,3,3,3,3,3,3,3,3,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,
5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
];
#[rustfmt::skip]
static SQQ_TABLE: [i32; 256] = [
0, 16, 22, 27, 32, 35, 39, 42, 45, 48, 50, 53, 55, 57, 59, 61,
64, 65, 67, 69, 71, 73, 75, 76, 78, 80, 81, 83, 84, 86, 87, 89,
90, 91, 93, 94, 96, 97, 98, 99, 101, 102, 103, 104, 106, 107, 108, 109,
110, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126,
128, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 144, 145, 146, 147, 148, 149, 150, 150, 151, 152, 153, 154, 155, 155,
156, 157, 158, 159, 160, 160, 161, 162, 163, 163, 164, 165, 166, 167, 167, 168,
169, 170, 170, 171, 172, 173, 173, 174, 175, 176, 176, 177, 178, 178, 179, 180,
181, 181, 182, 183, 183, 184, 185, 185, 186, 187, 187, 188, 189, 189, 190, 191,
192, 192, 193, 193, 194, 195, 195, 196, 197, 197, 198, 199, 199, 200, 201, 201,
202, 203, 203, 204, 204, 205, 206, 206, 207, 208, 208, 209, 209, 210, 211, 211,
212, 212, 213, 214, 214, 215, 215, 216, 217, 217, 218, 218, 219, 219, 220, 221,
221, 222, 222, 223, 224, 224, 225, 225, 226, 226, 227, 227, 228, 229, 229, 230,
230, 231, 231, 232, 232, 233, 234, 234, 235, 235, 236, 236, 237, 237, 238, 238,
239, 240, 240, 241, 241, 242, 242, 243, 243, 244, 244, 245, 245, 246, 246, 247,
247, 248, 248, 249, 249, 250, 250, 251, 251, 252, 252, 253, 253, 254, 254, 255,
];
/* `ss_ilg` in its `256 <= SS_BLOCKSIZE` configuration. */
#[inline]
fn ss_ilg(n: isize) -> i32 {
let n = n as i32;
if n & 0xff00 != 0 {
8 + LG_TABLE[((n >> 8) & 0xff) as usize]
} else {
LG_TABLE[(n & 0xff) as usize]
}
}
#[inline]
fn ss_isqrt(x: isize) -> isize {
if x >= SS_BLOCKSIZE * SS_BLOCKSIZE {
return SS_BLOCKSIZE;
}
let x = x as i32;
let e = if (x as u32) & 0xffff_0000 != 0 {
if (x as u32) & 0xff00_0000 != 0 {
24 + LG_TABLE[((x >> 24) & 0xff) as usize]
} else {
16 + LG_TABLE[((x >> 16) & 0xff) as usize]
}
} else if x & 0xff00 != 0 {
8 + LG_TABLE[((x >> 8) & 0xff) as usize]
} else {
LG_TABLE[(x & 0xff) as usize]
};
let mut y;
if e >= 16 {
y = SQQ_TABLE[(x >> ((e - 6) - (e & 1))) as usize] << ((e >> 1) - 7);
if e >= 24 {
y = (y + 1 + x / y) >> 1;
}
y = (y + 1 + x / y) >> 1;
} else if e >= 8 {
y = (SQQ_TABLE[(x >> ((e - 6) - (e & 1))) as usize] >> (7 - (e >> 1))) + 1;
} else {
return (SQQ_TABLE[x as usize] >> 4) as isize;
}
(if x < y * y { y - 1 } else { y }) as isize
}
/* --------------------------------------------------------------------- */
/// Compares two suffixes. `(p10, p11)` and `(p20, p21)` are the `p[0]`/`p[1]`
/// pairs the C routine reads through its `const int*` arguments; passing the
/// values directly also serves `sssort()`'s local two-element `PAi` array.
#[inline]
fn ss_compare(t: &[u8], p10: i32, p11: i32, p20: i32, p21: i32, depth: i32) -> i32 {
let mut u1 = (depth + p10) as isize;
let mut u2 = (depth + p20) as isize;
let u1n = (p11 + 2) as isize;
let u2n = (p21 + 2) as isize;
while u1 < u1n && u2 < u2n && t[u1 as usize] == t[u2 as usize] {
u1 += 1;
u2 += 1;
}
if u1 < u1n {
if u2 < u2n {
t[u1 as usize] as i32 - t[u2 as usize] as i32
} else {
1
}
} else if u2 < u2n {
-1
} else {
0
}
}
/// `ss_compare(T, p1, p2, depth)` for pointers `p1`/`p2` into the SA buffer.
#[inline]
fn ss_compare_pa(t: &[u8], sa: &[i32], p1: isize, p2: isize, depth: i32) -> i32 {
ss_compare(
t,
sa[p1 as usize],
sa[(p1 + 1) as usize],
sa[p2 as usize],
sa[(p2 + 1) as usize],
depth,
)
}
/* --------------------------------------------------------------------- */
/* Insertionsort for small size groups */
fn ss_insertionsort(t: &[u8], sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) {
let mut i = last - 2;
while first <= i {
let t0 = sa[i as usize];
let mut j = i + 1;
let mut r;
loop {
r = ss_compare_pa(t, sa, pa + t0 as isize, pa + sa[j as usize] as isize, depth);
if r <= 0 {
break;
}
loop {
sa[(j - 1) as usize] = sa[j as usize];
j += 1;
if !(j < last && sa[j as usize] < 0) {
break;
}
}
if last <= j {
break;
}
}
if r == 0 {
sa[j as usize] = !sa[j as usize];
}
sa[(j - 1) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/// `Td[PA[SA[p]]]` — the depth-`td` sorting key of the suffix stored at `p`.
#[inline(always)]
fn ss_key(t: &[u8], sa: &[i32], td: isize, pa: isize, p: isize) -> i32 {
t[(td + sa[(pa + sa[p as usize] as isize) as usize] as isize) as usize] as i32
}
/// `Td[v]` for an already-loaded SA element `v` (`Td[PA[v]]` in C).
#[inline(always)]
fn ss_key_of(t: &[u8], sa: &[i32], td: isize, pa: isize, v: i32) -> i32 {
t[(td + sa[(pa + v as isize) as usize] as isize) as usize] as i32
}
/// `Td[PA[SA[p]] - 1]` — the character preceding the depth-`td` key.
#[inline(always)]
fn ss_key_pred(t: &[u8], sa: &[i32], td: isize, pa: isize, p: isize) -> i32 {
t[(td + sa[(pa + sa[p as usize] as isize) as usize] as isize - 1) as usize] as i32
}
fn ss_fixdown(t: &[u8], td: isize, sa: &mut [i32], pa: isize, base: isize, i: isize, size: isize) {
let mut i = i;
let v = sa[(base + i) as usize];
let c = ss_key_of(t, sa, td, pa, v);
loop {
let mut j = 2 * i + 1;
if j >= size {
break;
}
let mut k = j;
j += 1;
let mut d = ss_key(t, sa, td, pa, base + k);
let e = ss_key(t, sa, td, pa, base + j);
if d < e {
k = j;
d = e;
}
if d <= c {
break;
}
sa[(base + i) as usize] = sa[(base + k) as usize];
i = k;
}
sa[(base + i) as usize] = v;
}
/* Simple top-down heapsort. */
fn ss_heapsort(t: &[u8], td: isize, sa: &mut [i32], pa: isize, base: isize, size: isize) {
let mut m = size;
if size % 2 == 0 {
m -= 1;
if ss_key(t, sa, td, pa, base + m / 2) < ss_key(t, sa, td, pa, base + m) {
sa.swap((base + m) as usize, (base + m / 2) as usize);
}
}
let mut i = m / 2 - 1;
while 0 <= i {
ss_fixdown(t, td, sa, pa, base, i, m);
i -= 1;
}
if size % 2 == 0 {
sa.swap(base as usize, (base + m) as usize);
ss_fixdown(t, td, sa, pa, base, 0, m);
}
let mut i = m - 1;
while 0 < i {
let t0 = sa[base as usize];
sa[base as usize] = sa[(base + i) as usize];
ss_fixdown(t, td, sa, pa, base, 0, i);
sa[(base + i) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/* Returns the median of three elements. */
#[inline]
fn ss_median3(
t: &[u8],
sa: &[i32],
td: isize,
pa: isize,
v1: isize,
v2: isize,
v3: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v2) {
std::mem::swap(&mut v1, &mut v2);
}
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v3) {
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v3) {
return v1;
}
return v3;
}
v2
}
/* Returns the median of five elements. */
#[inline]
fn ss_median5(
t: &[u8],
sa: &[i32],
td: isize,
pa: isize,
v1: isize,
v2: isize,
v3: isize,
v4: isize,
v5: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
let mut v3 = v3;
let mut v4 = v4;
let mut v5 = v5;
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v3) {
std::mem::swap(&mut v2, &mut v3);
}
if ss_key(t, sa, td, pa, v4) > ss_key(t, sa, td, pa, v5) {
std::mem::swap(&mut v4, &mut v5);
}
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v4) {
std::mem::swap(&mut v2, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v3) {
std::mem::swap(&mut v1, &mut v3);
}
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v4) {
std::mem::swap(&mut v1, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if ss_key(t, sa, td, pa, v3) > ss_key(t, sa, td, pa, v4) {
return v4;
}
v3
}
/* Returns the pivot element. */
#[inline]
fn ss_pivot(t: &[u8], sa: &[i32], td: isize, pa: isize, first: isize, last: isize) -> isize {
let mut t0 = last - first;
let middle = first + t0 / 2;
if t0 <= 512 {
if t0 <= 32 {
return ss_median3(t, sa, td, pa, first, middle, last - 1);
}
t0 >>= 2;
return ss_median5(
t,
sa,
td,
pa,
first,
first + t0,
middle,
last - 1 - t0,
last - 1,
);
}
t0 >>= 3;
let first = ss_median3(t, sa, td, pa, first, first + t0, first + (t0 << 1));
let middle = ss_median3(t, sa, td, pa, middle - t0, middle, middle + t0);
let last = ss_median3(t, sa, td, pa, last - 1 - (t0 << 1), last - 1 - t0, last - 1);
ss_median3(t, sa, td, pa, first, middle, last)
}
/* --------------------------------------------------------------------- */
/* Binary partition for substrings. */
/* The `>= x + 1` comparison deliberately mirrors the C expression shape. */
#[allow(clippy::int_plus_one)]
fn ss_partition(sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) -> isize {
let mut a = first - 1;
let mut b = last;
loop {
loop {
a += 1;
if !(a < b) {
break;
}
if !(sa[(pa + sa[a as usize] as isize) as usize] + depth
>= sa[(pa + sa[a as usize] as isize + 1) as usize] + 1)
{
break;
}
sa[a as usize] = !sa[a as usize];
}
loop {
b -= 1;
if !(a < b) {
break;
}
if !(sa[(pa + sa[b as usize] as isize) as usize] + depth
< sa[(pa + sa[b as usize] as isize + 1) as usize] + 1)
{
break;
}
}
if b <= a {
break;
}
let t0 = !sa[b as usize];
sa[b as usize] = sa[a as usize];
sa[a as usize] = t0;
}
if first < a {
sa[first as usize] = !sa[first as usize];
}
a
}
/* Multikey introsort for medium size groups. */
fn ss_mintrosort(t: &[u8], sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) {
let mut stack = [(0isize, 0isize, 0i32, 0i32); SS_MISORT_STACKSIZE];
let mut ssize = 0usize;
let mut first = first;
let mut last = last;
let mut depth = depth;
let mut limit = ss_ilg(last - first);
let mut x: i32 = 0;
loop {
if last - first <= SS_INSERTIONSORT_THRESHOLD {
if 1 < last - first {
ss_insertionsort(t, sa, pa, first, last, depth);
}
/* STACK_POP */
if ssize == 0 {
return;
}
ssize -= 1;
(first, last, depth, limit) = stack[ssize];
continue;
}
let td = depth as isize;
if limit == 0 {
ss_heapsort(t, td, sa, pa, first, last - first);
}
limit -= 1;
if limit < 0 {
let mut a = first + 1;
let mut v = ss_key(t, sa, td, pa, first);
while a < last {
x = ss_key(t, sa, td, pa, a);
if x != v {
if 1 < a - first {
break;
}
v = x;
first = a;
}
a += 1;
}
if ss_key_pred(t, sa, td, pa, first) < v {
first = ss_partition(sa, pa, first, a, depth);
}
if a - first <= last - a {
if 1 < a - first {
stack[ssize] = (a, last, depth, -1);
ssize += 1;
last = a;
depth += 1;
limit = ss_ilg(a - first);
} else {
first = a;
limit = -1;
}
} else if 1 < last - a {
stack[ssize] = (first, a, depth + 1, ss_ilg(a - first));
ssize += 1;
first = a;
limit = -1;
} else {
last = a;
depth += 1;
limit = ss_ilg(a - first);
}
continue;
}
/* choose pivot */
let mut a = ss_pivot(t, sa, td, pa, first, last);
let v = ss_key(t, sa, td, pa, a);
sa.swap(first as usize, a as usize);
/* partition */
let mut b = first;
loop {
b += 1;
if !(b < last) {
break;
}
x = ss_key(t, sa, td, pa, b);
if x != v {
break;
}
}
a = b;
if a < last && x < v {
loop {
b += 1;
if !(b < last) {
break;
}
x = ss_key(t, sa, td, pa, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
}
let mut c = last;
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if x != v {
break;
}
}
let mut d = c;
if b < d && x > v {
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
while b < c {
sa.swap(b as usize, c as usize);
loop {
b += 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
if a <= d {
c = b - 1;
let mut s = a - first;
let t0 = b - a;
if s > t0 {
s = t0;
}
let mut e = first;
let mut f = b - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
let mut s = d - c;
let t0 = last - d - 1;
if s > t0 {
s = t0;
}
let mut e = b;
let mut f = last - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
a = first + (b - a);
c = last - (d - c);
b = if v <= ss_key_pred(t, sa, td, pa, a) {
a
} else {
ss_partition(sa, pa, a, c, depth)
};
if a - first <= last - c {
if last - c <= c - b {
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
stack[ssize] = (c, last, depth, limit);
ssize += 1;
last = a;
} else if a - first <= c - b {
stack[ssize] = (c, last, depth, limit);
ssize += 1;
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
last = a;
} else {
stack[ssize] = (c, last, depth, limit);
ssize += 1;
stack[ssize] = (first, a, depth, limit);
ssize += 1;
first = b;
last = c;
depth += 1;
limit = ss_ilg(c - b);
}
} else if a - first <= c - b {
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
stack[ssize] = (first, a, depth, limit);
ssize += 1;
first = c;
} else if last - c <= c - b {
stack[ssize] = (first, a, depth, limit);
ssize += 1;
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
first = c;
} else {
stack[ssize] = (first, a, depth, limit);
ssize += 1;
stack[ssize] = (c, last, depth, limit);
ssize += 1;
first = b;
last = c;
depth += 1;
limit = ss_ilg(c - b);
}
} else {
limit += 1;
if ss_key_pred(t, sa, td, pa, first) < v {
first = ss_partition(sa, pa, first, last, depth);
limit = ss_ilg(last - first);
}
depth += 1;
}
}
}
/* --------------------------------------------------------------------- */
#[inline]
fn ss_blockswap(sa: &mut [i32], a: isize, b: isize, n: isize) {
let mut a = a;
let mut b = b;
let mut n = n;
while 0 < n {
sa.swap(a as usize, b as usize);
n -= 1;
a += 1;
b += 1;
}
}
#[inline]
fn ss_rotate(sa: &mut [i32], first: isize, middle: isize, last: isize) {
let mut first = first;
let mut last = last;
let mut l = middle - first;
let mut r = last - middle;
while 0 < l && 0 < r {
if l == r {
ss_blockswap(sa, first, middle, l);
break;
}
if l < r {
let mut a = last - 1;
let mut b = middle - 1;
let mut t0 = sa[a as usize];
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if b < first {
sa[a as usize] = t0;
last = a;
r -= l + 1;
if r <= l {
break;
}
a -= 1;
b = middle - 1;
t0 = sa[a as usize];
}
}
} else {
let mut a = first;
let mut b = middle;
let mut t0 = sa[a as usize];
loop {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
if last <= b {
sa[a as usize] = t0;
first = a + 1;
l -= r + 1;
if l <= r {
break;
}
a += 1;
b = middle;
t0 = sa[a as usize];
}
}
}
}
}
/* --------------------------------------------------------------------- */
fn ss_inplacemerge(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
depth: i32,
) {
let mut middle = middle;
let mut last = last;
loop {
let x: i32;
let p: isize;
if sa[(last - 1) as usize] < 0 {
x = 1;
p = pa + (!sa[(last - 1) as usize]) as isize;
} else {
x = 0;
p = pa + sa[(last - 1) as usize] as isize;
}
let mut a = first;
let mut len = middle - first;
let mut half = len >> 1;
let mut r: i32 = -1;
while 0 < len {
let b = a + half;
let bv = sa[b as usize];
let q = ss_compare_pa(
t,
sa,
pa + (if 0 <= bv { bv } else { !bv }) as isize,
p,
depth,
);
if q < 0 {
a = b + 1;
half -= (len & 1) ^ 1;
} else {
r = q;
}
len = half;
half >>= 1;
}
if a < middle {
if r == 0 {
sa[a as usize] = !sa[a as usize];
}
ss_rotate(sa, a, middle, last);
last -= middle - a;
middle = a;
if first == middle {
break;
}
}
last -= 1;
if x != 0 {
loop {
last -= 1;
if !(sa[last as usize] < 0) {
break;
}
}
}
if middle == last {
break;
}
}
}
/* --------------------------------------------------------------------- */
/* Merge-forward with internal buffer. */
fn ss_mergeforward(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
depth: i32,
) {
let bufend = buf + (middle - first) - 1;
ss_blockswap(sa, buf, first, middle - first);
let mut a = first;
let t0 = sa[a as usize];
let mut b = buf;
let mut c = middle;
loop {
let r = ss_compare_pa(
t,
sa,
pa + sa[b as usize] as isize,
pa + sa[c as usize] as isize,
depth,
);
if r < 0 {
loop {
sa[a as usize] = sa[b as usize];
a += 1;
if bufend <= b {
sa[bufend as usize] = t0;
return;
}
sa[b as usize] = sa[a as usize];
b += 1;
if !(sa[b as usize] < 0) {
break;
}
}
} else if r > 0 {
loop {
sa[a as usize] = sa[c as usize];
a += 1;
sa[c as usize] = sa[a as usize];
c += 1;
if last <= c {
while b < bufend {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
return;
}
if !(sa[c as usize] < 0) {
break;
}
}
} else {
sa[c as usize] = !sa[c as usize];
loop {
sa[a as usize] = sa[b as usize];
a += 1;
if bufend <= b {
sa[bufend as usize] = t0;
return;
}
sa[b as usize] = sa[a as usize];
b += 1;
if !(sa[b as usize] < 0) {
break;
}
}
loop {
sa[a as usize] = sa[c as usize];
a += 1;
sa[c as usize] = sa[a as usize];
c += 1;
if last <= c {
while b < bufend {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
return;
}
if !(sa[c as usize] < 0) {
break;
}
}
}
}
}
/* Merge-backward with internal buffer. */
fn ss_mergebackward(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
depth: i32,
) {
let bufend = buf + (last - middle) - 1;
ss_blockswap(sa, buf, middle, last - middle);
let mut x = 0i32;
let mut p1: isize;
let mut p2: isize;
if sa[bufend as usize] < 0 {
p1 = pa + (!sa[bufend as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[bufend as usize] as isize;
}
if sa[(middle - 1) as usize] < 0 {
p2 = pa + (!sa[(middle - 1) as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[(middle - 1) as usize] as isize;
}
let mut a = last - 1;
let t0 = sa[a as usize];
let mut b = bufend;
let mut c = middle - 1;
loop {
let r = ss_compare_pa(t, sa, p1, p2, depth);
if 0 < r {
if x & 1 != 0 {
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if !(sa[b as usize] < 0) {
break;
}
}
x ^= 1;
}
sa[a as usize] = sa[b as usize];
a -= 1;
if b <= buf {
sa[buf as usize] = t0;
break;
}
sa[b as usize] = sa[a as usize];
b -= 1;
if sa[b as usize] < 0 {
p1 = pa + (!sa[b as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[b as usize] as isize;
}
} else if r < 0 {
if x & 2 != 0 {
loop {
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if !(sa[c as usize] < 0) {
break;
}
}
x ^= 2;
}
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if c < first {
while buf < b {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
break;
}
if sa[c as usize] < 0 {
p2 = pa + (!sa[c as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[c as usize] as isize;
}
} else {
if x & 1 != 0 {
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if !(sa[b as usize] < 0) {
break;
}
}
x ^= 1;
}
sa[a as usize] = !sa[b as usize];
a -= 1;
if b <= buf {
sa[buf as usize] = t0;
break;
}
sa[b as usize] = sa[a as usize];
b -= 1;
if x & 2 != 0 {
loop {
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if !(sa[c as usize] < 0) {
break;
}
}
x ^= 2;
}
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if c < first {
while buf < b {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
break;
}
if sa[b as usize] < 0 {
p1 = pa + (!sa[b as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[b as usize] as isize;
}
if sa[c as usize] < 0 {
p2 = pa + (!sa[c as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[c as usize] as isize;
}
}
}
}
/// `GETIDX` — undoes the "already merged" complement marking.
#[inline(always)]
fn getidx(a: i32) -> i32 {
if 0 <= a {
a
} else {
!a
}
}
/// `MERGE_CHECK` — restores or sets the complement marks after a merge.
#[inline]
fn ss_merge_check(t: &[u8], sa: &mut [i32], pa: isize, a: isize, b: isize, c: i32, depth: i32) {
if (c & 1) != 0
|| ((c & 2) != 0
&& ss_compare_pa(
t,
sa,
pa + getidx(sa[(a - 1) as usize]) as isize,
pa + sa[a as usize] as isize,
depth,
) == 0)
{
sa[a as usize] = !sa[a as usize];
}
if (c & 4) != 0
&& ss_compare_pa(
t,
sa,
pa + getidx(sa[(b - 1) as usize]) as isize,
pa + sa[b as usize] as isize,
depth,
) == 0
{
sa[b as usize] = !sa[b as usize];
}
}
/* D&C based merge. */
fn ss_swapmerge(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
bufsize: isize,
depth: i32,
) {
let mut stack = [(0isize, 0isize, 0isize, 0i32); SS_SMERGE_STACKSIZE];
let mut ssize = 0usize;
let mut first = first;
let mut middle = middle;
let mut last = last;
let mut check = 0i32;
loop {
if last - middle <= bufsize {
if first < middle && middle < last {
ss_mergebackward(t, sa, pa, first, middle, last, buf, depth);
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
continue;
}
if middle - first <= bufsize {
if first < middle {
ss_mergeforward(t, sa, pa, first, middle, last, buf, depth);
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
continue;
}
let mut m: isize = 0;
let mut len = std::cmp::min(middle - first, last - middle);
let mut half = len >> 1;
while 0 < len {
if ss_compare_pa(
t,
sa,
pa + getidx(sa[(middle + m + half) as usize]) as isize,
pa + getidx(sa[(middle - m - half - 1) as usize]) as isize,
depth,
) < 0
{
m += half + 1;
half -= (len & 1) ^ 1;
}
len = half;
half >>= 1;
}
if 0 < m {
let lm = middle - m;
let rm = middle + m;
ss_blockswap(sa, lm, middle, m);
let mut l = middle;
let mut r = middle;
let mut next = 0i32;
if rm < last {
if sa[rm as usize] < 0 {
sa[rm as usize] = !sa[rm as usize];
if first < lm {
loop {
l -= 1;
if !(sa[l as usize] < 0) {
break;
}
}
next |= 4;
}
next |= 1;
} else if first < lm {
while sa[r as usize] < 0 {
r += 1;
}
next |= 2;
}
}
if l - first <= last - r {
stack[ssize] = (r, rm, last, (next & 3) | (check & 4));
ssize += 1;
middle = lm;
last = l;
check = (check & 3) | (next & 4);
} else {
if (next & 2) != 0 && r == middle {
next ^= 6;
}
stack[ssize] = (first, lm, l, (check & 3) | (next & 4));
ssize += 1;
first = r;
middle = rm;
check = (next & 3) | (check & 4);
}
} else {
if ss_compare_pa(
t,
sa,
pa + getidx(sa[(middle - 1) as usize]) as isize,
pa + sa[middle as usize] as isize,
depth,
) == 0
{
sa[middle as usize] = !sa[middle as usize];
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
}
}
}
/* --------------------------------------------------------------------- */
/* Substring sort */
fn sssort(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
last: isize,
buf: isize,
bufsize: isize,
depth: i32,
n: isize,
lastsuffix: bool,
) {
let mut first = first;
let mut buf = buf;
let mut bufsize = bufsize;
if lastsuffix {
first += 1;
}
let mut limit: isize = 0;
let mut middle = last;
if bufsize < SS_BLOCKSIZE && bufsize < last - first {
limit = ss_isqrt(last - first);
if bufsize < limit {
if SS_BLOCKSIZE < limit {
limit = SS_BLOCKSIZE;
}
middle = last - limit;
buf = middle;
bufsize = limit;
} else {
limit = 0;
}
}
let mut a = first;
let mut i: isize = 0;
while SS_BLOCKSIZE < middle - a {
ss_mintrosort(t, sa, pa, a, a + SS_BLOCKSIZE, depth);
let mut curbufsize = last - (a + SS_BLOCKSIZE);
let mut curbuf = a + SS_BLOCKSIZE;
if curbufsize <= bufsize {
curbufsize = bufsize;
curbuf = buf;
}
let mut b = a;
let mut k = SS_BLOCKSIZE;
let mut j = i;
while j & 1 != 0 {
ss_swapmerge(t, sa, pa, b - k, b, b + k, curbuf, curbufsize, depth);
b -= k;
k <<= 1;
j >>= 1;
}
a += SS_BLOCKSIZE;
i += 1;
}
ss_mintrosort(t, sa, pa, a, middle, depth);
let mut k = SS_BLOCKSIZE;
while i != 0 {
if i & 1 != 0 {
ss_swapmerge(t, sa, pa, a - k, a, middle, buf, bufsize, depth);
a -= k;
}
k <<= 1;
i >>= 1;
}
if limit != 0 {
ss_mintrosort(t, sa, pa, middle, last, depth);
ss_inplacemerge(t, sa, pa, first, middle, last, depth);
}
if lastsuffix {
/* Insert last type B* suffix. */
let pai0 = sa[(pa + sa[(first - 1) as usize] as isize) as usize];
let pai1 = (n - 2) as i32;
let i0 = sa[(first - 1) as usize];
let mut a = first;
while a < last {
let av = sa[a as usize];
if !(av < 0
|| 0 < ss_compare(
t,
pai0,
pai1,
sa[(pa + av as isize) as usize],
sa[(pa + av as isize + 1) as usize],
depth,
))
{
break;
}
sa[(a - 1) as usize] = av;
a += 1;
}
sa[(a - 1) as usize] = i0;
}
}
/* --------------------------------------------------------------------- */
#[inline]
fn tr_ilg(n: isize) -> i32 {
let n = n as i32;
if (n as u32) & 0xffff_0000 != 0 {
if (n as u32) & 0xff00_0000 != 0 {
24 + LG_TABLE[((n >> 24) & 0xff) as usize]
} else {
16 + LG_TABLE[((n >> 16) & 0xff) as usize]
}
} else if n & 0xff00 != 0 {
8 + LG_TABLE[((n >> 8) & 0xff) as usize]
} else {
LG_TABLE[(n & 0xff) as usize]
}
}
/* --------------------------------------------------------------------- */
/// `ISAd[SA[p]]` — the depth-offset rank of the suffix stored at `p`.
#[inline(always)]
fn tr_key(sa: &[i32], isad: isize, p: isize) -> i32 {
sa[(isad + sa[p as usize] as isize) as usize]
}
/* Simple insertionsort for small size groups. */
fn tr_insertionsort(sa: &mut [i32], isad: isize, first: isize, last: isize) {
let mut a = first + 1;
while a < last {
let t0 = sa[a as usize];
let mut b = a - 1;
let mut r;
loop {
r = sa[(isad + t0 as isize) as usize] - tr_key(sa, isad, b);
if !(0 > r) {
break;
}
loop {
sa[(b + 1) as usize] = sa[b as usize];
b -= 1;
if !(first <= b && sa[b as usize] < 0) {
break;
}
}
if b < first {
break;
}
}
if r == 0 {
sa[b as usize] = !sa[b as usize];
}
sa[(b + 1) as usize] = t0;
a += 1;
}
}
/* --------------------------------------------------------------------- */
fn tr_fixdown(sa: &mut [i32], isad: isize, base: isize, i: isize, size: isize) {
let mut i = i;
let v = sa[(base + i) as usize];
let c = sa[(isad + v as isize) as usize];
loop {
let mut j = 2 * i + 1;
if j >= size {
break;
}
let mut k = j;
j += 1;
let mut d = tr_key(sa, isad, base + k);
let e = tr_key(sa, isad, base + j);
if d < e {
k = j;
d = e;
}
if d <= c {
break;
}
sa[(base + i) as usize] = sa[(base + k) as usize];
i = k;
}
sa[(base + i) as usize] = v;
}
/* Simple top-down heapsort. */
fn tr_heapsort(sa: &mut [i32], isad: isize, base: isize, size: isize) {
let mut m = size;
if size % 2 == 0 {
m -= 1;
if tr_key(sa, isad, base + m / 2) < tr_key(sa, isad, base + m) {
sa.swap((base + m) as usize, (base + m / 2) as usize);
}
}
let mut i = m / 2 - 1;
while 0 <= i {
tr_fixdown(sa, isad, base, i, m);
i -= 1;
}
if size % 2 == 0 {
sa.swap(base as usize, (base + m) as usize);
tr_fixdown(sa, isad, base, 0, m);
}
let mut i = m - 1;
while 0 < i {
let t0 = sa[base as usize];
sa[base as usize] = sa[(base + i) as usize];
tr_fixdown(sa, isad, base, 0, i);
sa[(base + i) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/* Returns the median of three elements. */
#[inline]
fn tr_median3(sa: &[i32], isad: isize, v1: isize, v2: isize, v3: isize) -> isize {
let mut v1 = v1;
let mut v2 = v2;
if tr_key(sa, isad, v1) > tr_key(sa, isad, v2) {
std::mem::swap(&mut v1, &mut v2);
}
if tr_key(sa, isad, v2) > tr_key(sa, isad, v3) {
if tr_key(sa, isad, v1) > tr_key(sa, isad, v3) {
return v1;
}
return v3;
}
v2
}
/* Returns the median of five elements. */
#[inline]
fn tr_median5(
sa: &[i32],
isad: isize,
v1: isize,
v2: isize,
v3: isize,
v4: isize,
v5: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
let mut v3 = v3;
let mut v4 = v4;
let mut v5 = v5;
if tr_key(sa, isad, v2) > tr_key(sa, isad, v3) {
std::mem::swap(&mut v2, &mut v3);
}
if tr_key(sa, isad, v4) > tr_key(sa, isad, v5) {
std::mem::swap(&mut v4, &mut v5);
}
if tr_key(sa, isad, v2) > tr_key(sa, isad, v4) {
std::mem::swap(&mut v2, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if tr_key(sa, isad, v1) > tr_key(sa, isad, v3) {
std::mem::swap(&mut v1, &mut v3);
}
if tr_key(sa, isad, v1) > tr_key(sa, isad, v4) {
std::mem::swap(&mut v1, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if tr_key(sa, isad, v3) > tr_key(sa, isad, v4) {
return v4;
}
v3
}
/* Returns the pivot element. */
#[inline]
fn tr_pivot(sa: &[i32], isad: isize, first: isize, last: isize) -> isize {
let mut t0 = last - first;
let middle = first + t0 / 2;
if t0 <= 512 {
if t0 <= 32 {
return tr_median3(sa, isad, first, middle, last - 1);
}
t0 >>= 2;
return tr_median5(sa, isad, first, first + t0, middle, last - 1 - t0, last - 1);
}
t0 >>= 3;
let first = tr_median3(sa, isad, first, first + t0, first + (t0 << 1));
let middle = tr_median3(sa, isad, middle - t0, middle, middle + t0);
let last = tr_median3(sa, isad, last - 1 - (t0 << 1), last - 1 - t0, last - 1);
tr_median3(sa, isad, first, middle, last)
}
/* --------------------------------------------------------------------- */
struct TrBudget {
chance: i32,
remain: i32,
incval: i32,
count: i32,
}
impl TrBudget {
fn new(chance: i32, incval: i32) -> Self {
TrBudget {
chance,
remain: incval,
incval,
count: 0,
}
}
fn check(&mut self, size: isize) -> bool {
let size = size as i32;
if size <= self.remain {
self.remain -= size;
return true;
}
if self.chance == 0 {
self.count += size;
return false;
}
self.remain += self.incval - size;
self.chance -= 1;
true
}
}
/* --------------------------------------------------------------------- */
fn tr_partition(
sa: &mut [i32],
isad: isize,
first: isize,
middle: isize,
last: isize,
v: i32,
) -> (isize, isize) {
let mut first = first;
let mut last = last;
let mut x: i32 = 0;
let mut b = middle - 1;
loop {
b += 1;
if !(b < last) {
break;
}
x = tr_key(sa, isad, b);
if x != v {
break;
}
}
let mut a = b;
if a < last && x < v {
loop {
b += 1;
if !(b < last) {
break;
}
x = tr_key(sa, isad, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
}
let mut c = last;
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if x != v {
break;
}
}
let mut d = c;
if b < d && x > v {
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
while b < c {
sa.swap(b as usize, c as usize);
loop {
b += 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
if a <= d {
c = b - 1;
let mut s = a - first;
let t0 = b - a;
if s > t0 {
s = t0;
}
let mut e = first;
let mut f = b - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
let mut s = d - c;
let t0 = last - d - 1;
if s > t0 {
s = t0;
}
let mut e = b;
let mut f = last - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
first += b - a;
last -= d - c;
}
(first, last)
}
/* sort suffixes of middle partition by using sorted order of suffixes of
* left and right partition. */
fn tr_copy(
sa: &mut [i32],
isa: isize,
first: isize,
a: isize,
b: isize,
last: isize,
depth: isize,
) {
/* All cursor arithmetic is relative to the slice start, which is the C
* routine's `SA` pointer, so `x - SA` becomes plain `x`. */
let v = (b - 1) as i32;
let mut c = first;
let mut d = a - 1;
while c <= d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d += 1;
sa[d as usize] = s;
sa[(isa + s as isize) as usize] = d as i32;
}
c += 1;
}
let mut c = last - 1;
let e = d + 1;
let mut d = b;
while e < d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d -= 1;
sa[d as usize] = s;
sa[(isa + s as isize) as usize] = d as i32;
}
c -= 1;
}
}
fn tr_partialcopy(
sa: &mut [i32],
isa: isize,
first: isize,
a: isize,
b: isize,
last: isize,
depth: isize,
) {
let v = (b - 1) as i32;
let mut newrank: i32 = -1;
let mut lastrank: i32 = -1;
let mut c = first;
let mut d = a - 1;
while c <= d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d += 1;
sa[d as usize] = s;
let rank = sa[(isa + s as isize + depth) as usize];
if lastrank != rank {
lastrank = rank;
newrank = d as i32;
}
sa[(isa + s as isize) as usize] = newrank;
}
c += 1;
}
let mut lastrank: i32 = -1;
let mut e = d;
while first <= e {
let rank = sa[(isa + sa[e as usize] as isize) as usize];
if lastrank != rank {
lastrank = rank;
newrank = e as i32;
}
if newrank != rank {
sa[(isa + sa[e as usize] as isize) as usize] = newrank;
}
e -= 1;
}
let mut lastrank: i32 = -1;
let mut c = last - 1;
let e = d + 1;
let mut d = b;
while e < d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d -= 1;
sa[d as usize] = s;
let rank = sa[(isa + s as isize + depth) as usize];
if lastrank != rank {
lastrank = rank;
newrank = d as i32;
}
sa[(isa + s as isize) as usize] = newrank;
}
c -= 1;
}
}
fn tr_introsort(
sa: &mut [i32],
isa: isize,
isad: isize,
first: isize,
last: isize,
budget: &mut TrBudget,
) {
/* Stack frames are (ISAd, first, last, limit, trlink); the tandem-repeat
* copy frame stores its `(a, b)` pair in the pointer fields with a zero
* placeholder where C pushes a NULL ISAd. */
let mut stack = [(0isize, 0isize, 0isize, 0i32, 0i32); TR_STACKSIZE];
let mut ssize = 0usize;
let mut trlink: i32 = -1;
let mut isad = isad;
let mut first = first;
let mut last = last;
let incr = isad - isa;
let mut limit = tr_ilg(last - first);
loop {
if limit < 0 {
if limit == -1 {
/* tandem repeat partition */
let (a, b) = tr_partition(sa, isad - incr, first, first, last, (last - 1) as i32);
/* update ranks */
if a < last {
let v = (a - 1) as i32;
let mut c = first;
while c < a {
sa[(isa + sa[c as usize] as isize) as usize] = v;
c += 1;
}
}
if b < last {
let v = (b - 1) as i32;
let mut c = a;
while c < b {
sa[(isa + sa[c as usize] as isize) as usize] = v;
c += 1;
}
}
/* push */
if 1 < b - a {
stack[ssize] = (0, a, b, 0, 0);
ssize += 1;
stack[ssize] = (isad - incr, first, last, -2, trlink);
ssize += 1;
trlink = ssize as i32 - 2;
}
if a - first <= last - b {
if 1 < a - first {
stack[ssize] = (isad, b, last, tr_ilg(last - b), trlink);
ssize += 1;
last = a;
limit = tr_ilg(a - first);
} else if 1 < last - b {
first = b;
limit = tr_ilg(last - b);
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if 1 < last - b {
stack[ssize] = (isad, first, a, tr_ilg(a - first), trlink);
ssize += 1;
first = b;
limit = tr_ilg(last - b);
} else if 1 < a - first {
last = a;
limit = tr_ilg(a - first);
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if limit == -2 {
/* tandem repeat copy */
ssize -= 1;
let a = stack[ssize].1;
let b = stack[ssize].2;
if stack[ssize].3 == 0 {
tr_copy(sa, isa, first, a, b, last, isad - isa);
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
tr_partialcopy(sa, isa, first, a, b, last, isad - isa);
}
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
} else {
/* sorted partition */
if 0 <= sa[first as usize] {
let mut a = first;
loop {
sa[(isa + sa[a as usize] as isize) as usize] = a as i32;
a += 1;
if !(a < last && 0 <= sa[a as usize]) {
break;
}
}
first = a;
}
if first < last {
let mut a = first;
loop {
sa[a as usize] = !sa[a as usize];
a += 1;
if !(sa[a as usize] < 0) {
break;
}
}
let next =
if sa[(isa + sa[a as usize] as isize) as usize] != tr_key(sa, isad, a) {
tr_ilg(a - first + 1)
} else {
-1
};
a += 1;
if a < last {
let v = (a - 1) as i32;
let mut b = first;
while b < a {
sa[(isa + sa[b as usize] as isize) as usize] = v;
b += 1;
}
}
/* push */
if budget.check(a - first) {
if a - first <= last - a {
stack[ssize] = (isad, a, last, -3, trlink);
ssize += 1;
isad += incr;
last = a;
limit = next;
} else if 1 < last - a {
stack[ssize] = (isad + incr, first, a, next, trlink);
ssize += 1;
first = a;
limit = -3;
} else {
isad += incr;
last = a;
limit = next;
}
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if 1 < last - a {
first = a;
limit = -3;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
continue;
}
if last - first <= TR_INSERTIONSORT_THRESHOLD {
tr_insertionsort(sa, isad, first, last);
limit = -3;
continue;
}
/* C decrements `limit` here (`limit-- == 0`); the decrement is
* observable only on the not-taken path because the taken path
* overwrites `limit` with -3. */
if limit == 0 {
tr_heapsort(sa, isad, first, last - first);
let mut a = last - 1;
while first < a {
let x = tr_key(sa, isad, a);
let mut b = a - 1;
while first <= b && tr_key(sa, isad, b) == x {
sa[b as usize] = !sa[b as usize];
b -= 1;
}
a = b;
}
limit = -3;
continue;
}
limit -= 1;
/* choose pivot */
let a = tr_pivot(sa, isad, first, last);
sa.swap(first as usize, a as usize);
let v = tr_key(sa, isad, first);
/* partition */
let (a, b) = tr_partition(sa, isad, first, first + 1, last, v);
if last - first != b - a {
let next = if sa[(isa + sa[a as usize] as isize) as usize] != v {
tr_ilg(b - a)
} else {
-1
};
/* update ranks */
{
let vv = (a - 1) as i32;
let mut c = first;
while c < a {
sa[(isa + sa[c as usize] as isize) as usize] = vv;
c += 1;
}
}
if b < last {
let vv = (b - 1) as i32;
let mut c = a;
while c < b {
sa[(isa + sa[c as usize] as isize) as usize] = vv;
c += 1;
}
}
/* push */
if 1 < b - a && budget.check(b - a) {
if a - first <= last - b {
if last - b <= b - a {
if 1 < a - first {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
last = a;
} else if 1 < last - b {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
first = b;
} else {
isad += incr;
first = a;
last = b;
limit = next;
}
} else if a - first <= b - a {
if 1 < a - first {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
last = a;
} else {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else if a - first <= b - a {
if 1 < last - b {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
first = b;
} else if 1 < a - first {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
last = a;
} else {
isad += incr;
first = a;
last = b;
limit = next;
}
} else if last - b <= b - a {
if 1 < last - b {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
first = b;
} else {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
if 1 < b - a && 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if a - first <= last - b {
if 1 < a - first {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
last = a;
} else if 1 < last - b {
first = b;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if 1 < last - b {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
first = b;
} else if 1 < a - first {
last = a;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
} else if budget.check(last - first) {
limit = tr_ilg(last - first);
isad += incr;
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
}
/* --------------------------------------------------------------------- */
/* Tandem repeat sort */
fn trsort(sa: &mut [i32], isa: isize, n: isize, depth: isize) {
let mut budget = TrBudget::new(tr_ilg(n) * 2 / 3, n as i32);
/* trbudget_init(&budget, tr_ilg(n) * 3 / 4, n); */
let mut isad = isa + depth;
while -(n as i32) < sa[0] {
let mut first: isize = 0;
let mut skip: isize = 0;
let mut unsorted: i32 = 0;
loop {
let t0 = sa[first as usize];
if t0 < 0 {
first -= t0 as isize;
skip += t0 as isize;
} else {
if skip != 0 {
sa[(first + skip) as usize] = skip as i32;
skip = 0;
}
let last = sa[(isa + t0 as isize) as usize] as isize + 1;
if 1 < last - first {
budget.count = 0;
tr_introsort(sa, isa, isad, first, last, &mut budget);
if budget.count != 0 {
unsorted += budget.count;
} else {
skip = first - last;
}
} else if last - first == 1 {
skip = -1;
}
first = last;
}
if !(first < n) {
break;
}
}
if skip != 0 {
sa[(first + skip) as usize] = skip as i32;
}
if unsorted == 0 {
break;
}
isad += isad - isa;
}
}
/* --------------------------------------------------------------------- */
/// `BUCKET_B(c0, c1)` for the 256-symbol alphabet.
#[inline(always)]
fn bb(c0: i32, c1: i32) -> usize {
(((c1 as u32) << 8) | c0 as u32) as usize
}
/// `BUCKET_BSTAR(c0, c1)` for the 256-symbol alphabet.
#[inline(always)]
fn bstar(c0: i32, c1: i32) -> usize {
(((c0 as u32) << 8) | c1 as u32) as usize
}
/* Sorts suffixes of type B*. */
fn sort_type_bstar(
t: &[u8],
sa: &mut [i32],
bucket_a: &mut [i32],
bucket_b: &mut [i32],
n: isize,
) -> isize {
/* Initialize bucket arrays. */
for slot in bucket_a.iter_mut() {
*slot = 0;
}
for slot in bucket_b.iter_mut() {
*slot = 0;
}
/* Count the number of occurrences of the first one or two characters of
each type A, B and B* suffix. Moreover, store the beginning position of
all type B* suffixes into the array SA. */
let mut i = n - 1;
let mut m = n;
let mut c0 = t[(n - 1) as usize] as i32;
let mut c1;
while 0 <= i {
/* type A suffix. */
loop {
c1 = c0;
bucket_a[c1 as usize] += 1;
i -= 1;
if 0 <= i {
c0 = t[i as usize] as i32;
if c0 >= c1 {
continue;
}
}
break;
}
if 0 <= i {
/* type B* suffix. */
bucket_b[bstar(c0, c1)] += 1;
m -= 1;
sa[m as usize] = i as i32;
/* type B suffix. */
i -= 1;
c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 <= c1) {
break;
}
bucket_b[bb(c0, c1)] += 1;
i -= 1;
c1 = c0;
}
}
}
let m = n - m;
/*
note:
A type B* suffix is lexicographically smaller than a type B suffix that
begins with the same first two characters.
*/
/* Calculate the index of start/end point of each bucket. */
{
let mut i: i32 = 0;
let mut j: i32 = 0;
for c0 in 0..ALPHABET_SIZE {
let t0 = i + bucket_a[c0 as usize];
bucket_a[c0 as usize] = i + j; /* start point */
i = t0 + bucket_b[bb(c0, c0)];
for c1 in (c0 + 1)..ALPHABET_SIZE {
j += bucket_b[bstar(c0, c1)];
bucket_b[bstar(c0, c1)] = j; /* end point */
i += bucket_b[bb(c0, c1)];
}
}
}
if 0 < m {
/* Sort the type B* suffixes by their first two characters. */
let pab = n - m;
let isab = m;
let mut i = m - 2;
while 0 <= i {
let t0 = sa[(pab + i) as usize];
let c0 = t[t0 as usize] as i32;
let c1 = t[(t0 + 1) as usize] as i32;
bucket_b[bstar(c0, c1)] -= 1;
sa[bucket_b[bstar(c0, c1)] as usize] = i as i32;
i -= 1;
}
{
let t0 = sa[(pab + m - 1) as usize];
let c0 = t[t0 as usize] as i32;
let c1 = t[(t0 + 1) as usize] as i32;
bucket_b[bstar(c0, c1)] -= 1;
sa[bucket_b[bstar(c0, c1)] as usize] = (m - 1) as i32;
}
/* Sort the type B* substrings using sssort. */
let buf = m;
let bufsize = n - 2 * m;
let mut c0 = ALPHABET_SIZE - 2;
let mut j = m;
while 0 < j {
let mut c1 = ALPHABET_SIZE - 1;
while c0 < c1 {
let i = bucket_b[bstar(c0, c1)] as isize;
if 1 < j - i {
sssort(
t,
sa,
pab,
i,
j,
buf,
bufsize,
2,
n,
sa[i as usize] == (m - 1) as i32,
);
}
j = i;
c1 -= 1;
}
c0 -= 1;
}
/* Compute ranks of type B* substrings. */
let mut i = m - 1;
while 0 <= i {
if 0 <= sa[i as usize] {
let j = i;
loop {
sa[(isab + sa[i as usize] as isize) as usize] = i as i32;
i -= 1;
if !(0 <= i && 0 <= sa[i as usize]) {
break;
}
}
sa[(i + 1) as usize] = (i - j) as i32;
if i <= 0 {
break;
}
}
let j = i;
loop {
sa[i as usize] = !sa[i as usize];
sa[(isab + sa[i as usize] as isize) as usize] = j as i32;
i -= 1;
if !(sa[i as usize] < 0) {
break;
}
}
sa[(isab + sa[i as usize] as isize) as usize] = j as i32;
i -= 1;
}
/* Construct the inverse suffix array of type B* suffixes using
trsort. */
trsort(sa, isab, m, 1);
/* Set the sorted order of type B* suffixes. */
let mut i = n - 1;
let mut j = m;
let mut c0 = t[(n - 1) as usize] as i32;
while 0 <= i {
i -= 1;
let mut c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 >= c1) {
break;
}
i -= 1;
c1 = c0;
}
if 0 <= i {
let t0 = i;
i -= 1;
c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 <= c1) {
break;
}
i -= 1;
c1 = c0;
}
j -= 1;
sa[sa[(isab + j) as usize] as usize] = if t0 == 0 || 1 < t0 - i {
t0 as i32
} else {
!(t0 as i32)
};
}
}
/* Calculate the index of start/end point of each bucket. */
bucket_b[bb(ALPHABET_SIZE - 1, ALPHABET_SIZE - 1)] = n as i32; /* end point */
let mut k = m - 1;
let mut c0 = ALPHABET_SIZE - 2;
while 0 <= c0 {
let mut i = bucket_a[(c0 + 1) as usize] as isize - 1;
let mut c1 = ALPHABET_SIZE - 1;
while c0 < c1 {
let t0 = i - bucket_b[bb(c0, c1)] as isize;
bucket_b[bb(c0, c1)] = i as i32; /* end point */
/* Move all type B* suffixes to the correct position. */
i = t0;
let j = bucket_b[bstar(c0, c1)] as isize;
while j <= k {
sa[i as usize] = sa[k as usize];
i -= 1;
k -= 1;
}
c1 -= 1;
}
bucket_b[bstar(c0, c0 + 1)] = (i - bucket_b[bb(c0, c0)] as isize + 1) as i32; /* start point */
bucket_b[bb(c0, c0)] = i as i32; /* end point */
c0 -= 1;
}
}
m
}
/* Constructs the suffix array by using the sorted order of type B*
* suffixes. */
fn construct_sa(
t: &[u8],
sa: &mut [i32],
bucket_a: &mut [i32],
bucket_b: &mut [i32],
n: isize,
m: isize,
) {
if 0 < m {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
let mut c1 = ALPHABET_SIZE - 2;
while 0 <= c1 {
/* Scan the suffix array from right to left. */
let i = bucket_b[bstar(c1, c1 + 1)] as isize;
let mut j = bucket_a[(c1 + 1) as usize] as isize - 1;
let mut k: isize = 0;
let mut c2: i32 = -1;
while i <= j {
let mut s = sa[j as usize];
if 0 < s {
debug_assert_eq!(t[s as usize] as i32, c1);
debug_assert!((s as isize + 1) < n && t[s as usize] <= t[(s + 1) as usize]);
debug_assert!(t[(s - 1) as usize] <= t[s as usize]);
sa[j as usize] = !s;
s -= 1;
let c0 = t[s as usize] as i32;
if 0 < s && (t[(s - 1) as usize] as i32) > c0 {
s = !s;
}
if c0 != c2 {
if 0 <= c2 {
bucket_b[bb(c2, c1)] = k as i32;
}
c2 = c0;
k = bucket_b[bb(c2, c1)] as isize;
}
debug_assert!(k < j);
sa[k as usize] = s;
k -= 1;
} else {
debug_assert!((s == 0 && t[s as usize] as i32 == c1) || s < 0);
sa[j as usize] = !s;
}
j -= 1;
}
c1 -= 1;
}
}
/* Construct the suffix array by using the sorted order of type B
suffixes. */
let mut c2 = t[(n - 1) as usize] as i32;
let mut k = bucket_a[c2 as usize] as isize;
sa[k as usize] = if (t[(n - 2) as usize] as i32) < c2 {
!((n - 1) as i32)
} else {
(n - 1) as i32
};
k += 1;
/* Scan the suffix array from left to right. */
let mut i: isize = 0;
let j = n;
while i < j {
let mut s = sa[i as usize];
if 0 < s {
debug_assert!(t[(s - 1) as usize] >= t[s as usize]);
s -= 1;
let c0 = t[s as usize] as i32;
if s == 0 || (t[(s - 1) as usize] as i32) < c0 {
s = !s;
}
if c0 != c2 {
bucket_a[c2 as usize] = k as i32;
c2 = c0;
k = bucket_a[c2 as usize] as isize;
}
debug_assert!(i < k);
sa[k as usize] = s;
k += 1;
} else {
debug_assert!(s < 0);
sa[i as usize] = !s;
}
i += 1;
}
}
/* --------------------------------------------------------------------- */
/// Rust implementation of the `divsufsort()` entry point used by
/// `ZDICT_trainFromBuffer_legacy()`.
///
/// Integration removes the C function body, so this direct export provides
/// the existing library symbol without a wrapper. The `open_mp` parameter is
/// accepted for signature compatibility only: zstd never defines
/// `LIBBSC_OPENMP`, so the C implementation ignored it as well.
///
/// Returns 0 on success, -1 for invalid arguments, and -2 when the bucket
/// work arrays cannot be allocated, exactly like the C routine.
#[no_mangle]
pub unsafe extern "C" fn divsufsort(
t: *const u8,
sa: *mut c_int,
n: c_int,
open_mp: c_int,
) -> c_int {
let _ = open_mp;
/* Check arguments. */
if t.is_null() || sa.is_null() || n < 0 {
return -1;
}
if n == 0 {
return 0;
}
let text = unsafe { slice::from_raw_parts(t, n as usize) };
let suffix = unsafe { slice::from_raw_parts_mut(sa, n as usize) };
if n == 1 {
suffix[0] = 0;
return 0;
}
if n == 2 {
let m = usize::from(text[0] < text[1]);
suffix[m ^ 1] = 0;
suffix[m] = 1;
return 0;
}
let mut bucket_a: Vec<i32> = Vec::new();
let mut bucket_b: Vec<i32> = Vec::new();
if bucket_a.try_reserve_exact(BUCKET_A_SIZE).is_err()
|| bucket_b.try_reserve_exact(BUCKET_B_SIZE).is_err()
{
/* Match the C implementation's -2 result when malloc fails. */
return -2;
}
bucket_a.resize(BUCKET_A_SIZE, 0);
bucket_b.resize(BUCKET_B_SIZE, 0);
/* Suffixsort. */
let m = sort_type_bstar(text, suffix, &mut bucket_a, &mut bucket_b, n as isize);
construct_sa(text, suffix, &mut bucket_a, &mut bucket_b, n as isize, m);
0
}
#[cfg(test)]
mod tests {
use super::*;
use std::ptr;
fn build_sa(text: &[u8]) -> Vec<i32> {
let mut sa = vec![0i32; text.len()];
let result = unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), text.len() as c_int, 0) };
assert_eq!(result, 0);
sa
}
/// Trivial O(n^2 log n) reference: sort the suffix start positions by the
/// suffixes themselves.
fn reference_sa(text: &[u8]) -> Vec<i32> {
let mut sa: Vec<i32> = (0..text.len() as i32).collect();
sa.sort_by(|&a, &b| text[a as usize..].cmp(&text[b as usize..]));
sa
}
/// Suffix-array invariants: a permutation of `0..n` whose suffixes are in
/// strictly increasing lexicographic order.
fn assert_valid_sa(text: &[u8], sa: &[i32]) {
assert_eq!(sa.len(), text.len());
let mut seen = vec![false; text.len()];
for &p in sa {
let p = usize::try_from(p).expect("suffix index must be non-negative");
assert!(p < text.len(), "suffix index {p} out of range");
assert!(!seen[p], "duplicate suffix index {p}");
seen[p] = true;
}
for pair in sa.windows(2) {
assert!(
text[pair[0] as usize..] < text[pair[1] as usize..],
"suffixes {} and {} are not in sorted order",
pair[0],
pair[1]
);
}
}
/// Fixed-seed numerical-recipes LCG, used to generate reproducible
/// pseudo-random sample buffers.
fn lcg_bytes(len: usize, seed: u32, alphabet: u32) -> Vec<u8> {
let mut state = seed;
(0..len)
.map(|_| {
state = state.wrapping_mul(1_664_525).wrapping_add(1_013_904_223);
((state >> 24) % alphabet) as u8
})
.collect()
}
#[test]
fn rejects_invalid_arguments() {
let text = [0u8; 1];
let mut sa = [0i32; 1];
assert_eq!(
unsafe { divsufsort(ptr::null(), sa.as_mut_ptr(), 1, 0) },
-1
);
assert_eq!(
unsafe { divsufsort(text.as_ptr(), ptr::null_mut(), 1, 0) },
-1
);
assert_eq!(
unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), -1, 0) },
-1
);
}
#[test]
fn sorts_trivial_inputs() {
/* empty */
let text = [0u8; 1];
let mut sa = [i32::MIN; 1];
assert_eq!(
unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), 0, 0) },
0
);
assert_eq!(sa[0], i32::MIN, "n == 0 must not touch the output");
/* single byte */
assert_eq!(build_sa(b"z"), [0]);
/* two bytes: ascending, descending, and equal */
assert_eq!(build_sa(b"ab"), [0, 1]);
assert_eq!(build_sa(b"ba"), [1, 0]);
assert_eq!(build_sa(b"aa"), [1, 0]);
}
#[test]
fn sorts_all_equal_bytes() {
let text = vec![b'q'; 10_000];
let sa = build_sa(&text);
/* For a constant text the shortest suffix sorts first. */
let expected: Vec<i32> = (0..text.len() as i32).rev().collect();
assert_eq!(sa, expected);
}
#[test]
fn sorts_abracadabra_exactly() {
/* Hand-computed: a(10) abra(7) abracadabra(0) acadabra(3) adabra(5)
* bra(8) bracadabra(1) cadabra(4) dabra(6) ra(9) racadabra(2). */
assert_eq!(build_sa(b"abracadabra"), [10, 7, 0, 3, 5, 8, 1, 4, 6, 9, 2]);
}
#[test]
fn matches_reference_on_periodic_text() {
/* Tandem repeats exercise trsort's repeat partitioning. */
let text: Vec<u8> = b"ab".iter().copied().cycle().take(4096).collect();
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
#[test]
fn matches_reference_on_random_bytes() {
let text = lcg_bytes(8192, 0x0BAD_5EED, 256);
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
#[test]
fn matches_reference_on_low_alphabet_text() {
/* A four-symbol alphabet produces the large first-two-character
* buckets that reach sssort's block merging and the deeper trsort
* paths. */
let text = lcg_bytes(16_384, 0xDEAD_BEEF, 4);
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
}
+40
View File
@@ -0,0 +1,40 @@
//! Frozen decoders for the legacy zstd formats (v0.1 through v0.7).
//!
//! Frozen-decoder policy
//! =====================
//!
//! Each `lib/legacy/zstd_v0N.c` translation unit is a self-contained snapshot
//! of the entropy coders and frame logic of that historical release. The
//! Rust ports mirror that property:
//!
//! - Every version keeps its own frozen FSE/Huff0 and frame logic. The
//! modern `fse_decompress`, `huf_decompress`, `bitstream`, or `mem`
//! modules must NOT be reused here, and legacy versions must not share
//! code with each other, even where functions look identical. The legacy
//! formats are frozen; the modern modules keep evolving.
//! - Ports are line-by-line translations of the corresponding C file: same
//! table layouts, same arithmetic, same error codes. Outputs must be
//! byte-identical to the C implementation, including error behavior.
//! - The only shared dependency is `crate::errors`, because the C files
//! include `error_private.h` for the public `ZSTD_ErrorCode` values.
//!
//! Registration
//! ============
//!
//! Cargo features `legacy-v01` .. `legacy-v07` are all declared in
//! `Cargo.toml`. The build systems always pass the feature list derived
//! from the C configuration (`ZSTD_LEGACY_SUPPORT=N` enables versions N
//! and newer), so a feature may be enabled before its port exists. A version
//! without a Rust module simply stays implemented by its C file.
//!
//! To port version `v0N`: add `zstd_v0N.rs` next to this file, reduce
//! `lib/legacy/zstd_v0N.c` to a declaration-only shim, and register the
//! module here with exactly one line:
//!
//! ```text
//! #[cfg(feature = "legacy-v0N")]
//! pub mod zstd_v0N;
//! ```
#[cfg(feature = "legacy-v01")]
pub mod zstd_v01;
+2354
View File
@@ -0,0 +1,2354 @@
#![allow(non_snake_case)]
//! Frozen decoder for the zstd v0.1 format.
//!
//! This is a line-by-line port of `lib/legacy/zstd_v01.c`: the same table
//! layouts, the same arithmetic, and the same error codes. The C file is a
//! self-contained snapshot of the v0.1-era FSE and Huff0 coders, so this
//! module deliberately reimplements them instead of reusing the modern
//! `fse_decompress`/`huf_decompress` modules (see `legacy/mod.rs` for the
//! frozen-decoder policy). The only shared dependency is `crate::errors`,
//! mirroring the C file's `error_private.h` include.
//!
//! The streaming `ZSTDv01_Dctx` state lives entirely in Rust; C callers only
//! ever hold an opaque pointer to it. It is allocated with `libc::malloc`
//! and released with `libc::free`, exactly like the original C context.
use crate::errors::{ERR_isError, ZstdErrorCode, ERROR};
use std::os::raw::{c_uint, c_void};
use std::ptr;
/* ******************************************
* Error management (frozen v0.1 FSE codes)
********************************************/
/* FSE_LIST_ERRORS in zstd_v01.c; values are returned as `(size_t)-code`. */
const FSE_ERROR_GENERIC: usize = 1;
const FSE_ERROR_TABLELOG_TOO_LARGE: usize = 2;
const FSE_ERROR_MAX_SYMBOL_VALUE_TOO_LARGE: usize = 3;
const FSE_ERROR_MAX_SYMBOL_VALUE_TOO_SMALL: usize = 4;
const FSE_ERROR_DST_SIZE_TOO_SMALL: usize = 5;
const FSE_ERROR_SRC_SIZE_WRONG: usize = 6;
const FSE_ERROR_CORRUPTION_DETECTED: usize = 7;
const FSE_ERROR_MAX_CODE: usize = 8;
#[inline]
fn fse_error(code: usize) -> usize {
code.wrapping_neg()
}
#[inline]
fn fse_is_error(code: usize) -> bool {
code > fse_error(FSE_ERROR_MAX_CODE)
}
/* ******************************************
* Tuning parameters (frozen)
********************************************/
const FSE_MAX_MEMORY_USAGE: u32 = 14;
const FSE_MAX_SYMBOL_VALUE: u32 = 255;
const FSE_MAX_TABLELOG: u32 = FSE_MAX_MEMORY_USAGE - 2;
const FSE_MIN_TABLELOG: u32 = 5;
const FSE_TABLELOG_ABSOLUTE_MAX: u32 = 15;
const HUF_MAX_SYMBOL_VALUE: u32 = 255;
const HUF_MAX_TABLELOG: u32 = 12;
const HUF_ABSOLUTEMAX_TABLELOG: u32 = 16;
const IS_32BITS: bool = std::mem::size_of::<usize>() == 4;
const USIZE_BITS: u32 = usize::BITS;
/* ******************************************
* Memory I/O (FSE_read* / ZSTD_read* helpers)
********************************************/
#[inline]
unsafe fn fse_read_le16(mem_ptr: *const u8) -> u16 {
u16::from_le_bytes(ptr::read_unaligned(mem_ptr as *const [u8; 2]))
}
#[inline]
unsafe fn fse_read_le32(mem_ptr: *const u8) -> u32 {
u32::from_le_bytes(ptr::read_unaligned(mem_ptr as *const [u8; 4]))
}
#[inline]
unsafe fn fse_read_le64(mem_ptr: *const u8) -> u64 {
u64::from_le_bytes(ptr::read_unaligned(mem_ptr as *const [u8; 8]))
}
#[inline]
unsafe fn fse_read_lest(mem_ptr: *const u8) -> usize {
if IS_32BITS {
fse_read_le32(mem_ptr) as usize
} else {
fse_read_le64(mem_ptr) as usize
}
}
/// `FSE_highbit32`; the caller guarantees `val != 0`, as in C.
#[inline]
fn fse_highbit32(val: u32) -> u32 {
val.leading_zeros() ^ 31
}
/* ******************************************
* FSE structures
********************************************/
#[repr(C)]
#[derive(Clone, Copy)]
struct FseDecode {
new_state: u16,
symbol: u8,
nb_bits: u8,
}
#[repr(C)]
struct FseDTableHeader {
table_log: u16,
fast_mode: u16,
}
struct FseDStream {
bit_container: usize,
bits_consumed: u32,
ptr: *const u8,
start: *const u8,
}
struct FseDState {
state: usize,
table: *const FseDecode,
}
const FSE_DSTREAM_UNFINISHED: u32 = 0;
const FSE_DSTREAM_END_OF_BUFFER: u32 = 1;
const FSE_DSTREAM_COMPLETED: u32 = 2;
const FSE_DSTREAM_TOO_FAR: u32 = 3;
#[inline]
fn fse_table_step(table_size: u32) -> u32 {
(table_size >> 1) + (table_size >> 3) + 3
}
/* An FSE_DTable is an opaque u32 array: one header word followed by
* `1 << tableLog` FseDecode entries, exactly as in C. */
unsafe fn fse_build_dtable(
dt: *mut u32,
normalized_counter: *const i16,
max_symbol_value: u32,
table_log: u32,
) -> usize {
let dtable_h = dt as *mut FseDTableHeader;
let table_decode = dt.add(1) as *mut FseDecode;
/* Sanity checks */
if max_symbol_value > FSE_MAX_SYMBOL_VALUE {
return fse_error(FSE_ERROR_MAX_SYMBOL_VALUE_TOO_LARGE);
}
if table_log > FSE_MAX_TABLELOG {
return fse_error(FSE_ERROR_TABLELOG_TOO_LARGE);
}
let table_size: u32 = 1 << table_log;
let table_mask = table_size - 1;
let step = fse_table_step(table_size);
let mut symbol_next = [0u16; (FSE_MAX_SYMBOL_VALUE + 1) as usize];
let mut position: u32 = 0;
let mut high_threshold = table_size - 1;
let large_limit = (1i32 << (table_log - 1)) as i16;
let mut no_large: u32 = 1;
/* Init, lay down lowprob symbols */
(*dtable_h).table_log = table_log as u16;
for s in 0..=max_symbol_value {
let count = *normalized_counter.add(s as usize);
if count == -1 {
(*table_decode.add(high_threshold as usize)).symbol = s as u8;
high_threshold = high_threshold.wrapping_sub(1);
symbol_next[s as usize] = 1;
} else {
if count >= large_limit {
no_large = 0;
}
symbol_next[s as usize] = count as u16;
}
}
/* Spread symbols */
for s in 0..=max_symbol_value {
let count = *normalized_counter.add(s as usize);
let mut i = 0i32;
while i < count as i32 {
(*table_decode.add(position as usize)).symbol = s as u8;
position = (position + step) & table_mask;
while position > high_threshold {
position = (position + step) & table_mask; /* lowprob area */
}
i += 1;
}
}
if position != 0 {
/* position must reach all cells once, otherwise normalizedCounter is incorrect */
return fse_error(FSE_ERROR_GENERIC);
}
/* Build Decoding table */
for i in 0..table_size as usize {
let symbol = (*table_decode.add(i)).symbol;
let next_state = symbol_next[symbol as usize];
symbol_next[symbol as usize] = next_state.wrapping_add(1);
let nb_bits = (table_log - fse_highbit32(next_state as u32)) as u8;
(*table_decode.add(i)).nb_bits = nb_bits;
(*table_decode.add(i)).new_state =
(((next_state as u32) << nb_bits).wrapping_sub(table_size)) as u16;
}
(*dtable_h).fast_mode = no_large as u16;
0
}
/* ******************************************
* FSE header bitstream (FSE_readNCount)
********************************************/
unsafe fn fse_read_ncount(
normalized_counter: *mut i16,
max_sv_ptr: &mut u32,
table_log_ptr: &mut u32,
header_buffer: *const u8,
hb_size: usize,
) -> usize {
let istart = header_buffer;
let iend_addr = (istart as usize).wrapping_add(hb_size);
let mut ip = istart;
let mut charnum: u32 = 0;
let mut previous0 = false;
if hb_size < 4 {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
let mut bit_stream = fse_read_le32(ip);
let mut nb_bits: i32 = ((bit_stream & 0xF) + FSE_MIN_TABLELOG) as i32; /* extract tableLog */
if nb_bits > FSE_TABLELOG_ABSOLUTE_MAX as i32 {
return fse_error(FSE_ERROR_TABLELOG_TOO_LARGE);
}
bit_stream >>= 4;
let mut bit_count: i32 = 4;
*table_log_ptr = nb_bits as u32;
let mut remaining: i32 = (1 << nb_bits) + 1;
let mut threshold: i32 = 1 << nb_bits;
nb_bits += 1;
while remaining > 1 && charnum <= *max_sv_ptr {
if previous0 {
let mut n0 = charnum;
while (bit_stream & 0xFFFF) == 0xFFFF {
n0 += 24;
if (ip as usize) < iend_addr.wrapping_sub(5) {
ip = ip.add(2);
bit_stream = fse_read_le32(ip).wrapping_shr(bit_count as u32);
} else {
bit_stream >>= 16;
bit_count += 16;
}
}
while (bit_stream & 3) == 3 {
n0 += 3;
bit_stream >>= 2;
bit_count += 2;
}
n0 += bit_stream & 3;
bit_count += 2;
if n0 > *max_sv_ptr {
return fse_error(FSE_ERROR_MAX_SYMBOL_VALUE_TOO_SMALL);
}
while charnum < n0 {
*normalized_counter.add(charnum as usize) = 0;
charnum += 1;
}
if (ip as usize) <= iend_addr.wrapping_sub(7)
|| (ip as usize).wrapping_add((bit_count >> 3) as usize)
<= iend_addr.wrapping_sub(4)
{
ip = ip.add((bit_count >> 3) as usize);
bit_count &= 7;
bit_stream = fse_read_le32(ip).wrapping_shr(bit_count as u32);
} else {
bit_stream >>= 2;
}
}
{
let max: i16 = ((2 * threshold - 1) - remaining) as i16;
let mut count: i16;
if (bit_stream & (threshold - 1) as u32) < max as i32 as u32 {
count = (bit_stream & (threshold - 1) as u32) as u16 as i16;
bit_count += nb_bits - 1;
} else {
count = (bit_stream & (2 * threshold - 1) as u32) as u16 as i16;
if count as i32 >= threshold {
count = ((count as i32) - (max as i32)) as i16;
}
bit_count += nb_bits;
}
count = count.wrapping_sub(1); /* extra accuracy */
remaining -= (count as i32).abs();
*normalized_counter.add(charnum as usize) = count;
charnum += 1;
previous0 = count == 0;
while remaining < threshold {
nb_bits -= 1;
threshold >>= 1;
}
if (ip as usize) <= iend_addr.wrapping_sub(7)
|| (ip as usize).wrapping_add((bit_count >> 3) as usize)
<= iend_addr.wrapping_sub(4)
{
ip = ip.add((bit_count >> 3) as usize);
bit_count &= 7;
} else {
bit_count -=
(8 * (iend_addr.wrapping_sub(4) as isize - ip as usize as isize)) as i32;
ip = (iend_addr - 4) as *const u8;
}
bit_stream = fse_read_le32(ip).wrapping_shr((bit_count & 31) as u32);
}
}
if remaining != 1 {
return fse_error(FSE_ERROR_GENERIC);
}
*max_sv_ptr = charnum - 1;
ip = ip.wrapping_offset(((bit_count + 7) >> 3) as isize);
if (ip as usize).wrapping_sub(istart as usize) > hb_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
(ip as usize).wrapping_sub(istart as usize)
}
/* ******************************************
* FSE decompression, byte symbols
********************************************/
unsafe fn fse_build_dtable_rle(dt: *mut u32, symbol_value: u8) -> usize {
let dtable_h = dt as *mut FseDTableHeader;
let cell = dt.add(1) as *mut FseDecode;
(*dtable_h).table_log = 0;
(*dtable_h).fast_mode = 0;
(*cell).new_state = 0;
(*cell).symbol = symbol_value;
(*cell).nb_bits = 0;
0
}
unsafe fn fse_build_dtable_raw(dt: *mut u32, nb_bits: u32) -> usize {
let dtable_h = dt as *mut FseDTableHeader;
let dinfo = dt.add(1) as *mut FseDecode;
/* Sanity checks */
if nb_bits < 1 {
return fse_error(FSE_ERROR_GENERIC); /* min size */
}
let table_size: u32 = 1 << nb_bits;
let table_mask = table_size - 1;
let max_symbol_value = table_mask;
(*dtable_h).table_log = nb_bits as u16;
(*dtable_h).fast_mode = 1;
for s in 0..=max_symbol_value {
let cell = dinfo.add(s as usize);
(*cell).new_state = 0;
(*cell).symbol = s as u8;
(*cell).nb_bits = nb_bits as u8;
}
0
}
unsafe fn fse_init_dstream(
bit_d: &mut FseDStream,
src_buffer: *const u8,
src_size: usize,
) -> usize {
let word = std::mem::size_of::<usize>();
if src_size < 1 {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
if src_size >= word {
bit_d.start = src_buffer;
bit_d.ptr = src_buffer.add(src_size - word);
bit_d.bit_container = fse_read_lest(bit_d.ptr);
let contain32 = *src_buffer.add(src_size - 1) as u32;
if contain32 == 0 {
return fse_error(FSE_ERROR_GENERIC); /* stop bit not present */
}
bit_d.bits_consumed = 8 - fse_highbit32(contain32);
} else {
bit_d.start = src_buffer;
bit_d.ptr = bit_d.start;
bit_d.bit_container = *bit_d.start as usize;
/* C switch with fallthrough over srcSize 7..2 */
if src_size >= 7 {
bit_d.bit_container += (*src_buffer.add(6) as usize) << (USIZE_BITS as usize - 16);
}
if src_size >= 6 {
bit_d.bit_container += (*src_buffer.add(5) as usize) << (USIZE_BITS as usize - 24);
}
if src_size >= 5 {
bit_d.bit_container += (*src_buffer.add(4) as usize) << (USIZE_BITS as usize - 32);
}
if src_size >= 4 {
bit_d.bit_container += (*src_buffer.add(3) as usize) << 24;
}
if src_size >= 3 {
bit_d.bit_container += (*src_buffer.add(2) as usize) << 16;
}
if src_size >= 2 {
bit_d.bit_container += (*src_buffer.add(1) as usize) << 8;
}
let contain32 = *src_buffer.add(src_size - 1) as u32;
if contain32 == 0 {
return fse_error(FSE_ERROR_GENERIC); /* stop bit not present */
}
bit_d.bits_consumed = 8 - fse_highbit32(contain32);
bit_d.bits_consumed += ((word - src_size) * 8) as u32;
}
src_size
}
#[inline]
unsafe fn fse_look_bits(bit_d: &FseDStream, nb_bits: u32) -> usize {
let bit_mask = USIZE_BITS - 1;
((bit_d.bit_container << (bit_d.bits_consumed & bit_mask)) >> 1)
>> (bit_mask.wrapping_sub(nb_bits) & bit_mask)
}
#[inline]
unsafe fn fse_look_bits_fast(bit_d: &FseDStream, nb_bits: u32) -> usize {
/* only if nb_bits >= 1 */
let bit_mask = USIZE_BITS - 1;
(bit_d.bit_container << (bit_d.bits_consumed & bit_mask))
>> ((bit_mask + 1).wrapping_sub(nb_bits) & bit_mask)
}
#[inline]
fn fse_skip_bits(bit_d: &mut FseDStream, nb_bits: u32) {
bit_d.bits_consumed = bit_d.bits_consumed.wrapping_add(nb_bits);
}
#[inline]
unsafe fn fse_read_bits(bit_d: &mut FseDStream, nb_bits: u32) -> usize {
let value = fse_look_bits(bit_d, nb_bits);
fse_skip_bits(bit_d, nb_bits);
value
}
#[inline]
unsafe fn fse_read_bits_fast(bit_d: &mut FseDStream, nb_bits: u32) -> usize {
/* only if nb_bits >= 1 */
let value = fse_look_bits_fast(bit_d, nb_bits);
fse_skip_bits(bit_d, nb_bits);
value
}
unsafe fn fse_reload_dstream(bit_d: &mut FseDStream) -> u32 {
let word = std::mem::size_of::<usize>();
if bit_d.bits_consumed > (word * 8) as u32 {
/* should never happen */
return FSE_DSTREAM_TOO_FAR;
}
if (bit_d.ptr as usize) >= (bit_d.start as usize).wrapping_add(word) {
bit_d.ptr = bit_d.ptr.sub((bit_d.bits_consumed >> 3) as usize);
bit_d.bits_consumed &= 7;
bit_d.bit_container = fse_read_lest(bit_d.ptr);
return FSE_DSTREAM_UNFINISHED;
}
if bit_d.ptr == bit_d.start {
if bit_d.bits_consumed < (word * 8) as u32 {
return FSE_DSTREAM_END_OF_BUFFER;
}
return FSE_DSTREAM_COMPLETED;
}
{
let mut nb_bytes = bit_d.bits_consumed >> 3;
let mut result = FSE_DSTREAM_UNFINISHED;
if (bit_d.ptr as usize).wrapping_sub(nb_bytes as usize) < (bit_d.start as usize) {
nb_bytes = ((bit_d.ptr as usize) - (bit_d.start as usize)) as u32; /* ptr > start */
result = FSE_DSTREAM_END_OF_BUFFER;
}
bit_d.ptr = bit_d.ptr.sub(nb_bytes as usize);
bit_d.bits_consumed -= nb_bytes * 8;
bit_d.bit_container = fse_read_lest(bit_d.ptr); /* reminder : srcSize > sizeof(bitD) */
result
}
}
unsafe fn fse_init_dstate(d_state: &mut FseDState, bit_d: &mut FseDStream, dt: *const u32) {
let dtable_h = dt as *const FseDTableHeader;
d_state.state = fse_read_bits(bit_d, (*dtable_h).table_log as u32);
fse_reload_dstream(bit_d);
d_state.table = dt.add(1) as *const FseDecode;
}
unsafe fn fse_decode_symbol(d_state: &mut FseDState, bit_d: &mut FseDStream) -> u8 {
let d_info = *d_state.table.add(d_state.state);
let low_bits = fse_read_bits(bit_d, d_info.nb_bits as u32);
d_state.state = (d_info.new_state as usize).wrapping_add(low_bits);
d_info.symbol
}
unsafe fn fse_decode_symbol_fast(d_state: &mut FseDState, bit_d: &mut FseDStream) -> u8 {
let d_info = *d_state.table.add(d_state.state);
let low_bits = fse_read_bits_fast(bit_d, d_info.nb_bits as u32);
d_state.state = (d_info.new_state as usize).wrapping_add(low_bits);
d_info.symbol
}
#[inline]
fn fse_end_of_dstream(bit_d: &FseDStream) -> bool {
bit_d.ptr == bit_d.start && bit_d.bits_consumed == USIZE_BITS
}
#[inline]
fn fse_end_of_dstate(d_state: &FseDState) -> bool {
d_state.state == 0
}
unsafe fn fse_decompress_using_dtable_generic(
dst: *mut u8,
max_dst_size: usize,
c_src: *const u8,
c_src_size: usize,
dt: *const u32,
fast: bool,
) -> usize {
let ostart = dst;
let mut op = ostart;
let omax_addr = (op as usize).wrapping_add(max_dst_size);
let olimit_addr = omax_addr.wrapping_sub(3);
let mut bit_d = FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
};
let mut state1 = FseDState {
state: 0,
table: ptr::null(),
};
let mut state2 = FseDState {
state: 0,
table: ptr::null(),
};
/* Init */
let error_code = fse_init_dstream(&mut bit_d, c_src, c_src_size);
if fse_is_error(error_code) {
return error_code;
}
fse_init_dstate(&mut state1, &mut bit_d, dt);
fse_init_dstate(&mut state2, &mut bit_d, dt);
macro_rules! fse_getsymbol {
($state:expr) => {
if fast {
fse_decode_symbol_fast($state, &mut bit_d)
} else {
fse_decode_symbol($state, &mut bit_d)
}
};
}
/* Constant conditions from the C source; on 64-bit both are false, on
* 32-bit only the *4 variant reloads. */
const RELOAD_2: bool = FSE_MAX_TABLELOG * 2 + 7 > USIZE_BITS;
const RELOAD_4: bool = FSE_MAX_TABLELOG * 4 + 7 > USIZE_BITS;
/* 4 symbols per loop */
while fse_reload_dstream(&mut bit_d) == FSE_DSTREAM_UNFINISHED && (op as usize) < olimit_addr {
*op = fse_getsymbol!(&mut state1);
if RELOAD_2 {
/* This test must be static */
fse_reload_dstream(&mut bit_d);
}
*op.add(1) = fse_getsymbol!(&mut state2);
if RELOAD_4 {
/* This test must be static */
if fse_reload_dstream(&mut bit_d) > FSE_DSTREAM_UNFINISHED {
op = op.add(2);
break;
}
}
*op.add(2) = fse_getsymbol!(&mut state1);
if RELOAD_2 {
/* This test must be static */
fse_reload_dstream(&mut bit_d);
}
*op.add(3) = fse_getsymbol!(&mut state2);
op = op.add(4);
}
/* tail */
loop {
if fse_reload_dstream(&mut bit_d) > FSE_DSTREAM_COMPLETED
|| (op as usize) == omax_addr
|| (fse_end_of_dstream(&bit_d) && (fast || fse_end_of_dstate(&state1)))
{
break;
}
*op = fse_getsymbol!(&mut state1);
op = op.add(1);
if fse_reload_dstream(&mut bit_d) > FSE_DSTREAM_COMPLETED
|| (op as usize) == omax_addr
|| (fse_end_of_dstream(&bit_d) && (fast || fse_end_of_dstate(&state2)))
{
break;
}
*op = fse_getsymbol!(&mut state2);
op = op.add(1);
}
/* end ? */
if fse_end_of_dstream(&bit_d) && fse_end_of_dstate(&state1) && fse_end_of_dstate(&state2) {
return (op as usize) - (ostart as usize);
}
if (op as usize) == omax_addr {
/* dst buffer is full, but cSrc unfinished */
return fse_error(FSE_ERROR_DST_SIZE_TOO_SMALL);
}
fse_error(FSE_ERROR_CORRUPTION_DETECTED)
}
unsafe fn fse_decompress_using_dtable(
dst: *mut u8,
original_size: usize,
c_src: *const u8,
c_src_size: usize,
dt: *const u32,
) -> usize {
let fast_mode = (*(dt as *const FseDTableHeader)).fast_mode;
/* select fast mode (static) */
if fast_mode != 0 {
return fse_decompress_using_dtable_generic(
dst,
original_size,
c_src,
c_src_size,
dt,
true,
);
}
fse_decompress_using_dtable_generic(dst, original_size, c_src, c_src_size, dt, false)
}
unsafe fn fse_decompress(
dst: *mut u8,
max_dst_size: usize,
c_src: *const u8,
c_src_size: usize,
) -> usize {
let istart = c_src;
let mut ip = istart;
let mut counting = [0i16; (FSE_MAX_SYMBOL_VALUE + 1) as usize];
let mut dt = [0u32; 1 + (1 << FSE_MAX_TABLELOG)]; /* DTable_max_t */
let mut table_log: u32 = 0;
let mut max_symbol_value: u32 = FSE_MAX_SYMBOL_VALUE;
let mut remaining_size = c_src_size;
if c_src_size < 2 {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG); /* too small input size */
}
/* normal FSE decoding mode */
let error_code = fse_read_ncount(
counting.as_mut_ptr(),
&mut max_symbol_value,
&mut table_log,
istart,
c_src_size,
);
if fse_is_error(error_code) {
return error_code;
}
if error_code >= c_src_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG); /* too small input size */
}
ip = ip.add(error_code);
remaining_size -= error_code;
let error_code = fse_build_dtable(
dt.as_mut_ptr(),
counting.as_ptr(),
max_symbol_value,
table_log,
);
if fse_is_error(error_code) {
return error_code;
}
/* always return, even if it is an error code */
fse_decompress_using_dtable(dst, max_dst_size, ip, remaining_size, dt.as_ptr())
}
/* ******************************************
* Huff0 : Huffman block decompression
********************************************/
#[repr(C)]
#[derive(Clone, Copy)]
struct HufDElt {
byte: u8,
nb_bits: u8,
}
/* Loop shapes and arithmetic below intentionally mirror the frozen C. */
#[allow(clippy::needless_range_loop, clippy::manual_div_ceil)]
unsafe fn huf_read_dtable(dtable: *mut u16, src: *const u8, src_size: usize) -> usize {
let mut huff_weight = [0u8; (HUF_MAX_SYMBOL_VALUE + 1) as usize];
let mut rank_val = [0u32; (HUF_ABSOLUTEMAX_TABLELOG + 1) as usize];
let ip = src;
let dt = dtable.add(1) as *mut HufDElt;
if src_size == 0 {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
let mut i_size = *ip as usize;
let o_size: usize;
if i_size >= 128 {
/* special header */
if i_size >= 242 {
/* RLE */
const L: [usize; 14] = [1, 2, 3, 4, 7, 8, 15, 16, 31, 32, 63, 64, 127, 128];
o_size = L[i_size - 242];
huff_weight = [1u8; (HUF_MAX_SYMBOL_VALUE + 1) as usize];
i_size = 0;
} else {
/* Incompressible */
o_size = i_size - 127;
i_size = (o_size + 1) / 2;
if i_size + 1 > src_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
let ip = ip.add(1);
let mut n = 0usize;
while n < o_size {
huff_weight[n] = *ip.add(n / 2) >> 4;
huff_weight[n + 1] = *ip.add(n / 2) & 15;
n += 2;
}
}
} else {
/* header compressed with FSE (normal case) */
if i_size + 1 > src_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
/* max 255 values decoded, last one is implied */
let decoded = fse_decompress(
huff_weight.as_mut_ptr(),
HUF_MAX_SYMBOL_VALUE as usize,
ip.add(1),
i_size,
);
if fse_is_error(decoded) {
return decoded;
}
o_size = decoded;
}
/* collect weight stats */
let mut weight_total: u32 = 0;
for n in 0..o_size {
if huff_weight[n] >= HUF_ABSOLUTEMAX_TABLELOG as u8 {
return fse_error(FSE_ERROR_CORRUPTION_DETECTED);
}
rank_val[huff_weight[n] as usize] += 1;
weight_total += (1u32 << huff_weight[n]) >> 1;
}
if weight_total == 0 {
return fse_error(FSE_ERROR_CORRUPTION_DETECTED);
}
/* get last non-null symbol weight (implied, total must be 2^n) */
let max_bits = fse_highbit32(weight_total) + 1;
if max_bits > *dtable as u32 {
return fse_error(FSE_ERROR_TABLELOG_TOO_LARGE); /* DTable is too small */
}
*dtable = max_bits as u16;
{
let total = 1u32 << max_bits;
let rest = total - weight_total;
let verif = 1u32 << fse_highbit32(rest);
let last_weight = fse_highbit32(rest) + 1;
if verif != rest {
return fse_error(FSE_ERROR_CORRUPTION_DETECTED); /* last value must be a clean power of 2 */
}
huff_weight[o_size] = last_weight as u8;
rank_val[last_weight as usize] += 1;
}
/* check tree construction validity */
if rank_val[1] < 2 || (rank_val[1] & 1) != 0 {
/* by construction : at least 2 elts of rank 1, must be even */
return fse_error(FSE_ERROR_CORRUPTION_DETECTED);
}
/* Prepare ranks */
let mut next_rank_start: u32 = 0;
for n in 1..=(max_bits as usize) {
let current = next_rank_start;
next_rank_start += rank_val[n] << (n - 1);
rank_val[n] = current;
}
/* fill DTable */
for n in 0..=o_size {
let w = huff_weight[n] as usize;
let length = (1u32 << w) >> 1;
let d = HufDElt {
byte: n as u8,
nb_bits: (max_bits + 1 - w as u32) as u8,
};
for i in rank_val[w]..(rank_val[w] + length) {
*dt.add(i as usize) = d;
}
rank_val[w] += length;
}
i_size + 1
}
unsafe fn huf_decode_symbol(d_stream: &mut FseDStream, dt: *const HufDElt, dt_log: u32) -> u8 {
let val = fse_look_bits_fast(d_stream, dt_log); /* note : dtLog >= 1 */
let entry = *dt.add(val);
fse_skip_bits(d_stream, entry.nb_bits as u32);
entry.byte
}
/* The C decode macros: SYMBOL_1 reloads on 32-bit only when
* HUF_MAX_TABLELOG > 12 (never here); SYMBOL_2 reloads on 32-bit. */
#[inline]
unsafe fn huf_decode_symbol_1(
op: *mut u8,
d_stream: &mut FseDStream,
dt: *const HufDElt,
dt_log: u32,
) {
*op = huf_decode_symbol(d_stream, dt, dt_log);
if IS_32BITS && HUF_MAX_TABLELOG > 12 {
fse_reload_dstream(d_stream);
}
}
#[inline]
unsafe fn huf_decode_symbol_2(
op: *mut u8,
d_stream: &mut FseDStream,
dt: *const HufDElt,
dt_log: u32,
) {
*op = huf_decode_symbol(d_stream, dt, dt_log);
if IS_32BITS {
fse_reload_dstream(d_stream);
}
}
unsafe fn huf_decompress_using_dtable(
dst: *mut u8,
max_dst_size: usize,
c_src: *const u8,
c_src_size: usize,
dtable: *const u16,
) -> usize {
if c_src_size < 6 {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
let ostart = dst;
let mut op = ostart;
let omax_addr = (op as usize).wrapping_add(max_dst_size);
let olimit_addr = if max_dst_size < 15 {
op as usize
} else {
omax_addr - 15
};
let dt = dtable.add(1) as *const HufDElt;
let dt_log = *dtable as u32;
/* Init */
let length1 = fse_read_le16(c_src) as usize;
let length2 = fse_read_le16(c_src.add(2)) as usize;
let length3 = fse_read_le16(c_src.add(4)) as usize;
let length4 = c_src_size
.wrapping_sub(6)
.wrapping_sub(length1)
.wrapping_sub(length2)
.wrapping_sub(length3); /* check coherency !! */
let start1 = c_src.add(6);
let start2 = start1.wrapping_add(length1);
let start3 = start2.wrapping_add(length2);
let start4 = start3.wrapping_add(length3);
if length1 + length2 + length3 + 6 >= c_src_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
let mut bit_d1 = FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
};
let mut bit_d2 = FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
};
let mut bit_d3 = FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
};
let mut bit_d4 = FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
};
let error_code = fse_init_dstream(&mut bit_d1, start1, length1);
if fse_is_error(error_code) {
return error_code;
}
let error_code = fse_init_dstream(&mut bit_d2, start2, length2);
if fse_is_error(error_code) {
return error_code;
}
let error_code = fse_init_dstream(&mut bit_d3, start3, length3);
if fse_is_error(error_code) {
return error_code;
}
let error_code = fse_init_dstream(&mut bit_d4, start4, length4);
if fse_is_error(error_code) {
return error_code;
}
let mut reload_status = fse_reload_dstream(&mut bit_d2);
/* 16 symbols per loop; D2-3-4 are supposed to be synchronized and finish together */
while reload_status < FSE_DSTREAM_COMPLETED && (op as usize) < olimit_addr {
huf_decode_symbol_1(op, &mut bit_d1, dt, dt_log);
huf_decode_symbol_1(op.add(1), &mut bit_d2, dt, dt_log);
huf_decode_symbol_1(op.add(2), &mut bit_d3, dt, dt_log);
huf_decode_symbol_1(op.add(3), &mut bit_d4, dt, dt_log);
huf_decode_symbol_2(op.add(4), &mut bit_d1, dt, dt_log);
huf_decode_symbol_2(op.add(5), &mut bit_d2, dt, dt_log);
huf_decode_symbol_2(op.add(6), &mut bit_d3, dt, dt_log);
huf_decode_symbol_2(op.add(7), &mut bit_d4, dt, dt_log);
huf_decode_symbol_1(op.add(8), &mut bit_d1, dt, dt_log);
huf_decode_symbol_1(op.add(9), &mut bit_d2, dt, dt_log);
huf_decode_symbol_1(op.add(10), &mut bit_d3, dt, dt_log);
huf_decode_symbol_1(op.add(11), &mut bit_d4, dt, dt_log);
*op.add(12) = huf_decode_symbol(&mut bit_d1, dt, dt_log);
*op.add(13) = huf_decode_symbol(&mut bit_d2, dt, dt_log);
*op.add(14) = huf_decode_symbol(&mut bit_d3, dt, dt_log);
*op.add(15) = huf_decode_symbol(&mut bit_d4, dt, dt_log);
op = op.add(16);
reload_status = fse_reload_dstream(&mut bit_d2)
| fse_reload_dstream(&mut bit_d3)
| fse_reload_dstream(&mut bit_d4);
fse_reload_dstream(&mut bit_d1);
}
if reload_status != FSE_DSTREAM_COMPLETED {
/* not complete : some bitStream might be FSE_DStream_unfinished */
return fse_error(FSE_ERROR_CORRUPTION_DETECTED);
}
/* tail */
{
let mut bit_tail = FseDStream {
bit_container: bit_d1.bit_container, /* required in case of FSE_DStream_endOfBuffer */
bits_consumed: bit_d1.bits_consumed,
ptr: bit_d1.ptr,
start: start1,
};
while fse_reload_dstream(&mut bit_tail) < FSE_DSTREAM_COMPLETED && (op as usize) < omax_addr
{
*op = huf_decode_symbol(&mut bit_tail, dt, dt_log);
op = op.add(1);
}
if fse_end_of_dstream(&bit_tail) {
return (op as usize) - (ostart as usize);
}
}
if (op as usize) == omax_addr {
return fse_error(FSE_ERROR_DST_SIZE_TOO_SMALL); /* dst buffer is full, but cSrc unfinished */
}
fse_error(FSE_ERROR_CORRUPTION_DETECTED)
}
unsafe fn huf_decompress(
dst: *mut u8,
max_dst_size: usize,
c_src: *const u8,
c_src_size: usize,
) -> usize {
/* HUF_CREATE_STATIC_DTABLE(DTable, HUF_MAX_TABLELOG) */
let mut dtable = [0u16; 1 + (1 << HUF_MAX_TABLELOG)];
dtable[0] = HUF_MAX_TABLELOG as u16;
let mut ip = c_src;
let error_code = huf_read_dtable(dtable.as_mut_ptr(), c_src, c_src_size);
if fse_is_error(error_code) {
return error_code;
}
if error_code >= c_src_size {
return fse_error(FSE_ERROR_SRC_SIZE_WRONG);
}
ip = ip.add(error_code);
huf_decompress_using_dtable(
dst,
max_dst_size,
ip,
c_src_size - error_code,
dtable.as_ptr(),
)
}
/* ******************************************
* zstd v0.1 frame decoding
********************************************/
const ZSTD_MAGIC_NUMBER: u32 = 0xFD2FB51E; /* 3rd version : seqNb header */
const KB: usize = 1 << 10;
const BLOCKSIZE: usize = 128 * KB; /* define, for static allocation */
const MINMATCH: usize = 4;
const MLBITS: u32 = 7;
const LLBITS: u32 = 6;
const OFFBITS: u32 = 5;
const MAX_ML: u32 = (1 << MLBITS) - 1;
const MAX_LL: u32 = (1 << LLBITS) - 1;
const MAX_OFF: u32 = (1 << OFFBITS) - 1;
#[allow(dead_code)] /* part of the frozen v0.1 constant set; used only by the compressor */
const LIT_FSE_LOG: u32 = 11;
const ML_FSE_LOG: u32 = 10;
const LL_FSE_LOG: u32 = 10;
const OFF_FSE_LOG: u32 = 9;
const ZSTD_CONTENTSIZE_ERROR: u64 = 0u64.wrapping_sub(2);
const ZSTD_BLOCK_HEADER_SIZE: usize = 3;
const ZSTD_FRAME_HEADER_SIZE: usize = 4;
/* FSE_DTABLE_SIZE_U32(maxTableLog) == 1 + (1 << maxTableLog) */
const LL_DTABLE_SIZE_U32: usize = 1 + (1 << LL_FSE_LOG as usize);
const OFF_DTABLE_SIZE_U32: usize = 1 + (1 << OFF_FSE_LOG as usize);
const ML_DTABLE_SIZE_U32: usize = 1 + (1 << ML_FSE_LOG as usize);
#[inline]
unsafe fn zstd_copy4(dst: *mut u8, src: *const u8) {
ptr::copy_nonoverlapping(src, dst, 4);
}
#[inline]
unsafe fn zstd_copy8(dst: *mut u8, src: *const u8) {
ptr::copy_nonoverlapping(src, dst, 8);
}
unsafe fn zstd_wildcopy(dst: *mut u8, src: *const u8, length: isize) {
let mut ip = src;
let mut op = dst;
let oend_addr = (op as usize).wrapping_add(length as usize);
while (op as usize) < oend_addr {
zstd_copy8(op, ip);
op = op.add(8);
ip = ip.add(8);
}
}
#[inline]
unsafe fn zstd_read_le16(mem_ptr: *const u8) -> u16 {
u16::from_le_bytes(ptr::read_unaligned(mem_ptr as *const [u8; 2]))
}
#[inline]
unsafe fn zstd_read_le24(mem_ptr: *const u8) -> u32 {
zstd_read_le16(mem_ptr) as u32 + ((*mem_ptr.add(2) as u32) << 16)
}
#[inline]
unsafe fn zstd_read_be32(mem_ptr: *const u8) -> u32 {
u32::from_be_bytes(ptr::read_unaligned(mem_ptr as *const [u8; 4]))
}
/* blockType_t */
const BT_COMPRESSED: u32 = 0;
const BT_RAW: u32 = 1;
const BT_RLE: u32 = 2;
const BT_END: u32 = 3;
struct BlockProperties {
block_type: u32,
orig_size: u32,
}
/// The v0.1 streaming decompression context (`dctx_t` in C). The struct is
/// opaque to C: `zstd_v01.h` only forward-declares it, so the definition now
/// lives here.
#[repr(C)]
pub struct ZSTDv01_Dctx {
ll_table: [u32; LL_DTABLE_SIZE_U32],
off_table: [u32; OFF_DTABLE_SIZE_U32],
ml_table: [u32; ML_DTABLE_SIZE_U32],
previous_dst_end: *const u8,
base: *const u8,
expected: usize,
b_type: u32,
phase: u32,
}
unsafe fn zstdv01_getc_block_size(
src: *const u8,
src_size: usize,
bp_ptr: &mut BlockProperties,
) -> usize {
if src_size < 3 {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
let header_flags = *src;
let c_size =
(*src.add(2) as u32) + ((*src.add(1) as u32) << 8) + (((header_flags & 7) as u32) << 16);
bp_ptr.block_type = (header_flags >> 6) as u32;
bp_ptr.orig_size = if bp_ptr.block_type == BT_RLE {
c_size
} else {
0
};
if bp_ptr.block_type == BT_END {
return 0;
}
if bp_ptr.block_type == BT_RLE {
return 1;
}
c_size as usize
}
unsafe fn zstd_copy_uncompressed_block(
dst: *mut u8,
max_dst_size: usize,
src: *const u8,
src_size: usize,
) -> usize {
if src_size > max_dst_size {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
if src_size > 0 {
ptr::copy_nonoverlapping(src, dst, src_size);
}
src_size
}
unsafe fn zstd_decompress_literals(
dst: *mut u8,
max_dst_size: usize,
src: *const u8,
src_size: usize,
) -> usize {
let oend = dst.add(max_dst_size);
let ip = src;
/* check : minimum 2, for litSize, +1, for content */
if src_size <= 3 {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
let mut lit_size = (*ip.add(1) as usize) + ((*ip as usize) << 8);
lit_size += (((*ip.offset(-3) as usize) >> 3) & 7) << 16; /* mmmmh.... */
if lit_size > max_dst_size {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
let op = oend.sub(lit_size);
let error_code = huf_decompress(op, lit_size, ip.add(2), src_size - 2);
if fse_is_error(error_code) {
return ERROR(ZstdErrorCode::Generic);
}
lit_size
}
unsafe fn zstdv01_decode_literals_block(
dst: *mut u8,
max_dst_size: usize,
lit_start: &mut *const u8,
lit_size: &mut usize,
src: *const u8,
src_size: usize,
) -> usize {
let istart = src;
let mut ip = istart;
let ostart = dst;
let oend = ostart.add(max_dst_size);
let mut litbp = BlockProperties {
block_type: 0,
orig_size: 0,
};
let litc_size = zstdv01_getc_block_size(src, src_size, &mut litbp);
if ERR_isError(litc_size) {
return litc_size;
}
if litc_size > src_size - ZSTD_BLOCK_HEADER_SIZE {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
ip = ip.add(ZSTD_BLOCK_HEADER_SIZE);
match litbp.block_type {
BT_RAW => {
*lit_start = ip;
ip = ip.add(litc_size);
*lit_size = litc_size;
}
BT_RLE => {
let rle_size = litbp.orig_size as usize;
if rle_size > max_dst_size {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
if src_size == 0 {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
if rle_size > 0 {
ptr::write_bytes(oend.sub(rle_size), *ip, rle_size);
}
*lit_start = oend.sub(rle_size);
*lit_size = rle_size;
ip = ip.add(1);
}
BT_COMPRESSED => {
let decoded_lit_size = zstd_decompress_literals(dst, max_dst_size, ip, litc_size);
if ERR_isError(decoded_lit_size) {
return decoded_lit_size;
}
*lit_start = oend.sub(decoded_lit_size);
*lit_size = decoded_lit_size;
ip = ip.add(litc_size);
}
_ => {
/* bt_end and impossible values */
return ERROR(ZstdErrorCode::Generic);
}
}
(ip as usize) - (istart as usize)
}
#[allow(clippy::too_many_arguments)]
unsafe fn zstdv01_decode_seq_headers(
nb_seq: &mut i32,
dumps_ptr: &mut *const u8,
dumps_length_ptr: &mut usize,
dtable_ll: *mut u32,
dtable_ml: *mut u32,
dtable_offb: *mut u32,
src: *const u8,
src_size: usize,
) -> usize {
let istart = src;
let mut ip = istart;
let iend_addr = (istart as usize).wrapping_add(src_size);
/* check */
if src_size < 5 {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
/* SeqHead */
*nb_seq = zstd_read_le16(ip) as i32;
ip = ip.add(2);
let ll_type = (*ip >> 6) as u32;
let off_type = ((*ip >> 4) & 3) as u32;
let ml_type = ((*ip >> 2) & 3) as u32;
let dumps_length: usize;
if (*ip & 2) != 0 {
dumps_length = (*ip.add(2) as usize) + ((*ip.add(1) as usize) << 8);
ip = ip.add(3);
} else {
dumps_length = (*ip.add(1) as usize) + (((*ip as usize) & 1) << 8);
ip = ip.add(2);
}
*dumps_ptr = ip;
ip = ip.wrapping_add(dumps_length);
*dumps_length_ptr = dumps_length;
/* check */
if (ip as usize) > iend_addr.wrapping_sub(3) {
/* min : all 3 are "raw", hence no header, but at least xxLog bits per type */
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
/* sequences */
{
let mut norm = [0i16; (MAX_ML + 1) as usize]; /* assumption : MaxML >= MaxLL and MaxOff */
/* Build DTables */
match ll_type {
BT_RLE => {
fse_build_dtable_rle(dtable_ll, *ip);
ip = ip.add(1);
}
BT_RAW => {
fse_build_dtable_raw(dtable_ll, LLBITS);
}
_ => {
let mut max = MAX_LL;
let mut ll_log = 0u32;
let header_size = fse_read_ncount(
norm.as_mut_ptr(),
&mut max,
&mut ll_log,
ip,
iend_addr - ip as usize,
);
if fse_is_error(header_size) {
return ERROR(ZstdErrorCode::Generic);
}
if ll_log > LL_FSE_LOG {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
ip = ip.add(header_size);
fse_build_dtable(dtable_ll, norm.as_ptr(), max, ll_log);
}
}
match off_type {
BT_RLE => {
if (ip as usize) > iend_addr.wrapping_sub(2) {
/* min : "raw", hence no header, but at least xxLog bits */
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
fse_build_dtable_rle(dtable_offb, *ip);
ip = ip.add(1);
}
BT_RAW => {
fse_build_dtable_raw(dtable_offb, OFFBITS);
}
_ => {
let mut max = MAX_OFF;
let mut off_log = 0u32;
let header_size = fse_read_ncount(
norm.as_mut_ptr(),
&mut max,
&mut off_log,
ip,
iend_addr - ip as usize,
);
if fse_is_error(header_size) {
return ERROR(ZstdErrorCode::Generic);
}
if off_log > OFF_FSE_LOG {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
ip = ip.add(header_size);
fse_build_dtable(dtable_offb, norm.as_ptr(), max, off_log);
}
}
match ml_type {
BT_RLE => {
if (ip as usize) > iend_addr.wrapping_sub(2) {
/* min : "raw", hence no header, but at least xxLog bits */
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
fse_build_dtable_rle(dtable_ml, *ip);
ip = ip.add(1);
}
BT_RAW => {
fse_build_dtable_raw(dtable_ml, MLBITS);
}
_ => {
let mut max = MAX_ML;
let mut ml_log = 0u32;
let header_size = fse_read_ncount(
norm.as_mut_ptr(),
&mut max,
&mut ml_log,
ip,
iend_addr - ip as usize,
);
if fse_is_error(header_size) {
return ERROR(ZstdErrorCode::Generic);
}
if ml_log > ML_FSE_LOG {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
ip = ip.add(header_size);
fse_build_dtable(dtable_ml, norm.as_ptr(), max, ml_log);
}
}
}
(ip as usize) - (istart as usize)
}
#[derive(Clone, Copy)]
struct Seq {
lit_length: usize,
offset: usize,
match_length: usize,
}
struct SeqState {
d_stream: FseDStream,
state_ll: FseDState,
state_offb: FseDState,
state_ml: FseDState,
prev_offset: usize,
dumps: *const u8,
dumps_end: *const u8,
}
unsafe fn zstd_decode_sequence(seq: &mut Seq, seq_state: &mut SeqState) {
let mut dumps = seq_state.dumps;
let de = seq_state.dumps_end;
/* Literal length */
let mut lit_length =
fse_decode_symbol(&mut seq_state.state_ll, &mut seq_state.d_stream) as usize;
let prev_offset = if lit_length != 0 {
seq.offset
} else {
seq_state.prev_offset
};
seq_state.prev_offset = seq.offset;
if lit_length == MAX_LL as usize {
let add = if (dumps as usize) < (de as usize) {
let v = *dumps as u32;
dumps = dumps.add(1);
v
} else {
0
};
if add < 255 {
lit_length += add as usize;
} else if (dumps as usize) <= (de as usize).wrapping_sub(3) {
lit_length = zstd_read_le24(dumps) as usize;
dumps = dumps.add(3);
}
}
/* Offset */
let mut offset: usize;
{
let offset_code =
fse_decode_symbol(&mut seq_state.state_offb, &mut seq_state.d_stream) as u32;
if IS_32BITS {
fse_reload_dstream(&mut seq_state.d_stream);
}
let mut nb_bits = offset_code.wrapping_sub(1);
if offset_code == 0 {
nb_bits = 0; /* cmove */
}
offset = (1usize << (nb_bits & (USIZE_BITS - 1)))
.wrapping_add(fse_read_bits(&mut seq_state.d_stream, nb_bits));
if IS_32BITS {
fse_reload_dstream(&mut seq_state.d_stream);
}
if offset_code == 0 {
offset = prev_offset;
}
}
/* MatchLength */
let mut match_length =
fse_decode_symbol(&mut seq_state.state_ml, &mut seq_state.d_stream) as usize;
if match_length == MAX_ML as usize {
let add = if (dumps as usize) < (de as usize) {
let v = *dumps as u32;
dumps = dumps.add(1);
v
} else {
0
};
if add < 255 {
match_length += add as usize;
} else if (dumps as usize) <= (de as usize).wrapping_sub(3) {
match_length = zstd_read_le24(dumps) as usize;
dumps = dumps.add(3);
}
}
match_length += MINMATCH;
/* save result */
seq.lit_length = lit_length;
seq.offset = offset;
seq.match_length = match_length;
seq_state.dumps = dumps;
}
unsafe fn zstd_exec_sequence(
op: *mut u8,
sequence: Seq,
lit_ptr: &mut *const u8,
lit_limit: *const u8,
base: *const u8,
oend: *mut u8,
) -> usize {
static DEC32TABLE: [usize; 8] = [0, 1, 2, 1, 4, 4, 4, 4]; /* added */
static DEC64TABLE: [usize; 8] = [8, 8, 8, 7, 8, 9, 10, 11]; /* subtracted */
let ostart = op;
let mut op = op;
let o_lit_end = op.wrapping_add(sequence.lit_length);
let lit_length = sequence.lit_length;
/* risk : address space overflow (32-bits) */
let end_match = op
.wrapping_add(lit_length)
.wrapping_add(sequence.match_length);
let lit_end = (*lit_ptr).wrapping_add(lit_length);
/* checks */
let seq_length = sequence.lit_length.wrapping_add(sequence.match_length);
if seq_length > (oend as usize).wrapping_sub(op as usize) {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
if sequence.lit_length > (lit_limit as usize).wrapping_sub(*lit_ptr as usize) {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
/* Now we know there are no overflow in literal nor match lengths, can use pointer checks */
if sequence.offset > ((o_lit_end as usize).wrapping_sub(base as usize)) as u32 as usize {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
if (end_match as usize) > (oend as usize) {
/* overwrite beyond dst buffer */
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
if (lit_end as usize) > (lit_limit as usize) {
/* overRead beyond lit buffer */
return ERROR(ZstdErrorCode::CorruptionDetected);
}
if sequence.match_length > (*lit_ptr as usize).wrapping_sub(op as usize) {
/* overwrite literal segment */
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
/* copy Literals */
/* note : v0.1 seems to allow scenarios where output or input are close to end of buffer */
ptr::copy(*lit_ptr, op, sequence.lit_length);
op = op.add(lit_length);
*lit_ptr = lit_end; /* update for next sequence */
/* check : last match must be at a minimum distance of 8 from end of dest buffer */
if (oend as usize) - (op as usize) < 8 {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
/* copy Match */
{
let overlap_risk = (lit_end as usize).wrapping_sub(end_match as usize) < 12;
/* possible underflow at op - offset ? */
let mut match_ptr = (op as usize).wrapping_sub(sequence.offset) as *const u8;
let mut qutt: usize = 12;
let mut saved = [0u8; 16]; /* U64 saved[2] */
/* check */
if (match_ptr as usize) < (base as usize) {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
if sequence.offset > base as usize {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
/* save beginning of literal sequence, in case of write overlap */
if overlap_risk {
if (end_match as usize).wrapping_add(qutt) > (oend as usize) {
qutt = (oend as usize) - (end_match as usize);
}
ptr::copy_nonoverlapping(end_match as *const u8, saved.as_mut_ptr(), qutt);
}
if sequence.offset < 8 {
let dec64 = DEC64TABLE[sequence.offset];
*op = *match_ptr;
*op.add(1) = *match_ptr.add(1);
*op.add(2) = *match_ptr.add(2);
*op.add(3) = *match_ptr.add(3);
match_ptr = match_ptr.add(DEC32TABLE[sequence.offset]);
zstd_copy4(op.add(4), match_ptr);
match_ptr = match_ptr.sub(dec64);
} else {
zstd_copy8(op, match_ptr);
}
op = op.add(8);
match_ptr = match_ptr.add(8);
if (end_match as usize) > (oend as usize).wrapping_sub(16 - MINMATCH) {
if (op as usize) < (oend as usize).wrapping_sub(8) {
let dist = ((oend as usize) - 8) - (op as usize);
zstd_wildcopy(op, match_ptr, dist as isize);
match_ptr = match_ptr.add(dist);
op = oend.sub(8);
}
while (op as usize) < (end_match as usize) {
*op = *match_ptr;
op = op.add(1);
match_ptr = match_ptr.add(1);
}
} else {
/* works even if matchLength < 8 */
zstd_wildcopy(op, match_ptr, sequence.match_length as isize - 8);
}
/* restore, in case of overlap */
if overlap_risk {
ptr::copy_nonoverlapping(saved.as_ptr(), end_match, qutt);
}
}
(end_match as usize) - (ostart as usize)
}
unsafe fn zstd_decompress_sequences(
ctx: *mut ZSTDv01_Dctx,
dst: *mut u8,
max_dst_size: usize,
seq_start: *const u8,
seq_size: usize,
lit_start: *const u8,
lit_size: usize,
) -> usize {
let dctx = ctx;
let mut ip = seq_start;
let iend_addr = (ip as usize).wrapping_add(seq_size);
let ostart = dst;
let mut op = ostart;
let oend = ostart.add(max_dst_size);
let mut lit_ptr = lit_start;
let lit_end = lit_start.add(lit_size);
let mut nb_seq: i32 = 0;
let mut dumps: *const u8 = ptr::null();
let mut dumps_length: usize = 0;
let dtable_ll = (*dctx).ll_table.as_mut_ptr();
let dtable_ml = (*dctx).ml_table.as_mut_ptr();
let dtable_offb = (*dctx).off_table.as_mut_ptr();
let base = (*dctx).base;
/* Build Decoding Tables */
let error_code = zstdv01_decode_seq_headers(
&mut nb_seq,
&mut dumps,
&mut dumps_length,
dtable_ll,
dtable_ml,
dtable_offb,
ip,
iend_addr - ip as usize,
);
if ERR_isError(error_code) {
return error_code;
}
ip = ip.add(error_code);
/* Regen sequences */
{
let mut sequence = Seq {
lit_length: 0,
offset: 0,
match_length: 0,
};
let mut seq_state = SeqState {
d_stream: FseDStream {
bit_container: 0,
bits_consumed: 0,
ptr: ptr::null(),
start: ptr::null(),
},
state_ll: FseDState {
state: 0,
table: ptr::null(),
},
state_offb: FseDState {
state: 0,
table: ptr::null(),
},
state_ml: FseDState {
state: 0,
table: ptr::null(),
},
prev_offset: 1,
dumps,
dumps_end: dumps.wrapping_add(dumps_length),
};
let error_code = fse_init_dstream(&mut seq_state.d_stream, ip, iend_addr - ip as usize);
if fse_is_error(error_code) {
return ERROR(ZstdErrorCode::CorruptionDetected);
}
fse_init_dstate(&mut seq_state.state_ll, &mut seq_state.d_stream, dtable_ll);
fse_init_dstate(
&mut seq_state.state_offb,
&mut seq_state.d_stream,
dtable_offb,
);
fse_init_dstate(&mut seq_state.state_ml, &mut seq_state.d_stream, dtable_ml);
while fse_reload_dstream(&mut seq_state.d_stream) <= FSE_DSTREAM_COMPLETED && nb_seq > 0 {
nb_seq -= 1;
zstd_decode_sequence(&mut sequence, &mut seq_state);
let one_seq_size = zstd_exec_sequence(op, sequence, &mut lit_ptr, lit_end, base, oend);
if ERR_isError(one_seq_size) {
return one_seq_size;
}
op = op.add(one_seq_size);
}
/* check if reached exact end */
if !fse_end_of_dstream(&seq_state.d_stream) {
/* requested too much : data is corrupted */
return ERROR(ZstdErrorCode::CorruptionDetected);
}
if nb_seq < 0 {
/* requested too many sequences : data is corrupted */
return ERROR(ZstdErrorCode::CorruptionDetected);
}
/* last literal segment */
{
let last_ll_size = (lit_end as usize) - (lit_ptr as usize);
if (op as usize).wrapping_add(last_ll_size) > (oend as usize) {
return ERROR(ZstdErrorCode::DstSizeTooSmall);
}
if last_ll_size > 0 {
if !std::ptr::eq(op as *const u8, lit_ptr) {
ptr::copy(lit_ptr, op, last_ll_size);
}
op = op.add(last_ll_size);
}
}
}
(op as usize) - (ostart as usize)
}
unsafe fn zstd_decompress_block(
ctx: *mut ZSTDv01_Dctx,
dst: *mut u8,
max_dst_size: usize,
src: *const u8,
src_size: usize,
) -> usize {
/* blockType == blockCompressed, srcSize is trusted */
let mut ip = src;
let mut lit_ptr: *const u8 = ptr::null();
let mut lit_size: usize = 0;
/* Decode literals sub-block */
let error_code = zstdv01_decode_literals_block(
dst,
max_dst_size,
&mut lit_ptr,
&mut lit_size,
src,
src_size,
);
if ERR_isError(error_code) {
return error_code;
}
ip = ip.add(error_code);
let src_size = src_size - error_code;
zstd_decompress_sequences(ctx, dst, max_dst_size, ip, src_size, lit_ptr, lit_size)
}
unsafe fn zstdv01_decompress_dctx(
ctx: *mut ZSTDv01_Dctx,
dst: *mut u8,
max_dst_size: usize,
src: *const u8,
src_size: usize,
) -> usize {
let mut ip = src;
let iend_addr = (ip as usize).wrapping_add(src_size);
let ostart = dst;
let mut op = ostart;
let oend_addr = (ostart as usize).wrapping_add(max_dst_size);
let mut remaining_size = src_size;
let mut error_code: usize = 0;
/* Frame Header */
if src_size < ZSTD_FRAME_HEADER_SIZE + ZSTD_BLOCK_HEADER_SIZE {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
let magic_number = zstd_read_be32(src);
if magic_number != ZSTD_MAGIC_NUMBER {
return ERROR(ZstdErrorCode::PrefixUnknown);
}
ip = ip.add(ZSTD_FRAME_HEADER_SIZE);
remaining_size -= ZSTD_FRAME_HEADER_SIZE;
/* Loop on each block */
loop {
let mut block_properties = BlockProperties {
block_type: 0,
orig_size: 0,
};
let block_size =
zstdv01_getc_block_size(ip, iend_addr - ip as usize, &mut block_properties);
if ERR_isError(block_size) {
return block_size;
}
ip = ip.add(ZSTD_BLOCK_HEADER_SIZE);
remaining_size -= ZSTD_BLOCK_HEADER_SIZE;
if block_size > remaining_size {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
match block_properties.block_type {
BT_COMPRESSED => {
error_code =
zstd_decompress_block(ctx, op, oend_addr - op as usize, ip, block_size);
}
BT_RAW => {
error_code =
zstd_copy_uncompressed_block(op, oend_addr - op as usize, ip, block_size);
}
BT_RLE => {
return ERROR(ZstdErrorCode::Generic); /* not yet supported */
}
BT_END => {
/* end of frame */
if remaining_size != 0 {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
}
_ => {
return ERROR(ZstdErrorCode::Generic);
}
}
if block_size == 0 {
break; /* bt_end */
}
if ERR_isError(error_code) {
return error_code;
}
op = op.add(error_code);
ip = ip.add(block_size);
remaining_size -= block_size;
}
(op as usize) - (ostart as usize)
}
/* ZSTD_errorFrameSizeInfoLegacy() :
assumes `cSize` and `dBound` are _not_ NULL */
unsafe fn zstd_error_frame_size_info_legacy(c_size: *mut usize, d_bound: *mut u64, ret: usize) {
*c_size = ret;
*d_bound = ZSTD_CONTENTSIZE_ERROR;
}
/* ******************************************
* Exported C ABI (zstd_v01.h)
********************************************/
/// C ABI: `ZSTDv01_isError`.
#[no_mangle]
pub extern "C" fn ZSTDv01_isError(code: usize) -> c_uint {
ERR_isError(code) as c_uint
}
/// C ABI: `ZSTDv01_decompressDCtx`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_decompressDCtx(
ctx: *mut c_void,
dst: *mut c_void,
max_original_size: usize,
src: *const c_void,
compressed_size: usize,
) -> usize {
zstdv01_decompress_dctx(
ctx as *mut ZSTDv01_Dctx,
dst as *mut u8,
max_original_size,
src as *const u8,
compressed_size,
)
}
/// C ABI: `ZSTDv01_decompress`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_decompress(
dst: *mut c_void,
max_original_size: usize,
src: *const c_void,
compressed_size: usize,
) -> usize {
/* The C version uses an uninitialized on-stack dctx_t; only `base` is
* read before being written, so a zeroed context is equivalent. */
let mut ctx = std::mem::MaybeUninit::<ZSTDv01_Dctx>::zeroed();
let ctx_ptr = ctx.as_mut_ptr();
(*ctx_ptr).base = dst as *const u8;
zstdv01_decompress_dctx(
ctx_ptr,
dst as *mut u8,
max_original_size,
src as *const u8,
compressed_size,
)
}
/// C ABI: `ZSTDv01_findFrameSizeInfoLegacy`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_findFrameSizeInfoLegacy(
src: *const c_void,
src_size: usize,
c_size: *mut usize,
d_bound: *mut u64,
) {
let mut ip = src as *const u8;
let mut remaining_size = src_size;
let mut nb_blocks: usize = 0;
/* Frame Header */
if src_size < ZSTD_FRAME_HEADER_SIZE + ZSTD_BLOCK_HEADER_SIZE {
zstd_error_frame_size_info_legacy(c_size, d_bound, ERROR(ZstdErrorCode::SrcSizeWrong));
return;
}
let magic_number = zstd_read_be32(src as *const u8);
if magic_number != ZSTD_MAGIC_NUMBER {
zstd_error_frame_size_info_legacy(c_size, d_bound, ERROR(ZstdErrorCode::PrefixUnknown));
return;
}
ip = ip.add(ZSTD_FRAME_HEADER_SIZE);
remaining_size -= ZSTD_FRAME_HEADER_SIZE;
/* Loop on each block */
loop {
let mut block_properties = BlockProperties {
block_type: 0,
orig_size: 0,
};
let block_size = zstdv01_getc_block_size(ip, remaining_size, &mut block_properties);
if ERR_isError(block_size) {
zstd_error_frame_size_info_legacy(c_size, d_bound, block_size);
return;
}
ip = ip.add(ZSTD_BLOCK_HEADER_SIZE);
remaining_size -= ZSTD_BLOCK_HEADER_SIZE;
if block_size > remaining_size {
zstd_error_frame_size_info_legacy(c_size, d_bound, ERROR(ZstdErrorCode::SrcSizeWrong));
return;
}
if block_size == 0 {
break; /* bt_end */
}
ip = ip.add(block_size);
remaining_size -= block_size;
nb_blocks += 1;
}
*c_size = (ip as usize) - (src as usize);
*d_bound = (nb_blocks * BLOCKSIZE) as u64;
}
/* ******************************************
* Streaming Decompression API
********************************************/
/// C ABI: `ZSTDv01_resetDCtx`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_resetDCtx(dctx: *mut ZSTDv01_Dctx) -> usize {
(*dctx).expected = ZSTD_FRAME_HEADER_SIZE;
(*dctx).phase = 0;
(*dctx).previous_dst_end = ptr::null();
(*dctx).base = ptr::null();
0
}
/// C ABI: `ZSTDv01_createDCtx`. Allocated with `malloc` exactly like the C
/// implementation, so create/free may be paired across the C/Rust boundary.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_createDCtx() -> *mut ZSTDv01_Dctx {
let dctx = libc::malloc(std::mem::size_of::<ZSTDv01_Dctx>()) as *mut ZSTDv01_Dctx;
if dctx.is_null() {
return ptr::null_mut();
}
ZSTDv01_resetDCtx(dctx);
dctx
}
/// C ABI: `ZSTDv01_freeDCtx`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_freeDCtx(dctx: *mut ZSTDv01_Dctx) -> usize {
libc::free(dctx as *mut c_void);
0
}
/// C ABI: `ZSTDv01_nextSrcSizeToDecompress`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_nextSrcSizeToDecompress(dctx: *mut ZSTDv01_Dctx) -> usize {
(*dctx).expected
}
/// C ABI: `ZSTDv01_decompressContinue`.
#[no_mangle]
pub unsafe extern "C" fn ZSTDv01_decompressContinue(
dctx: *mut ZSTDv01_Dctx,
dst: *mut c_void,
max_dst_size: usize,
src: *const c_void,
src_size: usize,
) -> usize {
let ctx = dctx;
let src = src as *const u8;
let dst = dst as *mut u8;
/* Sanity check */
if src_size != (*ctx).expected {
return ERROR(ZstdErrorCode::SrcSizeWrong);
}
if !std::ptr::eq(dst as *const u8, (*ctx).previous_dst_end) {
/* not contiguous */
(*ctx).base = dst as *const u8;
}
/* Decompress : frame header */
if (*ctx).phase == 0 {
/* Check frame magic header */
let magic_number = zstd_read_be32(src);
if magic_number != ZSTD_MAGIC_NUMBER {
return ERROR(ZstdErrorCode::PrefixUnknown);
}
(*ctx).phase = 1;
(*ctx).expected = ZSTD_BLOCK_HEADER_SIZE;
return 0;
}
/* Decompress : block header */
if (*ctx).phase == 1 {
let mut bp = BlockProperties {
block_type: 0,
orig_size: 0,
};
let block_size = zstdv01_getc_block_size(src, ZSTD_BLOCK_HEADER_SIZE, &mut bp);
if ERR_isError(block_size) {
return block_size;
}
if bp.block_type == BT_END {
(*ctx).expected = 0;
(*ctx).phase = 0;
} else {
(*ctx).expected = block_size;
(*ctx).b_type = bp.block_type;
(*ctx).phase = 2;
}
return 0;
}
/* Decompress : block content */
{
let r_size = match (*ctx).b_type {
BT_COMPRESSED => zstd_decompress_block(ctx, dst, max_dst_size, src, src_size),
BT_RAW => zstd_copy_uncompressed_block(dst, max_dst_size, src, src_size),
BT_RLE => {
return ERROR(ZstdErrorCode::Generic); /* not yet handled */
}
BT_END => {
/* should never happen (filtered at phase 1) */
0
}
_ => {
return ERROR(ZstdErrorCode::Generic);
}
};
(*ctx).phase = 1;
(*ctx).expected = ZSTD_BLOCK_HEADER_SIZE;
if ERR_isError(r_size) {
return r_size;
}
(*ctx).previous_dst_end = (dst as usize + r_size) as *const u8;
r_size
}
}
#[cfg(test)]
mod tests {
use super::*;
/// 237-byte base text; the fixtures compress repetitions of it.
const SAMPLE: &str = "snowden is snowed in / he's now then in his snow den / when does the snow end?\ngoodbye little dog / you dug some holes in your day / they'll be hard to fill.\nwhen life shuts a door, / just open it. it's a door. / that is how doors work.\n";
/// `SAMPLE` repeated 3 times (711 bytes), compressed by a zstd binary
/// built from the v0.1.0 tag. Single compressed block (FSE + Huff0),
/// verified byte-identical against the pristine C decoder.
const FRAME_ENTROPY: &[u8] = &[
0xFD, 0x2F, 0xB5, 0x1E, 0x00, 0x00, 0xB0, 0x00, 0x00, 0x93, 0x00, 0xD6, 0x19, 0x80, 0x87,
0xB1, 0x00, 0xC0, 0x73, 0xC5, 0x8A, 0xA5, 0x6C, 0x6F, 0x4B, 0xF2, 0x9F, 0x94, 0xC4, 0x2B,
0xE4, 0x6B, 0x2C, 0x96, 0xAE, 0x5F, 0xC8, 0x9C, 0x1F, 0x00, 0x1C, 0x00, 0x1B, 0x00, 0x43,
0x31, 0x70, 0xE4, 0xC9, 0xA3, 0x46, 0xD9, 0xAD, 0xFD, 0x78, 0x37, 0x66, 0xA6, 0x4D, 0x15,
0xCA, 0x33, 0xF5, 0xA7, 0x31, 0x83, 0x02, 0x96, 0x7E, 0xD0, 0x5E, 0xAC, 0xE9, 0x48, 0x03,
0xC7, 0xEF, 0x44, 0x81, 0x0B, 0x00, 0x90, 0xEA, 0x55, 0x41, 0xD8, 0x4E, 0x36, 0x10, 0xAC,
0x95, 0x92, 0x42, 0x5D, 0x92, 0xF4, 0x64, 0x7F, 0x96, 0x44, 0x59, 0x0B, 0x03, 0x4A, 0xF6,
0x16, 0x9F, 0xC8, 0xB9, 0x11, 0x70, 0xB8, 0x35, 0x06, 0xC4, 0x26, 0xC7, 0x6A, 0x89, 0xE4,
0x6F, 0x65, 0xC0, 0x9B, 0x55, 0x08, 0xEB, 0xE8, 0x37, 0x66, 0x67, 0x9E, 0x7D, 0x57, 0xA0,
0x25, 0x0D, 0xA4, 0x3E, 0x26, 0x8F, 0x32, 0xF6, 0xC9, 0x49, 0x50, 0x40, 0x0A, 0xED, 0x34,
0x94, 0x15, 0x5E, 0xE3, 0x1B, 0xD6, 0x27, 0x05, 0x00, 0x54, 0x05, 0x41, 0xFF, 0xD6, 0x01,
0x00, 0x80, 0x36, 0x00, 0xFE, 0x8D, 0x38, 0x42, 0x51, 0x7E, 0x40, 0x38, 0xD5, 0x60, 0x46,
0x03, 0x90, 0x25, 0xC0, 0x00, 0x00,
];
/// 48 random bytes compressed by v0.1.0: a raw (uncompressed) block.
const FRAME_RAW: &[u8] = &[
0xFD, 0x2F, 0xB5, 0x1E, 0x40, 0x00, 0x30, 0x77, 0xE9, 0x9A, 0x70, 0x2F, 0x77, 0x16, 0x9A,
0x92, 0xD4, 0xEE, 0x72, 0xFD, 0xDD, 0x0F, 0x39, 0x33, 0x8B, 0x8B, 0x3C, 0x6A, 0x91, 0xB2,
0x2B, 0x3A, 0xF0, 0x0E, 0xA6, 0x4C, 0x51, 0xA0, 0xC9, 0x5F, 0x69, 0x1A, 0xCE, 0x93, 0x5C,
0x7F, 0x43, 0x0C, 0xD9, 0x25, 0xBC, 0x91, 0xDD, 0x6E, 0xD7, 0xC0, 0x00, 0x00,
];
const RAW_PAYLOAD: &[u8] = &[
0x77, 0xE9, 0x9A, 0x70, 0x2F, 0x77, 0x16, 0x9A, 0x92, 0xD4, 0xEE, 0x72, 0xFD, 0xDD, 0x0F,
0x39, 0x33, 0x8B, 0x8B, 0x3C, 0x6A, 0x91, 0xB2, 0x2B, 0x3A, 0xF0, 0x0E, 0xA6, 0x4C, 0x51,
0xA0, 0xC9, 0x5F, 0x69, 0x1A, 0xCE, 0x93, 0x5C, 0x7F, 0x43, 0x0C, 0xD9, 0x25, 0xBC, 0x91,
0xDD, 0x6E, 0xD7,
];
/// `SAMPLE` repeated 1800 times (426600 bytes) compressed by v0.1.0:
/// four data blocks, so it exercises multi-block frames and repeated
/// offsets.
const FRAME_MULTI_BLOCK: &[u8] = &[
0xFD, 0x2F, 0xB5, 0x1E, 0x00, 0x00, 0xB0, 0x00, 0x00, 0x93, 0x00, 0xD6, 0x19, 0x80, 0x87,
0xB1, 0x00, 0xC0, 0x73, 0xC5, 0x8A, 0xA5, 0x6C, 0x6F, 0x4B, 0xF2, 0x9F, 0x94, 0xC4, 0x2B,
0xE4, 0x6B, 0x2C, 0x96, 0xAE, 0x5F, 0xC8, 0x9C, 0x1F, 0x00, 0x1C, 0x00, 0x1B, 0x00, 0x43,
0x31, 0x70, 0xE4, 0xC9, 0xA3, 0x46, 0xD9, 0xAD, 0xFD, 0x78, 0x37, 0x66, 0xA6, 0x4D, 0x15,
0xCA, 0x33, 0xF5, 0xA7, 0x31, 0x83, 0x02, 0x96, 0x7E, 0xD0, 0x5E, 0xAC, 0xE9, 0x48, 0x03,
0xC7, 0xEF, 0x44, 0x81, 0x0B, 0x00, 0x90, 0xEA, 0x55, 0x41, 0xD8, 0x4E, 0x36, 0x10, 0xAC,
0x95, 0x92, 0x42, 0x5D, 0x92, 0xF4, 0x64, 0x7F, 0x96, 0x44, 0x59, 0x0B, 0x03, 0x4A, 0xF6,
0x16, 0x9F, 0xC8, 0xB9, 0x11, 0x70, 0xB8, 0x35, 0x06, 0xC4, 0x26, 0xC7, 0x6A, 0x89, 0xE4,
0x6F, 0x65, 0xC0, 0x9B, 0x55, 0x08, 0xEB, 0xE8, 0x37, 0x66, 0x67, 0x9E, 0x7D, 0x57, 0xA0,
0x25, 0x0D, 0xA4, 0x3E, 0x26, 0x8F, 0x32, 0xF6, 0xC9, 0x49, 0x50, 0x40, 0x0A, 0xED, 0x34,
0x94, 0x15, 0x5E, 0xE3, 0x1B, 0xD6, 0x27, 0x05, 0x00, 0x54, 0x05, 0x41, 0xFF, 0x0F, 0xFF,
0x01, 0x80, 0x36, 0x00, 0xFE, 0x8D, 0x38, 0x42, 0x51, 0x7E, 0x40, 0x38, 0xD5, 0x60, 0x46,
0x03, 0x90, 0x25, 0x00, 0x00, 0x12, 0x40, 0x00, 0x00, 0x01, 0x00, 0x54, 0x04, 0xFF, 0xFC,
0xFF, 0x01, 0x80, 0xFA, 0x7F, 0x00, 0xFC, 0x23, 0x10, 0x00, 0x00, 0x12, 0x40, 0x00, 0x00,
0x01, 0x00, 0x54, 0x04, 0xFF, 0xFC, 0xFF, 0x01, 0x00, 0xF5, 0xFF, 0x00, 0xF8, 0x4B, 0x20,
0x00, 0x00, 0x12, 0x40, 0x00, 0x00, 0x01, 0x00, 0x54, 0x04, 0xFF, 0x64, 0x82, 0x00, 0x80,
0xEF, 0xFF, 0x00, 0xF0, 0x9F, 0x40, 0xC0, 0x00, 0x00,
];
fn decompress(frame: &[u8], capacity: usize) -> Result<Vec<u8>, usize> {
let mut out = vec![0u8; capacity];
let code = unsafe {
ZSTDv01_decompress(
out.as_mut_ptr() as *mut c_void,
capacity,
frame.as_ptr() as *const c_void,
frame.len(),
)
};
if ZSTDv01_isError(code) != 0 {
return Err(code);
}
out.truncate(code);
Ok(out)
}
fn frame_size_info(frame: &[u8]) -> (usize, u64) {
let mut c_size = 0usize;
let mut d_bound = 0u64;
unsafe {
ZSTDv01_findFrameSizeInfoLegacy(
frame.as_ptr() as *const c_void,
frame.len(),
&mut c_size,
&mut d_bound,
);
}
(c_size, d_bound)
}
#[test]
fn decodes_entropy_frame_byte_identically() {
let expected = SAMPLE.repeat(3).into_bytes();
let out = decompress(FRAME_ENTROPY, expected.len()).unwrap();
assert_eq!(out, expected);
}
#[test]
fn decodes_raw_block_frame_byte_identically() {
let out = decompress(FRAME_RAW, RAW_PAYLOAD.len()).unwrap();
assert_eq!(out, RAW_PAYLOAD);
}
#[test]
fn decodes_multi_block_frame_byte_identically() {
let expected = SAMPLE.repeat(1800).into_bytes();
assert_eq!(expected.len(), 426600);
let out = decompress(FRAME_MULTI_BLOCK, expected.len()).unwrap();
assert_eq!(out, expected);
}
#[test]
fn reports_frame_size_info_like_c() {
/* Values verified against the pristine C implementation. */
assert_eq!(
frame_size_info(FRAME_ENTROPY),
(FRAME_ENTROPY.len(), 131072)
);
assert_eq!(frame_size_info(FRAME_RAW), (FRAME_RAW.len(), 131072));
assert_eq!(
frame_size_info(FRAME_MULTI_BLOCK),
(FRAME_MULTI_BLOCK.len(), 524288)
);
}
#[test]
fn rejects_bad_magic_with_prefix_unknown() {
let mut frame = FRAME_ENTROPY.to_vec();
frame[0] ^= 0x55;
assert_eq!(
decompress(&frame, 1024).unwrap_err(),
ERROR(ZstdErrorCode::PrefixUnknown)
);
let (c_size, d_bound) = frame_size_info(&frame);
assert_eq!(c_size, ERROR(ZstdErrorCode::PrefixUnknown));
assert_eq!(d_bound, ZSTD_CONTENTSIZE_ERROR);
}
#[test]
fn rejects_truncated_frames_with_src_size_wrong() {
for frame in [FRAME_ENTROPY, FRAME_RAW, FRAME_MULTI_BLOCK] {
let truncated = &frame[..frame.len() - 1];
assert_eq!(
decompress(truncated, 1 << 20).unwrap_err(),
ERROR(ZstdErrorCode::SrcSizeWrong)
);
assert_eq!(
decompress(&frame[..5], 1 << 20).unwrap_err(),
ERROR(ZstdErrorCode::SrcSizeWrong)
);
}
}
#[test]
fn rejects_small_destination_with_dst_size_too_small() {
/* Error codes verified against the pristine C implementation. */
let expected_len = SAMPLE.len() * 3;
assert_eq!(
decompress(FRAME_ENTROPY, expected_len - 1).unwrap_err(),
ERROR(ZstdErrorCode::DstSizeTooSmall)
);
assert_eq!(
decompress(FRAME_ENTROPY, 0).unwrap_err(),
ERROR(ZstdErrorCode::DstSizeTooSmall)
);
assert_eq!(
decompress(FRAME_RAW, RAW_PAYLOAD.len() - 1).unwrap_err(),
ERROR(ZstdErrorCode::DstSizeTooSmall)
);
}
#[test]
fn streaming_api_decodes_the_frame() {
let expected = SAMPLE.repeat(3).into_bytes();
let mut out = vec![0u8; expected.len()];
let mut in_pos = 0usize;
let mut out_pos = 0usize;
unsafe {
let dctx = ZSTDv01_createDCtx();
assert!(!dctx.is_null());
loop {
let needed = ZSTDv01_nextSrcSizeToDecompress(dctx);
if needed == 0 {
break;
}
assert!(in_pos + needed <= FRAME_ENTROPY.len());
let produced = ZSTDv01_decompressContinue(
dctx,
out.as_mut_ptr().add(out_pos) as *mut c_void,
out.len() - out_pos,
FRAME_ENTROPY.as_ptr().add(in_pos) as *const c_void,
needed,
);
assert_eq!(ZSTDv01_isError(produced), 0);
in_pos += needed;
out_pos += produced;
}
assert_eq!(ZSTDv01_freeDCtx(dctx), 0);
}
assert_eq!(in_pos, FRAME_ENTROPY.len());
assert_eq!(out_pos, expected.len());
assert_eq!(out, expected);
}
#[test]
fn streaming_context_reset_and_null_free() {
unsafe {
let dctx = ZSTDv01_createDCtx();
assert_eq!(ZSTDv01_nextSrcSizeToDecompress(dctx), 4);
/* bad magic through the streaming entry point */
let bad = [0u8; 4];
assert_eq!(
ZSTDv01_decompressContinue(
dctx,
ptr::null_mut(),
0,
bad.as_ptr() as *const c_void,
4
),
ERROR(ZstdErrorCode::PrefixUnknown)
);
/* wrong srcSize */
assert_eq!(
ZSTDv01_decompressContinue(
dctx,
ptr::null_mut(),
0,
bad.as_ptr() as *const c_void,
3
),
ERROR(ZstdErrorCode::SrcSizeWrong)
);
assert_eq!(ZSTDv01_resetDCtx(dctx), 0);
assert_eq!(ZSTDv01_nextSrcSizeToDecompress(dctx), 4);
assert_eq!(ZSTDv01_freeDCtx(dctx), 0);
assert_eq!(ZSTDv01_freeDCtx(ptr::null_mut()), 0);
}
}
}
+3
View File
@@ -5,6 +5,8 @@ pub mod bitstream;
pub mod common;
pub mod cpu;
pub mod debug;
#[cfg(feature = "dict-builder")]
pub mod divsufsort;
pub mod entropy_common;
pub mod errors;
#[cfg(feature = "compression")]
@@ -16,6 +18,7 @@ pub mod hist;
pub mod huf_compress;
#[cfg(feature = "decompression")]
pub mod huf_decompress;
pub mod legacy;
pub mod mem;
pub mod pool;
pub mod threading;
+198
View File
@@ -0,0 +1,198 @@
#![allow(non_camel_case_types)]
#![allow(non_snake_case)]
//! Precise monotonic time measurement for the command-line programs.
//!
//! Port of `programs/timefn.c`. `UTIL_time_t` is a plain nanosecond counter
//! whose absolute value is meaningless; only spans between two measurements
//! are valid. The struct crosses the C ABI by value, so it stays `repr(C)`
//! with the exact `timefn.h` layout.
//!
//! Platform selection mirrors the C preprocessor structure: Windows uses the
//! performance counter, Apple systems use the Mach absolute clock, and other
//! POSIX systems use `clock_gettime(CLOCK_MONOTONIC)`. The C90 `clock()`
//! fallback is never needed on targets Rust supports, so multi-threaded
//! measurements are always supported.
use std::os::raw::c_int;
/// Precise Time (`PTime` in timefn.h): an unsigned 64-bit nanosecond count.
pub type PTime = u64;
/// Nanosecond time counter with the `timefn.h` `UTIL_time_t` layout.
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct UTIL_time_t {
pub t: PTime,
}
const _: () = assert!(std::mem::size_of::<PTime>() == 8);
const _: () = assert!(std::mem::size_of::<UTIL_time_t>() == std::mem::size_of::<PTime>());
#[cfg(windows)]
mod platform {
use super::PTime;
use std::sync::OnceLock;
#[link(name = "kernel32")]
unsafe extern "system" {
/// Takes a `LARGE_INTEGER*`; the union is ABI-identical to `i64*`.
fn QueryPerformanceCounter(count: *mut i64) -> i32;
fn QueryPerformanceFrequency(frequency: *mut i64) -> i32;
}
pub fn monotonic_ns() -> PTime {
static TICKS_PER_SECOND: OnceLock<i64> = OnceLock::new();
let ticks_per_second = *TICKS_PER_SECOND.get_or_init(|| {
let mut frequency = 0i64;
if unsafe { QueryPerformanceFrequency(&mut frequency) } == 0 {
eprintln!(
"timefn::QueryPerformanceFrequency: {}",
std::io::Error::last_os_error()
);
std::process::abort();
}
frequency
});
let mut counter = 0i64;
unsafe { QueryPerformanceCounter(&mut counter) };
(counter as PTime).wrapping_mul(1_000_000_000) / ticks_per_second as PTime
}
}
#[cfg(all(unix, target_vendor = "apple"))]
mod platform {
use super::PTime;
use std::sync::OnceLock;
pub fn monotonic_ns() -> PTime {
static RATE: OnceLock<(PTime, PTime)> = OnceLock::new();
let (numer, denom) = *RATE.get_or_init(|| {
let mut rate = libc::mach_timebase_info { numer: 0, denom: 0 };
unsafe { libc::mach_timebase_info(&mut rate) };
(PTime::from(rate.numer), PTime::from(rate.denom))
});
unsafe { libc::mach_absolute_time() }.wrapping_mul(numer) / denom
}
}
#[cfg(all(unix, not(target_vendor = "apple")))]
mod platform {
use super::PTime;
pub fn monotonic_ns() -> PTime {
// Zero-initialized like the C source, which works around timespec_get
// msan limitations on some targets.
let mut time: libc::timespec = unsafe { std::mem::zeroed() };
if unsafe { libc::clock_gettime(libc::CLOCK_MONOTONIC, &mut time) } != 0 {
eprintln!(
"timefn::clock_gettime(CLOCK_MONOTONIC): {}",
std::io::Error::last_os_error()
);
std::process::abort();
}
(time.tv_sec as PTime)
.wrapping_mul(1_000_000_000)
.wrapping_add(time.tv_nsec as PTime)
}
}
/// Returns the current value of the platform's monotonic nanosecond clock.
#[no_mangle]
pub extern "C" fn UTIL_getTime() -> UTIL_time_t {
UTIL_time_t {
t: platform::monotonic_ns(),
}
}
/// Nanoseconds elapsed between two measurements, with C unsigned wrap-around.
#[no_mangle]
pub extern "C" fn UTIL_getSpanTimeNano(clockStart: UTIL_time_t, clockEnd: UTIL_time_t) -> PTime {
clockEnd.t.wrapping_sub(clockStart.t)
}
/// Microseconds elapsed between two measurements, truncated like C division.
#[no_mangle]
pub extern "C" fn UTIL_getSpanTimeMicro(begin: UTIL_time_t, end: UTIL_time_t) -> PTime {
UTIL_getSpanTimeNano(begin, end) / 1000
}
/// Microseconds elapsed since `clockStart`.
#[no_mangle]
pub extern "C" fn UTIL_clockSpanMicro(clockStart: UTIL_time_t) -> PTime {
UTIL_getSpanTimeMicro(clockStart, UTIL_getTime())
}
/// Nanoseconds elapsed since `clockStart`.
#[no_mangle]
pub extern "C" fn UTIL_clockSpanNano(clockStart: UTIL_time_t) -> PTime {
UTIL_getSpanTimeNano(clockStart, UTIL_getTime())
}
/// Busy-waits until the clock produces a new tick, improving measurement
/// accuracy on platforms with a low timer resolution.
#[no_mangle]
pub extern "C" fn UTIL_waitForNextTick() {
let clockStart = UTIL_getTime();
loop {
let clockEnd = UTIL_getTime();
if UTIL_getSpanTimeNano(clockStart, clockEnd) != 0 {
return;
}
}
}
/// All clock sources used by the Rust port are valid under multi-threaded
/// workloads; only the C90 `clock()` fallback of the C source was not.
#[no_mangle]
pub extern "C" fn UTIL_support_MT_measurements() -> c_int {
1
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn nanosecond_spans_subtract_with_unsigned_wrap_around() {
let start = UTIL_time_t { t: 100 };
let end = UTIL_time_t { t: 350 };
assert_eq!(UTIL_getSpanTimeNano(start, end), 250);
assert_eq!(UTIL_getSpanTimeNano(end, start), u64::MAX - 249);
assert_eq!(UTIL_getSpanTimeNano(start, start), 0);
}
#[test]
fn microsecond_spans_truncate_sub_tick_remainders() {
let start = UTIL_time_t { t: 0 };
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 999 }), 0);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 1_000 }), 1);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 1_999 }), 1);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 2_000 }), 2);
}
#[test]
fn clock_is_monotonic_across_measurements() {
let first = UTIL_getTime();
let second = UTIL_getTime();
assert!(second.t >= first.t);
assert!(UTIL_clockSpanNano(first) >= UTIL_getSpanTimeNano(first, second));
}
#[test]
fn waiting_for_the_next_tick_advances_the_clock() {
let before = UTIL_getTime();
UTIL_waitForNextTick();
let after = UTIL_getTime();
assert!(UTIL_getSpanTimeNano(before, after) > 0);
}
#[test]
fn multi_threaded_measurements_are_supported() {
assert_eq!(UTIL_support_MT_measurements(), 1);
}
}
+147 -5
View File
@@ -10,9 +10,15 @@
//! writes, dictionary loading, streaming, and metadata preservation remain in
//! `programs/fileio.c` for this first migration step.
//!
//! Benchmark mode (`-b`) parses here and dispatches through the
//! `ZSTD_rust_cli_bench` bridge in `programs/zstdcli.c`: the run/timing loop
//! (benchfn, timefn) is Rust, while orchestration and result formatting
//! (`benchzstd.c`) remain C behind the preprocessor-gated bridge, so builds
//! with `ZSTD_NOBENCH` never reference benchmark symbols.
//!
//! Remaining C-only CLI boundaries are called out in `unsupported()` below:
//! benchmark execution, dictionary training, recursive/file-list expansion,
//! tracing, alternate-format selection, and the advanced directory modes.
//! dictionary training, recursive/file-list expansion, tracing,
//! alternate-format selection, and the advanced directory modes.
use std::env;
use std::ffi::{CStr, CString, OsStr, OsString};
@@ -30,6 +36,7 @@ use std::os::unix::fs::FileTypeExt;
const DEFAULT_CLEVEL: i32 = 3;
#[cfg(feature = "compression")]
const DEFAULT_MAX_CLEVEL: i32 = 19;
const DEFAULT_BENCH_NB_SECONDS: u32 = 3;
const DEFAULT_MEM_LIMIT: u32 = 1 << 27;
const DEFAULT_LONG_WINDOW_LOG: u32 = 27;
const MAX_FAST_ACCELERATION: i32 = 128 << 10;
@@ -173,6 +180,22 @@ unsafe extern "C" {
output: *const c_char,
dict: *const c_char,
) -> c_int;
/// Benchmark bridge implemented by the `programs/zstdcli.c` shim, which
/// owns the `ZSTD_NOBENCH` preprocessor decision. Returns the benchmark
/// result (>= 0), or -1 when benchmarking is compiled out.
fn ZSTD_rust_cli_bench(
file_names: *const *const c_char,
nb_files: c_uint,
dict_file_name: *const c_char,
start_level: c_int,
end_level: c_int,
compression_params: *const ZSTD_compressionParameters,
display_level: c_int,
nb_seconds: c_uint,
block_size: usize,
nb_workers: c_int,
) -> c_int;
}
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
@@ -180,6 +203,7 @@ enum Operation {
Compress,
Decompress,
Test,
Bench,
}
#[derive(Debug)]
@@ -230,6 +254,8 @@ struct Cli {
row_match_finder: i32,
exclude_compressed: bool,
compression_params: ZSTD_compressionParameters,
bench_end_level: Option<i32>,
bench_nb_seconds: Option<u32>,
unsupported_program: Option<String>,
}
@@ -275,6 +301,8 @@ impl Cli {
row_match_finder: ZSTD_PS_AUTO,
exclude_compressed: false,
compression_params: ZSTD_compressionParameters::default(),
bench_end_level: None,
bench_nb_seconds: None,
unsupported_program: None,
};
@@ -404,9 +432,22 @@ fn usage(advanced: bool) {
out,
" --adapt[=min=#,max=#], --rsyncable, --[no-]row-match-finder"
);
let _ = writeln!(out, "\nBenchmark options:");
let _ = writeln!(
out,
"\nNot yet migrated: benchmark, dictionary training, recursive/file-list expansion,"
" -b# Benchmark file(s) at compression level #"
);
let _ = writeln!(
out,
" -e# Test all levels from -b# up to # included"
);
let _ = writeln!(
out,
" -i# Set the minimum evaluation time to # seconds"
);
let _ = writeln!(
out,
"\nNot yet migrated: dictionary training, recursive/file-list expansion,"
);
let _ = writeln!(out, "trace, alternate formats, and output-directory modes.");
}
@@ -905,6 +946,33 @@ fn parse_short_options(
'd' => cli.operation = Operation::Decompress,
'z' => cli.operation = Operation::Compress,
't' => cli.operation = Operation::Test,
'b' => cli.operation = Operation::Bench,
'e' | 'i' => {
// Benchmark range end (-e#) and duration (-i#): like the C
// parser, digits attach directly and default to 0.
let mut digits_end = offset + 1;
while digits_end < bytes.len() && bytes[digits_end].is_ascii_digit() {
digits_end += 1;
}
let digits = &value[offset + 1..digits_end];
if option == 'e' {
cli.bench_end_level = Some(if digits.is_empty() {
0
} else {
parse_i32(digits, "benchmark end level")?
});
} else {
cli.bench_nb_seconds = Some(if digits.is_empty() {
0
} else {
digits
.parse::<u32>()
.map_err(|_| format!("invalid benchmark duration: {digits:?}"))?
});
}
offset = digits_end;
continue;
}
'c' => {
cli.output = Some(cstring(STDOUT_MARK)?);
cli.force_stdout = true;
@@ -940,7 +1008,7 @@ fn parse_short_options(
}
break;
}
'b' | 'e' | 'i' | 'l' | 'p' | 'P' | 'r' | 's' | 'S' => {
'l' | 'p' | 'P' | 'r' | 's' | 'S' => {
unsupported(&format!("-{option}"))?;
}
_ => return Err(format!("unknown option -{option}")),
@@ -1231,16 +1299,52 @@ unsafe fn run_decompress(
dictionary,
)
},
Operation::Compress => unreachable!("compression is dispatched separately"),
Operation::Compress | Operation::Bench => {
unreachable!("compression and benchmark are dispatched separately")
}
}
}
/// Runs benchmark mode through the C bridge. Level clamping against
/// `ZSTD_maxCLevel()` happens on the C side, where the symbol is always
/// available when benchmarking is compiled in. No input file means a
/// synthetic-sample benchmark, matching the C CLI.
fn run_bench(cli: &Cli) -> Result<i32, String> {
let inputs: Vec<*const c_char> = cli.inputs.iter().map(|value| value.as_ptr()).collect();
let dictionary = cli
.dictionary
.as_ref()
.map_or(ptr::null(), |value| value.as_ptr());
let result = unsafe {
ZSTD_rust_cli_bench(
inputs.as_ptr(),
inputs.len() as c_uint,
dictionary,
cli.level,
cli.bench_end_level.unwrap_or(cli.level),
&cli.compression_params,
cli.display_level,
cli.bench_nb_seconds.unwrap_or(DEFAULT_BENCH_NB_SECONDS),
cli.block_size.unwrap_or(0),
// The C CLI benchmarks single-threaded unless -T was given.
cli.workers.unwrap_or(1),
)
};
if result < 0 {
return Err("benchmark mode is not available in this build".to_owned());
}
Ok(result)
}
fn run_cli(mut cli: Cli) -> Result<i32, String> {
if let Some(program_name) = &cli.unsupported_program {
return Err(format!(
"{program_name} compatibility mode is not yet implemented by the Rust CLI frontend"
));
}
if cli.operation == Operation::Bench {
return run_bench(&cli);
}
let explicit_input_count = cli.inputs.len();
filter_symlink_inputs(&mut cli);
if explicit_input_count > 0 && cli.inputs.is_empty() {
@@ -1347,6 +1451,7 @@ fn run_cli(mut cli: Cli) -> Result<i32, String> {
#[cfg(not(feature = "decompression"))]
unreachable!("unsupported decompression was rejected above")
}
Operation::Bench => unreachable!("benchmark mode was dispatched earlier"),
}
};
@@ -1581,4 +1686,41 @@ mod tests {
assert!(error.contains("not yet implemented"));
}
#[test]
fn bench_mode_parses_level_duration_and_defaults() {
let cli = parse(&["zstd", "-b1", "-i0", "input"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, 1);
assert_eq!(cli.bench_nb_seconds, Some(0));
assert_eq!(cli.bench_end_level, None);
assert_eq!(
cli.inputs
.iter()
.map(|input| input.as_bytes())
.collect::<Vec<_>>(),
vec![&b"input"[..]]
);
}
#[test]
fn bench_range_aggregates_within_a_single_argument() {
let cli = parse(&["zstd", "-b5e6i2", "input"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, 5);
assert_eq!(cli.bench_end_level, Some(6));
assert_eq!(cli.bench_nb_seconds, Some(2));
}
#[test]
fn bench_duration_without_digits_defaults_to_zero() {
let cli = parse(&["zstd", "-b", "-e", "-i"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, DEFAULT_CLEVEL);
assert_eq!(cli.bench_end_level, Some(0));
assert_eq!(cli.bench_nb_seconds, Some(0));
}
}
+51 -7
View File
@@ -65,25 +65,36 @@ endif
endif
RUST_HUF_FEATURE :=
RUST_BUILD_CONFIG := default
RUST_HUF_MODE := default
ifneq ($(RUST_HUF_FORCE_X1),0)
RUST_HUF_FEATURE := huf-force-decompress-x1
RUST_BUILD_CONFIG := huf-force-decompress-x1
RUST_HUF_MODE := huf-force-decompress-x1
endif
ifneq ($(RUST_HUF_FORCE_X2),0)
RUST_HUF_FEATURE := huf-force-decompress-x2
RUST_BUILD_CONFIG := huf-force-decompress-x2
RUST_HUF_MODE := huf-force-decompress-x2
endif
# Test binaries compile every lib/legacy/*.c file regardless of the dispatch
# level (ZSTDLEGACY_FILES is a plain wildcard below), so the Rust archive must
# always carry all ported legacy decoders too. The build configuration still
# encodes the level because the flat C objects bake -DZSTD_LEGACY_SUPPORT into
# the dispatch code and must never outlive a level change.
RUST_LEGACY_FEATURES := legacy-v01,legacy-v02,legacy-v03,legacy-v04,legacy-v05,legacy-v06,legacy-v07
RUST_BUILD_CONFIG := $(RUST_HUF_MODE)-legacy$(ZSTD_LEGACY_SUPPORT)
RUST_TARGET_DIR := $(RUST_DIR)/target/$(RUST_BUILD_CONFIG)
RUST_STATICLIB := $(RUST_TARGET_DIR)/release/libzstd_rs.a
RUST_TARGET_32 ?= i686-unknown-linux-gnu
RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a
# Tests build every library module, so they use the crate's default feature
# set (compression, decompression, and dict-builder) plus any forced HUF mode.
RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_TARGET_DIR)
ifneq ($(RUST_HUF_FEATURE),)
RUST_CARGO_FLAGS += --features $(RUST_HUF_FEATURE)
endif
RUST_CARGO_FLAGS += --features $(RUST_LEGACY_FEATURES)
$(RUST_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_CARGO_FLAGS)
@@ -91,6 +102,28 @@ $(RUST_STATICLIB): $(RUST_SOURCES)
$(RUST_STATICLIB_32): $(RUST_SOURCES)
$(CARGO) build $(RUST_CARGO_FLAGS) --target $(RUST_TARGET_32)
# Program-only helpers that were C sources shared with the tests (timefn,
# benchfn) now live in the Rust CLI package. The tests link a helpers-only
# archive, built without the `cli` feature: the parser/dispatch layer needs
# the C fileio backend, which test binaries do not provide.
RUST_CLI_DIR := $(RUST_DIR)/cli
RUST_CLI_MANIFEST := $(RUST_CLI_DIR)/Cargo.toml
RUST_CLI_HELPER_SOURCES := $(RUST_CLI_MANIFEST) $(RUST_CLI_DIR)/Cargo.lock \
$(RUST_CLI_DIR)/src/lib.rs \
$(RUST_DIR)/src/timefn.rs $(RUST_DIR)/src/benchfn.rs
RUST_CLI_HELPERS_TARGET_DIR := $(RUST_DIR)/target/cli-helpers
RUST_CLI_HELPERS_STATICLIB := $(RUST_CLI_HELPERS_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_CLI_HELPERS_STATICLIB_32 := $(RUST_CLI_HELPERS_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_cli_rs.a
RUST_CLI_HELPERS_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_CLI_HELPERS_TARGET_DIR) \
--no-default-features
$(RUST_CLI_HELPERS_STATICLIB): $(RUST_CLI_HELPER_SOURCES)
$(CARGO) build $(RUST_CLI_HELPERS_CARGO_FLAGS)
$(RUST_CLI_HELPERS_STATICLIB_32): $(RUST_CLI_HELPER_SOURCES)
$(CARGO) build $(RUST_CLI_HELPERS_CARGO_FLAGS) --target $(RUST_TARGET_32)
# These test objects have flat filenames, unlike the configuration-hashed
# program objects. Track the HUF mode separately so a C object set compiled
# for one decoder is never relinked with a Rust archive for another decoder.
@@ -250,8 +283,8 @@ fuzzer32 : $(ZSTD_FILES)
$(LINK.c) $^ -o $@$(EXT)
# note : broken : requires symbols unavailable from dynamic library
fuzzer-dll : $(LIB_SRCDIR)/common/xxhash.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c $(PRGDIR)/datagen.c fuzzer.c
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(LDFLAGS) -o $@$(EXT)
fuzzer-dll : $(LIB_SRCDIR)/common/xxhash.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c $(PRGDIR)/datagen.c fuzzer.c $(RUST_CLI_HELPERS_STATICLIB)
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(RUST_CLI_HELPERS_STATICLIB) $(LDFLAGS) -o $@$(EXT)
CLEAN += zstreamtest zstreamtest32
ZSTREAM_LOCAL_FILES := $(PRGDIR)/datagen.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c seqgen.c zstreamtest.c external_matchfinder.c
@@ -282,8 +315,8 @@ zstreamtest_ubsan : $(ZSTREAMFILES)
# note : broken : requires symbols unavailable from dynamic library
zstreamtest-dll : $(LIB_SRCDIR)/common/xxhash.c # xxh symbols not exposed from dll
zstreamtest-dll : $(ZSTREAM_LOCAL_FILES)
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(LDFLAGS) -o $@$(EXT)
zstreamtest-dll : $(ZSTREAM_LOCAL_FILES) $(RUST_CLI_HELPERS_STATICLIB)
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(RUST_CLI_HELPERS_STATICLIB) $(LDFLAGS) -o $@$(EXT)
CLEAN += paramgrill
paramgrill : DEBUGFLAGS = # turn off debug for speed measurements
@@ -342,6 +375,17 @@ $(RUST_LINK_TARGETS_32): $(RUST_STATICLIB_32)
$(RUST_LINK_TARGETS) $(RUST_LINK_TARGETS_32): $(RUST_HUF_C_MODE_STAMP)
# Tests that compile the timefn/benchfn C shims also link the Rust CLI
# helpers archive, which owns those implementations. The archive is a
# prerequisite so `$^` places it after every C object referencing its symbols.
RUST_CLI_LINK_TARGETS := fullbench fullbench-lib fullbench-dll fuzzer \
zstreamtest zstreamtest_asan zstreamtest_tsan \
zstreamtest_ubsan paramgrill decodecorpus poolTests
$(RUST_CLI_LINK_TARGETS): $(RUST_CLI_HELPERS_STATICLIB)
RUST_CLI_LINK_TARGETS_32 := fullbench32 fuzzer32 zstreamtest32
$(RUST_CLI_LINK_TARGETS_32): $(RUST_CLI_HELPERS_STATICLIB_32)
.PHONY: versionsTest
versionsTest: clean
$(PYTHON) test-zstd-versions.py