Compare commits

..
Author SHA1 Message Date
ddidderr 24488e4eaa feat(rust): port legacy v0.1 decoder
Port lib/legacy/zstd_v01.c (the frozen zstd v0.1 decoder) to
rust/src/legacy/zstd_v01.rs as the first legacy-format port on the new
scaffolding, and reduce the C file to a declaration-only shim that
keeps its header includes for configuration and platform preprocessor
behavior.

Frozen-decoder policy: zstd_v01.c embeds its own v0.1-era FSE and
Huff0 snapshot, distinct from every other release. The Rust port is a
line-by-line translation with the same table layouts (FSE_DTable as a
u32 header word plus packed newState/symbol/nbBits entries, the Huff0
u16 DTable with byte/nbBits pairs), the same arithmetic including
wrap-around and pointer-comparison quirks (e.g. the offset-vs-base
address check in ZSTD_execSequence), the same internal FSE error space
(size_t)-1..-7, and the same public ZSTD error codes. It reuses no
modern Rust entropy module; its only crate dependency is `errors`,
matching the C file's error_private.h include. The 32-bit-only reload
points are kept as compile-time conditions on usize::BITS.

Symbol takeover boundary: all nine ZSTDv01_* entry points from
zstd_v01.h now come from Rust as context-free #[no_mangle] extern "C"
functions (isError, decompress, decompressDCtx,
findFrameSizeInfoLegacy, createDCtx, freeDCtx, resetDCtx,
nextSrcSizeToDecompress, decompressContinue). zstd_legacy.h only uses
the first four for v0.1; streaming for v0.1-v0.3 intentionally returns
version_unsupported there, unchanged. The ZSTDv01_Dctx struct
definition moves entirely into Rust: C code only ever holds an opaque
pointer (zstd_v01.h forward-declares the type), and the context is
malloc/free-allocated exactly like the C version so create/free may
pair across the language boundary.

Byte-identity verification against the pristine pre-migration C build
(f8745da6, pure C, ZSTD_LEGACY_SUPPORT=1):

- Real v0.1 frames were generated by building the v0.1.0 git tag and
  compressing text, random, and 426 KB multi-block inputs. A one-shot
  ZSTD_decompress harness linked once against the pristine C libzstd.a
  and once against the Rust-backed libzstd.a produced bit-identical
  outputs for all frames.
- A direct ZSTDv01_* probe (one-shot decode, dst-too-small, truncated
  input, bad magic, findFrameSizeInfoLegacy, and the streaming
  continue loop) printed identical results, including exact error
  codes (-70 dstSize_tooSmall, -72 srcSize_wrong, -10 prefix_unknown)
  and identical dBound values.
- zstd -l -v on v0.1 files matches the pristine binary; CLI streaming
  decode of v0.1 fails with the same "Version not supported" in both,
  by design of zstd_legacy.h.

Unit tests embed three v0.1.0-generated fixtures (entropy-coded,
raw-block, and four-block frames) plus the truncation, bad-magic,
small-destination, and streaming-API cases, all asserting the exact C
error codes above. Note that `make -C tests test-legacy` only covers
v0.4+ frames, so the embedded fixtures and the harness comparison are
the actual v0.1 coverage.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets
  --features legacy-v01 -- -D warnings && cargo test --all-targets
  --features legacy-v01 (127 tests, 9 for v0.1)
- cargo clippy/test --no-default-features --features
  decompression,legacy-v01 (module builds standalone)
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests, also with
  ZSTD_LEGACY_SUPPORT=1 (mixed Rust v0.1 + C v0.2-0.7 link)
- make -C tests test-rust-lib-smoke && make -C tests test-legacy
- make -C programs zstd (default and ZSTD_LEGACY_SUPPORT=1); nm shows
  the nine ZSTDv01_* symbols provided by Rust at level 1
- make -C lib libzstd.a ZSTD_LEGACY_SUPPORT=0 (no legacy symbols) and
  meson -Dlegacy_level=1 shared library exporting all nine
2026-07-11 14:25:23 +02:00
ddidderr c0a1b0bae1 build(rust): add legacy feature scaffolding
The legacy decoders (lib/legacy/zstd_v01.c .. zstd_v07.c) are next in
the Rust migration. Each of those files is a frozen snapshot of the
FSE/Huff0 entropy coders and frame logic of one historical release, so
their ports must not reuse the modern Rust entropy modules and must not
share code with each other: outputs and error codes have to stay
byte-identical to the frozen C forever. This commit installs the
build-system scaffolding so seven per-version ports can land
independently, each adding only its own module file plus a one-line
registration in rust/src/legacy/mod.rs.

Cargo grows features legacy-v01 .. legacy-v07. They are never default
features: the C build defaults differ per build system, so each build
system passes the list explicitly, derived from its own legacy
configuration:

- lib/Makefile and programs/Makefile map ZSTD_LEGACY_SUPPORT=N to the
  features for versions N..7 (0 disables legacy), mirroring the
  ZSTD_LEGACY_FILES selection in lib/libzstd.mk.
- tests/Makefile always enables all seven features because its
  ZSTDLEGACY_FILES wildcard compiles every lib/legacy/*.c regardless of
  the dispatch level.
- build/meson maps legacy_level exactly like the makefiles; build/cmake
  enables all seven whenever ZSTD_LEGACY_SUPPORT is ON because it
  always compiles all seven C files (ZSTD_LEGACY_LEVEL only selects the
  C dispatch).

Every build system also encodes the legacy selection in the Rust target
directory name (e.g. c1-d1-default-legacy5), for the same reason the
HUF mode is encoded there: a cached archive built for one configuration
must never be linked into a build expecting another. In tests/Makefile
the legacy level additionally flows into the existing HUF C-mode stamp,
so the flat C test objects (which bake -DZSTD_LEGACY_SUPPORT into the
dispatch) are rebuilt whenever the level changes. In programs/Makefile
the compress-only, decompress-only, and CLI archives keep
level-independent directories (RUST_HUF_MODE) because they are only
linked into ZSTD_LEGACY_SUPPORT=0 program variants and carry no legacy
features.

A feature whose version has not been ported yet gates nothing: the
module registration in rust/src/legacy/mod.rs is added by each port,
so enabling e.g. legacy-v05 today simply leaves that decoder in C.
This is what makes mixed C/Rust legacy levels link cleanly while the
seven ports land in any order.

Test plan:
- cd rust && cargo fmt --check && cargo clippy --all-targets
  -- -D warnings && cargo test --all-targets
- cargo clippy with --no-default-features --features
  decompression,legacy-v01 and with all seven legacy features
- make -C tests fuzzer && ./tests/fuzzer -i1 --no-big-tests
- make -C tests test-rust-lib-smoke; make -C tests test-legacy
- make -C lib libzstd.a with ZSTD_LEGACY_SUPPORT=0, 1 and default (5)
- cmake configure and meson setup (including -Dlegacy_level=1) emit the
  expected --features lists and legacy-suffixed target directories
2026-07-11 14:24:47 +02:00
22 changed files with 2590 additions and 5091 deletions
-1
View File
@@ -34,7 +34,6 @@ install/
# Build artefacts # Build artefacts
/rust/target/ /rust/target/
/rust/cli/target/
contrib/linux-kernel/linux/ contrib/linux-kernel/linux/
projects/ projects/
bin/ bin/
+1 -6
View File
@@ -142,7 +142,6 @@ endif()
set(_zstd_rust_features) set(_zstd_rust_features)
set(_zstd_rust_compression 0) set(_zstd_rust_compression 0)
set(_zstd_rust_decompression 0) set(_zstd_rust_decompression 0)
set(_zstd_rust_dictbuilder 0)
if(ZSTD_BUILD_COMPRESSION) if(ZSTD_BUILD_COMPRESSION)
list(APPEND _zstd_rust_features compression) list(APPEND _zstd_rust_features compression)
set(_zstd_rust_compression 1) set(_zstd_rust_compression 1)
@@ -151,10 +150,6 @@ if(ZSTD_BUILD_DECOMPRESSION)
list(APPEND _zstd_rust_features decompression) list(APPEND _zstd_rust_features decompression)
set(_zstd_rust_decompression 1) set(_zstd_rust_decompression 1)
endif() endif()
if(ZSTD_BUILD_DICTBUILDER)
list(APPEND _zstd_rust_features dict-builder)
set(_zstd_rust_dictbuilder 1)
endif()
set(_zstd_rust_huf_mode default) set(_zstd_rust_huf_mode default)
if(_zstd_huf_force_x1) if(_zstd_huf_force_x1)
@@ -207,7 +202,7 @@ if(_zstd_rust_features)
endif() endif()
set(_zstd_rust_build_config set(_zstd_rust_build_config
"c${_zstd_rust_compression}-d${_zstd_rust_decompression}-b${_zstd_rust_dictbuilder}-${_zstd_rust_huf_mode}-legacy${_zstd_rust_legacy}") "c${_zstd_rust_compression}-d${_zstd_rust_decompression}-${_zstd_rust_huf_mode}-legacy${_zstd_rust_legacy}")
set(ZSTD_RUST_MANIFEST "${ZSTD_SOURCE_DIR}/rust/Cargo.toml") set(ZSTD_RUST_MANIFEST "${ZSTD_SOURCE_DIR}/rust/Cargo.toml")
set(ZSTD_RUST_TARGET_DIR set(ZSTD_RUST_TARGET_DIR
"${CMAKE_CURRENT_BINARY_DIR}/rust-target/${_zstd_rust_build_config}") "${CMAKE_CURRENT_BINARY_DIR}/rust-target/${_zstd_rust_build_config}")
+2 -4
View File
@@ -73,9 +73,7 @@ if rust_huf_force_x1 and rust_huf_force_x2
error('HUF_FORCE_DECOMPRESS_X1 and HUF_FORCE_DECOMPRESS_X2 are mutually exclusive') error('HUF_FORCE_DECOMPRESS_X1 and HUF_FORCE_DECOMPRESS_X2 are mutually exclusive')
endif endif
# Meson always compiles the dictBuilder sources above, so the Rust archive rust_features = ['compression', 'decompression']
# must always carry the matching dict-builder module set.
rust_features = ['compression', 'decompression', 'dict-builder']
rust_huf_mode = 'default' rust_huf_mode = 'default'
rust_huf_c_args = [] rust_huf_c_args = []
if rust_huf_force_x1 if rust_huf_force_x1
@@ -108,7 +106,7 @@ if rust_target == ''
endif endif
endif endif
rust_build_config = 'c1-d1-b1-' + rust_huf_mode + '-legacy@0@'.format(legacy_level) rust_build_config = 'c1-d1-' + rust_huf_mode + '-legacy@0@'.format(legacy_level)
rust_target_dir = join_paths(meson.current_build_dir(), 'rust-target', rust_build_config) rust_target_dir = join_paths(meson.current_build_dir(), 'rust-target', rust_build_config)
is_msvc = cc_id == compiler_msvc or cc_id == 'clang-cl' is_msvc = cc_id == compiler_msvc or cc_id == 'clang-cl'
rust_staticlib_name = is_msvc ? 'zstd_rs.lib' : 'libzstd_rs.a' rust_staticlib_name = is_msvc ? 'zstd_rs.lib' : 'libzstd_rs.a'
+3 -7
View File
@@ -84,20 +84,16 @@ endif
ifneq ($(ZSTD_LIB_DECOMPRESSION),0) ifneq ($(ZSTD_LIB_DECOMPRESSION),0)
RUST_CARGO_FEATURES += decompression RUST_CARGO_FEATURES += decompression
endif endif
ifneq ($(ZSTD_LIB_DICTBUILDER),0)
RUST_CARGO_FEATURES += dict-builder
endif
RUST_MODULE_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-b$(ZSTD_LIB_DICTBUILDER)
RUST_HUF_FEATURE := RUST_HUF_FEATURE :=
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-default RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-default
ifneq ($(RUST_HUF_FORCE_X1),0) ifneq ($(RUST_HUF_FORCE_X1),0)
RUST_HUF_FEATURE := huf-force-decompress-x1 RUST_HUF_FEATURE := huf-force-decompress-x1
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-huf-force-decompress-x1 RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-huf-force-decompress-x1
endif endif
ifneq ($(RUST_HUF_FORCE_X2),0) ifneq ($(RUST_HUF_FORCE_X2),0)
RUST_HUF_FEATURE := huf-force-decompress-x2 RUST_HUF_FEATURE := huf-force-decompress-x2
RUST_BUILD_CONFIG := $(RUST_MODULE_CONFIG)-huf-force-decompress-x2 RUST_BUILD_CONFIG := c$(ZSTD_LIB_COMPRESSION)-d$(ZSTD_LIB_DECOMPRESSION)-huf-force-decompress-x2
endif endif
ifneq ($(RUST_HUF_FEATURE),) ifneq ($(RUST_HUF_FEATURE),)
ifneq ($(ZSTD_LIB_DECOMPRESSION),0) ifneq ($(ZSTD_LIB_DECOMPRESSION),0)
+271 -106
View File
@@ -38,62 +38,6 @@ size_t ZSTD_rust_writeFrameHeader(void* dst, size_t dstCapacity,
U32 windowLog, U64 pledgedSrcSize, U32 windowLog, U64 pledgedSrcSize,
U32 dictID); U32 dictID);
/* Context-free compression-parameter selection and sizing leaves live in
* Rust (rust/src/zstd_compress_params.rs). This file retains
* configuration-sensitive policy: the excluded-block-compressor strategy
* cascade, private ZSTD_CCtx_params handling, and sanitizer workspace
* policy, which it feeds to the leaves as explicit scalar inputs. The
* ZSTD_CParamMode_e and ZSTD_ParamSwitch_e enums are passed as int; the
* Rust side mirrors their values. */
int ZSTD_rust_params_maxCLevel(void);
int ZSTD_rust_params_minCLevel(void);
int ZSTD_rust_params_defaultCLevel(void);
ZSTD_bounds ZSTD_rust_params_getBounds(int param);
size_t ZSTD_rust_params_checkCParams(ZSTD_compressionParameters cParams);
ZSTD_compressionParameters
ZSTD_rust_params_clampCParams(ZSTD_compressionParameters cParams);
U32 ZSTD_rust_params_cycleLog(U32 hashLog, int strategy);
ZSTD_compressionParameters
ZSTD_rust_params_selectCParams(int compressionLevel, U64 srcSizeHint,
size_t dictSize, int mode);
ZSTD_compressionParameters
ZSTD_rust_params_adjustCParams(ZSTD_compressionParameters cParams, U64 srcSize,
size_t dictSize, int mode,
int useRowMatchFinder);
ZSTD_parameters ZSTD_rust_params_makeParams(ZSTD_compressionParameters cParams);
size_t ZSTD_rust_params_maxNbSeq(size_t blockSize, U32 minMatch,
int useSequenceProducer);
size_t ZSTD_rust_params_resolveMaxBlockSize(size_t maxBlockSize);
typedef struct {
U32 hashLog3Max;
size_t matchTSize;
size_t optimalTSize;
size_t asanRedzoneSize;
} ZSTD_rustMatchStateSizing;
size_t ZSTD_rust_params_estimateMatchStateSize(
ZSTD_compressionParameters cParams, int useRowMatchFinder,
int enableDedicatedDictSearch, U32 forCCtx,
const ZSTD_rustMatchStateSizing* sizing);
typedef struct {
size_t cdictSize;
size_t hufWorkspaceSize;
U32 hashLog3Max;
size_t matchTSize;
size_t optimalTSize;
size_t asanRedzoneSize;
} ZSTD_rustCDictSizing;
size_t ZSTD_rust_params_estimateCDictSizeFromCParams(
size_t dictSize, ZSTD_compressionParameters cParams,
int dictLoadMethod, const ZSTD_rustCDictSizing* sizing);
#if ZSTD_ADDRESS_SANITIZER && !defined (ZSTD_ASAN_DONT_POISON_WORKSPACE)
# define ZSTD_RUST_ASAN_REDZONE_SIZE ((size_t)ZSTD_CWKSP_ASAN_REDZONE_SIZE)
#else
# define ZSTD_RUST_ASAN_REDZONE_SIZE ((size_t)0)
#endif
/* *************************************************************** /* ***************************************************************
* Tuning parameters * Tuning parameters
*****************************************************************/ *****************************************************************/
@@ -332,7 +276,11 @@ static int ZSTD_resolveExternalSequenceValidation(int mode) {
/* Resolves maxBlockSize to the default if no value is present. */ /* Resolves maxBlockSize to the default if no value is present. */
static size_t ZSTD_resolveMaxBlockSize(size_t maxBlockSize) { static size_t ZSTD_resolveMaxBlockSize(size_t maxBlockSize) {
return ZSTD_rust_params_resolveMaxBlockSize(maxBlockSize); if (maxBlockSize == 0) {
return ZSTD_BLOCKSIZE_MAX;
} else {
return maxBlockSize;
}
} }
static ZSTD_ParamSwitch_e ZSTD_resolveExternalRepcodeSearch(ZSTD_ParamSwitch_e value, int cLevel) { static ZSTD_ParamSwitch_e ZSTD_resolveExternalRepcodeSearch(ZSTD_ParamSwitch_e value, int cLevel) {
@@ -472,17 +420,45 @@ ZSTD_bounds ZSTD_cParam_getBounds(ZSTD_cParameter param)
switch(param) switch(param)
{ {
/* The compression level and the seven core compression parameters are
* bounded by the Rust leaf. */
case ZSTD_c_compressionLevel: case ZSTD_c_compressionLevel:
bounds.lowerBound = ZSTD_minCLevel();
bounds.upperBound = ZSTD_maxCLevel();
return bounds;
case ZSTD_c_windowLog: case ZSTD_c_windowLog:
bounds.lowerBound = ZSTD_WINDOWLOG_MIN;
bounds.upperBound = ZSTD_WINDOWLOG_MAX;
return bounds;
case ZSTD_c_hashLog: case ZSTD_c_hashLog:
bounds.lowerBound = ZSTD_HASHLOG_MIN;
bounds.upperBound = ZSTD_HASHLOG_MAX;
return bounds;
case ZSTD_c_chainLog: case ZSTD_c_chainLog:
bounds.lowerBound = ZSTD_CHAINLOG_MIN;
bounds.upperBound = ZSTD_CHAINLOG_MAX;
return bounds;
case ZSTD_c_searchLog: case ZSTD_c_searchLog:
bounds.lowerBound = ZSTD_SEARCHLOG_MIN;
bounds.upperBound = ZSTD_SEARCHLOG_MAX;
return bounds;
case ZSTD_c_minMatch: case ZSTD_c_minMatch:
bounds.lowerBound = ZSTD_MINMATCH_MIN;
bounds.upperBound = ZSTD_MINMATCH_MAX;
return bounds;
case ZSTD_c_targetLength: case ZSTD_c_targetLength:
bounds.lowerBound = ZSTD_TARGETLENGTH_MIN;
bounds.upperBound = ZSTD_TARGETLENGTH_MAX;
return bounds;
case ZSTD_c_strategy: case ZSTD_c_strategy:
return ZSTD_rust_params_getBounds((int)param); bounds.lowerBound = ZSTD_STRATEGY_MIN;
bounds.upperBound = ZSTD_STRATEGY_MAX;
return bounds;
case ZSTD_c_contentSizeFlag: case ZSTD_c_contentSizeFlag:
bounds.lowerBound = 0; bounds.lowerBound = 0;
@@ -1409,7 +1385,14 @@ size_t ZSTD_CCtx_reset(ZSTD_CCtx* cctx, ZSTD_ResetDirective reset)
@return : 0, or an error code if one value is beyond authorized range */ @return : 0, or an error code if one value is beyond authorized range */
size_t ZSTD_checkCParams(ZSTD_compressionParameters cParams) size_t ZSTD_checkCParams(ZSTD_compressionParameters cParams)
{ {
return ZSTD_rust_params_checkCParams(cParams); BOUNDCHECK(ZSTD_c_windowLog, (int)cParams.windowLog);
BOUNDCHECK(ZSTD_c_chainLog, (int)cParams.chainLog);
BOUNDCHECK(ZSTD_c_hashLog, (int)cParams.hashLog);
BOUNDCHECK(ZSTD_c_searchLog, (int)cParams.searchLog);
BOUNDCHECK(ZSTD_c_minMatch, (int)cParams.minMatch);
BOUNDCHECK(ZSTD_c_targetLength,(int)cParams.targetLength);
BOUNDCHECK(ZSTD_c_strategy, (int)cParams.strategy);
return 0;
} }
/** ZSTD_clampCParams() : /** ZSTD_clampCParams() :
@@ -1418,14 +1401,63 @@ size_t ZSTD_checkCParams(ZSTD_compressionParameters cParams)
static ZSTD_compressionParameters static ZSTD_compressionParameters
ZSTD_clampCParams(ZSTD_compressionParameters cParams) ZSTD_clampCParams(ZSTD_compressionParameters cParams)
{ {
return ZSTD_rust_params_clampCParams(cParams); # define CLAMP_TYPE(cParam, val, type) \
do { \
ZSTD_bounds const bounds = ZSTD_cParam_getBounds(cParam); \
if ((int)val<bounds.lowerBound) val=(type)bounds.lowerBound; \
else if ((int)val>bounds.upperBound) val=(type)bounds.upperBound; \
} while (0)
# define CLAMP(cParam, val) CLAMP_TYPE(cParam, val, unsigned)
CLAMP(ZSTD_c_windowLog, cParams.windowLog);
CLAMP(ZSTD_c_chainLog, cParams.chainLog);
CLAMP(ZSTD_c_hashLog, cParams.hashLog);
CLAMP(ZSTD_c_searchLog, cParams.searchLog);
CLAMP(ZSTD_c_minMatch, cParams.minMatch);
CLAMP(ZSTD_c_targetLength,cParams.targetLength);
CLAMP_TYPE(ZSTD_c_strategy,cParams.strategy, ZSTD_strategy);
return cParams;
} }
/** ZSTD_cycleLog() : /** ZSTD_cycleLog() :
* condition for correct operation : hashLog > 1 */ * condition for correct operation : hashLog > 1 */
U32 ZSTD_cycleLog(U32 hashLog, ZSTD_strategy strat) U32 ZSTD_cycleLog(U32 hashLog, ZSTD_strategy strat)
{ {
return ZSTD_rust_params_cycleLog(hashLog, (int)strat); U32 const btScale = ((U32)strat >= (U32)ZSTD_btlazy2);
return hashLog - btScale;
}
/** ZSTD_dictAndWindowLog() :
* Returns an adjusted window log that is large enough to fit the source and the dictionary.
* The zstd format says that the entire dictionary is valid if one byte of the dictionary
* is within the window. So the hashLog and chainLog should be large enough to reference both
* the dictionary and the window. So we must use this adjusted dictAndWindowLog when downsizing
* the hashLog and windowLog.
* NOTE: srcSize must not be ZSTD_CONTENTSIZE_UNKNOWN.
*/
static U32 ZSTD_dictAndWindowLog(U32 windowLog, U64 srcSize, U64 dictSize)
{
const U64 maxWindowSize = 1ULL << ZSTD_WINDOWLOG_MAX;
/* No dictionary ==> No change */
if (dictSize == 0) {
return windowLog;
}
assert(windowLog <= ZSTD_WINDOWLOG_MAX);
assert(srcSize != ZSTD_CONTENTSIZE_UNKNOWN); /* Handled in ZSTD_adjustCParams_internal() */
{
U64 const windowSize = 1ULL << windowLog;
U64 const dictAndWindowSize = dictSize + windowSize;
/* If the window size is already large enough to fit both the source and the dictionary
* then just use the window size. Otherwise adjust so that it fits the dictionary and
* the window.
*/
if (windowSize >= dictSize + srcSize) {
return windowLog; /* Window size large enough already */
} else if (dictAndWindowSize >= maxWindowSize) {
return ZSTD_WINDOWLOG_MAX; /* Larger than max window log */
} else {
return ZSTD_highbit32((U32)dictAndWindowSize - 1) + 1;
}
}
} }
/** ZSTD_adjustCParams_internal() : /** ZSTD_adjustCParams_internal() :
@@ -1442,6 +1474,8 @@ ZSTD_adjustCParams_internal(ZSTD_compressionParameters cPar,
ZSTD_CParamMode_e mode, ZSTD_CParamMode_e mode,
ZSTD_ParamSwitch_e useRowMatchFinder) ZSTD_ParamSwitch_e useRowMatchFinder)
{ {
const U64 minSrcSize = 513; /* (1<<9) + 1 */
const U64 maxWindowResize = 1ULL << (ZSTD_WINDOWLOG_MAX-1);
assert(ZSTD_checkCParams(cPar)==0); assert(ZSTD_checkCParams(cPar)==0);
/* Cascade the selected strategy down to the next-highest one built into /* Cascade the selected strategy down to the next-highest one built into
@@ -1486,13 +1520,90 @@ ZSTD_adjustCParams_internal(ZSTD_compressionParameters cPar,
} }
#endif #endif
/* The remaining adjustment logic is context-free and lives in Rust. The switch (mode) {
* short-cache and row-hash tag widths the leaf assumes are fixed case ZSTD_cpm_unknown:
* private-header constants; keep them checked here. */ case ZSTD_cpm_noAttachDict:
ZSTD_STATIC_ASSERT(ZSTD_SHORT_CACHE_TAG_BITS == 8); /* If we don't know the source size, don't make any
ZSTD_STATIC_ASSERT(ZSTD_ROW_HASH_TAG_BITS == 8); * assumptions about it. We will already have selected
return ZSTD_rust_params_adjustCParams(cPar, (U64)srcSize, dictSize, * smaller parameters if a dictionary is in use.
(int)mode, (int)useRowMatchFinder); */
break;
case ZSTD_cpm_createCDict:
/* Assume a small source size when creating a dictionary
* with an unknown source size.
*/
if (dictSize && srcSize == ZSTD_CONTENTSIZE_UNKNOWN)
srcSize = minSrcSize;
break;
case ZSTD_cpm_attachDict:
/* Dictionary has its own dedicated parameters which have
* already been selected. We are selecting parameters
* for only the source.
*/
dictSize = 0;
break;
default:
assert(0);
break;
}
/* resize windowLog if input is small enough, to use less memory */
if ( (srcSize <= maxWindowResize)
&& (dictSize <= maxWindowResize) ) {
U32 const tSize = (U32)(srcSize + dictSize);
static U32 const hashSizeMin = 1 << ZSTD_HASHLOG_MIN;
U32 const srcLog = (tSize < hashSizeMin) ? ZSTD_HASHLOG_MIN :
ZSTD_highbit32(tSize-1) + 1;
if (cPar.windowLog > srcLog) cPar.windowLog = srcLog;
}
if (srcSize != ZSTD_CONTENTSIZE_UNKNOWN) {
U32 const dictAndWindowLog = ZSTD_dictAndWindowLog(cPar.windowLog, (U64)srcSize, (U64)dictSize);
U32 const cycleLog = ZSTD_cycleLog(cPar.chainLog, cPar.strategy);
if (cPar.hashLog > dictAndWindowLog+1) cPar.hashLog = dictAndWindowLog+1;
if (cycleLog > dictAndWindowLog)
cPar.chainLog -= (cycleLog - dictAndWindowLog);
}
if (cPar.windowLog < ZSTD_WINDOWLOG_ABSOLUTEMIN)
cPar.windowLog = ZSTD_WINDOWLOG_ABSOLUTEMIN; /* minimum wlog required for valid frame header */
/* We can't use more than 32 bits of hash in total, so that means that we require:
* (hashLog + 8) <= 32 && (chainLog + 8) <= 32
*/
if (mode == ZSTD_cpm_createCDict && ZSTD_CDictIndicesAreTagged(&cPar)) {
U32 const maxShortCacheHashLog = 32 - ZSTD_SHORT_CACHE_TAG_BITS;
if (cPar.hashLog > maxShortCacheHashLog) {
cPar.hashLog = maxShortCacheHashLog;
}
if (cPar.chainLog > maxShortCacheHashLog) {
cPar.chainLog = maxShortCacheHashLog;
}
}
/* At this point, we aren't 100% sure if we are using the row match finder.
* Unless it is explicitly disabled, conservatively assume that it is enabled.
* In this case it will only be disabled for small sources, so shrinking the
* hash log a little bit shouldn't result in any ratio loss.
*/
if (useRowMatchFinder == ZSTD_ps_auto)
useRowMatchFinder = ZSTD_ps_enable;
/* We can't hash more than 32-bits in total. So that means that we require:
* (hashLog - rowLog + 8) <= 32
*/
if (ZSTD_rowMatchFinderUsed(cPar.strategy, useRowMatchFinder)) {
/* Switch to 32-entry rows if searchLog is 5 (or more) */
U32 const rowLog = BOUNDED(4, cPar.searchLog, 6);
U32 const maxRowHashLog = 32 - ZSTD_ROW_HASH_TAG_BITS;
U32 const maxHashLog = maxRowHashLog + rowLog;
assert(cPar.hashLog >= rowLog);
if (cPar.hashLog > maxHashLog) {
cPar.hashLog = maxHashLog;
}
}
return cPar;
} }
ZSTD_compressionParameters ZSTD_compressionParameters
@@ -1543,31 +1654,47 @@ ZSTD_sizeof_matchState(const ZSTD_compressionParameters* const cParams,
const int enableDedicatedDictSearch, const int enableDedicatedDictSearch,
const U32 forCCtx) const U32 forCCtx)
{ {
ZSTD_rustMatchStateSizing sizing; /* chain table size should be 0 for fast or row-hash strategies */
sizing.hashLog3Max = ZSTD_HASHLOG3_MAX; size_t const chainSize = ZSTD_allocateChainTable(cParams->strategy, useRowMatchFinder, enableDedicatedDictSearch && !forCCtx)
sizing.matchTSize = sizeof(ZSTD_match_t); ? ((size_t)1 << cParams->chainLog)
sizing.optimalTSize = sizeof(ZSTD_optimal_t); : 0;
sizing.asanRedzoneSize = ZSTD_RUST_ASAN_REDZONE_SIZE; size_t const hSize = ((size_t)1) << cParams->hashLog;
U32 const hashLog3 = (forCCtx && cParams->minMatch==3) ? MIN(ZSTD_HASHLOG3_MAX, cParams->windowLog) : 0;
size_t const h3Size = hashLog3 ? ((size_t)1) << hashLog3 : 0;
/* We don't use ZSTD_cwksp_alloc_size() here because the tables aren't
* surrounded by redzones in ASAN. */
size_t const tableSpace = chainSize * sizeof(U32)
+ hSize * sizeof(U32)
+ h3Size * sizeof(U32);
size_t const optPotentialSpace =
ZSTD_cwksp_aligned64_alloc_size((MaxML+1) * sizeof(U32))
+ ZSTD_cwksp_aligned64_alloc_size((MaxLL+1) * sizeof(U32))
+ ZSTD_cwksp_aligned64_alloc_size((MaxOff+1) * sizeof(U32))
+ ZSTD_cwksp_aligned64_alloc_size((1<<Litbits) * sizeof(U32))
+ ZSTD_cwksp_aligned64_alloc_size(ZSTD_OPT_SIZE * sizeof(ZSTD_match_t))
+ ZSTD_cwksp_aligned64_alloc_size(ZSTD_OPT_SIZE * sizeof(ZSTD_optimal_t));
size_t const lazyAdditionalSpace = ZSTD_rowMatchFinderUsed(cParams->strategy, useRowMatchFinder)
? ZSTD_cwksp_aligned64_alloc_size(hSize)
: 0;
size_t const optSpace = (forCCtx && (cParams->strategy >= ZSTD_btopt))
? optPotentialSpace
: 0;
size_t const slackSpace = ZSTD_cwksp_slack_space_required();
/* The Rust leaf hardcodes the frozen format bounds and workspace rules;
* keep them checked against the private headers here. */
ZSTD_STATIC_ASSERT(MaxML == 52 && MaxLL == 35 && MaxOff == 31);
ZSTD_STATIC_ASSERT(Litbits == 8 && ZSTD_OPT_SIZE == 4099);
ZSTD_STATIC_ASSERT(ZSTD_CWKSP_ALIGNMENT_BYTES == 64);
/* tables are guaranteed to be sized in multiples of 64 bytes (or 16 uint32_t) */ /* tables are guaranteed to be sized in multiples of 64 bytes (or 16 uint32_t) */
ZSTD_STATIC_ASSERT(ZSTD_HASHLOG_MIN >= 4 && ZSTD_WINDOWLOG_MIN >= 4 && ZSTD_CHAINLOG_MIN >= 4); ZSTD_STATIC_ASSERT(ZSTD_HASHLOG_MIN >= 4 && ZSTD_WINDOWLOG_MIN >= 4 && ZSTD_CHAINLOG_MIN >= 4);
assert(useRowMatchFinder != ZSTD_ps_auto); assert(useRowMatchFinder != ZSTD_ps_auto);
return ZSTD_rust_params_estimateMatchStateSize(*cParams, DEBUGLOG(4, "chainSize: %u - hSize: %u - h3Size: %u",
(int)useRowMatchFinder, (U32)chainSize, (U32)hSize, (U32)h3Size);
enableDedicatedDictSearch, return tableSpace + optSpace + slackSpace + lazyAdditionalSpace;
forCCtx, &sizing);
} }
/* Helper function for calculating memory requirements. /* Helper function for calculating memory requirements.
* Gives a tighter bound than ZSTD_sequenceBound() by taking minMatch into account. */ * Gives a tighter bound than ZSTD_sequenceBound() by taking minMatch into account. */
static size_t ZSTD_maxNbSeq(size_t blockSize, unsigned minMatch, int useSequenceProducer) { static size_t ZSTD_maxNbSeq(size_t blockSize, unsigned minMatch, int useSequenceProducer) {
return ZSTD_rust_params_maxNbSeq(blockSize, minMatch, useSequenceProducer); U32 const divider = (minMatch==3 || useSequenceProducer) ? 3 : 4;
return blockSize / divider;
} }
static size_t ZSTD_estimateCCtxSize_usingCCtxParams_internal( static size_t ZSTD_estimateCCtxSize_usingCCtxParams_internal(
@@ -5303,15 +5430,15 @@ size_t ZSTD_estimateCDictSize_advanced(
size_t dictSize, ZSTD_compressionParameters cParams, size_t dictSize, ZSTD_compressionParameters cParams,
ZSTD_dictLoadMethod_e dictLoadMethod) ZSTD_dictLoadMethod_e dictLoadMethod)
{ {
ZSTD_rustCDictSizing sizing; DEBUGLOG(5, "sizeof(ZSTD_CDict) : %u", (unsigned)sizeof(ZSTD_CDict));
sizing.cdictSize = sizeof(ZSTD_CDict); return ZSTD_cwksp_alloc_size(sizeof(ZSTD_CDict))
sizing.hufWorkspaceSize = HUF_WORKSPACE_SIZE; + ZSTD_cwksp_alloc_size(HUF_WORKSPACE_SIZE)
sizing.hashLog3Max = ZSTD_HASHLOG3_MAX; /* enableDedicatedDictSearch == 1 ensures that CDict estimation will not be too small
sizing.matchTSize = sizeof(ZSTD_match_t); * in case we are using DDS with row-hash. */
sizing.optimalTSize = sizeof(ZSTD_optimal_t); + ZSTD_sizeof_matchState(&cParams, ZSTD_resolveRowMatchFinderMode(ZSTD_ps_auto, &cParams),
sizing.asanRedzoneSize = ZSTD_RUST_ASAN_REDZONE_SIZE; /* enableDedicatedDictSearch */ 1, /* forCCtx */ 0)
return ZSTD_rust_params_estimateCDictSizeFromCParams( + (dictLoadMethod == ZSTD_dlm_byRef ? 0
dictSize, cParams, (int)dictLoadMethod, &sizing); : ZSTD_cwksp_alloc_size(ZSTD_cwksp_align(dictSize, sizeof(void *))));
} }
size_t ZSTD_estimateCDictSize(size_t dictSize, int compressionLevel) size_t ZSTD_estimateCDictSize(size_t dictSize, int compressionLevel)
@@ -7449,12 +7576,11 @@ size_t ZSTD_endStream(ZSTD_CStream* zcs, ZSTD_outBuffer* output)
/*-===== Pre-defined compression levels =====-*/ /*-===== Pre-defined compression levels =====-*/
/* The compression-level tables (formerly included from clevels.h) live in #include "clevels.h"
* rust/src/zstd_compress_params.rs. */
int ZSTD_maxCLevel(void) { return ZSTD_rust_params_maxCLevel(); } int ZSTD_maxCLevel(void) { return ZSTD_MAX_CLEVEL; }
int ZSTD_minCLevel(void) { return ZSTD_rust_params_minCLevel(); } int ZSTD_minCLevel(void) { return (int)-ZSTD_TARGETLENGTH_MAX; }
int ZSTD_defaultCLevel(void) { return ZSTD_rust_params_defaultCLevel(); } int ZSTD_defaultCLevel(void) { return ZSTD_CLEVEL_DEFAULT; }
static ZSTD_compressionParameters ZSTD_dedicatedDictSearch_getCParams(int const compressionLevel, size_t const dictSize) static ZSTD_compressionParameters ZSTD_dedicatedDictSearch_getCParams(int const compressionLevel, size_t const dictSize)
{ {
@@ -7513,6 +7639,26 @@ static void ZSTD_dedicatedDictSearch_revertCParams(
} }
} }
static U64 ZSTD_getCParamRowSize(U64 srcSizeHint, size_t dictSize, ZSTD_CParamMode_e mode)
{
switch (mode) {
case ZSTD_cpm_unknown:
case ZSTD_cpm_noAttachDict:
case ZSTD_cpm_createCDict:
break;
case ZSTD_cpm_attachDict:
dictSize = 0;
break;
default:
assert(0);
break;
}
{ int const unknown = srcSizeHint == ZSTD_CONTENTSIZE_UNKNOWN;
size_t const addedSize = unknown && dictSize > 0 ? 500 : 0;
return unknown && dictSize == 0 ? ZSTD_CONTENTSIZE_UNKNOWN : srcSizeHint+dictSize+addedSize;
}
}
/*! ZSTD_getCParams_internal() : /*! ZSTD_getCParams_internal() :
* @return ZSTD_compressionParameters structure for a selected compression level, srcSize and dictSize. * @return ZSTD_compressionParameters structure for a selected compression level, srcSize and dictSize.
* Note: srcSizeHint 0 means 0, use ZSTD_CONTENTSIZE_UNKNOWN for unknown. * Note: srcSizeHint 0 means 0, use ZSTD_CONTENTSIZE_UNKNOWN for unknown.
@@ -7520,12 +7666,27 @@ static void ZSTD_dedicatedDictSearch_revertCParams(
* Note: `mode` controls how we treat the `dictSize`. See docs for `ZSTD_CParamMode_e`. */ * Note: `mode` controls how we treat the `dictSize`. See docs for `ZSTD_CParamMode_e`. */
static ZSTD_compressionParameters ZSTD_getCParams_internal(int compressionLevel, unsigned long long srcSizeHint, size_t dictSize, ZSTD_CParamMode_e mode) static ZSTD_compressionParameters ZSTD_getCParams_internal(int compressionLevel, unsigned long long srcSizeHint, size_t dictSize, ZSTD_CParamMode_e mode)
{ {
/* Table selection is context-free and lives in Rust; the adjustment step U64 const rSize = ZSTD_getCParamRowSize(srcSizeHint, dictSize, mode);
* stays behind ZSTD_adjustCParams_internal() so this build's strategy U32 const tableID = (rSize <= 256 KB) + (rSize <= 128 KB) + (rSize <= 16 KB);
* cascade applies. */ int row;
ZSTD_compressionParameters const cp = ZSTD_rust_params_selectCParams( DEBUGLOG(5, "ZSTD_getCParams_internal (cLevel=%i)", compressionLevel);
compressionLevel, srcSizeHint, dictSize, (int)mode);
return ZSTD_adjustCParams_internal(cp, srcSizeHint, dictSize, mode, ZSTD_ps_auto); /* row */
if (compressionLevel == 0) row = ZSTD_CLEVEL_DEFAULT; /* 0 == default */
else if (compressionLevel < 0) row = 0; /* entry 0 is baseline for fast mode */
else if (compressionLevel > ZSTD_MAX_CLEVEL) row = ZSTD_MAX_CLEVEL;
else row = compressionLevel;
{ ZSTD_compressionParameters cp = ZSTD_defaultCParameters[tableID][row];
DEBUGLOG(5, "ZSTD_getCParams_internal selected tableID: %u row: %u strat: %u", tableID, row, (U32)cp.strategy);
/* acceleration factor */
if (compressionLevel < 0) {
int const clampedCompressionLevel = MAX(ZSTD_minCLevel(), compressionLevel);
cp.targetLength = (unsigned)(-clampedCompressionLevel);
}
/* refine parameters based on srcSize & dictSize */
return ZSTD_adjustCParams_internal(cp, srcSizeHint, dictSize, mode, ZSTD_ps_auto);
}
} }
/*! ZSTD_getCParams() : /*! ZSTD_getCParams() :
@@ -7544,9 +7705,13 @@ ZSTD_compressionParameters ZSTD_getCParams(int compressionLevel, unsigned long l
static ZSTD_parameters static ZSTD_parameters
ZSTD_getParams_internal(int compressionLevel, unsigned long long srcSizeHint, size_t dictSize, ZSTD_CParamMode_e mode) ZSTD_getParams_internal(int compressionLevel, unsigned long long srcSizeHint, size_t dictSize, ZSTD_CParamMode_e mode)
{ {
ZSTD_parameters params;
ZSTD_compressionParameters const cParams = ZSTD_getCParams_internal(compressionLevel, srcSizeHint, dictSize, mode); ZSTD_compressionParameters const cParams = ZSTD_getCParams_internal(compressionLevel, srcSizeHint, dictSize, mode);
DEBUGLOG(5, "ZSTD_getParams (cLevel=%i)", compressionLevel); DEBUGLOG(5, "ZSTD_getParams (cLevel=%i)", compressionLevel);
return ZSTD_rust_params_makeParams(cParams); ZSTD_memset(&params, 0, sizeof(params));
params.cParams = cParams;
params.fParams.contentSizeFlag = 1;
return params;
} }
/*! ZSTD_getParams() : /*! ZSTD_getParams() :
+1886 -5
View File
@@ -24,9 +24,1890 @@
* OTHER DEALINGS IN THE SOFTWARE. * OTHER DEALINGS IN THE SOFTWARE.
*/ */
/* divsufsort() is implemented in rust/src/divsufsort.rs, which provides the /*- Compiler specifics -*/
* symbol directly. This translation unit keeps the header's prototypes in #ifdef __clang__
* the build so the dictionary builder continues to compile against the #pragma clang diagnostic ignored "-Wshorten-64-to-32"
* original interface. divbwt() has no callers in zstd and is declaration- #endif
* only; it moves to Rust if a user ever appears. */
#if defined(_MSC_VER)
# pragma warning(disable : 4244)
# pragma warning(disable : 4127) /* C4127 : Condition expression is constant */
#endif
/*- Dependencies -*/
#include <assert.h>
#include <stdio.h>
#include <stdlib.h>
#include "divsufsort.h" #include "divsufsort.h"
/*- Constants -*/
#if defined(INLINE)
# undef INLINE
#endif
#if !defined(INLINE)
# define INLINE __inline
#endif
#if defined(ALPHABET_SIZE) && (ALPHABET_SIZE < 1)
# undef ALPHABET_SIZE
#endif
#if !defined(ALPHABET_SIZE)
# define ALPHABET_SIZE (256)
#endif
#define BUCKET_A_SIZE (ALPHABET_SIZE)
#define BUCKET_B_SIZE (ALPHABET_SIZE * ALPHABET_SIZE)
#if defined(SS_INSERTIONSORT_THRESHOLD)
# if SS_INSERTIONSORT_THRESHOLD < 1
# undef SS_INSERTIONSORT_THRESHOLD
# define SS_INSERTIONSORT_THRESHOLD (1)
# endif
#else
# define SS_INSERTIONSORT_THRESHOLD (8)
#endif
#if defined(SS_BLOCKSIZE)
# if SS_BLOCKSIZE < 0
# undef SS_BLOCKSIZE
# define SS_BLOCKSIZE (0)
# elif 32768 <= SS_BLOCKSIZE
# undef SS_BLOCKSIZE
# define SS_BLOCKSIZE (32767)
# endif
#else
# define SS_BLOCKSIZE (1024)
#endif
/* minstacksize = log(SS_BLOCKSIZE) / log(3) * 2 */
#if SS_BLOCKSIZE == 0
# define SS_MISORT_STACKSIZE (96)
#elif SS_BLOCKSIZE <= 4096
# define SS_MISORT_STACKSIZE (16)
#else
# define SS_MISORT_STACKSIZE (24)
#endif
#define SS_SMERGE_STACKSIZE (32)
#define TR_INSERTIONSORT_THRESHOLD (8)
#define TR_STACKSIZE (64)
/*- Macros -*/
#ifndef SWAP
# define SWAP(_a, _b) do { t = (_a); (_a) = (_b); (_b) = t; } while(0)
#endif /* SWAP */
#ifndef MIN
# define MIN(_a, _b) (((_a) < (_b)) ? (_a) : (_b))
#endif /* MIN */
#ifndef MAX
# define MAX(_a, _b) (((_a) > (_b)) ? (_a) : (_b))
#endif /* MAX */
#define STACK_PUSH(_a, _b, _c, _d)\
do {\
assert(ssize < STACK_SIZE);\
stack[ssize].a = (_a), stack[ssize].b = (_b),\
stack[ssize].c = (_c), stack[ssize++].d = (_d);\
} while(0)
#define STACK_PUSH5(_a, _b, _c, _d, _e)\
do {\
assert(ssize < STACK_SIZE);\
stack[ssize].a = (_a), stack[ssize].b = (_b),\
stack[ssize].c = (_c), stack[ssize].d = (_d), stack[ssize++].e = (_e);\
} while(0)
#define STACK_POP(_a, _b, _c, _d)\
do {\
assert(0 <= ssize);\
if(ssize == 0) { return; }\
(_a) = stack[--ssize].a, (_b) = stack[ssize].b,\
(_c) = stack[ssize].c, (_d) = stack[ssize].d;\
} while(0)
#define STACK_POP5(_a, _b, _c, _d, _e)\
do {\
assert(0 <= ssize);\
if(ssize == 0) { return; }\
(_a) = stack[--ssize].a, (_b) = stack[ssize].b,\
(_c) = stack[ssize].c, (_d) = stack[ssize].d, (_e) = stack[ssize].e;\
} while(0)
#define BUCKET_A(_c0) bucket_A[(_c0)]
#if ALPHABET_SIZE == 256
#define BUCKET_B(_c0, _c1) (bucket_B[((_c1) << 8) | (_c0)])
#define BUCKET_BSTAR(_c0, _c1) (bucket_B[((_c0) << 8) | (_c1)])
#else
#define BUCKET_B(_c0, _c1) (bucket_B[(_c1) * ALPHABET_SIZE + (_c0)])
#define BUCKET_BSTAR(_c0, _c1) (bucket_B[(_c0) * ALPHABET_SIZE + (_c1)])
#endif
/*- Private Functions -*/
static const int lg_table[256]= {
-1,0,1,1,2,2,2,2,3,3,3,3,3,3,3,3,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,
5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7
};
#if (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE)
static INLINE
int
ss_ilg(int n) {
#if SS_BLOCKSIZE == 0
return (n & 0xffff0000) ?
((n & 0xff000000) ?
24 + lg_table[(n >> 24) & 0xff] :
16 + lg_table[(n >> 16) & 0xff]) :
((n & 0x0000ff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff]);
#elif SS_BLOCKSIZE < 256
return lg_table[n];
#else
return (n & 0xff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff];
#endif
}
#endif /* (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE) */
#if SS_BLOCKSIZE != 0
static const int sqq_table[256] = {
0, 16, 22, 27, 32, 35, 39, 42, 45, 48, 50, 53, 55, 57, 59, 61,
64, 65, 67, 69, 71, 73, 75, 76, 78, 80, 81, 83, 84, 86, 87, 89,
90, 91, 93, 94, 96, 97, 98, 99, 101, 102, 103, 104, 106, 107, 108, 109,
110, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126,
128, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 144, 145, 146, 147, 148, 149, 150, 150, 151, 152, 153, 154, 155, 155,
156, 157, 158, 159, 160, 160, 161, 162, 163, 163, 164, 165, 166, 167, 167, 168,
169, 170, 170, 171, 172, 173, 173, 174, 175, 176, 176, 177, 178, 178, 179, 180,
181, 181, 182, 183, 183, 184, 185, 185, 186, 187, 187, 188, 189, 189, 190, 191,
192, 192, 193, 193, 194, 195, 195, 196, 197, 197, 198, 199, 199, 200, 201, 201,
202, 203, 203, 204, 204, 205, 206, 206, 207, 208, 208, 209, 209, 210, 211, 211,
212, 212, 213, 214, 214, 215, 215, 216, 217, 217, 218, 218, 219, 219, 220, 221,
221, 222, 222, 223, 224, 224, 225, 225, 226, 226, 227, 227, 228, 229, 229, 230,
230, 231, 231, 232, 232, 233, 234, 234, 235, 235, 236, 236, 237, 237, 238, 238,
239, 240, 240, 241, 241, 242, 242, 243, 243, 244, 244, 245, 245, 246, 246, 247,
247, 248, 248, 249, 249, 250, 250, 251, 251, 252, 252, 253, 253, 254, 254, 255
};
static INLINE
int
ss_isqrt(int x) {
int y, e;
if(x >= (SS_BLOCKSIZE * SS_BLOCKSIZE)) { return SS_BLOCKSIZE; }
e = (x & 0xffff0000) ?
((x & 0xff000000) ?
24 + lg_table[(x >> 24) & 0xff] :
16 + lg_table[(x >> 16) & 0xff]) :
((x & 0x0000ff00) ?
8 + lg_table[(x >> 8) & 0xff] :
0 + lg_table[(x >> 0) & 0xff]);
if(e >= 16) {
y = sqq_table[x >> ((e - 6) - (e & 1))] << ((e >> 1) - 7);
if(e >= 24) { y = (y + 1 + x / y) >> 1; }
y = (y + 1 + x / y) >> 1;
} else if(e >= 8) {
y = (sqq_table[x >> ((e - 6) - (e & 1))] >> (7 - (e >> 1))) + 1;
} else {
return sqq_table[x] >> 4;
}
return (x < (y * y)) ? y - 1 : y;
}
#endif /* SS_BLOCKSIZE != 0 */
/*---------------------------------------------------------------------------*/
/* Compares two suffixes. */
static INLINE
int
ss_compare(const unsigned char *T,
const int *p1, const int *p2,
int depth) {
const unsigned char *U1, *U2, *U1n, *U2n;
for(U1 = T + depth + *p1,
U2 = T + depth + *p2,
U1n = T + *(p1 + 1) + 2,
U2n = T + *(p2 + 1) + 2;
(U1 < U1n) && (U2 < U2n) && (*U1 == *U2);
++U1, ++U2) {
}
return U1 < U1n ?
(U2 < U2n ? *U1 - *U2 : 1) :
(U2 < U2n ? -1 : 0);
}
/*---------------------------------------------------------------------------*/
#if (SS_BLOCKSIZE != 1) && (SS_INSERTIONSORT_THRESHOLD != 1)
/* Insertionsort for small size groups */
static
void
ss_insertionsort(const unsigned char *T, const int *PA,
int *first, int *last, int depth) {
int *i, *j;
int t;
int r;
for(i = last - 2; first <= i; --i) {
for(t = *i, j = i + 1; 0 < (r = ss_compare(T, PA + t, PA + *j, depth));) {
do { *(j - 1) = *j; } while((++j < last) && (*j < 0));
if(last <= j) { break; }
}
if(r == 0) { *j = ~*j; }
*(j - 1) = t;
}
}
#endif /* (SS_BLOCKSIZE != 1) && (SS_INSERTIONSORT_THRESHOLD != 1) */
/*---------------------------------------------------------------------------*/
#if (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE)
static INLINE
void
ss_fixdown(const unsigned char *Td, const int *PA,
int *SA, int i, int size) {
int j, k;
int v;
int c, d, e;
for(v = SA[i], c = Td[PA[v]]; (j = 2 * i + 1) < size; SA[i] = SA[k], i = k) {
d = Td[PA[SA[k = j++]]];
if(d < (e = Td[PA[SA[j]]])) { k = j; d = e; }
if(d <= c) { break; }
}
SA[i] = v;
}
/* Simple top-down heapsort. */
static
void
ss_heapsort(const unsigned char *Td, const int *PA, int *SA, int size) {
int i, m;
int t;
m = size;
if((size % 2) == 0) {
m--;
if(Td[PA[SA[m / 2]]] < Td[PA[SA[m]]]) { SWAP(SA[m], SA[m / 2]); }
}
for(i = m / 2 - 1; 0 <= i; --i) { ss_fixdown(Td, PA, SA, i, m); }
if((size % 2) == 0) { SWAP(SA[0], SA[m]); ss_fixdown(Td, PA, SA, 0, m); }
for(i = m - 1; 0 < i; --i) {
t = SA[0], SA[0] = SA[i];
ss_fixdown(Td, PA, SA, 0, i);
SA[i] = t;
}
}
/*---------------------------------------------------------------------------*/
/* Returns the median of three elements. */
static INLINE
int *
ss_median3(const unsigned char *Td, const int *PA,
int *v1, int *v2, int *v3) {
int *t;
if(Td[PA[*v1]] > Td[PA[*v2]]) { SWAP(v1, v2); }
if(Td[PA[*v2]] > Td[PA[*v3]]) {
if(Td[PA[*v1]] > Td[PA[*v3]]) { return v1; }
else { return v3; }
}
return v2;
}
/* Returns the median of five elements. */
static INLINE
int *
ss_median5(const unsigned char *Td, const int *PA,
int *v1, int *v2, int *v3, int *v4, int *v5) {
int *t;
if(Td[PA[*v2]] > Td[PA[*v3]]) { SWAP(v2, v3); }
if(Td[PA[*v4]] > Td[PA[*v5]]) { SWAP(v4, v5); }
if(Td[PA[*v2]] > Td[PA[*v4]]) { SWAP(v2, v4); SWAP(v3, v5); }
if(Td[PA[*v1]] > Td[PA[*v3]]) { SWAP(v1, v3); }
if(Td[PA[*v1]] > Td[PA[*v4]]) { SWAP(v1, v4); SWAP(v3, v5); }
if(Td[PA[*v3]] > Td[PA[*v4]]) { return v4; }
return v3;
}
/* Returns the pivot element. */
static INLINE
int *
ss_pivot(const unsigned char *Td, const int *PA, int *first, int *last) {
int *middle;
int t;
t = last - first;
middle = first + t / 2;
if(t <= 512) {
if(t <= 32) {
return ss_median3(Td, PA, first, middle, last - 1);
} else {
t >>= 2;
return ss_median5(Td, PA, first, first + t, middle, last - 1 - t, last - 1);
}
}
t >>= 3;
first = ss_median3(Td, PA, first, first + t, first + (t << 1));
middle = ss_median3(Td, PA, middle - t, middle, middle + t);
last = ss_median3(Td, PA, last - 1 - (t << 1), last - 1 - t, last - 1);
return ss_median3(Td, PA, first, middle, last);
}
/*---------------------------------------------------------------------------*/
/* Binary partition for substrings. */
static INLINE
int *
ss_partition(const int *PA,
int *first, int *last, int depth) {
int *a, *b;
int t;
for(a = first - 1, b = last;;) {
for(; (++a < b) && ((PA[*a] + depth) >= (PA[*a + 1] + 1));) { *a = ~*a; }
for(; (a < --b) && ((PA[*b] + depth) < (PA[*b + 1] + 1));) { }
if(b <= a) { break; }
t = ~*b;
*b = *a;
*a = t;
}
if(first < a) { *first = ~*first; }
return a;
}
/* Multikey introsort for medium size groups. */
static
void
ss_mintrosort(const unsigned char *T, const int *PA,
int *first, int *last,
int depth) {
#define STACK_SIZE SS_MISORT_STACKSIZE
struct { int *a, *b, c; int d; } stack[STACK_SIZE];
const unsigned char *Td;
int *a, *b, *c, *d, *e, *f;
int s, t;
int ssize;
int limit;
int v, x = 0;
for(ssize = 0, limit = ss_ilg(last - first);;) {
if((last - first) <= SS_INSERTIONSORT_THRESHOLD) {
#if 1 < SS_INSERTIONSORT_THRESHOLD
if(1 < (last - first)) { ss_insertionsort(T, PA, first, last, depth); }
#endif
STACK_POP(first, last, depth, limit);
continue;
}
Td = T + depth;
if(limit-- == 0) { ss_heapsort(Td, PA, first, last - first); }
if(limit < 0) {
for(a = first + 1, v = Td[PA[*first]]; a < last; ++a) {
if((x = Td[PA[*a]]) != v) {
if(1 < (a - first)) { break; }
v = x;
first = a;
}
}
if(Td[PA[*first] - 1] < v) {
first = ss_partition(PA, first, a, depth);
}
if((a - first) <= (last - a)) {
if(1 < (a - first)) {
STACK_PUSH(a, last, depth, -1);
last = a, depth += 1, limit = ss_ilg(a - first);
} else {
first = a, limit = -1;
}
} else {
if(1 < (last - a)) {
STACK_PUSH(first, a, depth + 1, ss_ilg(a - first));
first = a, limit = -1;
} else {
last = a, depth += 1, limit = ss_ilg(a - first);
}
}
continue;
}
/* choose pivot */
a = ss_pivot(Td, PA, first, last);
v = Td[PA[*a]];
SWAP(*first, *a);
/* partition */
for(b = first; (++b < last) && ((x = Td[PA[*b]]) == v);) { }
if(((a = b) < last) && (x < v)) {
for(; (++b < last) && ((x = Td[PA[*b]]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
}
for(c = last; (b < --c) && ((x = Td[PA[*c]]) == v);) { }
if((b < (d = c)) && (x > v)) {
for(; (b < --c) && ((x = Td[PA[*c]]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
for(; b < c;) {
SWAP(*b, *c);
for(; (++b < c) && ((x = Td[PA[*b]]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
for(; (b < --c) && ((x = Td[PA[*c]]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
if(a <= d) {
c = b - 1;
if((s = a - first) > (t = b - a)) { s = t; }
for(e = first, f = b - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
if((s = d - c) > (t = last - d - 1)) { s = t; }
for(e = b, f = last - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
a = first + (b - a), c = last - (d - c);
b = (v <= Td[PA[*a] - 1]) ? a : ss_partition(PA, a, c, depth);
if((a - first) <= (last - c)) {
if((last - c) <= (c - b)) {
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
STACK_PUSH(c, last, depth, limit);
last = a;
} else if((a - first) <= (c - b)) {
STACK_PUSH(c, last, depth, limit);
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
last = a;
} else {
STACK_PUSH(c, last, depth, limit);
STACK_PUSH(first, a, depth, limit);
first = b, last = c, depth += 1, limit = ss_ilg(c - b);
}
} else {
if((a - first) <= (c - b)) {
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
STACK_PUSH(first, a, depth, limit);
first = c;
} else if((last - c) <= (c - b)) {
STACK_PUSH(first, a, depth, limit);
STACK_PUSH(b, c, depth + 1, ss_ilg(c - b));
first = c;
} else {
STACK_PUSH(first, a, depth, limit);
STACK_PUSH(c, last, depth, limit);
first = b, last = c, depth += 1, limit = ss_ilg(c - b);
}
}
} else {
limit += 1;
if(Td[PA[*first] - 1] < v) {
first = ss_partition(PA, first, last, depth);
limit = ss_ilg(last - first);
}
depth += 1;
}
}
#undef STACK_SIZE
}
#endif /* (SS_BLOCKSIZE == 0) || (SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE) */
/*---------------------------------------------------------------------------*/
#if SS_BLOCKSIZE != 0
static INLINE
void
ss_blockswap(int *a, int *b, int n) {
int t;
for(; 0 < n; --n, ++a, ++b) {
t = *a, *a = *b, *b = t;
}
}
static INLINE
void
ss_rotate(int *first, int *middle, int *last) {
int *a, *b, t;
int l, r;
l = middle - first, r = last - middle;
for(; (0 < l) && (0 < r);) {
if(l == r) { ss_blockswap(first, middle, l); break; }
if(l < r) {
a = last - 1, b = middle - 1;
t = *a;
do {
*a-- = *b, *b-- = *a;
if(b < first) {
*a = t;
last = a;
if((r -= l + 1) <= l) { break; }
a -= 1, b = middle - 1;
t = *a;
}
} while(1);
} else {
a = first, b = middle;
t = *a;
do {
*a++ = *b, *b++ = *a;
if(last <= b) {
*a = t;
first = a + 1;
if((l -= r + 1) <= r) { break; }
a += 1, b = middle;
t = *a;
}
} while(1);
}
}
}
/*---------------------------------------------------------------------------*/
static
void
ss_inplacemerge(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int depth) {
const int *p;
int *a, *b;
int len, half;
int q, r;
int x;
for(;;) {
if(*(last - 1) < 0) { x = 1; p = PA + ~*(last - 1); }
else { x = 0; p = PA + *(last - 1); }
for(a = first, len = middle - first, half = len >> 1, r = -1;
0 < len;
len = half, half >>= 1) {
b = a + half;
q = ss_compare(T, PA + ((0 <= *b) ? *b : ~*b), p, depth);
if(q < 0) {
a = b + 1;
half -= (len & 1) ^ 1;
} else {
r = q;
}
}
if(a < middle) {
if(r == 0) { *a = ~*a; }
ss_rotate(a, middle, last);
last -= middle - a;
middle = a;
if(first == middle) { break; }
}
--last;
if(x != 0) { while(*--last < 0) { } }
if(middle == last) { break; }
}
}
/*---------------------------------------------------------------------------*/
/* Merge-forward with internal buffer. */
static
void
ss_mergeforward(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int depth) {
int *a, *b, *c, *bufend;
int t;
int r;
bufend = buf + (middle - first) - 1;
ss_blockswap(buf, first, middle - first);
for(t = *(a = first), b = buf, c = middle;;) {
r = ss_compare(T, PA + *b, PA + *c, depth);
if(r < 0) {
do {
*a++ = *b;
if(bufend <= b) { *bufend = t; return; }
*b++ = *a;
} while(*b < 0);
} else if(r > 0) {
do {
*a++ = *c, *c++ = *a;
if(last <= c) {
while(b < bufend) { *a++ = *b, *b++ = *a; }
*a = *b, *b = t;
return;
}
} while(*c < 0);
} else {
*c = ~*c;
do {
*a++ = *b;
if(bufend <= b) { *bufend = t; return; }
*b++ = *a;
} while(*b < 0);
do {
*a++ = *c, *c++ = *a;
if(last <= c) {
while(b < bufend) { *a++ = *b, *b++ = *a; }
*a = *b, *b = t;
return;
}
} while(*c < 0);
}
}
}
/* Merge-backward with internal buffer. */
static
void
ss_mergebackward(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int depth) {
const int *p1, *p2;
int *a, *b, *c, *bufend;
int t;
int r;
int x;
bufend = buf + (last - middle) - 1;
ss_blockswap(buf, middle, last - middle);
x = 0;
if(*bufend < 0) { p1 = PA + ~*bufend; x |= 1; }
else { p1 = PA + *bufend; }
if(*(middle - 1) < 0) { p2 = PA + ~*(middle - 1); x |= 2; }
else { p2 = PA + *(middle - 1); }
for(t = *(a = last - 1), b = bufend, c = middle - 1;;) {
r = ss_compare(T, p1, p2, depth);
if(0 < r) {
if(x & 1) { do { *a-- = *b, *b-- = *a; } while(*b < 0); x ^= 1; }
*a-- = *b;
if(b <= buf) { *buf = t; break; }
*b-- = *a;
if(*b < 0) { p1 = PA + ~*b; x |= 1; }
else { p1 = PA + *b; }
} else if(r < 0) {
if(x & 2) { do { *a-- = *c, *c-- = *a; } while(*c < 0); x ^= 2; }
*a-- = *c, *c-- = *a;
if(c < first) {
while(buf < b) { *a-- = *b, *b-- = *a; }
*a = *b, *b = t;
break;
}
if(*c < 0) { p2 = PA + ~*c; x |= 2; }
else { p2 = PA + *c; }
} else {
if(x & 1) { do { *a-- = *b, *b-- = *a; } while(*b < 0); x ^= 1; }
*a-- = ~*b;
if(b <= buf) { *buf = t; break; }
*b-- = *a;
if(x & 2) { do { *a-- = *c, *c-- = *a; } while(*c < 0); x ^= 2; }
*a-- = *c, *c-- = *a;
if(c < first) {
while(buf < b) { *a-- = *b, *b-- = *a; }
*a = *b, *b = t;
break;
}
if(*b < 0) { p1 = PA + ~*b; x |= 1; }
else { p1 = PA + *b; }
if(*c < 0) { p2 = PA + ~*c; x |= 2; }
else { p2 = PA + *c; }
}
}
}
/* D&C based merge. */
static
void
ss_swapmerge(const unsigned char *T, const int *PA,
int *first, int *middle, int *last,
int *buf, int bufsize, int depth) {
#define STACK_SIZE SS_SMERGE_STACKSIZE
#define GETIDX(a) ((0 <= (a)) ? (a) : (~(a)))
#define MERGE_CHECK(a, b, c)\
do {\
if(((c) & 1) ||\
(((c) & 2) && (ss_compare(T, PA + GETIDX(*((a) - 1)), PA + *(a), depth) == 0))) {\
*(a) = ~*(a);\
}\
if(((c) & 4) && ((ss_compare(T, PA + GETIDX(*((b) - 1)), PA + *(b), depth) == 0))) {\
*(b) = ~*(b);\
}\
} while(0)
struct { int *a, *b, *c; int d; } stack[STACK_SIZE];
int *l, *r, *lm, *rm;
int m, len, half;
int ssize;
int check, next;
for(check = 0, ssize = 0;;) {
if((last - middle) <= bufsize) {
if((first < middle) && (middle < last)) {
ss_mergebackward(T, PA, first, middle, last, buf, depth);
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
continue;
}
if((middle - first) <= bufsize) {
if(first < middle) {
ss_mergeforward(T, PA, first, middle, last, buf, depth);
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
continue;
}
for(m = 0, len = MIN(middle - first, last - middle), half = len >> 1;
0 < len;
len = half, half >>= 1) {
if(ss_compare(T, PA + GETIDX(*(middle + m + half)),
PA + GETIDX(*(middle - m - half - 1)), depth) < 0) {
m += half + 1;
half -= (len & 1) ^ 1;
}
}
if(0 < m) {
lm = middle - m, rm = middle + m;
ss_blockswap(lm, middle, m);
l = r = middle, next = 0;
if(rm < last) {
if(*rm < 0) {
*rm = ~*rm;
if(first < lm) { for(; *--l < 0;) { } next |= 4; }
next |= 1;
} else if(first < lm) {
for(; *r < 0; ++r) { }
next |= 2;
}
}
if((l - first) <= (last - r)) {
STACK_PUSH(r, rm, last, (next & 3) | (check & 4));
middle = lm, last = l, check = (check & 3) | (next & 4);
} else {
if((next & 2) && (r == middle)) { next ^= 6; }
STACK_PUSH(first, lm, l, (check & 3) | (next & 4));
first = r, middle = rm, check = (next & 3) | (check & 4);
}
} else {
if(ss_compare(T, PA + GETIDX(*(middle - 1)), PA + *middle, depth) == 0) {
*middle = ~*middle;
}
MERGE_CHECK(first, last, check);
STACK_POP(first, middle, last, check);
}
}
#undef STACK_SIZE
}
#endif /* SS_BLOCKSIZE != 0 */
/*---------------------------------------------------------------------------*/
/* Substring sort */
static
void
sssort(const unsigned char *T, const int *PA,
int *first, int *last,
int *buf, int bufsize,
int depth, int n, int lastsuffix) {
int *a;
#if SS_BLOCKSIZE != 0
int *b, *middle, *curbuf;
int j, k, curbufsize, limit;
#endif
int i;
if(lastsuffix != 0) { ++first; }
#if SS_BLOCKSIZE == 0
ss_mintrosort(T, PA, first, last, depth);
#else
if((bufsize < SS_BLOCKSIZE) &&
(bufsize < (last - first)) &&
(bufsize < (limit = ss_isqrt(last - first)))) {
if(SS_BLOCKSIZE < limit) { limit = SS_BLOCKSIZE; }
buf = middle = last - limit, bufsize = limit;
} else {
middle = last, limit = 0;
}
for(a = first, i = 0; SS_BLOCKSIZE < (middle - a); a += SS_BLOCKSIZE, ++i) {
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, a, a + SS_BLOCKSIZE, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, a, a + SS_BLOCKSIZE, depth);
#endif
curbufsize = last - (a + SS_BLOCKSIZE);
curbuf = a + SS_BLOCKSIZE;
if(curbufsize <= bufsize) { curbufsize = bufsize, curbuf = buf; }
for(b = a, k = SS_BLOCKSIZE, j = i; j & 1; b -= k, k <<= 1, j >>= 1) {
ss_swapmerge(T, PA, b - k, b, b + k, curbuf, curbufsize, depth);
}
}
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, a, middle, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, a, middle, depth);
#endif
for(k = SS_BLOCKSIZE; i != 0; k <<= 1, i >>= 1) {
if(i & 1) {
ss_swapmerge(T, PA, a - k, a, middle, buf, bufsize, depth);
a -= k;
}
}
if(limit != 0) {
#if SS_INSERTIONSORT_THRESHOLD < SS_BLOCKSIZE
ss_mintrosort(T, PA, middle, last, depth);
#elif 1 < SS_BLOCKSIZE
ss_insertionsort(T, PA, middle, last, depth);
#endif
ss_inplacemerge(T, PA, first, middle, last, depth);
}
#endif
if(lastsuffix != 0) {
/* Insert last type B* suffix. */
int PAi[2]; PAi[0] = PA[*(first - 1)], PAi[1] = n - 2;
for(a = first, i = *(first - 1);
(a < last) && ((*a < 0) || (0 < ss_compare(T, &(PAi[0]), PA + *a, depth)));
++a) {
*(a - 1) = *a;
}
*(a - 1) = i;
}
}
/*---------------------------------------------------------------------------*/
static INLINE
int
tr_ilg(int n) {
return (n & 0xffff0000) ?
((n & 0xff000000) ?
24 + lg_table[(n >> 24) & 0xff] :
16 + lg_table[(n >> 16) & 0xff]) :
((n & 0x0000ff00) ?
8 + lg_table[(n >> 8) & 0xff] :
0 + lg_table[(n >> 0) & 0xff]);
}
/*---------------------------------------------------------------------------*/
/* Simple insertionsort for small size groups. */
static
void
tr_insertionsort(const int *ISAd, int *first, int *last) {
int *a, *b;
int t, r;
for(a = first + 1; a < last; ++a) {
for(t = *a, b = a - 1; 0 > (r = ISAd[t] - ISAd[*b]);) {
do { *(b + 1) = *b; } while((first <= --b) && (*b < 0));
if(b < first) { break; }
}
if(r == 0) { *b = ~*b; }
*(b + 1) = t;
}
}
/*---------------------------------------------------------------------------*/
static INLINE
void
tr_fixdown(const int *ISAd, int *SA, int i, int size) {
int j, k;
int v;
int c, d, e;
for(v = SA[i], c = ISAd[v]; (j = 2 * i + 1) < size; SA[i] = SA[k], i = k) {
d = ISAd[SA[k = j++]];
if(d < (e = ISAd[SA[j]])) { k = j; d = e; }
if(d <= c) { break; }
}
SA[i] = v;
}
/* Simple top-down heapsort. */
static
void
tr_heapsort(const int *ISAd, int *SA, int size) {
int i, m;
int t;
m = size;
if((size % 2) == 0) {
m--;
if(ISAd[SA[m / 2]] < ISAd[SA[m]]) { SWAP(SA[m], SA[m / 2]); }
}
for(i = m / 2 - 1; 0 <= i; --i) { tr_fixdown(ISAd, SA, i, m); }
if((size % 2) == 0) { SWAP(SA[0], SA[m]); tr_fixdown(ISAd, SA, 0, m); }
for(i = m - 1; 0 < i; --i) {
t = SA[0], SA[0] = SA[i];
tr_fixdown(ISAd, SA, 0, i);
SA[i] = t;
}
}
/*---------------------------------------------------------------------------*/
/* Returns the median of three elements. */
static INLINE
int *
tr_median3(const int *ISAd, int *v1, int *v2, int *v3) {
int *t;
if(ISAd[*v1] > ISAd[*v2]) { SWAP(v1, v2); }
if(ISAd[*v2] > ISAd[*v3]) {
if(ISAd[*v1] > ISAd[*v3]) { return v1; }
else { return v3; }
}
return v2;
}
/* Returns the median of five elements. */
static INLINE
int *
tr_median5(const int *ISAd,
int *v1, int *v2, int *v3, int *v4, int *v5) {
int *t;
if(ISAd[*v2] > ISAd[*v3]) { SWAP(v2, v3); }
if(ISAd[*v4] > ISAd[*v5]) { SWAP(v4, v5); }
if(ISAd[*v2] > ISAd[*v4]) { SWAP(v2, v4); SWAP(v3, v5); }
if(ISAd[*v1] > ISAd[*v3]) { SWAP(v1, v3); }
if(ISAd[*v1] > ISAd[*v4]) { SWAP(v1, v4); SWAP(v3, v5); }
if(ISAd[*v3] > ISAd[*v4]) { return v4; }
return v3;
}
/* Returns the pivot element. */
static INLINE
int *
tr_pivot(const int *ISAd, int *first, int *last) {
int *middle;
int t;
t = last - first;
middle = first + t / 2;
if(t <= 512) {
if(t <= 32) {
return tr_median3(ISAd, first, middle, last - 1);
} else {
t >>= 2;
return tr_median5(ISAd, first, first + t, middle, last - 1 - t, last - 1);
}
}
t >>= 3;
first = tr_median3(ISAd, first, first + t, first + (t << 1));
middle = tr_median3(ISAd, middle - t, middle, middle + t);
last = tr_median3(ISAd, last - 1 - (t << 1), last - 1 - t, last - 1);
return tr_median3(ISAd, first, middle, last);
}
/*---------------------------------------------------------------------------*/
typedef struct _trbudget_t trbudget_t;
struct _trbudget_t {
int chance;
int remain;
int incval;
int count;
};
static INLINE
void
trbudget_init(trbudget_t *budget, int chance, int incval) {
budget->chance = chance;
budget->remain = budget->incval = incval;
}
static INLINE
int
trbudget_check(trbudget_t *budget, int size) {
if(size <= budget->remain) { budget->remain -= size; return 1; }
if(budget->chance == 0) { budget->count += size; return 0; }
budget->remain += budget->incval - size;
budget->chance -= 1;
return 1;
}
/*---------------------------------------------------------------------------*/
static INLINE
void
tr_partition(const int *ISAd,
int *first, int *middle, int *last,
int **pa, int **pb, int v) {
int *a, *b, *c, *d, *e, *f;
int t, s;
int x = 0;
for(b = middle - 1; (++b < last) && ((x = ISAd[*b]) == v);) { }
if(((a = b) < last) && (x < v)) {
for(; (++b < last) && ((x = ISAd[*b]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
}
for(c = last; (b < --c) && ((x = ISAd[*c]) == v);) { }
if((b < (d = c)) && (x > v)) {
for(; (b < --c) && ((x = ISAd[*c]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
for(; b < c;) {
SWAP(*b, *c);
for(; (++b < c) && ((x = ISAd[*b]) <= v);) {
if(x == v) { SWAP(*b, *a); ++a; }
}
for(; (b < --c) && ((x = ISAd[*c]) >= v);) {
if(x == v) { SWAP(*c, *d); --d; }
}
}
if(a <= d) {
c = b - 1;
if((s = a - first) > (t = b - a)) { s = t; }
for(e = first, f = b - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
if((s = d - c) > (t = last - d - 1)) { s = t; }
for(e = b, f = last - s; 0 < s; --s, ++e, ++f) { SWAP(*e, *f); }
first += (b - a), last -= (d - c);
}
*pa = first, *pb = last;
}
static
void
tr_copy(int *ISA, const int *SA,
int *first, int *a, int *b, int *last,
int depth) {
/* sort suffixes of middle partition
by using sorted order of suffixes of left and right partition. */
int *c, *d, *e;
int s, v;
v = b - SA - 1;
for(c = first, d = a - 1; c <= d; ++c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*++d = s;
ISA[s] = d - SA;
}
}
for(c = last - 1, e = d + 1, d = b; e < d; --c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*--d = s;
ISA[s] = d - SA;
}
}
}
static
void
tr_partialcopy(int *ISA, const int *SA,
int *first, int *a, int *b, int *last,
int depth) {
int *c, *d, *e;
int s, v;
int rank, lastrank, newrank = -1;
v = b - SA - 1;
lastrank = -1;
for(c = first, d = a - 1; c <= d; ++c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*++d = s;
rank = ISA[s + depth];
if(lastrank != rank) { lastrank = rank; newrank = d - SA; }
ISA[s] = newrank;
}
}
lastrank = -1;
for(e = d; first <= e; --e) {
rank = ISA[*e];
if(lastrank != rank) { lastrank = rank; newrank = e - SA; }
if(newrank != rank) { ISA[*e] = newrank; }
}
lastrank = -1;
for(c = last - 1, e = d + 1, d = b; e < d; --c) {
if((0 <= (s = *c - depth)) && (ISA[s] == v)) {
*--d = s;
rank = ISA[s + depth];
if(lastrank != rank) { lastrank = rank; newrank = d - SA; }
ISA[s] = newrank;
}
}
}
static
void
tr_introsort(int *ISA, const int *ISAd,
int *SA, int *first, int *last,
trbudget_t *budget) {
#define STACK_SIZE TR_STACKSIZE
struct { const int *a; int *b, *c; int d, e; }stack[STACK_SIZE];
int *a, *b, *c;
int t;
int v, x = 0;
int incr = ISAd - ISA;
int limit, next;
int ssize, trlink = -1;
for(ssize = 0, limit = tr_ilg(last - first);;) {
if(limit < 0) {
if(limit == -1) {
/* tandem repeat partition */
tr_partition(ISAd - incr, first, first, last, &a, &b, last - SA - 1);
/* update ranks */
if(a < last) {
for(c = first, v = a - SA - 1; c < a; ++c) { ISA[*c] = v; }
}
if(b < last) {
for(c = a, v = b - SA - 1; c < b; ++c) { ISA[*c] = v; }
}
/* push */
if(1 < (b - a)) {
STACK_PUSH5(NULL, a, b, 0, 0);
STACK_PUSH5(ISAd - incr, first, last, -2, trlink);
trlink = ssize - 2;
}
if((a - first) <= (last - b)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, tr_ilg(last - b), trlink);
last = a, limit = tr_ilg(a - first);
} else if(1 < (last - b)) {
first = b, limit = tr_ilg(last - b);
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
} else {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, tr_ilg(a - first), trlink);
first = b, limit = tr_ilg(last - b);
} else if(1 < (a - first)) {
last = a, limit = tr_ilg(a - first);
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
} else if(limit == -2) {
/* tandem repeat copy */
a = stack[--ssize].b, b = stack[ssize].c;
if(stack[ssize].d == 0) {
tr_copy(ISA, SA, first, a, b, last, ISAd - ISA);
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
tr_partialcopy(ISA, SA, first, a, b, last, ISAd - ISA);
}
STACK_POP5(ISAd, first, last, limit, trlink);
} else {
/* sorted partition */
if(0 <= *first) {
a = first;
do { ISA[*a] = a - SA; } while((++a < last) && (0 <= *a));
first = a;
}
if(first < last) {
a = first; do { *a = ~*a; } while(*++a < 0);
next = (ISA[*a] != ISAd[*a]) ? tr_ilg(a - first + 1) : -1;
if(++a < last) { for(b = first, v = a - SA - 1; b < a; ++b) { ISA[*b] = v; } }
/* push */
if(trbudget_check(budget, a - first)) {
if((a - first) <= (last - a)) {
STACK_PUSH5(ISAd, a, last, -3, trlink);
ISAd += incr, last = a, limit = next;
} else {
if(1 < (last - a)) {
STACK_PUSH5(ISAd + incr, first, a, next, trlink);
first = a, limit = -3;
} else {
ISAd += incr, last = a, limit = next;
}
}
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
if(1 < (last - a)) {
first = a, limit = -3;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
continue;
}
if((last - first) <= TR_INSERTIONSORT_THRESHOLD) {
tr_insertionsort(ISAd, first, last);
limit = -3;
continue;
}
if(limit-- == 0) {
tr_heapsort(ISAd, first, last - first);
for(a = last - 1; first < a; a = b) {
for(x = ISAd[*a], b = a - 1; (first <= b) && (ISAd[*b] == x); --b) { *b = ~*b; }
}
limit = -3;
continue;
}
/* choose pivot */
a = tr_pivot(ISAd, first, last);
SWAP(*first, *a);
v = ISAd[*first];
/* partition */
tr_partition(ISAd, first, first + 1, last, &a, &b, v);
if((last - first) != (b - a)) {
next = (ISA[*a] != v) ? tr_ilg(b - a) : -1;
/* update ranks */
for(c = first, v = a - SA - 1; c < a; ++c) { ISA[*c] = v; }
if(b < last) { for(c = a, v = b - SA - 1; c < b; ++c) { ISA[*c] = v; } }
/* push */
if((1 < (b - a)) && (trbudget_check(budget, b - a))) {
if((a - first) <= (last - b)) {
if((last - b) <= (b - a)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
STACK_PUSH5(ISAd, b, last, limit, trlink);
last = a;
} else if(1 < (last - b)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
first = b;
} else {
ISAd += incr, first = a, last = b, limit = next;
}
} else if((a - first) <= (b - a)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, limit, trlink);
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
last = a;
} else {
STACK_PUSH5(ISAd, b, last, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
STACK_PUSH5(ISAd, b, last, limit, trlink);
STACK_PUSH5(ISAd, first, a, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
if((a - first) <= (b - a)) {
if(1 < (last - b)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
STACK_PUSH5(ISAd, first, a, limit, trlink);
first = b;
} else if(1 < (a - first)) {
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
last = a;
} else {
ISAd += incr, first = a, last = b, limit = next;
}
} else if((last - b) <= (b - a)) {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, limit, trlink);
STACK_PUSH5(ISAd + incr, a, b, next, trlink);
first = b;
} else {
STACK_PUSH5(ISAd, first, a, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
} else {
STACK_PUSH5(ISAd, first, a, limit, trlink);
STACK_PUSH5(ISAd, b, last, limit, trlink);
ISAd += incr, first = a, last = b, limit = next;
}
}
} else {
if((1 < (b - a)) && (0 <= trlink)) { stack[trlink].d = -1; }
if((a - first) <= (last - b)) {
if(1 < (a - first)) {
STACK_PUSH5(ISAd, b, last, limit, trlink);
last = a;
} else if(1 < (last - b)) {
first = b;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
} else {
if(1 < (last - b)) {
STACK_PUSH5(ISAd, first, a, limit, trlink);
first = b;
} else if(1 < (a - first)) {
last = a;
} else {
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
}
} else {
if(trbudget_check(budget, last - first)) {
limit = tr_ilg(last - first), ISAd += incr;
} else {
if(0 <= trlink) { stack[trlink].d = -1; }
STACK_POP5(ISAd, first, last, limit, trlink);
}
}
}
#undef STACK_SIZE
}
/*---------------------------------------------------------------------------*/
/* Tandem repeat sort */
static
void
trsort(int *ISA, int *SA, int n, int depth) {
int *ISAd;
int *first, *last;
trbudget_t budget;
int t, skip, unsorted;
trbudget_init(&budget, tr_ilg(n) * 2 / 3, n);
/* trbudget_init(&budget, tr_ilg(n) * 3 / 4, n); */
for(ISAd = ISA + depth; -n < *SA; ISAd += ISAd - ISA) {
first = SA;
skip = 0;
unsorted = 0;
do {
if((t = *first) < 0) { first -= t; skip += t; }
else {
if(skip != 0) { *(first + skip) = skip; skip = 0; }
last = SA + ISA[t] + 1;
if(1 < (last - first)) {
budget.count = 0;
tr_introsort(ISA, ISAd, SA, first, last, &budget);
if(budget.count != 0) { unsorted += budget.count; }
else { skip = first - last; }
} else if((last - first) == 1) {
skip = -1;
}
first = last;
}
} while(first < (SA + n));
if(skip != 0) { *(first + skip) = skip; }
if(unsorted == 0) { break; }
}
}
/*---------------------------------------------------------------------------*/
/* Sorts suffixes of type B*. */
static
int
sort_typeBstar(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int openMP) {
int *PAb, *ISAb, *buf;
#ifdef LIBBSC_OPENMP
int *curbuf;
int l;
#endif
int i, j, k, t, m, bufsize;
int c0, c1;
#ifdef LIBBSC_OPENMP
int d0, d1;
#endif
(void)openMP;
/* Initialize bucket arrays. */
for(i = 0; i < BUCKET_A_SIZE; ++i) { bucket_A[i] = 0; }
for(i = 0; i < BUCKET_B_SIZE; ++i) { bucket_B[i] = 0; }
/* Count the number of occurrences of the first one or two characters of each
type A, B and B* suffix. Moreover, store the beginning position of all
type B* suffixes into the array SA. */
for(i = n - 1, m = n, c0 = T[n - 1]; 0 <= i;) {
/* type A suffix. */
do { ++BUCKET_A(c1 = c0); } while((0 <= --i) && ((c0 = T[i]) >= c1));
if(0 <= i) {
/* type B* suffix. */
++BUCKET_BSTAR(c0, c1);
SA[--m] = i;
/* type B suffix. */
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) <= c1); --i, c1 = c0) {
++BUCKET_B(c0, c1);
}
}
}
m = n - m;
/*
note:
A type B* suffix is lexicographically smaller than a type B suffix that
begins with the same first two characters.
*/
/* Calculate the index of start/end point of each bucket. */
for(c0 = 0, i = 0, j = 0; c0 < ALPHABET_SIZE; ++c0) {
t = i + BUCKET_A(c0);
BUCKET_A(c0) = i + j; /* start point */
i = t + BUCKET_B(c0, c0);
for(c1 = c0 + 1; c1 < ALPHABET_SIZE; ++c1) {
j += BUCKET_BSTAR(c0, c1);
BUCKET_BSTAR(c0, c1) = j; /* end point */
i += BUCKET_B(c0, c1);
}
}
if(0 < m) {
/* Sort the type B* suffixes by their first two characters. */
PAb = SA + n - m; ISAb = SA + m;
for(i = m - 2; 0 <= i; --i) {
t = PAb[i], c0 = T[t], c1 = T[t + 1];
SA[--BUCKET_BSTAR(c0, c1)] = i;
}
t = PAb[m - 1], c0 = T[t], c1 = T[t + 1];
SA[--BUCKET_BSTAR(c0, c1)] = m - 1;
/* Sort the type B* substrings using sssort. */
#ifdef LIBBSC_OPENMP
if (openMP)
{
buf = SA + m;
c0 = ALPHABET_SIZE - 2, c1 = ALPHABET_SIZE - 1, j = m;
#pragma omp parallel default(shared) private(bufsize, curbuf, k, l, d0, d1)
{
bufsize = (n - (2 * m)) / omp_get_num_threads();
curbuf = buf + omp_get_thread_num() * bufsize;
k = 0;
for(;;) {
#pragma omp critical(sssort_lock)
{
if(0 < (l = j)) {
d0 = c0, d1 = c1;
do {
k = BUCKET_BSTAR(d0, d1);
if(--d1 <= d0) {
d1 = ALPHABET_SIZE - 1;
if(--d0 < 0) { break; }
}
} while(((l - k) <= 1) && (0 < (l = k)));
c0 = d0, c1 = d1, j = k;
}
}
if(l == 0) { break; }
sssort(T, PAb, SA + k, SA + l,
curbuf, bufsize, 2, n, *(SA + k) == (m - 1));
}
}
}
else
{
buf = SA + m, bufsize = n - (2 * m);
for(c0 = ALPHABET_SIZE - 2, j = m; 0 < j; --c0) {
for(c1 = ALPHABET_SIZE - 1; c0 < c1; j = i, --c1) {
i = BUCKET_BSTAR(c0, c1);
if(1 < (j - i)) {
sssort(T, PAb, SA + i, SA + j,
buf, bufsize, 2, n, *(SA + i) == (m - 1));
}
}
}
}
#else
buf = SA + m, bufsize = n - (2 * m);
for(c0 = ALPHABET_SIZE - 2, j = m; 0 < j; --c0) {
for(c1 = ALPHABET_SIZE - 1; c0 < c1; j = i, --c1) {
i = BUCKET_BSTAR(c0, c1);
if(1 < (j - i)) {
sssort(T, PAb, SA + i, SA + j,
buf, bufsize, 2, n, *(SA + i) == (m - 1));
}
}
}
#endif
/* Compute ranks of type B* substrings. */
for(i = m - 1; 0 <= i; --i) {
if(0 <= SA[i]) {
j = i;
do { ISAb[SA[i]] = i; } while((0 <= --i) && (0 <= SA[i]));
SA[i + 1] = i - j;
if(i <= 0) { break; }
}
j = i;
do { ISAb[SA[i] = ~SA[i]] = j; } while(SA[--i] < 0);
ISAb[SA[i]] = j;
}
/* Construct the inverse suffix array of type B* suffixes using trsort. */
trsort(ISAb, SA, m, 1);
/* Set the sorted order of type B* suffixes. */
for(i = n - 1, j = m, c0 = T[n - 1]; 0 <= i;) {
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) >= c1); --i, c1 = c0) { }
if(0 <= i) {
t = i;
for(--i, c1 = c0; (0 <= i) && ((c0 = T[i]) <= c1); --i, c1 = c0) { }
SA[ISAb[--j]] = ((t == 0) || (1 < (t - i))) ? t : ~t;
}
}
/* Calculate the index of start/end point of each bucket. */
BUCKET_B(ALPHABET_SIZE - 1, ALPHABET_SIZE - 1) = n; /* end point */
for(c0 = ALPHABET_SIZE - 2, k = m - 1; 0 <= c0; --c0) {
i = BUCKET_A(c0 + 1) - 1;
for(c1 = ALPHABET_SIZE - 1; c0 < c1; --c1) {
t = i - BUCKET_B(c0, c1);
BUCKET_B(c0, c1) = i; /* end point */
/* Move all type B* suffixes to the correct position. */
for(i = t, j = BUCKET_BSTAR(c0, c1);
j <= k;
--i, --k) { SA[i] = SA[k]; }
}
BUCKET_BSTAR(c0, c0 + 1) = i - BUCKET_B(c0, c0) + 1; /* start point */
BUCKET_B(c0, c0) = i; /* end point */
}
}
return m;
}
/* Constructs the suffix array by using the sorted order of type B* suffixes. */
static
void
construct_SA(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m) {
int *i, *j, *k;
int s;
int c0, c1, c2;
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
*j = ~s;
c0 = T[--s];
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else {
assert(((s == 0) && (T[s] == c1)) || (s < 0));
*j = ~s;
}
}
}
}
/* Construct the suffix array by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
*k++ = (T[n - 2] < c2) ? ~(n - 1) : (n - 1);
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
c0 = T[--s];
if((s == 0) || (T[s - 1] < c0)) { s = ~s; }
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
*k++ = s;
} else {
assert(s < 0);
*i = ~s;
}
}
}
/* Constructs the burrows-wheeler transformed string directly
by using the sorted order of type B* suffixes. */
static
int
construct_BWT(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m) {
int *i, *j, *k, *orig;
int s;
int c0, c1, c2;
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
c0 = T[--s];
*j = ~((int)c0);
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else if(s != 0) {
*j = ~s;
#ifndef NDEBUG
} else {
assert(T[s] == c1);
#endif
}
}
}
}
/* Construct the BWTed string by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
*k++ = (T[n - 2] < c2) ? ~((int)T[n - 2]) : (n - 1);
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n, orig = SA; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
c0 = T[--s];
*i = c0;
if((0 < s) && (T[s - 1] < c0)) { s = ~((int)T[s - 1]); }
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
*k++ = s;
} else if(s != 0) {
*i = ~s;
} else {
orig = i;
}
}
return orig - SA;
}
/* Constructs the burrows-wheeler transformed string directly
by using the sorted order of type B* suffixes. */
static
int
construct_BWT_indexes(const unsigned char *T, int *SA,
int *bucket_A, int *bucket_B,
int n, int m,
unsigned char * num_indexes, int * indexes) {
int *i, *j, *k, *orig;
int s;
int c0, c1, c2;
int mod = n / 8;
{
mod |= mod >> 1; mod |= mod >> 2;
mod |= mod >> 4; mod |= mod >> 8;
mod |= mod >> 16; mod >>= 1;
*num_indexes = (unsigned char)((n - 1) / (mod + 1));
}
if(0 < m) {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
for(c1 = ALPHABET_SIZE - 2; 0 <= c1; --c1) {
/* Scan the suffix array from right to left. */
for(i = SA + BUCKET_BSTAR(c1, c1 + 1),
j = SA + BUCKET_A(c1 + 1) - 1, k = NULL, c2 = -1;
i <= j;
--j) {
if(0 < (s = *j)) {
assert(T[s] == c1);
assert(((s + 1) < n) && (T[s] <= T[s + 1]));
assert(T[s - 1] <= T[s]);
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = j - SA;
c0 = T[--s];
*j = ~((int)c0);
if((0 < s) && (T[s - 1] > c0)) { s = ~s; }
if(c0 != c2) {
if(0 <= c2) { BUCKET_B(c2, c1) = k - SA; }
k = SA + BUCKET_B(c2 = c0, c1);
}
assert(k < j); assert(k != NULL);
*k-- = s;
} else if(s != 0) {
*j = ~s;
#ifndef NDEBUG
} else {
assert(T[s] == c1);
#endif
}
}
}
}
/* Construct the BWTed string by using
the sorted order of type B suffixes. */
k = SA + BUCKET_A(c2 = T[n - 1]);
if (T[n - 2] < c2) {
if (((n - 1) & mod) == 0) indexes[(n - 1) / (mod + 1) - 1] = k - SA;
*k++ = ~((int)T[n - 2]);
}
else {
*k++ = n - 1;
}
/* Scan the suffix array from left to right. */
for(i = SA, j = SA + n, orig = SA; i < j; ++i) {
if(0 < (s = *i)) {
assert(T[s - 1] >= T[s]);
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = i - SA;
c0 = T[--s];
*i = c0;
if(c0 != c2) {
BUCKET_A(c2) = k - SA;
k = SA + BUCKET_A(c2 = c0);
}
assert(i < k);
if((0 < s) && (T[s - 1] < c0)) {
if ((s & mod) == 0) indexes[s / (mod + 1) - 1] = k - SA;
*k++ = ~((int)T[s - 1]);
} else
*k++ = s;
} else if(s != 0) {
*i = ~s;
} else {
orig = i;
}
}
return orig - SA;
}
/*---------------------------------------------------------------------------*/
/*- Function -*/
int
divsufsort(const unsigned char *T, int *SA, int n, int openMP) {
int *bucket_A, *bucket_B;
int m;
int err = 0;
/* Check arguments. */
if((T == NULL) || (SA == NULL) || (n < 0)) { return -1; }
else if(n == 0) { return 0; }
else if(n == 1) { SA[0] = 0; return 0; }
else if(n == 2) { m = (T[0] < T[1]); SA[m ^ 1] = 0, SA[m] = 1; return 0; }
bucket_A = (int *)malloc(BUCKET_A_SIZE * sizeof(int));
bucket_B = (int *)malloc(BUCKET_B_SIZE * sizeof(int));
/* Suffixsort. */
if((bucket_A != NULL) && (bucket_B != NULL)) {
m = sort_typeBstar(T, SA, bucket_A, bucket_B, n, openMP);
construct_SA(T, SA, bucket_A, bucket_B, n, m);
} else {
err = -2;
}
free(bucket_B);
free(bucket_A);
return err;
}
int
divbwt(const unsigned char *T, unsigned char *U, int *A, int n, unsigned char * num_indexes, int * indexes, int openMP) {
int *B;
int *bucket_A, *bucket_B;
int m, pidx, i;
/* Check arguments. */
if((T == NULL) || (U == NULL) || (n < 0)) { return -1; }
else if(n <= 1) { if(n == 1) { U[0] = T[0]; } return n; }
if((B = A) == NULL) { B = (int *)malloc((size_t)(n + 1) * sizeof(int)); }
bucket_A = (int *)malloc(BUCKET_A_SIZE * sizeof(int));
bucket_B = (int *)malloc(BUCKET_B_SIZE * sizeof(int));
/* Burrows-Wheeler Transform. */
if((B != NULL) && (bucket_A != NULL) && (bucket_B != NULL)) {
m = sort_typeBstar(T, B, bucket_A, bucket_B, n, openMP);
if (num_indexes == NULL || indexes == NULL) {
pidx = construct_BWT(T, B, bucket_A, bucket_B, n, m);
} else {
pidx = construct_BWT_indexes(T, B, bucket_A, bucket_B, n, m, num_indexes, indexes);
}
/* Copy to output string. */
U[0] = T[n - 1];
for(i = 0; i < pidx; ++i) { U[i + 1] = (unsigned char)B[i]; }
for(i += 1; i < n; ++i) { U[i] = (unsigned char)B[i]; }
pidx += 1;
} else {
pidx = -2;
}
free(bucket_B);
free(bucket_A);
if(A == NULL) { free(B); }
return pidx;
}
+8 -20
View File
@@ -32,8 +32,7 @@ RUST_CLI_MANIFEST := $(RUST_CLI_DIR)/Cargo.toml
RUST_SOURCES := $(RUST_MANIFEST) $(RUST_DIR)/Cargo.lock \ RUST_SOURCES := $(RUST_MANIFEST) $(RUST_DIR)/Cargo.lock \
$(shell find $(RUST_DIR)/src -type f -name '*.rs' -print) $(shell find $(RUST_DIR)/src -type f -name '*.rs' -print)
RUST_CLI_SOURCES := $(RUST_CLI_MANIFEST) $(RUST_CLI_DIR)/Cargo.lock \ RUST_CLI_SOURCES := $(RUST_CLI_MANIFEST) $(RUST_CLI_DIR)/Cargo.lock \
$(RUST_CLI_DIR)/src/lib.rs $(RUST_DIR)/src/zstd_cli.rs \ $(RUST_CLI_DIR)/src/lib.rs $(RUST_DIR)/src/zstd_cli.rs
$(RUST_DIR)/src/timefn.rs $(RUST_DIR)/src/benchfn.rs
# Keep Rust's HUF implementation in lockstep with libzstd.mk's C selection. # Keep Rust's HUF implementation in lockstep with libzstd.mk's C selection.
# Forced modes may arrive as libzstd.mk variables or as direct -D flags in # Forced modes may arrive as libzstd.mk variables or as direct -D flags in
@@ -89,7 +88,7 @@ RUST_TARGET_32 ?= i686-unknown-linux-gnu
RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a
RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \ RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_TARGET_DIR) --no-default-features --target-dir $(RUST_TARGET_DIR) --no-default-features
RUST_CARGO_FLAGS += --features compression,decompression,dict-builder RUST_CARGO_FLAGS += --features compression,decompression
ifneq ($(RUST_HUF_FEATURE),) ifneq ($(RUST_HUF_FEATURE),)
RUST_CARGO_FLAGS += --features $(RUST_HUF_FEATURE) RUST_CARGO_FLAGS += --features $(RUST_HUF_FEATURE)
endif endif
@@ -109,7 +108,7 @@ RUST_CLI_STATICLIB := $(RUST_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_CLI_STATICLIB_32 := $(RUST_CLI_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_cli_rs.a RUST_CLI_STATICLIB_32 := $(RUST_CLI_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_cli_rs.a
RUST_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \ RUST_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_CLI_TARGET_DIR) \ --target-dir $(RUST_CLI_TARGET_DIR) \
--no-default-features --features cli,compression,decompression --no-default-features --features compression,decompression
$(RUST_CLI_STATICLIB): $(RUST_CLI_SOURCES) $(RUST_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_CLI_CARGO_FLAGS) $(CARGO) build $(RUST_CLI_CARGO_FLAGS)
@@ -117,7 +116,7 @@ $(RUST_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(RUST_CLI_STATICLIB_32): $(RUST_CLI_SOURCES) $(RUST_CLI_STATICLIB_32): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_CLI_CARGO_FLAGS) --target $(RUST_TARGET_32) $(CARGO) build $(RUST_CLI_CARGO_FLAGS) --target $(RUST_TARGET_32)
RUST_DECOMPRESS_BUILD_CONFIG := lib-c0-d1-b0-$(RUST_HUF_MODE) RUST_DECOMPRESS_BUILD_CONFIG := lib-c0-d1-$(RUST_HUF_MODE)
RUST_DECOMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DECOMPRESS_BUILD_CONFIG) RUST_DECOMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DECOMPRESS_BUILD_CONFIG)
RUST_DECOMPRESS_STATICLIB := $(RUST_DECOMPRESS_TARGET_DIR)/release/libzstd_rs.a RUST_DECOMPRESS_STATICLIB := $(RUST_DECOMPRESS_TARGET_DIR)/release/libzstd_rs.a
RUST_DECOMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \ RUST_DECOMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
@@ -135,12 +134,12 @@ RUST_DECOMPRESS_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DECOMPRESS_CLI_BUILD
RUST_DECOMPRESS_CLI_STATICLIB := $(RUST_DECOMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a RUST_DECOMPRESS_CLI_STATICLIB := $(RUST_DECOMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_DECOMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \ RUST_DECOMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_DECOMPRESS_CLI_TARGET_DIR) \ --target-dir $(RUST_DECOMPRESS_CLI_TARGET_DIR) \
--no-default-features --features cli,decompression --no-default-features --features decompression
$(RUST_DECOMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES) $(RUST_DECOMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_DECOMPRESS_CLI_CARGO_FLAGS) $(CARGO) build $(RUST_DECOMPRESS_CLI_CARGO_FLAGS)
RUST_COMPRESS_BUILD_CONFIG := lib-c1-d0-b0-$(RUST_HUF_MODE) RUST_COMPRESS_BUILD_CONFIG := lib-c1-d0-$(RUST_HUF_MODE)
RUST_COMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_BUILD_CONFIG) RUST_COMPRESS_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_BUILD_CONFIG)
RUST_COMPRESS_STATICLIB := $(RUST_COMPRESS_TARGET_DIR)/release/libzstd_rs.a RUST_COMPRESS_STATICLIB := $(RUST_COMPRESS_TARGET_DIR)/release/libzstd_rs.a
RUST_COMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \ RUST_COMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
@@ -150,23 +149,12 @@ RUST_COMPRESS_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
$(RUST_COMPRESS_STATICLIB): $(RUST_SOURCES) $(RUST_COMPRESS_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_COMPRESS_CARGO_FLAGS) $(CARGO) build $(RUST_COMPRESS_CARGO_FLAGS)
RUST_DICTBUILDER_BUILD_CONFIG := lib-c1-d0-b1-$(RUST_HUF_MODE)
RUST_DICTBUILDER_TARGET_DIR := $(RUST_DIR)/target/$(RUST_DICTBUILDER_BUILD_CONFIG)
RUST_DICTBUILDER_STATICLIB := $(RUST_DICTBUILDER_TARGET_DIR)/release/libzstd_rs.a
RUST_DICTBUILDER_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_DICTBUILDER_TARGET_DIR) \
--no-default-features \
--features compression,dict-builder
$(RUST_DICTBUILDER_STATICLIB): $(RUST_SOURCES)
$(CARGO) build $(RUST_DICTBUILDER_CARGO_FLAGS)
RUST_COMPRESS_CLI_BUILD_CONFIG := cli-c1-d0-$(RUST_HUF_MODE) RUST_COMPRESS_CLI_BUILD_CONFIG := cli-c1-d0-$(RUST_HUF_MODE)
RUST_COMPRESS_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_CLI_BUILD_CONFIG) RUST_COMPRESS_CLI_TARGET_DIR := $(RUST_DIR)/target/$(RUST_COMPRESS_CLI_BUILD_CONFIG)
RUST_COMPRESS_CLI_STATICLIB := $(RUST_COMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a RUST_COMPRESS_CLI_STATICLIB := $(RUST_COMPRESS_CLI_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_COMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \ RUST_COMPRESS_CLI_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_COMPRESS_CLI_TARGET_DIR) \ --target-dir $(RUST_COMPRESS_CLI_TARGET_DIR) \
--no-default-features --features cli,compression --no-default-features --features compression
$(RUST_COMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES) $(RUST_COMPRESS_CLI_STATICLIB): $(RUST_CLI_SOURCES)
$(CARGO) build $(RUST_COMPRESS_CLI_CARGO_FLAGS) $(CARGO) build $(RUST_COMPRESS_CLI_CARGO_FLAGS)
@@ -444,7 +432,7 @@ zstd-compress: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) zstdcli.c util.c ti
## zstd-dictBuilder: executable supporting dictionary creation and compression (only) ## zstd-dictBuilder: executable supporting dictionary creation and compression (only)
CLEAN += zstd-dictBuilder CLEAN += zstd-dictBuilder
zstd-dictBuilder: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) $(ZDICT_SRC) zstdcli.c util.c timefn.c fileio.c fileio_asyncio.c dibio.c $(RUST_DICTBUILDER_STATICLIB) $(RUST_COMPRESS_CLI_STATICLIB) zstd-dictBuilder: $(ZSTDLIB_COMMON_SRC) $(ZSTDLIB_COMPRESS_SRC) $(ZDICT_SRC) zstdcli.c util.c timefn.c fileio.c fileio_asyncio.c dibio.c $(RUST_COMPRESS_STATICLIB) $(RUST_COMPRESS_CLI_STATICLIB)
$(CC) $(FLAGS) -DZSTD_NOBENCH -DZSTD_NODECOMPRESS -DZSTD_NOTRACE $^ -o $@$(EXT) $(CC) $(FLAGS) -DZSTD_NOBENCH -DZSTD_NODECOMPRESS -DZSTD_NOTRACE $^ -o $@$(EXT)
RUST_DIRECT_LINK_TARGETS := zstd32 zstd-nolegacy zstd-small zstd-frugal \ RUST_DIRECT_LINK_TARGETS := zstd32 zstd-nolegacy zstd-small zstd-frugal \
+243 -17
View File
@@ -8,23 +8,249 @@
* You may select, at your option, one of the above-listed licenses. * You may select, at your option, one of the above-listed licenses.
*/ */
/* The implementation lives in rust/src/benchfn.rs, built into the Rust CLI
* static archive. This translation unit stays in the original source lists so
* build configuration keeps working while the implementation is in Rust. */
#include <stddef.h> /* size_t, offsetof */
/* *************************************
* Includes
***************************************/
#include <stdlib.h> /* malloc, free */
#include <string.h> /* memset */
#include <assert.h> /* assert */
#include "timefn.h" /* UTIL_time_t, UTIL_getTime */
#include "benchfn.h" #include "benchfn.h"
/* BMK_runTime_t and BMK_runOutcome_t are returned by value across the C/Rust
* boundary, and BMK_benchParams_t is passed by value. The Rust #[repr(C)] /* *************************************
* definitions mirror the offsets pinned here. */ * Constants
typedef char BMK_staticAssert_runTimeSumOffset[ ***************************************/
(offsetof(BMK_runTime_t, sumOfReturn) == sizeof(double)) ? 1 : -1]; #define TIMELOOP_MICROSEC SEC_TO_MICRO /* 1 second */
typedef char BMK_staticAssert_outcomeResultOffset[ #define TIMELOOP_NANOSEC (1*1000000000ULL) /* 1 second */
(offsetof(BMK_runOutcome_t, error_result_never_ever_use_directly)
== sizeof(BMK_runTime_t)) ? 1 : -1]; #define KB *(1 <<10)
typedef char BMK_staticAssert_outcomeTagOffset[ #define MB *(1 <<20)
(offsetof(BMK_runOutcome_t, error_tag_never_ever_use_directly) #define GB *(1U<<30)
== sizeof(BMK_runTime_t) + sizeof(size_t)) ? 1 : -1];
typedef char BMK_staticAssert_shellAlignment[
(sizeof(BMK_timedFnState_shell) == BMK_TIMEDFNSTATE_SIZE) ? 1 : -1]; /* *************************************
* Debug errors
***************************************/
#if defined(DEBUG) && (DEBUG >= 1)
# include <stdio.h> /* fprintf */
# define DISPLAY(...) fprintf(stderr, __VA_ARGS__)
# define DEBUGOUTPUT(...) { if (DEBUG) DISPLAY(__VA_ARGS__); }
#else
# define DEBUGOUTPUT(...)
#endif
/* error without displaying */
#define RETURN_QUIET_ERROR(retValue, ...) { \
DEBUGOUTPUT("%s: %i: \n", __FILE__, __LINE__); \
DEBUGOUTPUT("Error : "); \
DEBUGOUTPUT(__VA_ARGS__); \
DEBUGOUTPUT(" \n"); \
return retValue; \
}
/* Abort execution if a condition is not met */
#define CONTROL(c) { if (!(c)) { DEBUGOUTPUT("error: %s \n", #c); abort(); } }
/* *************************************
* Benchmarking an arbitrary function
***************************************/
int BMK_isSuccessful_runOutcome(BMK_runOutcome_t outcome)
{
return outcome.error_tag_never_ever_use_directly == 0;
}
/* warning : this function will stop program execution if outcome is invalid !
* check outcome validity first, using BMK_isValid_runResult() */
BMK_runTime_t BMK_extract_runTime(BMK_runOutcome_t outcome)
{
CONTROL(outcome.error_tag_never_ever_use_directly == 0);
return outcome.internal_never_ever_use_directly;
}
size_t BMK_extract_errorResult(BMK_runOutcome_t outcome)
{
CONTROL(outcome.error_tag_never_ever_use_directly != 0);
return outcome.error_result_never_ever_use_directly;
}
static BMK_runOutcome_t BMK_runOutcome_error(size_t errorResult)
{
BMK_runOutcome_t b;
memset(&b, 0, sizeof(b));
b.error_tag_never_ever_use_directly = 1;
b.error_result_never_ever_use_directly = errorResult;
return b;
}
static BMK_runOutcome_t BMK_setValid_runTime(BMK_runTime_t runTime)
{
BMK_runOutcome_t outcome;
outcome.error_tag_never_ever_use_directly = 0;
outcome.internal_never_ever_use_directly = runTime;
return outcome;
}
/* initFn will be measured once, benchFn will be measured `nbLoops` times */
/* initFn is optional, provide NULL if none */
/* benchFn must return a size_t value that errorFn can interpret */
/* takes # of blocks and list of size & stuff for each. */
/* can report result of benchFn for each block into blockResult. */
/* blockResult is optional, provide NULL if this information is not required */
/* note : time per loop can be reported as zero if run time < timer resolution */
BMK_runOutcome_t BMK_benchFunction(BMK_benchParams_t p,
unsigned nbLoops)
{
nbLoops += !nbLoops; /* minimum nbLoops is 1 */
/* init */
{ size_t i;
for(i = 0; i < p.blockCount; i++) {
memset(p.dstBuffers[i], 0xE5, p.dstCapacities[i]); /* warm up and erase result buffer */
} }
/* benchmark */
{ size_t dstSize = 0;
UTIL_time_t const clockStart = UTIL_getTime();
unsigned loopNb, blockNb;
if (p.initFn != NULL) p.initFn(p.initPayload);
for (loopNb = 0; loopNb < nbLoops; loopNb++) {
for (blockNb = 0; blockNb < p.blockCount; blockNb++) {
size_t const res = p.benchFn(p.srcBuffers[blockNb], p.srcSizes[blockNb],
p.dstBuffers[blockNb], p.dstCapacities[blockNb],
p.benchPayload);
if (loopNb == 0) {
if (p.blockResults != NULL) p.blockResults[blockNb] = res;
if ((p.errorFn != NULL) && (p.errorFn(res))) {
RETURN_QUIET_ERROR(BMK_runOutcome_error(res),
"Function benchmark failed on block %u (of size %u) with error %i",
blockNb, (unsigned)p.srcSizes[blockNb], (int)res);
}
dstSize += res;
} }
} /* for (loopNb = 0; loopNb < nbLoops; loopNb++) */
{ PTime const totalTime = UTIL_clockSpanNano(clockStart);
BMK_runTime_t rt;
rt.nanoSecPerRun = (double)totalTime / nbLoops;
rt.sumOfReturn = dstSize;
return BMK_setValid_runTime(rt);
} }
}
/* ==== Benchmarking any function, providing intermediate results ==== */
struct BMK_timedFnState_s {
PTime timeSpent_ns;
PTime timeBudget_ns;
PTime runBudget_ns;
BMK_runTime_t fastestRun;
unsigned nbLoops;
UTIL_time_t coolTime;
}; /* typedef'd to BMK_timedFnState_t within bench.h */
BMK_timedFnState_t* BMK_createTimedFnState(unsigned total_ms, unsigned run_ms)
{
BMK_timedFnState_t* const r = (BMK_timedFnState_t*)malloc(sizeof(*r));
if (r == NULL) return NULL; /* malloc() error */
BMK_resetTimedFnState(r, total_ms, run_ms);
return r;
}
void BMK_freeTimedFnState(BMK_timedFnState_t* state) { free(state); }
BMK_timedFnState_t*
BMK_initStatic_timedFnState(void* buffer, size_t size, unsigned total_ms, unsigned run_ms)
{
typedef char check_size[ 2 * (sizeof(BMK_timedFnState_shell) >= sizeof(struct BMK_timedFnState_s)) - 1]; /* static assert : a compilation failure indicates that BMK_timedFnState_shell is not large enough */
typedef struct { check_size c; BMK_timedFnState_t tfs; } tfs_align; /* force tfs to be aligned at its next best position */
size_t const tfs_alignment = offsetof(tfs_align, tfs); /* provides the minimal alignment restriction for BMK_timedFnState_t */
BMK_timedFnState_t* const r = (BMK_timedFnState_t*)buffer;
if (buffer == NULL) return NULL;
if (size < sizeof(struct BMK_timedFnState_s)) return NULL;
if ((size_t)buffer % tfs_alignment) return NULL; /* buffer must be properly aligned */
BMK_resetTimedFnState(r, total_ms, run_ms);
return r;
}
void BMK_resetTimedFnState(BMK_timedFnState_t* timedFnState, unsigned total_ms, unsigned run_ms)
{
if (!total_ms) total_ms = 1 ;
if (!run_ms) run_ms = 1;
if (run_ms > total_ms) run_ms = total_ms;
timedFnState->timeSpent_ns = 0;
timedFnState->timeBudget_ns = (PTime)total_ms * TIMELOOP_NANOSEC / 1000;
timedFnState->runBudget_ns = (PTime)run_ms * TIMELOOP_NANOSEC / 1000;
timedFnState->fastestRun.nanoSecPerRun = (double)TIMELOOP_NANOSEC * 2000000000; /* hopefully large enough : must be larger than any potential measurement */
timedFnState->fastestRun.sumOfReturn = (size_t)(-1LL);
timedFnState->nbLoops = 1;
timedFnState->coolTime = UTIL_getTime();
}
/* Tells if nb of seconds set in timedFnState for all runs is spent.
* note : this function will return 1 if BMK_benchFunctionTimed() has actually errored. */
int BMK_isCompleted_TimedFn(const BMK_timedFnState_t* timedFnState)
{
return (timedFnState->timeSpent_ns >= timedFnState->timeBudget_ns);
}
#undef MIN
#define MIN(a,b) ( (a) < (b) ? (a) : (b) )
#define MINUSABLETIME (TIMELOOP_NANOSEC / 2) /* 0.5 seconds */
BMK_runOutcome_t BMK_benchTimedFn(BMK_timedFnState_t* cont,
BMK_benchParams_t p)
{
PTime const runBudget_ns = cont->runBudget_ns;
PTime const runTimeMin_ns = runBudget_ns / 2;
int completed = 0;
BMK_runTime_t bestRunTime = cont->fastestRun;
while (!completed) {
BMK_runOutcome_t const runResult = BMK_benchFunction(p, cont->nbLoops);
if(!BMK_isSuccessful_runOutcome(runResult)) { /* error : move out */
return runResult;
}
{ BMK_runTime_t const newRunTime = BMK_extract_runTime(runResult);
double const loopDuration_ns = newRunTime.nanoSecPerRun * cont->nbLoops;
cont->timeSpent_ns += (unsigned long long)loopDuration_ns;
/* estimate nbLoops for next run to last approximately 1 second */
if (loopDuration_ns > ((double)runBudget_ns / 50)) {
double const fastestRun_ns = MIN(bestRunTime.nanoSecPerRun, newRunTime.nanoSecPerRun);
cont->nbLoops = (unsigned)((double)runBudget_ns / fastestRun_ns) + 1;
} else {
/* previous run was too short : blindly increase workload by x multiplier */
const unsigned multiplier = 10;
assert(cont->nbLoops < ((unsigned)-1) / multiplier); /* avoid overflow */
cont->nbLoops *= multiplier;
}
if(loopDuration_ns < (double)runTimeMin_ns) {
/* don't report results for which benchmark run time was too small : increased risks of rounding errors */
assert(completed == 0);
continue;
} else {
if(newRunTime.nanoSecPerRun < bestRunTime.nanoSecPerRun) {
bestRunTime = newRunTime;
}
completed = 1;
}
}
} /* while (!completed) */
return BMK_setValid_runTime(bestRunTime);
}
+155 -8
View File
@@ -8,14 +8,161 @@
* You may select, at your option, one of the above-listed licenses. * You may select, at your option, one of the above-listed licenses.
*/ */
/* The implementation lives in rust/src/timefn.rs, built into the Rust CLI
* static archive. This translation unit stays in the original source lists so /* === Dependencies === */
* build configuration keeps working while the implementation is in Rust. */
#include "timefn.h" #include "timefn.h"
#include "platform.h" /* set _POSIX_C_SOURCE */
#include <time.h> /* CLOCK_MONOTONIC, TIME_UTC */
/* The Rust port mirrors this exact ABI: UTIL_time_t is returned by value and /*-****************************************
* must remain a plain 64-bit nanosecond counter. */ * Time functions
typedef char UTIL_staticAssert_ptimeIs64Bit[(sizeof(PTime) == 8) ? 1 : -1]; ******************************************/
typedef char UTIL_staticAssert_timeIsPlainCounter[
(sizeof(UTIL_time_t) == sizeof(PTime)) ? 1 : -1]; #if defined(_WIN32) /* Windows */
#include <windows.h> /* LARGE_INTEGER */
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
static LARGE_INTEGER ticksPerSecond;
static int init = 0;
if (!init) {
if (!QueryPerformanceFrequency(&ticksPerSecond)) {
perror("timefn::QueryPerformanceFrequency");
abort();
}
init = 1;
}
{ UTIL_time_t r;
LARGE_INTEGER x;
QueryPerformanceCounter(&x);
r.t = (PTime)(x.QuadPart * 1000000000ULL / ticksPerSecond.QuadPart);
return r;
}
}
#elif defined(__APPLE__) && defined(__MACH__)
#include <mach/mach_time.h> /* mach_timebase_info_data_t, mach_timebase_info, mach_absolute_time */
UTIL_time_t UTIL_getTime(void)
{
static mach_timebase_info_data_t rate;
static int init = 0;
if (!init) {
mach_timebase_info(&rate);
init = 1;
}
{ UTIL_time_t r;
r.t = mach_absolute_time() * (PTime)rate.numer / (PTime)rate.denom;
return r;
}
}
/* POSIX.1-2001 (optional) */
#elif defined(CLOCK_MONOTONIC)
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
/* time must be initialized, othersize it may fail msan test.
* No good reason, likely a limitation of timespec_get() for some target */
struct timespec time = { 0, 0 };
if (clock_gettime(CLOCK_MONOTONIC, &time) != 0) {
perror("timefn::clock_gettime(CLOCK_MONOTONIC)");
abort();
}
{ UTIL_time_t r;
r.t = (PTime)time.tv_sec * 1000000000ULL + (PTime)time.tv_nsec;
return r;
}
}
/* C11 requires support of timespec_get().
* However, FreeBSD 11 claims C11 compliance while lacking timespec_get().
* Double confirm timespec_get() support by checking the definition of TIME_UTC.
* However, some versions of Android manage to simultaneously define TIME_UTC
* and lack timespec_get() support... */
#elif (defined (__STDC_VERSION__) && (__STDC_VERSION__ >= 201112L) /* C11 */) \
&& defined(TIME_UTC) && !defined(__ANDROID__)
#include <stdlib.h> /* abort */
#include <stdio.h> /* perror */
UTIL_time_t UTIL_getTime(void)
{
/* time must be initialized, othersize it may fail msan test.
* No good reason, likely a limitation of timespec_get() for some target */
struct timespec time = { 0, 0 };
if (timespec_get(&time, TIME_UTC) != TIME_UTC) {
perror("timefn::timespec_get(TIME_UTC)");
abort();
}
{ UTIL_time_t r;
r.t = (PTime)time.tv_sec * 1000000000ULL + (PTime)time.tv_nsec;
return r;
}
}
#else /* relies on standard C90 (note : clock_t produces wrong measurements for multi-threaded workloads) */
UTIL_time_t UTIL_getTime(void)
{
UTIL_time_t r;
r.t = (PTime)clock() * 1000000000ULL / CLOCKS_PER_SEC;
return r;
}
#define TIME_MT_MEASUREMENTS_NOT_SUPPORTED
#endif
/* ==== Common functions, valid for all time API ==== */
PTime UTIL_getSpanTimeNano(UTIL_time_t clockStart, UTIL_time_t clockEnd)
{
return clockEnd.t - clockStart.t;
}
PTime UTIL_getSpanTimeMicro(UTIL_time_t begin, UTIL_time_t end)
{
return UTIL_getSpanTimeNano(begin, end) / 1000ULL;
}
PTime UTIL_clockSpanMicro(UTIL_time_t clockStart )
{
UTIL_time_t const clockEnd = UTIL_getTime();
return UTIL_getSpanTimeMicro(clockStart, clockEnd);
}
PTime UTIL_clockSpanNano(UTIL_time_t clockStart )
{
UTIL_time_t const clockEnd = UTIL_getTime();
return UTIL_getSpanTimeNano(clockStart, clockEnd);
}
void UTIL_waitForNextTick(void)
{
UTIL_time_t const clockStart = UTIL_getTime();
UTIL_time_t clockEnd;
do {
clockEnd = UTIL_getTime();
} while (UTIL_getSpanTimeNano(clockStart, clockEnd) == 0);
}
int UTIL_support_MT_measurements(void)
{
# if defined(TIME_MT_MEASUREMENTS_NOT_SUPPORTED)
return 0;
# else
return 1;
# endif
}
-51
View File
@@ -10,67 +10,16 @@
/* The CLI parser and control flow live in rust/src/zstd_cli.rs. Keep this /* The CLI parser and control flow live in rust/src/zstd_cli.rs. Keep this
* translation unit as the stable C entry point used by program launchers. */ * translation unit as the stable C entry point used by program launchers. */
#include <stddef.h> /* size_t */
#define ZSTD_STATIC_LINKING_ONLY /* ZSTD_compressionParameters */
#include "../lib/zstd.h" #include "../lib/zstd.h"
#ifndef ZSTD_NOBENCH
# include "benchzstd.h" /* BMK_benchFilesAdvanced, BMK_syntheticTest */
#endif
int ZSTD_rust_cli_main(int argCount, const char* const argv[]); int ZSTD_rust_cli_main(int argCount, const char* const argv[]);
const char* ZSTD_rust_cli_expected_version(void); const char* ZSTD_rust_cli_expected_version(void);
int ZSTD_rust_cli_bench(const char* const* fileNames, unsigned nbFiles,
const char* dictFileName,
int startCLevel, int endCLevel,
const ZSTD_compressionParameters* compressionParams,
int displayLevel, unsigned nbSeconds,
size_t blockSize, int nbWorkers);
const char* ZSTD_rust_cli_expected_version(void) const char* ZSTD_rust_cli_expected_version(void)
{ {
return ZSTD_VERSION_STRING; return ZSTD_VERSION_STRING;
} }
/* Benchmark bridge for the Rust CLI. Whether benchmarking exists is a C
* preprocessor property (ZSTD_NOBENCH), so the decision stays in this shim:
* the Rust frontend calls in unconditionally, and stripped program variants
* never reference benchmark symbols.
* @return the benchmark result code (>= 0), or -1 when unavailable. */
int ZSTD_rust_cli_bench(const char* const* fileNames, unsigned nbFiles,
const char* dictFileName,
int startCLevel, int endCLevel,
const ZSTD_compressionParameters* compressionParams,
int displayLevel, unsigned nbSeconds,
size_t blockSize, int nbWorkers)
{
#ifndef ZSTD_NOBENCH
BMK_advancedParams_t advancedParams = BMK_initAdvancedParams();
int startLevel = startCLevel;
int endLevel = endCLevel;
advancedParams.nbSeconds = nbSeconds;
advancedParams.blockSize = blockSize;
advancedParams.nbWorkers = nbWorkers;
if (startLevel > ZSTD_maxCLevel()) startLevel = ZSTD_maxCLevel();
if (endLevel > ZSTD_maxCLevel()) endLevel = ZSTD_maxCLevel();
if (endLevel < startLevel) endLevel = startLevel;
if (nbFiles == 0) {
/* No input file: benchmark a synthetic sample (lorem generator). */
return BMK_syntheticTest(-1.0, startLevel, endLevel,
compressionParams, displayLevel,
&advancedParams);
}
return BMK_benchFilesAdvanced(fileNames, nbFiles, dictFileName,
startLevel, endLevel,
compressionParams, displayLevel,
&advancedParams);
#else
(void)fileNames; (void)nbFiles; (void)dictFileName;
(void)startCLevel; (void)endCLevel; (void)compressionParams;
(void)displayLevel; (void)nbSeconds; (void)blockSize; (void)nbWorkers;
return -1;
#endif
}
int main(int argCount, const char* argv[]) int main(int argCount, const char* argv[])
{ {
return ZSTD_rust_cli_main(argCount, argv); return ZSTD_rust_cli_main(argCount, argv);
+1 -2
View File
@@ -7,10 +7,9 @@ edition = "2021"
crate-type = ["staticlib"] crate-type = ["staticlib"]
[features] [features]
default = ["compression", "decompression", "dict-builder"] default = ["compression", "decompression"]
compression = [] compression = []
decompression = [] decompression = []
dict-builder = []
huf-force-decompress-x1 = [] huf-force-decompress-x1 = []
huf-force-decompress-x2 = [] huf-force-decompress-x2 = []
# Legacy-format decoders (zstd v0.1 .. v0.7). Never default features: the # Legacy-format decoders (zstd v0.1 .. v0.7). Never default features: the
+6 -28
View File
@@ -34,12 +34,6 @@ zstd ABI:
- `zstd_compress_frame` serializes frame headers, skippable frames, and the - `zstd_compress_frame` serializes frame headers, skippable frames, and the
last empty block; it takes scalar frame parameters so the C-owned last empty block; it takes scalar frame parameters so the C-owned
`ZSTD_CCtx_params` layout never crosses the language boundary. `ZSTD_CCtx_params` layout never crosses the language boundary.
- `zstd_compress_params` owns the compression-level tables (formerly
`clevels.h`), parameter bounds, clamping, validation, table selection,
source/dictionary adjustment, and match-state/CDict size estimation.
The C integration layer keeps the public `ZSTD_*` symbols and feeds the
leaves configuration-owned scalars: the excluded-block-compressor
strategy cascade, struct sizes, and sanitizer redzone policy.
- `zstd_fast` and `zstd_double_fast` implement the single- and two-table - `zstd_fast` and `zstd_double_fast` implement the single- and two-table
fast block match finders, including attached and external dictionary paths. fast block match finders, including attached and external dictionary paths.
- `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching, - `zstd_lazy` implements greedy, lazy, lazy2, and binary-tree matching,
@@ -48,10 +42,6 @@ zstd ABI:
the dynamic-programming optimal parser itself remains in C for now. the dynamic-programming optimal parser itself remains in C for now.
- `zstd_ldm` implements long-distance-match parameter selection, table - `zstd_ldm` implements long-distance-match parameter selection, table
maintenance, sequence generation, and sequence consumption. maintenance, sequence generation, and sequence consumption.
- Dictionary building
- `divsufsort` constructs the suffix array that drives the legacy `ZDICT`
trainer (`ZDICT_trainFromBuffer_legacy`). The sample analysis and
dictionary assembly in `zdict.c`, `cover.c`, and `fastcover.c` remain C.
- Runtime support - Runtime support
- `threading` provides platform pthread wrappers required by zstd headers. - `threading` provides platform pthread wrappers required by zstd headers.
- `pool` implements the bounded worker pool used by multithreaded compression. - `pool` implements the bounded worker pool used by multithreaded compression.
@@ -73,19 +63,9 @@ zstd ABI:
so library builds do not acquire program-only dependencies. The C so library builds do not acquire program-only dependencies. The C
`fileio` backend still owns file opening, safe replacement, sparse writes, `fileio` backend still owns file opening, safe replacement, sparse writes,
metadata, and streaming I/O. metadata, and streaming I/O.
- `timefn` provides the monotonic nanosecond clock behind `UTIL_time_t`,
and `benchfn` owns the benchmark run/timing loop (`BMK_benchFunction`,
`BMK_benchTimedFn`) used by the CLI benchmark mode and by C test tools.
Both live in the `cli/` package, but C test binaries (fullbench, fuzzer,
zstreamtest, paramgrill, ...) link a helpers-only build of that archive,
produced without the package's `cli` feature, because the parser layer
requires the C `fileio` backend that tests do not compile. Benchmark
orchestration and reporting (`benchzstd.c`) remain C, reached from the
Rust parser through the `ZSTD_NOBENCH`-gated bridge in `zstdcli.c`.
The optimal block matcher, high-level frame compression, dictionary-building The optimal block matcher, high-level frame compression, dictionary-building,
except suffix-array construction, the legacy v0.2-v0.7 decoders, benchmark the legacy v0.2-v0.7 decoders, and the CLI file-I/O backend are still C. They
orchestration (`benchzstd`), and the CLI file-I/O backend are still C. They
must move before the rewrite is complete. Keeping that boundary explicit must move before the rewrite is complete. Keeping that boundary explicit
prevents a passing hybrid build from being mistaken for the final all-Rust prevents a passing hybrid build from being mistaken for the final all-Rust
result. result.
@@ -135,9 +115,8 @@ makefile source list as a small shim so header configuration and platform
preprocessor behavior stay available during the transition. preprocessor behavior stay available during the transition.
The library, test, and program makefiles select an archive directory for the The library, test, and program makefiles select an archive directory for the
active C configuration: enabled compression/decompression/dictionary-builder active C configuration: enabled compression/decompression modules, default or
modules, default or forced HUF X1/X2, and the matching Rust target for 32-bit forced HUF X1/X2, and the matching Rust target for 32-bit C binaries. The
C binaries. The
native static archive flattens Rust object members rather than nesting a Rust native static archive flattens Rust object members rather than nesting a Rust
archive, while the native shared library retains all migrated Rust exports. archive, while the native shared library retains all migrated Rust exports.
When the HUF mode changes, the test and program paths also rebuild cached C When the HUF mode changes, the test and program paths also rebuild cached C
@@ -161,9 +140,8 @@ from `rust/cli` as well:
```sh ```sh
cargo clippy --all-targets -- -D warnings cargo clippy --all-targets -- -D warnings
cargo test --all-targets cargo test --all-targets
cargo test --no-default-features --features cli,compression --all-targets cargo test --no-default-features --features compression --all-targets
cargo test --no-default-features --features cli,decompression --all-targets cargo test --no-default-features --features decompression --all-targets
cargo test --no-default-features --all-targets
``` ```
Then run original compatibility tests from the repository root, starting with Then run original compatibility tests from the repository root, starting with
-9
View File
@@ -2,15 +2,6 @@
# It is not intended for manual editing. # It is not intended for manual editing.
version = 4 version = 4
[[package]]
name = "libc"
version = "0.2.186"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "68ab91017fe16c622486840e4c83c9a37afeff978bd239b5293d61ece587de66"
[[package]] [[package]]
name = "zstd-cli-rs" name = "zstd-cli-rs"
version = "0.1.0" version = "0.1.0"
dependencies = [
"libc",
]
+1 -8
View File
@@ -7,13 +7,6 @@ edition = "2021"
crate-type = ["staticlib"] crate-type = ["staticlib"]
[features] [features]
default = ["cli", "compression", "decompression"] default = ["compression", "decompression"]
# The command-line parser and dispatch layer, which requires the C fileio
# backend at link time. Program archives enable it; C test binaries link a
# helpers-only archive (timefn) built without it.
cli = []
compression = [] compression = []
decompression = [] decompression = []
[dependencies]
libc = "0.2"
-5
View File
@@ -1,7 +1,2 @@
#[path = "../../src/benchfn.rs"]
mod benchfn;
#[path = "../../src/timefn.rs"]
mod timefn;
#[cfg(feature = "cli")]
#[path = "../../src/zstd_cli.rs"] #[path = "../../src/zstd_cli.rs"]
mod zstd_cli; mod zstd_cli;
-575
View File
@@ -1,575 +0,0 @@
#![allow(non_camel_case_types)]
#![allow(non_snake_case)]
#![allow(clippy::missing_safety_doc)]
//! Benchmark loop for arbitrary functions over a set of blocks.
//!
//! Port of `programs/benchfn.c`. `BMK_benchFunction` measures one batch of
//! runs; `BMK_benchTimedFn` repeats batches, growing the loop count until a
//! run lasts long enough to be reported reliably against `run_ms`, within a
//! `total_ms` budget tracked by `BMK_timedFnState_t`.
//!
//! ABI notes: `BMK_runOutcome_t` and `BMK_runTime_t` are returned by value
//! across the C boundary and `BMK_benchParams_t` is passed by value, so all
//! three are `repr(C)` mirrors of the `benchfn.h` layout, pinned by asserts
//! here and in the C shim. `BMK_timedFnState_t` is opaque to C, but
//! `BMK_initStatic_timedFnState` guarantees it fits the 64-byte
//! `BMK_timedFnState_shell`, and `BMK_createTimedFnState` uses `malloc` so
//! creation and destruction stay interchangeable with C callers.
use std::os::raw::{c_int, c_uint, c_void};
use std::ptr;
use crate::timefn::{PTime, UTIL_clockSpanNano, UTIL_getTime, UTIL_time_t};
const TIMELOOP_NANOSEC: PTime = 1_000_000_000;
/// Valid benchmark result (`BMK_runTime_t` in benchfn.h).
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct BMK_runTime_t {
/// Time per iteration, over all blocks.
pub nanoSecPerRun: f64,
/// Sum of the benchmarked function's return values, first loop only.
pub sumOfReturn: usize,
}
/// Outcome variant of a benchmark run (`BMK_runOutcome_t` in benchfn.h):
/// either a valid `BMK_runTime_t` or an error result. C callers treat it as
/// opaque and use the accessor functions below.
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct BMK_runOutcome_t {
pub internal_never_ever_use_directly: BMK_runTime_t,
pub error_result_never_ever_use_directly: usize,
pub error_tag_never_ever_use_directly: c_int,
}
// These mirror the static asserts in the programs/benchfn.c shim: the structs
// cross the ABI by value, so field offsets must match the C header exactly.
const _: () =
assert!(std::mem::offset_of!(BMK_runTime_t, sumOfReturn) == std::mem::size_of::<f64>());
const _: () = assert!(
std::mem::offset_of!(BMK_runOutcome_t, error_result_never_ever_use_directly)
== std::mem::size_of::<BMK_runTime_t>()
);
const _: () = assert!(
std::mem::offset_of!(BMK_runOutcome_t, error_tag_never_ever_use_directly)
== std::mem::size_of::<BMK_runTime_t>() + std::mem::size_of::<usize>()
);
/// `size_t (*BMK_benchFn_t)(const void*, size_t, void*, size_t, void*)`
pub type BMK_benchFn_t = Option<
unsafe extern "C" fn(
src: *const c_void,
srcSize: usize,
dst: *mut c_void,
dstCapacity: usize,
customPayload: *mut c_void,
) -> usize,
>;
/// `size_t (*BMK_initFn_t)(void*)`
pub type BMK_initFn_t = Option<unsafe extern "C" fn(initPayload: *mut c_void) -> usize>;
/// `unsigned (*BMK_errorFn_t)(size_t)`
pub type BMK_errorFn_t = Option<unsafe extern "C" fn(result: usize) -> c_uint>;
/// Parameters of `BMK_benchFunction`, passed by value (`BMK_benchParams_t`).
#[repr(C)]
#[derive(Clone, Copy)]
pub struct BMK_benchParams_t {
pub benchFn: BMK_benchFn_t,
pub benchPayload: *mut c_void,
pub initFn: BMK_initFn_t,
pub initPayload: *mut c_void,
pub errorFn: BMK_errorFn_t,
pub blockCount: usize,
pub srcBuffers: *const *const c_void,
pub srcSizes: *const usize,
pub dstBuffers: *const *mut c_void,
pub dstCapacities: *const usize,
pub blockResults: *mut usize,
}
/// Aborts, like benchfn.c's `CONTROL`, when an accessor is used on the wrong
/// outcome variant.
fn control(condition: bool) {
if !condition {
std::process::abort();
}
}
fn error_outcome(errorResult: usize) -> BMK_runOutcome_t {
BMK_runOutcome_t {
internal_never_ever_use_directly: BMK_runTime_t {
nanoSecPerRun: 0.0,
sumOfReturn: 0,
},
error_result_never_ever_use_directly: errorResult,
error_tag_never_ever_use_directly: 1,
}
}
fn valid_outcome(runTime: BMK_runTime_t) -> BMK_runOutcome_t {
BMK_runOutcome_t {
internal_never_ever_use_directly: runTime,
error_result_never_ever_use_directly: 0,
error_tag_never_ever_use_directly: 0,
}
}
/// Tells if the outcome carries a valid measurement.
#[no_mangle]
pub extern "C" fn BMK_isSuccessful_runOutcome(outcome: BMK_runOutcome_t) -> c_int {
c_int::from(outcome.error_tag_never_ever_use_directly == 0)
}
/// Extracts the measurement; aborts if the outcome is an error, so validity
/// must be checked first with `BMK_isSuccessful_runOutcome`.
#[no_mangle]
pub extern "C" fn BMK_extract_runTime(outcome: BMK_runOutcome_t) -> BMK_runTime_t {
control(outcome.error_tag_never_ever_use_directly == 0);
outcome.internal_never_ever_use_directly
}
/// Extracts the faulty `benchFn` return value; aborts if the outcome is
/// valid, so failure must be checked first.
#[no_mangle]
pub extern "C" fn BMK_extract_errorResult(outcome: BMK_runOutcome_t) -> usize {
control(outcome.error_tag_never_ever_use_directly != 0);
outcome.error_result_never_ever_use_directly
}
/// Runs `initFn` once, then `benchFn` `nbLoops` times over every block, and
/// reports the mean time per loop. On the first loop, per-block results are
/// stored into `blockResults` (when provided) and checked with `errorFn`
/// (when provided); the first failing block aborts the measurement and
/// produces an error outcome carrying the faulty return value.
#[no_mangle]
pub unsafe extern "C" fn BMK_benchFunction(
p: BMK_benchParams_t,
mut nbLoops: c_uint,
) -> BMK_runOutcome_t {
// Minimum nbLoops is 1.
nbLoops += c_uint::from(nbLoops == 0);
// Warm up and erase the result buffers.
for blockNb in 0..p.blockCount {
unsafe {
let dst = *p.dstBuffers.add(blockNb);
ptr::write_bytes(dst.cast::<u8>(), 0xE5, *p.dstCapacities.add(blockNb));
}
}
let benchFn = p.benchFn.expect("benchFn is mandatory");
let mut dstSize = 0usize;
let clockStart = UTIL_getTime();
if let Some(initFn) = p.initFn {
unsafe { initFn(p.initPayload) };
}
for loopNb in 0..nbLoops {
for blockNb in 0..p.blockCount {
let res = unsafe {
benchFn(
*p.srcBuffers.add(blockNb),
*p.srcSizes.add(blockNb),
*p.dstBuffers.add(blockNb),
*p.dstCapacities.add(blockNb),
p.benchPayload,
)
};
if loopNb == 0 {
if !p.blockResults.is_null() {
unsafe { *p.blockResults.add(blockNb) = res };
}
if let Some(errorFn) = p.errorFn {
if unsafe { errorFn(res) } != 0 {
return error_outcome(res);
}
}
dstSize = dstSize.wrapping_add(res);
}
}
}
let totalTime = UTIL_clockSpanNano(clockStart);
valid_outcome(BMK_runTime_t {
nanoSecPerRun: totalTime as f64 / f64::from(nbLoops),
sumOfReturn: dstSize,
})
}
/// Benchmark session state (`struct BMK_timedFnState_s`), opaque to C.
#[repr(C)]
pub struct BMK_timedFnState_t {
timeSpent_ns: PTime,
timeBudget_ns: PTime,
runBudget_ns: PTime,
fastestRun: BMK_runTime_t,
nbLoops: c_uint,
coolTime: UTIL_time_t,
}
/// `BMK_TIMEDFNSTATE_SIZE` in benchfn.h: capacity of the caller-provided
/// `BMK_timedFnState_shell`, which the state must always fit.
const BMK_TIMEDFNSTATE_SIZE: usize = 64;
const _: () = assert!(std::mem::size_of::<BMK_timedFnState_t>() <= BMK_TIMEDFNSTATE_SIZE);
// The shell aligns via a `long long` member; the state must not need more.
const _: () = assert!(std::mem::align_of::<BMK_timedFnState_t>() <= std::mem::align_of::<u64>());
/// Allocates and initializes a benchmark session lasting a minimum of
/// `total_ms`, paced at intervals of approximately `run_ms`. Uses `malloc`
/// so ownership stays interchangeable with the original C implementation.
#[no_mangle]
pub extern "C" fn BMK_createTimedFnState(
total_ms: c_uint,
run_ms: c_uint,
) -> *mut BMK_timedFnState_t {
let state = unsafe { libc::malloc(std::mem::size_of::<BMK_timedFnState_t>()) }
.cast::<BMK_timedFnState_t>();
if state.is_null() {
return ptr::null_mut();
}
unsafe { BMK_resetTimedFnState(state, total_ms, run_ms) };
state
}
/// Releases a state obtained from `BMK_createTimedFnState`.
#[no_mangle]
pub unsafe extern "C" fn BMK_freeTimedFnState(state: *mut BMK_timedFnState_t) {
unsafe { libc::free(state.cast()) };
}
/// Places the session state into a caller-provided buffer, typically a
/// `BMK_timedFnState_shell`. Returns NULL when the buffer is missing, too
/// small, or misaligned.
#[no_mangle]
pub unsafe extern "C" fn BMK_initStatic_timedFnState(
buffer: *mut c_void,
size: usize,
total_ms: c_uint,
run_ms: c_uint,
) -> *mut BMK_timedFnState_t {
if buffer.is_null() {
return ptr::null_mut();
}
if size < std::mem::size_of::<BMK_timedFnState_t>() {
return ptr::null_mut();
}
if !(buffer as usize).is_multiple_of(std::mem::align_of::<BMK_timedFnState_t>()) {
return ptr::null_mut();
}
let state = buffer.cast::<BMK_timedFnState_t>();
unsafe { BMK_resetTimedFnState(state, total_ms, run_ms) };
state
}
/// Re-arms a session for a new benchmark of `total_ms`, paced at `run_ms`.
#[no_mangle]
pub unsafe extern "C" fn BMK_resetTimedFnState(
timedFnState: *mut BMK_timedFnState_t,
total_ms: c_uint,
run_ms: c_uint,
) {
let total_ms = if total_ms == 0 { 1 } else { total_ms };
let mut run_ms = if run_ms == 0 { 1 } else { run_ms };
if run_ms > total_ms {
run_ms = total_ms;
}
let state = BMK_timedFnState_t {
timeSpent_ns: 0,
timeBudget_ns: PTime::from(total_ms) * TIMELOOP_NANOSEC / 1000,
runBudget_ns: PTime::from(run_ms) * TIMELOOP_NANOSEC / 1000,
fastestRun: BMK_runTime_t {
// Must be larger than any potential measurement.
nanoSecPerRun: TIMELOOP_NANOSEC as f64 * 2_000_000_000.0,
sumOfReturn: usize::MAX,
},
nbLoops: 1,
coolTime: UTIL_getTime(),
};
unsafe { timedFnState.write(state) };
}
/// Tells if the total time budget of the session is spent. Also reports 1
/// after `BMK_benchTimedFn` returned an error.
#[no_mangle]
pub unsafe extern "C" fn BMK_isCompleted_TimedFn(timedFnState: *const BMK_timedFnState_t) -> c_int {
let state = unsafe { &*timedFnState };
c_int::from(state.timeSpent_ns >= state.timeBudget_ns)
}
/// Runs one measurement supposed to last about `run_ms`, automatically
/// scaling `nbLoops`. Runs shorter than half the run budget are re-tried
/// with a larger workload instead of being reported, limiting rounding-error
/// risks; the best (fastest) qualifying run is returned.
#[no_mangle]
pub unsafe extern "C" fn BMK_benchTimedFn(
cont: *mut BMK_timedFnState_t,
p: BMK_benchParams_t,
) -> BMK_runOutcome_t {
let cont = unsafe { &mut *cont };
let runBudget_ns = cont.runBudget_ns;
let runTimeMin_ns = runBudget_ns / 2;
let mut bestRunTime = cont.fastestRun;
loop {
let runResult = unsafe { BMK_benchFunction(p, cont.nbLoops) };
if BMK_isSuccessful_runOutcome(runResult) == 0 {
// Error: move out.
return runResult;
}
let newRunTime = BMK_extract_runTime(runResult);
let loopDuration_ns = newRunTime.nanoSecPerRun * f64::from(cont.nbLoops);
cont.timeSpent_ns = cont.timeSpent_ns.wrapping_add(loopDuration_ns as PTime);
// Estimate nbLoops for the next run to last approximately run_ms.
if loopDuration_ns > runBudget_ns as f64 / 50.0 {
let fastestRun_ns = bestRunTime.nanoSecPerRun.min(newRunTime.nanoSecPerRun);
cont.nbLoops = ((runBudget_ns as f64 / fastestRun_ns) as c_uint).wrapping_add(1);
} else {
// Previous run was too short: blindly increase workload by a
// x10 multiplier.
const MULTIPLIER: c_uint = 10;
debug_assert!(cont.nbLoops < c_uint::MAX / MULTIPLIER); // avoid overflow
cont.nbLoops = cont.nbLoops.wrapping_mul(MULTIPLIER);
}
if loopDuration_ns < runTimeMin_ns as f64 {
// Don't report results when the run time was too small, which
// increases the risk of rounding errors.
continue;
}
if newRunTime.nanoSecPerRun < bestRunTime.nanoSecPerRun {
bestRunTime = newRunTime;
}
return valid_outcome(bestRunTime);
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Test payload observed through `benchPayload`/`initPayload` pointers.
#[derive(Default)]
struct CallLog {
bench_calls: usize,
init_calls: usize,
}
/// Counts invocations and reports `srcSize`, like a size-preserving codec.
unsafe extern "C" fn counting_bench_fn(
_src: *const c_void,
srcSize: usize,
_dst: *mut c_void,
_dstCapacity: usize,
payload: *mut c_void,
) -> usize {
let log = unsafe { &mut *payload.cast::<CallLog>() };
log.bench_calls += 1;
srcSize
}
unsafe extern "C" fn counting_init_fn(payload: *mut c_void) -> usize {
let log = unsafe { &mut *payload.cast::<CallLog>() };
log.init_calls += 1;
0
}
/// Flags results of 5 bytes and above as errors.
unsafe extern "C" fn error_on_5(result: usize) -> c_uint {
c_uint::from(result >= 5)
}
struct Fixture {
srcs: Vec<Vec<u8>>,
dsts: Vec<Vec<u8>>,
src_ptrs: Vec<*const c_void>,
src_sizes: Vec<usize>,
dst_ptrs: Vec<*mut c_void>,
dst_capacities: Vec<usize>,
block_results: Vec<usize>,
log: CallLog,
}
impl Fixture {
fn new(block_sizes: &[usize]) -> Box<Self> {
let srcs: Vec<Vec<u8>> = block_sizes.iter().map(|size| vec![0u8; *size]).collect();
let mut dsts: Vec<Vec<u8>> = block_sizes.iter().map(|size| vec![0u8; *size]).collect();
let src_ptrs = srcs.iter().map(|src| src.as_ptr().cast()).collect();
let src_sizes = srcs.iter().map(Vec::len).collect();
let dst_ptrs = dsts.iter_mut().map(|dst| dst.as_mut_ptr().cast()).collect();
let dst_capacities = dsts.iter().map(Vec::len).collect();
let block_results = vec![0usize; block_sizes.len()];
Box::new(Self {
srcs,
dsts,
src_ptrs,
src_sizes,
dst_ptrs,
dst_capacities,
block_results,
log: CallLog::default(),
})
}
fn params(&mut self, errorFn: BMK_errorFn_t) -> BMK_benchParams_t {
let payload: *mut CallLog = &mut self.log;
BMK_benchParams_t {
benchFn: Some(counting_bench_fn),
benchPayload: payload.cast(),
initFn: Some(counting_init_fn),
initPayload: payload.cast(),
errorFn,
blockCount: self.srcs.len(),
srcBuffers: self.src_ptrs.as_ptr(),
srcSizes: self.src_sizes.as_ptr(),
dstBuffers: self.dst_ptrs.as_ptr(),
dstCapacities: self.dst_capacities.as_ptr(),
blockResults: self.block_results.as_mut_ptr(),
}
}
}
#[test]
fn bench_function_accounts_loops_blocks_and_first_loop_results() {
let mut fixture = Fixture::new(&[3, 8]);
let params = fixture.params(None);
let outcome = unsafe { BMK_benchFunction(params, 4) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
let run_time = BMK_extract_runTime(outcome);
// benchFn ran nbLoops times over each block; initFn ran once.
assert_eq!(fixture.log.bench_calls, 4 * 2);
assert_eq!(fixture.log.init_calls, 1);
// sumOfReturn and blockResults reflect the first loop only.
assert_eq!(run_time.sumOfReturn, 3 + 8);
assert_eq!(fixture.block_results, vec![3, 8]);
assert!(run_time.nanoSecPerRun >= 0.0);
}
#[test]
fn bench_function_treats_zero_loops_as_one_and_warms_up_buffers() {
let mut fixture = Fixture::new(&[4]);
let params = fixture.params(None);
let outcome = unsafe { BMK_benchFunction(params, 0) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
assert_eq!(fixture.log.bench_calls, 1);
// The result buffer was erased with the 0xE5 warm-up pattern.
assert_eq!(fixture.dsts[0], vec![0xE5; 4]);
}
#[test]
fn bench_function_reports_the_first_failing_block() {
let mut fixture = Fixture::new(&[3, 5, 7]);
let params = fixture.params(Some(error_on_5));
let outcome = unsafe { BMK_benchFunction(params, 10) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 0);
assert_eq!(BMK_extract_errorResult(outcome), 5);
// Execution stopped at the failing block, before the third one.
assert_eq!(fixture.log.bench_calls, 2);
// blockResults were recorded up to and including the failure.
assert_eq!(fixture.block_results[..2], [3, 5]);
}
#[test]
fn reset_clamps_budgets_and_rearms_the_loop_counter() {
let state = BMK_createTimedFnState(0, 7);
assert!(!state.is_null());
{
let state = unsafe { &*state };
// total_ms 0 becomes 1ms, and run_ms is clamped to total_ms.
assert_eq!(state.timeBudget_ns, 1_000_000);
assert_eq!(state.runBudget_ns, 1_000_000);
assert_eq!(state.nbLoops, 1);
assert_eq!(state.timeSpent_ns, 0);
assert_eq!(state.fastestRun.sumOfReturn, usize::MAX);
}
assert_eq!(unsafe { BMK_isCompleted_TimedFn(state) }, 0);
unsafe { BMK_resetTimedFnState(state, 2_000, 500) };
{
let state = unsafe { &*state };
assert_eq!(state.timeBudget_ns, 2_000_000_000);
assert_eq!(state.runBudget_ns, 500_000_000);
}
unsafe { BMK_freeTimedFnState(state) };
}
#[test]
fn static_state_initialization_validates_its_buffer() {
let mut shell = [0u64; BMK_TIMEDFNSTATE_SIZE / 8];
let buffer: *mut c_void = shell.as_mut_ptr().cast();
// A properly sized and aligned buffer is accepted.
let state = unsafe { BMK_initStatic_timedFnState(buffer, 64, 1_000, 100) };
assert!(!state.is_null());
assert_eq!(unsafe { BMK_isCompleted_TimedFn(state) }, 0);
// NULL, undersized, and misaligned buffers are rejected.
let too_small = std::mem::size_of::<BMK_timedFnState_t>() - 1;
unsafe {
assert!(BMK_initStatic_timedFnState(ptr::null_mut(), 64, 1, 1).is_null());
assert!(BMK_initStatic_timedFnState(buffer, too_small, 1, 1).is_null());
assert!(
BMK_initStatic_timedFnState(buffer.cast::<u8>().add(1).cast(), 63, 1, 1).is_null()
);
}
}
#[test]
fn timed_runs_grow_the_workload_and_spend_the_budget() {
let mut fixture = Fixture::new(&[16]);
let params = fixture.params(None);
let state = BMK_createTimedFnState(4, 2);
assert!(!state.is_null());
let mut rounds = 0usize;
while unsafe { BMK_isCompleted_TimedFn(state) } == 0 {
let outcome = unsafe { BMK_benchTimedFn(state, params) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 1);
let run_time = BMK_extract_runTime(outcome);
assert_eq!(run_time.sumOfReturn, 16);
rounds += 1;
assert!(rounds < 1_000, "the time budget must eventually be spent");
}
{
let state = unsafe { &*state };
// A reported run had to last at least runBudget/2, which is only
// reachable for this trivial function with a grown loop counter.
assert!(state.nbLoops > 1);
assert!(state.timeSpent_ns >= state.timeBudget_ns);
}
// Every reported outcome came from a run of >= runBudget/2, and the
// budget accounting matches BMK_isCompleted_TimedFn.
assert!(rounds >= 1);
assert!(fixture.log.bench_calls >= rounds);
unsafe { BMK_freeTimedFnState(state) };
}
#[test]
fn timed_runs_propagate_errors_without_aborting() {
let mut fixture = Fixture::new(&[9]);
let params = fixture.params(Some(error_on_5));
let state = BMK_createTimedFnState(1_000, 100);
let outcome = unsafe { BMK_benchTimedFn(state, params) };
assert_eq!(BMK_isSuccessful_runOutcome(outcome), 0);
assert_eq!(BMK_extract_errorResult(outcome), 9);
unsafe { BMK_freeTimedFnState(state) };
}
}
-2756
View File
@@ -1,2756 +0,0 @@
#![allow(clippy::missing_safety_doc)]
#![allow(clippy::too_many_arguments)]
//! Suffix-array construction for the dictionary builder.
//!
//! Port of `lib/dictBuilder/divsufsort.c` (libdivsufsort-lite, Copyright (c)
//! 2003-2008 Yuta Mori, MIT license) in the exact configuration zstd compiles
//! it with: `ALPHABET_SIZE = 256`, `SS_INSERTIONSORT_THRESHOLD = 8`,
//! `SS_BLOCKSIZE = 1024`, and no OpenMP. Only `divsufsort()` is exported;
//! `divbwt()` has no callers anywhere in zstd and was not ported.
//!
//! The C implementation walks raw `int*` cursors through the caller's SA
//! buffer, including transient one-before-the-range positions. Every such
//! cursor is translated to an `isize` index into one `&mut [i32]` slice
//! covering the whole buffer, so all arithmetic — including the
//! bitwise-complement rank marking and the C `int` value semantics — matches
//! the original exactly while staying bounds-checked.
use std::os::raw::c_int;
use std::slice;
const BUCKET_A_SIZE: usize = 256; /* ALPHABET_SIZE */
const BUCKET_B_SIZE: usize = 256 * 256; /* ALPHABET_SIZE * ALPHABET_SIZE */
const ALPHABET_SIZE: i32 = 256;
const SS_INSERTIONSORT_THRESHOLD: isize = 8;
const SS_BLOCKSIZE: isize = 1024;
/* minstacksize = log(SS_BLOCKSIZE) / log(3) * 2 */
const SS_MISORT_STACKSIZE: usize = 16;
const SS_SMERGE_STACKSIZE: usize = 32;
const TR_INSERTIONSORT_THRESHOLD: isize = 8;
const TR_STACKSIZE: usize = 64;
#[rustfmt::skip]
static LG_TABLE: [i32; 256] = [
-1,0,1,1,2,2,2,2,3,3,3,3,3,3,3,3,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,4,
5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,5,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,6,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,7,
];
#[rustfmt::skip]
static SQQ_TABLE: [i32; 256] = [
0, 16, 22, 27, 32, 35, 39, 42, 45, 48, 50, 53, 55, 57, 59, 61,
64, 65, 67, 69, 71, 73, 75, 76, 78, 80, 81, 83, 84, 86, 87, 89,
90, 91, 93, 94, 96, 97, 98, 99, 101, 102, 103, 104, 106, 107, 108, 109,
110, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126,
128, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 144, 145, 146, 147, 148, 149, 150, 150, 151, 152, 153, 154, 155, 155,
156, 157, 158, 159, 160, 160, 161, 162, 163, 163, 164, 165, 166, 167, 167, 168,
169, 170, 170, 171, 172, 173, 173, 174, 175, 176, 176, 177, 178, 178, 179, 180,
181, 181, 182, 183, 183, 184, 185, 185, 186, 187, 187, 188, 189, 189, 190, 191,
192, 192, 193, 193, 194, 195, 195, 196, 197, 197, 198, 199, 199, 200, 201, 201,
202, 203, 203, 204, 204, 205, 206, 206, 207, 208, 208, 209, 209, 210, 211, 211,
212, 212, 213, 214, 214, 215, 215, 216, 217, 217, 218, 218, 219, 219, 220, 221,
221, 222, 222, 223, 224, 224, 225, 225, 226, 226, 227, 227, 228, 229, 229, 230,
230, 231, 231, 232, 232, 233, 234, 234, 235, 235, 236, 236, 237, 237, 238, 238,
239, 240, 240, 241, 241, 242, 242, 243, 243, 244, 244, 245, 245, 246, 246, 247,
247, 248, 248, 249, 249, 250, 250, 251, 251, 252, 252, 253, 253, 254, 254, 255,
];
/* `ss_ilg` in its `256 <= SS_BLOCKSIZE` configuration. */
#[inline]
fn ss_ilg(n: isize) -> i32 {
let n = n as i32;
if n & 0xff00 != 0 {
8 + LG_TABLE[((n >> 8) & 0xff) as usize]
} else {
LG_TABLE[(n & 0xff) as usize]
}
}
#[inline]
fn ss_isqrt(x: isize) -> isize {
if x >= SS_BLOCKSIZE * SS_BLOCKSIZE {
return SS_BLOCKSIZE;
}
let x = x as i32;
let e = if (x as u32) & 0xffff_0000 != 0 {
if (x as u32) & 0xff00_0000 != 0 {
24 + LG_TABLE[((x >> 24) & 0xff) as usize]
} else {
16 + LG_TABLE[((x >> 16) & 0xff) as usize]
}
} else if x & 0xff00 != 0 {
8 + LG_TABLE[((x >> 8) & 0xff) as usize]
} else {
LG_TABLE[(x & 0xff) as usize]
};
let mut y;
if e >= 16 {
y = SQQ_TABLE[(x >> ((e - 6) - (e & 1))) as usize] << ((e >> 1) - 7);
if e >= 24 {
y = (y + 1 + x / y) >> 1;
}
y = (y + 1 + x / y) >> 1;
} else if e >= 8 {
y = (SQQ_TABLE[(x >> ((e - 6) - (e & 1))) as usize] >> (7 - (e >> 1))) + 1;
} else {
return (SQQ_TABLE[x as usize] >> 4) as isize;
}
(if x < y * y { y - 1 } else { y }) as isize
}
/* --------------------------------------------------------------------- */
/// Compares two suffixes. `(p10, p11)` and `(p20, p21)` are the `p[0]`/`p[1]`
/// pairs the C routine reads through its `const int*` arguments; passing the
/// values directly also serves `sssort()`'s local two-element `PAi` array.
#[inline]
fn ss_compare(t: &[u8], p10: i32, p11: i32, p20: i32, p21: i32, depth: i32) -> i32 {
let mut u1 = (depth + p10) as isize;
let mut u2 = (depth + p20) as isize;
let u1n = (p11 + 2) as isize;
let u2n = (p21 + 2) as isize;
while u1 < u1n && u2 < u2n && t[u1 as usize] == t[u2 as usize] {
u1 += 1;
u2 += 1;
}
if u1 < u1n {
if u2 < u2n {
t[u1 as usize] as i32 - t[u2 as usize] as i32
} else {
1
}
} else if u2 < u2n {
-1
} else {
0
}
}
/// `ss_compare(T, p1, p2, depth)` for pointers `p1`/`p2` into the SA buffer.
#[inline]
fn ss_compare_pa(t: &[u8], sa: &[i32], p1: isize, p2: isize, depth: i32) -> i32 {
ss_compare(
t,
sa[p1 as usize],
sa[(p1 + 1) as usize],
sa[p2 as usize],
sa[(p2 + 1) as usize],
depth,
)
}
/* --------------------------------------------------------------------- */
/* Insertionsort for small size groups */
fn ss_insertionsort(t: &[u8], sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) {
let mut i = last - 2;
while first <= i {
let t0 = sa[i as usize];
let mut j = i + 1;
let mut r;
loop {
r = ss_compare_pa(t, sa, pa + t0 as isize, pa + sa[j as usize] as isize, depth);
if r <= 0 {
break;
}
loop {
sa[(j - 1) as usize] = sa[j as usize];
j += 1;
if !(j < last && sa[j as usize] < 0) {
break;
}
}
if last <= j {
break;
}
}
if r == 0 {
sa[j as usize] = !sa[j as usize];
}
sa[(j - 1) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/// `Td[PA[SA[p]]]` — the depth-`td` sorting key of the suffix stored at `p`.
#[inline(always)]
fn ss_key(t: &[u8], sa: &[i32], td: isize, pa: isize, p: isize) -> i32 {
t[(td + sa[(pa + sa[p as usize] as isize) as usize] as isize) as usize] as i32
}
/// `Td[v]` for an already-loaded SA element `v` (`Td[PA[v]]` in C).
#[inline(always)]
fn ss_key_of(t: &[u8], sa: &[i32], td: isize, pa: isize, v: i32) -> i32 {
t[(td + sa[(pa + v as isize) as usize] as isize) as usize] as i32
}
/// `Td[PA[SA[p]] - 1]` — the character preceding the depth-`td` key.
#[inline(always)]
fn ss_key_pred(t: &[u8], sa: &[i32], td: isize, pa: isize, p: isize) -> i32 {
t[(td + sa[(pa + sa[p as usize] as isize) as usize] as isize - 1) as usize] as i32
}
fn ss_fixdown(t: &[u8], td: isize, sa: &mut [i32], pa: isize, base: isize, i: isize, size: isize) {
let mut i = i;
let v = sa[(base + i) as usize];
let c = ss_key_of(t, sa, td, pa, v);
loop {
let mut j = 2 * i + 1;
if j >= size {
break;
}
let mut k = j;
j += 1;
let mut d = ss_key(t, sa, td, pa, base + k);
let e = ss_key(t, sa, td, pa, base + j);
if d < e {
k = j;
d = e;
}
if d <= c {
break;
}
sa[(base + i) as usize] = sa[(base + k) as usize];
i = k;
}
sa[(base + i) as usize] = v;
}
/* Simple top-down heapsort. */
fn ss_heapsort(t: &[u8], td: isize, sa: &mut [i32], pa: isize, base: isize, size: isize) {
let mut m = size;
if size % 2 == 0 {
m -= 1;
if ss_key(t, sa, td, pa, base + m / 2) < ss_key(t, sa, td, pa, base + m) {
sa.swap((base + m) as usize, (base + m / 2) as usize);
}
}
let mut i = m / 2 - 1;
while 0 <= i {
ss_fixdown(t, td, sa, pa, base, i, m);
i -= 1;
}
if size % 2 == 0 {
sa.swap(base as usize, (base + m) as usize);
ss_fixdown(t, td, sa, pa, base, 0, m);
}
let mut i = m - 1;
while 0 < i {
let t0 = sa[base as usize];
sa[base as usize] = sa[(base + i) as usize];
ss_fixdown(t, td, sa, pa, base, 0, i);
sa[(base + i) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/* Returns the median of three elements. */
#[inline]
fn ss_median3(
t: &[u8],
sa: &[i32],
td: isize,
pa: isize,
v1: isize,
v2: isize,
v3: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v2) {
std::mem::swap(&mut v1, &mut v2);
}
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v3) {
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v3) {
return v1;
}
return v3;
}
v2
}
/* Returns the median of five elements. */
#[inline]
fn ss_median5(
t: &[u8],
sa: &[i32],
td: isize,
pa: isize,
v1: isize,
v2: isize,
v3: isize,
v4: isize,
v5: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
let mut v3 = v3;
let mut v4 = v4;
let mut v5 = v5;
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v3) {
std::mem::swap(&mut v2, &mut v3);
}
if ss_key(t, sa, td, pa, v4) > ss_key(t, sa, td, pa, v5) {
std::mem::swap(&mut v4, &mut v5);
}
if ss_key(t, sa, td, pa, v2) > ss_key(t, sa, td, pa, v4) {
std::mem::swap(&mut v2, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v3) {
std::mem::swap(&mut v1, &mut v3);
}
if ss_key(t, sa, td, pa, v1) > ss_key(t, sa, td, pa, v4) {
std::mem::swap(&mut v1, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if ss_key(t, sa, td, pa, v3) > ss_key(t, sa, td, pa, v4) {
return v4;
}
v3
}
/* Returns the pivot element. */
#[inline]
fn ss_pivot(t: &[u8], sa: &[i32], td: isize, pa: isize, first: isize, last: isize) -> isize {
let mut t0 = last - first;
let middle = first + t0 / 2;
if t0 <= 512 {
if t0 <= 32 {
return ss_median3(t, sa, td, pa, first, middle, last - 1);
}
t0 >>= 2;
return ss_median5(
t,
sa,
td,
pa,
first,
first + t0,
middle,
last - 1 - t0,
last - 1,
);
}
t0 >>= 3;
let first = ss_median3(t, sa, td, pa, first, first + t0, first + (t0 << 1));
let middle = ss_median3(t, sa, td, pa, middle - t0, middle, middle + t0);
let last = ss_median3(t, sa, td, pa, last - 1 - (t0 << 1), last - 1 - t0, last - 1);
ss_median3(t, sa, td, pa, first, middle, last)
}
/* --------------------------------------------------------------------- */
/* Binary partition for substrings. */
/* The `>= x + 1` comparison deliberately mirrors the C expression shape. */
#[allow(clippy::int_plus_one)]
fn ss_partition(sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) -> isize {
let mut a = first - 1;
let mut b = last;
loop {
loop {
a += 1;
if !(a < b) {
break;
}
if !(sa[(pa + sa[a as usize] as isize) as usize] + depth
>= sa[(pa + sa[a as usize] as isize + 1) as usize] + 1)
{
break;
}
sa[a as usize] = !sa[a as usize];
}
loop {
b -= 1;
if !(a < b) {
break;
}
if !(sa[(pa + sa[b as usize] as isize) as usize] + depth
< sa[(pa + sa[b as usize] as isize + 1) as usize] + 1)
{
break;
}
}
if b <= a {
break;
}
let t0 = !sa[b as usize];
sa[b as usize] = sa[a as usize];
sa[a as usize] = t0;
}
if first < a {
sa[first as usize] = !sa[first as usize];
}
a
}
/* Multikey introsort for medium size groups. */
fn ss_mintrosort(t: &[u8], sa: &mut [i32], pa: isize, first: isize, last: isize, depth: i32) {
let mut stack = [(0isize, 0isize, 0i32, 0i32); SS_MISORT_STACKSIZE];
let mut ssize = 0usize;
let mut first = first;
let mut last = last;
let mut depth = depth;
let mut limit = ss_ilg(last - first);
let mut x: i32 = 0;
loop {
if last - first <= SS_INSERTIONSORT_THRESHOLD {
if 1 < last - first {
ss_insertionsort(t, sa, pa, first, last, depth);
}
/* STACK_POP */
if ssize == 0 {
return;
}
ssize -= 1;
(first, last, depth, limit) = stack[ssize];
continue;
}
let td = depth as isize;
if limit == 0 {
ss_heapsort(t, td, sa, pa, first, last - first);
}
limit -= 1;
if limit < 0 {
let mut a = first + 1;
let mut v = ss_key(t, sa, td, pa, first);
while a < last {
x = ss_key(t, sa, td, pa, a);
if x != v {
if 1 < a - first {
break;
}
v = x;
first = a;
}
a += 1;
}
if ss_key_pred(t, sa, td, pa, first) < v {
first = ss_partition(sa, pa, first, a, depth);
}
if a - first <= last - a {
if 1 < a - first {
stack[ssize] = (a, last, depth, -1);
ssize += 1;
last = a;
depth += 1;
limit = ss_ilg(a - first);
} else {
first = a;
limit = -1;
}
} else if 1 < last - a {
stack[ssize] = (first, a, depth + 1, ss_ilg(a - first));
ssize += 1;
first = a;
limit = -1;
} else {
last = a;
depth += 1;
limit = ss_ilg(a - first);
}
continue;
}
/* choose pivot */
let mut a = ss_pivot(t, sa, td, pa, first, last);
let v = ss_key(t, sa, td, pa, a);
sa.swap(first as usize, a as usize);
/* partition */
let mut b = first;
loop {
b += 1;
if !(b < last) {
break;
}
x = ss_key(t, sa, td, pa, b);
if x != v {
break;
}
}
a = b;
if a < last && x < v {
loop {
b += 1;
if !(b < last) {
break;
}
x = ss_key(t, sa, td, pa, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
}
let mut c = last;
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if x != v {
break;
}
}
let mut d = c;
if b < d && x > v {
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
while b < c {
sa.swap(b as usize, c as usize);
loop {
b += 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
loop {
c -= 1;
if !(b < c) {
break;
}
x = ss_key(t, sa, td, pa, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
if a <= d {
c = b - 1;
let mut s = a - first;
let t0 = b - a;
if s > t0 {
s = t0;
}
let mut e = first;
let mut f = b - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
let mut s = d - c;
let t0 = last - d - 1;
if s > t0 {
s = t0;
}
let mut e = b;
let mut f = last - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
a = first + (b - a);
c = last - (d - c);
b = if v <= ss_key_pred(t, sa, td, pa, a) {
a
} else {
ss_partition(sa, pa, a, c, depth)
};
if a - first <= last - c {
if last - c <= c - b {
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
stack[ssize] = (c, last, depth, limit);
ssize += 1;
last = a;
} else if a - first <= c - b {
stack[ssize] = (c, last, depth, limit);
ssize += 1;
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
last = a;
} else {
stack[ssize] = (c, last, depth, limit);
ssize += 1;
stack[ssize] = (first, a, depth, limit);
ssize += 1;
first = b;
last = c;
depth += 1;
limit = ss_ilg(c - b);
}
} else if a - first <= c - b {
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
stack[ssize] = (first, a, depth, limit);
ssize += 1;
first = c;
} else if last - c <= c - b {
stack[ssize] = (first, a, depth, limit);
ssize += 1;
stack[ssize] = (b, c, depth + 1, ss_ilg(c - b));
ssize += 1;
first = c;
} else {
stack[ssize] = (first, a, depth, limit);
ssize += 1;
stack[ssize] = (c, last, depth, limit);
ssize += 1;
first = b;
last = c;
depth += 1;
limit = ss_ilg(c - b);
}
} else {
limit += 1;
if ss_key_pred(t, sa, td, pa, first) < v {
first = ss_partition(sa, pa, first, last, depth);
limit = ss_ilg(last - first);
}
depth += 1;
}
}
}
/* --------------------------------------------------------------------- */
#[inline]
fn ss_blockswap(sa: &mut [i32], a: isize, b: isize, n: isize) {
let mut a = a;
let mut b = b;
let mut n = n;
while 0 < n {
sa.swap(a as usize, b as usize);
n -= 1;
a += 1;
b += 1;
}
}
#[inline]
fn ss_rotate(sa: &mut [i32], first: isize, middle: isize, last: isize) {
let mut first = first;
let mut last = last;
let mut l = middle - first;
let mut r = last - middle;
while 0 < l && 0 < r {
if l == r {
ss_blockswap(sa, first, middle, l);
break;
}
if l < r {
let mut a = last - 1;
let mut b = middle - 1;
let mut t0 = sa[a as usize];
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if b < first {
sa[a as usize] = t0;
last = a;
r -= l + 1;
if r <= l {
break;
}
a -= 1;
b = middle - 1;
t0 = sa[a as usize];
}
}
} else {
let mut a = first;
let mut b = middle;
let mut t0 = sa[a as usize];
loop {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
if last <= b {
sa[a as usize] = t0;
first = a + 1;
l -= r + 1;
if l <= r {
break;
}
a += 1;
b = middle;
t0 = sa[a as usize];
}
}
}
}
}
/* --------------------------------------------------------------------- */
fn ss_inplacemerge(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
depth: i32,
) {
let mut middle = middle;
let mut last = last;
loop {
let x: i32;
let p: isize;
if sa[(last - 1) as usize] < 0 {
x = 1;
p = pa + (!sa[(last - 1) as usize]) as isize;
} else {
x = 0;
p = pa + sa[(last - 1) as usize] as isize;
}
let mut a = first;
let mut len = middle - first;
let mut half = len >> 1;
let mut r: i32 = -1;
while 0 < len {
let b = a + half;
let bv = sa[b as usize];
let q = ss_compare_pa(
t,
sa,
pa + (if 0 <= bv { bv } else { !bv }) as isize,
p,
depth,
);
if q < 0 {
a = b + 1;
half -= (len & 1) ^ 1;
} else {
r = q;
}
len = half;
half >>= 1;
}
if a < middle {
if r == 0 {
sa[a as usize] = !sa[a as usize];
}
ss_rotate(sa, a, middle, last);
last -= middle - a;
middle = a;
if first == middle {
break;
}
}
last -= 1;
if x != 0 {
loop {
last -= 1;
if !(sa[last as usize] < 0) {
break;
}
}
}
if middle == last {
break;
}
}
}
/* --------------------------------------------------------------------- */
/* Merge-forward with internal buffer. */
fn ss_mergeforward(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
depth: i32,
) {
let bufend = buf + (middle - first) - 1;
ss_blockswap(sa, buf, first, middle - first);
let mut a = first;
let t0 = sa[a as usize];
let mut b = buf;
let mut c = middle;
loop {
let r = ss_compare_pa(
t,
sa,
pa + sa[b as usize] as isize,
pa + sa[c as usize] as isize,
depth,
);
if r < 0 {
loop {
sa[a as usize] = sa[b as usize];
a += 1;
if bufend <= b {
sa[bufend as usize] = t0;
return;
}
sa[b as usize] = sa[a as usize];
b += 1;
if !(sa[b as usize] < 0) {
break;
}
}
} else if r > 0 {
loop {
sa[a as usize] = sa[c as usize];
a += 1;
sa[c as usize] = sa[a as usize];
c += 1;
if last <= c {
while b < bufend {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
return;
}
if !(sa[c as usize] < 0) {
break;
}
}
} else {
sa[c as usize] = !sa[c as usize];
loop {
sa[a as usize] = sa[b as usize];
a += 1;
if bufend <= b {
sa[bufend as usize] = t0;
return;
}
sa[b as usize] = sa[a as usize];
b += 1;
if !(sa[b as usize] < 0) {
break;
}
}
loop {
sa[a as usize] = sa[c as usize];
a += 1;
sa[c as usize] = sa[a as usize];
c += 1;
if last <= c {
while b < bufend {
sa[a as usize] = sa[b as usize];
a += 1;
sa[b as usize] = sa[a as usize];
b += 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
return;
}
if !(sa[c as usize] < 0) {
break;
}
}
}
}
}
/* Merge-backward with internal buffer. */
fn ss_mergebackward(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
depth: i32,
) {
let bufend = buf + (last - middle) - 1;
ss_blockswap(sa, buf, middle, last - middle);
let mut x = 0i32;
let mut p1: isize;
let mut p2: isize;
if sa[bufend as usize] < 0 {
p1 = pa + (!sa[bufend as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[bufend as usize] as isize;
}
if sa[(middle - 1) as usize] < 0 {
p2 = pa + (!sa[(middle - 1) as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[(middle - 1) as usize] as isize;
}
let mut a = last - 1;
let t0 = sa[a as usize];
let mut b = bufend;
let mut c = middle - 1;
loop {
let r = ss_compare_pa(t, sa, p1, p2, depth);
if 0 < r {
if x & 1 != 0 {
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if !(sa[b as usize] < 0) {
break;
}
}
x ^= 1;
}
sa[a as usize] = sa[b as usize];
a -= 1;
if b <= buf {
sa[buf as usize] = t0;
break;
}
sa[b as usize] = sa[a as usize];
b -= 1;
if sa[b as usize] < 0 {
p1 = pa + (!sa[b as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[b as usize] as isize;
}
} else if r < 0 {
if x & 2 != 0 {
loop {
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if !(sa[c as usize] < 0) {
break;
}
}
x ^= 2;
}
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if c < first {
while buf < b {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
break;
}
if sa[c as usize] < 0 {
p2 = pa + (!sa[c as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[c as usize] as isize;
}
} else {
if x & 1 != 0 {
loop {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
if !(sa[b as usize] < 0) {
break;
}
}
x ^= 1;
}
sa[a as usize] = !sa[b as usize];
a -= 1;
if b <= buf {
sa[buf as usize] = t0;
break;
}
sa[b as usize] = sa[a as usize];
b -= 1;
if x & 2 != 0 {
loop {
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if !(sa[c as usize] < 0) {
break;
}
}
x ^= 2;
}
sa[a as usize] = sa[c as usize];
a -= 1;
sa[c as usize] = sa[a as usize];
c -= 1;
if c < first {
while buf < b {
sa[a as usize] = sa[b as usize];
a -= 1;
sa[b as usize] = sa[a as usize];
b -= 1;
}
sa[a as usize] = sa[b as usize];
sa[b as usize] = t0;
break;
}
if sa[b as usize] < 0 {
p1 = pa + (!sa[b as usize]) as isize;
x |= 1;
} else {
p1 = pa + sa[b as usize] as isize;
}
if sa[c as usize] < 0 {
p2 = pa + (!sa[c as usize]) as isize;
x |= 2;
} else {
p2 = pa + sa[c as usize] as isize;
}
}
}
}
/// `GETIDX` — undoes the "already merged" complement marking.
#[inline(always)]
fn getidx(a: i32) -> i32 {
if 0 <= a {
a
} else {
!a
}
}
/// `MERGE_CHECK` — restores or sets the complement marks after a merge.
#[inline]
fn ss_merge_check(t: &[u8], sa: &mut [i32], pa: isize, a: isize, b: isize, c: i32, depth: i32) {
if (c & 1) != 0
|| ((c & 2) != 0
&& ss_compare_pa(
t,
sa,
pa + getidx(sa[(a - 1) as usize]) as isize,
pa + sa[a as usize] as isize,
depth,
) == 0)
{
sa[a as usize] = !sa[a as usize];
}
if (c & 4) != 0
&& ss_compare_pa(
t,
sa,
pa + getidx(sa[(b - 1) as usize]) as isize,
pa + sa[b as usize] as isize,
depth,
) == 0
{
sa[b as usize] = !sa[b as usize];
}
}
/* D&C based merge. */
fn ss_swapmerge(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
middle: isize,
last: isize,
buf: isize,
bufsize: isize,
depth: i32,
) {
let mut stack = [(0isize, 0isize, 0isize, 0i32); SS_SMERGE_STACKSIZE];
let mut ssize = 0usize;
let mut first = first;
let mut middle = middle;
let mut last = last;
let mut check = 0i32;
loop {
if last - middle <= bufsize {
if first < middle && middle < last {
ss_mergebackward(t, sa, pa, first, middle, last, buf, depth);
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
continue;
}
if middle - first <= bufsize {
if first < middle {
ss_mergeforward(t, sa, pa, first, middle, last, buf, depth);
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
continue;
}
let mut m: isize = 0;
let mut len = std::cmp::min(middle - first, last - middle);
let mut half = len >> 1;
while 0 < len {
if ss_compare_pa(
t,
sa,
pa + getidx(sa[(middle + m + half) as usize]) as isize,
pa + getidx(sa[(middle - m - half - 1) as usize]) as isize,
depth,
) < 0
{
m += half + 1;
half -= (len & 1) ^ 1;
}
len = half;
half >>= 1;
}
if 0 < m {
let lm = middle - m;
let rm = middle + m;
ss_blockswap(sa, lm, middle, m);
let mut l = middle;
let mut r = middle;
let mut next = 0i32;
if rm < last {
if sa[rm as usize] < 0 {
sa[rm as usize] = !sa[rm as usize];
if first < lm {
loop {
l -= 1;
if !(sa[l as usize] < 0) {
break;
}
}
next |= 4;
}
next |= 1;
} else if first < lm {
while sa[r as usize] < 0 {
r += 1;
}
next |= 2;
}
}
if l - first <= last - r {
stack[ssize] = (r, rm, last, (next & 3) | (check & 4));
ssize += 1;
middle = lm;
last = l;
check = (check & 3) | (next & 4);
} else {
if (next & 2) != 0 && r == middle {
next ^= 6;
}
stack[ssize] = (first, lm, l, (check & 3) | (next & 4));
ssize += 1;
first = r;
middle = rm;
check = (next & 3) | (check & 4);
}
} else {
if ss_compare_pa(
t,
sa,
pa + getidx(sa[(middle - 1) as usize]) as isize,
pa + sa[middle as usize] as isize,
depth,
) == 0
{
sa[middle as usize] = !sa[middle as usize];
}
ss_merge_check(t, sa, pa, first, last, check, depth);
if ssize == 0 {
return;
}
ssize -= 1;
(first, middle, last, check) = stack[ssize];
}
}
}
/* --------------------------------------------------------------------- */
/* Substring sort */
fn sssort(
t: &[u8],
sa: &mut [i32],
pa: isize,
first: isize,
last: isize,
buf: isize,
bufsize: isize,
depth: i32,
n: isize,
lastsuffix: bool,
) {
let mut first = first;
let mut buf = buf;
let mut bufsize = bufsize;
if lastsuffix {
first += 1;
}
let mut limit: isize = 0;
let mut middle = last;
if bufsize < SS_BLOCKSIZE && bufsize < last - first {
limit = ss_isqrt(last - first);
if bufsize < limit {
if SS_BLOCKSIZE < limit {
limit = SS_BLOCKSIZE;
}
middle = last - limit;
buf = middle;
bufsize = limit;
} else {
limit = 0;
}
}
let mut a = first;
let mut i: isize = 0;
while SS_BLOCKSIZE < middle - a {
ss_mintrosort(t, sa, pa, a, a + SS_BLOCKSIZE, depth);
let mut curbufsize = last - (a + SS_BLOCKSIZE);
let mut curbuf = a + SS_BLOCKSIZE;
if curbufsize <= bufsize {
curbufsize = bufsize;
curbuf = buf;
}
let mut b = a;
let mut k = SS_BLOCKSIZE;
let mut j = i;
while j & 1 != 0 {
ss_swapmerge(t, sa, pa, b - k, b, b + k, curbuf, curbufsize, depth);
b -= k;
k <<= 1;
j >>= 1;
}
a += SS_BLOCKSIZE;
i += 1;
}
ss_mintrosort(t, sa, pa, a, middle, depth);
let mut k = SS_BLOCKSIZE;
while i != 0 {
if i & 1 != 0 {
ss_swapmerge(t, sa, pa, a - k, a, middle, buf, bufsize, depth);
a -= k;
}
k <<= 1;
i >>= 1;
}
if limit != 0 {
ss_mintrosort(t, sa, pa, middle, last, depth);
ss_inplacemerge(t, sa, pa, first, middle, last, depth);
}
if lastsuffix {
/* Insert last type B* suffix. */
let pai0 = sa[(pa + sa[(first - 1) as usize] as isize) as usize];
let pai1 = (n - 2) as i32;
let i0 = sa[(first - 1) as usize];
let mut a = first;
while a < last {
let av = sa[a as usize];
if !(av < 0
|| 0 < ss_compare(
t,
pai0,
pai1,
sa[(pa + av as isize) as usize],
sa[(pa + av as isize + 1) as usize],
depth,
))
{
break;
}
sa[(a - 1) as usize] = av;
a += 1;
}
sa[(a - 1) as usize] = i0;
}
}
/* --------------------------------------------------------------------- */
#[inline]
fn tr_ilg(n: isize) -> i32 {
let n = n as i32;
if (n as u32) & 0xffff_0000 != 0 {
if (n as u32) & 0xff00_0000 != 0 {
24 + LG_TABLE[((n >> 24) & 0xff) as usize]
} else {
16 + LG_TABLE[((n >> 16) & 0xff) as usize]
}
} else if n & 0xff00 != 0 {
8 + LG_TABLE[((n >> 8) & 0xff) as usize]
} else {
LG_TABLE[(n & 0xff) as usize]
}
}
/* --------------------------------------------------------------------- */
/// `ISAd[SA[p]]` — the depth-offset rank of the suffix stored at `p`.
#[inline(always)]
fn tr_key(sa: &[i32], isad: isize, p: isize) -> i32 {
sa[(isad + sa[p as usize] as isize) as usize]
}
/* Simple insertionsort for small size groups. */
fn tr_insertionsort(sa: &mut [i32], isad: isize, first: isize, last: isize) {
let mut a = first + 1;
while a < last {
let t0 = sa[a as usize];
let mut b = a - 1;
let mut r;
loop {
r = sa[(isad + t0 as isize) as usize] - tr_key(sa, isad, b);
if !(0 > r) {
break;
}
loop {
sa[(b + 1) as usize] = sa[b as usize];
b -= 1;
if !(first <= b && sa[b as usize] < 0) {
break;
}
}
if b < first {
break;
}
}
if r == 0 {
sa[b as usize] = !sa[b as usize];
}
sa[(b + 1) as usize] = t0;
a += 1;
}
}
/* --------------------------------------------------------------------- */
fn tr_fixdown(sa: &mut [i32], isad: isize, base: isize, i: isize, size: isize) {
let mut i = i;
let v = sa[(base + i) as usize];
let c = sa[(isad + v as isize) as usize];
loop {
let mut j = 2 * i + 1;
if j >= size {
break;
}
let mut k = j;
j += 1;
let mut d = tr_key(sa, isad, base + k);
let e = tr_key(sa, isad, base + j);
if d < e {
k = j;
d = e;
}
if d <= c {
break;
}
sa[(base + i) as usize] = sa[(base + k) as usize];
i = k;
}
sa[(base + i) as usize] = v;
}
/* Simple top-down heapsort. */
fn tr_heapsort(sa: &mut [i32], isad: isize, base: isize, size: isize) {
let mut m = size;
if size % 2 == 0 {
m -= 1;
if tr_key(sa, isad, base + m / 2) < tr_key(sa, isad, base + m) {
sa.swap((base + m) as usize, (base + m / 2) as usize);
}
}
let mut i = m / 2 - 1;
while 0 <= i {
tr_fixdown(sa, isad, base, i, m);
i -= 1;
}
if size % 2 == 0 {
sa.swap(base as usize, (base + m) as usize);
tr_fixdown(sa, isad, base, 0, m);
}
let mut i = m - 1;
while 0 < i {
let t0 = sa[base as usize];
sa[base as usize] = sa[(base + i) as usize];
tr_fixdown(sa, isad, base, 0, i);
sa[(base + i) as usize] = t0;
i -= 1;
}
}
/* --------------------------------------------------------------------- */
/* Returns the median of three elements. */
#[inline]
fn tr_median3(sa: &[i32], isad: isize, v1: isize, v2: isize, v3: isize) -> isize {
let mut v1 = v1;
let mut v2 = v2;
if tr_key(sa, isad, v1) > tr_key(sa, isad, v2) {
std::mem::swap(&mut v1, &mut v2);
}
if tr_key(sa, isad, v2) > tr_key(sa, isad, v3) {
if tr_key(sa, isad, v1) > tr_key(sa, isad, v3) {
return v1;
}
return v3;
}
v2
}
/* Returns the median of five elements. */
#[inline]
fn tr_median5(
sa: &[i32],
isad: isize,
v1: isize,
v2: isize,
v3: isize,
v4: isize,
v5: isize,
) -> isize {
let mut v1 = v1;
let mut v2 = v2;
let mut v3 = v3;
let mut v4 = v4;
let mut v5 = v5;
if tr_key(sa, isad, v2) > tr_key(sa, isad, v3) {
std::mem::swap(&mut v2, &mut v3);
}
if tr_key(sa, isad, v4) > tr_key(sa, isad, v5) {
std::mem::swap(&mut v4, &mut v5);
}
if tr_key(sa, isad, v2) > tr_key(sa, isad, v4) {
std::mem::swap(&mut v2, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if tr_key(sa, isad, v1) > tr_key(sa, isad, v3) {
std::mem::swap(&mut v1, &mut v3);
}
if tr_key(sa, isad, v1) > tr_key(sa, isad, v4) {
std::mem::swap(&mut v1, &mut v4);
std::mem::swap(&mut v3, &mut v5);
}
if tr_key(sa, isad, v3) > tr_key(sa, isad, v4) {
return v4;
}
v3
}
/* Returns the pivot element. */
#[inline]
fn tr_pivot(sa: &[i32], isad: isize, first: isize, last: isize) -> isize {
let mut t0 = last - first;
let middle = first + t0 / 2;
if t0 <= 512 {
if t0 <= 32 {
return tr_median3(sa, isad, first, middle, last - 1);
}
t0 >>= 2;
return tr_median5(sa, isad, first, first + t0, middle, last - 1 - t0, last - 1);
}
t0 >>= 3;
let first = tr_median3(sa, isad, first, first + t0, first + (t0 << 1));
let middle = tr_median3(sa, isad, middle - t0, middle, middle + t0);
let last = tr_median3(sa, isad, last - 1 - (t0 << 1), last - 1 - t0, last - 1);
tr_median3(sa, isad, first, middle, last)
}
/* --------------------------------------------------------------------- */
struct TrBudget {
chance: i32,
remain: i32,
incval: i32,
count: i32,
}
impl TrBudget {
fn new(chance: i32, incval: i32) -> Self {
TrBudget {
chance,
remain: incval,
incval,
count: 0,
}
}
fn check(&mut self, size: isize) -> bool {
let size = size as i32;
if size <= self.remain {
self.remain -= size;
return true;
}
if self.chance == 0 {
self.count += size;
return false;
}
self.remain += self.incval - size;
self.chance -= 1;
true
}
}
/* --------------------------------------------------------------------- */
fn tr_partition(
sa: &mut [i32],
isad: isize,
first: isize,
middle: isize,
last: isize,
v: i32,
) -> (isize, isize) {
let mut first = first;
let mut last = last;
let mut x: i32 = 0;
let mut b = middle - 1;
loop {
b += 1;
if !(b < last) {
break;
}
x = tr_key(sa, isad, b);
if x != v {
break;
}
}
let mut a = b;
if a < last && x < v {
loop {
b += 1;
if !(b < last) {
break;
}
x = tr_key(sa, isad, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
}
let mut c = last;
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if x != v {
break;
}
}
let mut d = c;
if b < d && x > v {
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
while b < c {
sa.swap(b as usize, c as usize);
loop {
b += 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, b);
if !(x <= v) {
break;
}
if x == v {
sa.swap(b as usize, a as usize);
a += 1;
}
}
loop {
c -= 1;
if !(b < c) {
break;
}
x = tr_key(sa, isad, c);
if !(x >= v) {
break;
}
if x == v {
sa.swap(c as usize, d as usize);
d -= 1;
}
}
}
if a <= d {
c = b - 1;
let mut s = a - first;
let t0 = b - a;
if s > t0 {
s = t0;
}
let mut e = first;
let mut f = b - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
let mut s = d - c;
let t0 = last - d - 1;
if s > t0 {
s = t0;
}
let mut e = b;
let mut f = last - s;
while 0 < s {
sa.swap(e as usize, f as usize);
s -= 1;
e += 1;
f += 1;
}
first += b - a;
last -= d - c;
}
(first, last)
}
/* sort suffixes of middle partition by using sorted order of suffixes of
* left and right partition. */
fn tr_copy(
sa: &mut [i32],
isa: isize,
first: isize,
a: isize,
b: isize,
last: isize,
depth: isize,
) {
/* All cursor arithmetic is relative to the slice start, which is the C
* routine's `SA` pointer, so `x - SA` becomes plain `x`. */
let v = (b - 1) as i32;
let mut c = first;
let mut d = a - 1;
while c <= d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d += 1;
sa[d as usize] = s;
sa[(isa + s as isize) as usize] = d as i32;
}
c += 1;
}
let mut c = last - 1;
let e = d + 1;
let mut d = b;
while e < d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d -= 1;
sa[d as usize] = s;
sa[(isa + s as isize) as usize] = d as i32;
}
c -= 1;
}
}
fn tr_partialcopy(
sa: &mut [i32],
isa: isize,
first: isize,
a: isize,
b: isize,
last: isize,
depth: isize,
) {
let v = (b - 1) as i32;
let mut newrank: i32 = -1;
let mut lastrank: i32 = -1;
let mut c = first;
let mut d = a - 1;
while c <= d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d += 1;
sa[d as usize] = s;
let rank = sa[(isa + s as isize + depth) as usize];
if lastrank != rank {
lastrank = rank;
newrank = d as i32;
}
sa[(isa + s as isize) as usize] = newrank;
}
c += 1;
}
let mut lastrank: i32 = -1;
let mut e = d;
while first <= e {
let rank = sa[(isa + sa[e as usize] as isize) as usize];
if lastrank != rank {
lastrank = rank;
newrank = e as i32;
}
if newrank != rank {
sa[(isa + sa[e as usize] as isize) as usize] = newrank;
}
e -= 1;
}
let mut lastrank: i32 = -1;
let mut c = last - 1;
let e = d + 1;
let mut d = b;
while e < d {
let s = sa[c as usize] - depth as i32;
if 0 <= s && sa[(isa + s as isize) as usize] == v {
d -= 1;
sa[d as usize] = s;
let rank = sa[(isa + s as isize + depth) as usize];
if lastrank != rank {
lastrank = rank;
newrank = d as i32;
}
sa[(isa + s as isize) as usize] = newrank;
}
c -= 1;
}
}
fn tr_introsort(
sa: &mut [i32],
isa: isize,
isad: isize,
first: isize,
last: isize,
budget: &mut TrBudget,
) {
/* Stack frames are (ISAd, first, last, limit, trlink); the tandem-repeat
* copy frame stores its `(a, b)` pair in the pointer fields with a zero
* placeholder where C pushes a NULL ISAd. */
let mut stack = [(0isize, 0isize, 0isize, 0i32, 0i32); TR_STACKSIZE];
let mut ssize = 0usize;
let mut trlink: i32 = -1;
let mut isad = isad;
let mut first = first;
let mut last = last;
let incr = isad - isa;
let mut limit = tr_ilg(last - first);
loop {
if limit < 0 {
if limit == -1 {
/* tandem repeat partition */
let (a, b) = tr_partition(sa, isad - incr, first, first, last, (last - 1) as i32);
/* update ranks */
if a < last {
let v = (a - 1) as i32;
let mut c = first;
while c < a {
sa[(isa + sa[c as usize] as isize) as usize] = v;
c += 1;
}
}
if b < last {
let v = (b - 1) as i32;
let mut c = a;
while c < b {
sa[(isa + sa[c as usize] as isize) as usize] = v;
c += 1;
}
}
/* push */
if 1 < b - a {
stack[ssize] = (0, a, b, 0, 0);
ssize += 1;
stack[ssize] = (isad - incr, first, last, -2, trlink);
ssize += 1;
trlink = ssize as i32 - 2;
}
if a - first <= last - b {
if 1 < a - first {
stack[ssize] = (isad, b, last, tr_ilg(last - b), trlink);
ssize += 1;
last = a;
limit = tr_ilg(a - first);
} else if 1 < last - b {
first = b;
limit = tr_ilg(last - b);
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if 1 < last - b {
stack[ssize] = (isad, first, a, tr_ilg(a - first), trlink);
ssize += 1;
first = b;
limit = tr_ilg(last - b);
} else if 1 < a - first {
last = a;
limit = tr_ilg(a - first);
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if limit == -2 {
/* tandem repeat copy */
ssize -= 1;
let a = stack[ssize].1;
let b = stack[ssize].2;
if stack[ssize].3 == 0 {
tr_copy(sa, isa, first, a, b, last, isad - isa);
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
tr_partialcopy(sa, isa, first, a, b, last, isad - isa);
}
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
} else {
/* sorted partition */
if 0 <= sa[first as usize] {
let mut a = first;
loop {
sa[(isa + sa[a as usize] as isize) as usize] = a as i32;
a += 1;
if !(a < last && 0 <= sa[a as usize]) {
break;
}
}
first = a;
}
if first < last {
let mut a = first;
loop {
sa[a as usize] = !sa[a as usize];
a += 1;
if !(sa[a as usize] < 0) {
break;
}
}
let next =
if sa[(isa + sa[a as usize] as isize) as usize] != tr_key(sa, isad, a) {
tr_ilg(a - first + 1)
} else {
-1
};
a += 1;
if a < last {
let v = (a - 1) as i32;
let mut b = first;
while b < a {
sa[(isa + sa[b as usize] as isize) as usize] = v;
b += 1;
}
}
/* push */
if budget.check(a - first) {
if a - first <= last - a {
stack[ssize] = (isad, a, last, -3, trlink);
ssize += 1;
isad += incr;
last = a;
limit = next;
} else if 1 < last - a {
stack[ssize] = (isad + incr, first, a, next, trlink);
ssize += 1;
first = a;
limit = -3;
} else {
isad += incr;
last = a;
limit = next;
}
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if 1 < last - a {
first = a;
limit = -3;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
continue;
}
if last - first <= TR_INSERTIONSORT_THRESHOLD {
tr_insertionsort(sa, isad, first, last);
limit = -3;
continue;
}
/* C decrements `limit` here (`limit-- == 0`); the decrement is
* observable only on the not-taken path because the taken path
* overwrites `limit` with -3. */
if limit == 0 {
tr_heapsort(sa, isad, first, last - first);
let mut a = last - 1;
while first < a {
let x = tr_key(sa, isad, a);
let mut b = a - 1;
while first <= b && tr_key(sa, isad, b) == x {
sa[b as usize] = !sa[b as usize];
b -= 1;
}
a = b;
}
limit = -3;
continue;
}
limit -= 1;
/* choose pivot */
let a = tr_pivot(sa, isad, first, last);
sa.swap(first as usize, a as usize);
let v = tr_key(sa, isad, first);
/* partition */
let (a, b) = tr_partition(sa, isad, first, first + 1, last, v);
if last - first != b - a {
let next = if sa[(isa + sa[a as usize] as isize) as usize] != v {
tr_ilg(b - a)
} else {
-1
};
/* update ranks */
{
let vv = (a - 1) as i32;
let mut c = first;
while c < a {
sa[(isa + sa[c as usize] as isize) as usize] = vv;
c += 1;
}
}
if b < last {
let vv = (b - 1) as i32;
let mut c = a;
while c < b {
sa[(isa + sa[c as usize] as isize) as usize] = vv;
c += 1;
}
}
/* push */
if 1 < b - a && budget.check(b - a) {
if a - first <= last - b {
if last - b <= b - a {
if 1 < a - first {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
last = a;
} else if 1 < last - b {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
first = b;
} else {
isad += incr;
first = a;
last = b;
limit = next;
}
} else if a - first <= b - a {
if 1 < a - first {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
last = a;
} else {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else if a - first <= b - a {
if 1 < last - b {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
first = b;
} else if 1 < a - first {
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
last = a;
} else {
isad += incr;
first = a;
last = b;
limit = next;
}
} else if last - b <= b - a {
if 1 < last - b {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
stack[ssize] = (isad + incr, a, b, next, trlink);
ssize += 1;
first = b;
} else {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
isad += incr;
first = a;
last = b;
limit = next;
}
} else {
if 1 < b - a && 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if a - first <= last - b {
if 1 < a - first {
stack[ssize] = (isad, b, last, limit, trlink);
ssize += 1;
last = a;
} else if 1 < last - b {
first = b;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
} else if 1 < last - b {
stack[ssize] = (isad, first, a, limit, trlink);
ssize += 1;
first = b;
} else if 1 < a - first {
last = a;
} else {
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
} else if budget.check(last - first) {
limit = tr_ilg(last - first);
isad += incr;
} else {
if 0 <= trlink {
stack[trlink as usize].3 = -1;
}
if ssize == 0 {
return;
}
ssize -= 1;
(isad, first, last, limit, trlink) = stack[ssize];
}
}
}
/* --------------------------------------------------------------------- */
/* Tandem repeat sort */
fn trsort(sa: &mut [i32], isa: isize, n: isize, depth: isize) {
let mut budget = TrBudget::new(tr_ilg(n) * 2 / 3, n as i32);
/* trbudget_init(&budget, tr_ilg(n) * 3 / 4, n); */
let mut isad = isa + depth;
while -(n as i32) < sa[0] {
let mut first: isize = 0;
let mut skip: isize = 0;
let mut unsorted: i32 = 0;
loop {
let t0 = sa[first as usize];
if t0 < 0 {
first -= t0 as isize;
skip += t0 as isize;
} else {
if skip != 0 {
sa[(first + skip) as usize] = skip as i32;
skip = 0;
}
let last = sa[(isa + t0 as isize) as usize] as isize + 1;
if 1 < last - first {
budget.count = 0;
tr_introsort(sa, isa, isad, first, last, &mut budget);
if budget.count != 0 {
unsorted += budget.count;
} else {
skip = first - last;
}
} else if last - first == 1 {
skip = -1;
}
first = last;
}
if !(first < n) {
break;
}
}
if skip != 0 {
sa[(first + skip) as usize] = skip as i32;
}
if unsorted == 0 {
break;
}
isad += isad - isa;
}
}
/* --------------------------------------------------------------------- */
/// `BUCKET_B(c0, c1)` for the 256-symbol alphabet.
#[inline(always)]
fn bb(c0: i32, c1: i32) -> usize {
(((c1 as u32) << 8) | c0 as u32) as usize
}
/// `BUCKET_BSTAR(c0, c1)` for the 256-symbol alphabet.
#[inline(always)]
fn bstar(c0: i32, c1: i32) -> usize {
(((c0 as u32) << 8) | c1 as u32) as usize
}
/* Sorts suffixes of type B*. */
fn sort_type_bstar(
t: &[u8],
sa: &mut [i32],
bucket_a: &mut [i32],
bucket_b: &mut [i32],
n: isize,
) -> isize {
/* Initialize bucket arrays. */
for slot in bucket_a.iter_mut() {
*slot = 0;
}
for slot in bucket_b.iter_mut() {
*slot = 0;
}
/* Count the number of occurrences of the first one or two characters of
each type A, B and B* suffix. Moreover, store the beginning position of
all type B* suffixes into the array SA. */
let mut i = n - 1;
let mut m = n;
let mut c0 = t[(n - 1) as usize] as i32;
let mut c1;
while 0 <= i {
/* type A suffix. */
loop {
c1 = c0;
bucket_a[c1 as usize] += 1;
i -= 1;
if 0 <= i {
c0 = t[i as usize] as i32;
if c0 >= c1 {
continue;
}
}
break;
}
if 0 <= i {
/* type B* suffix. */
bucket_b[bstar(c0, c1)] += 1;
m -= 1;
sa[m as usize] = i as i32;
/* type B suffix. */
i -= 1;
c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 <= c1) {
break;
}
bucket_b[bb(c0, c1)] += 1;
i -= 1;
c1 = c0;
}
}
}
let m = n - m;
/*
note:
A type B* suffix is lexicographically smaller than a type B suffix that
begins with the same first two characters.
*/
/* Calculate the index of start/end point of each bucket. */
{
let mut i: i32 = 0;
let mut j: i32 = 0;
for c0 in 0..ALPHABET_SIZE {
let t0 = i + bucket_a[c0 as usize];
bucket_a[c0 as usize] = i + j; /* start point */
i = t0 + bucket_b[bb(c0, c0)];
for c1 in (c0 + 1)..ALPHABET_SIZE {
j += bucket_b[bstar(c0, c1)];
bucket_b[bstar(c0, c1)] = j; /* end point */
i += bucket_b[bb(c0, c1)];
}
}
}
if 0 < m {
/* Sort the type B* suffixes by their first two characters. */
let pab = n - m;
let isab = m;
let mut i = m - 2;
while 0 <= i {
let t0 = sa[(pab + i) as usize];
let c0 = t[t0 as usize] as i32;
let c1 = t[(t0 + 1) as usize] as i32;
bucket_b[bstar(c0, c1)] -= 1;
sa[bucket_b[bstar(c0, c1)] as usize] = i as i32;
i -= 1;
}
{
let t0 = sa[(pab + m - 1) as usize];
let c0 = t[t0 as usize] as i32;
let c1 = t[(t0 + 1) as usize] as i32;
bucket_b[bstar(c0, c1)] -= 1;
sa[bucket_b[bstar(c0, c1)] as usize] = (m - 1) as i32;
}
/* Sort the type B* substrings using sssort. */
let buf = m;
let bufsize = n - 2 * m;
let mut c0 = ALPHABET_SIZE - 2;
let mut j = m;
while 0 < j {
let mut c1 = ALPHABET_SIZE - 1;
while c0 < c1 {
let i = bucket_b[bstar(c0, c1)] as isize;
if 1 < j - i {
sssort(
t,
sa,
pab,
i,
j,
buf,
bufsize,
2,
n,
sa[i as usize] == (m - 1) as i32,
);
}
j = i;
c1 -= 1;
}
c0 -= 1;
}
/* Compute ranks of type B* substrings. */
let mut i = m - 1;
while 0 <= i {
if 0 <= sa[i as usize] {
let j = i;
loop {
sa[(isab + sa[i as usize] as isize) as usize] = i as i32;
i -= 1;
if !(0 <= i && 0 <= sa[i as usize]) {
break;
}
}
sa[(i + 1) as usize] = (i - j) as i32;
if i <= 0 {
break;
}
}
let j = i;
loop {
sa[i as usize] = !sa[i as usize];
sa[(isab + sa[i as usize] as isize) as usize] = j as i32;
i -= 1;
if !(sa[i as usize] < 0) {
break;
}
}
sa[(isab + sa[i as usize] as isize) as usize] = j as i32;
i -= 1;
}
/* Construct the inverse suffix array of type B* suffixes using
trsort. */
trsort(sa, isab, m, 1);
/* Set the sorted order of type B* suffixes. */
let mut i = n - 1;
let mut j = m;
let mut c0 = t[(n - 1) as usize] as i32;
while 0 <= i {
i -= 1;
let mut c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 >= c1) {
break;
}
i -= 1;
c1 = c0;
}
if 0 <= i {
let t0 = i;
i -= 1;
c1 = c0;
while 0 <= i {
c0 = t[i as usize] as i32;
if !(c0 <= c1) {
break;
}
i -= 1;
c1 = c0;
}
j -= 1;
sa[sa[(isab + j) as usize] as usize] = if t0 == 0 || 1 < t0 - i {
t0 as i32
} else {
!(t0 as i32)
};
}
}
/* Calculate the index of start/end point of each bucket. */
bucket_b[bb(ALPHABET_SIZE - 1, ALPHABET_SIZE - 1)] = n as i32; /* end point */
let mut k = m - 1;
let mut c0 = ALPHABET_SIZE - 2;
while 0 <= c0 {
let mut i = bucket_a[(c0 + 1) as usize] as isize - 1;
let mut c1 = ALPHABET_SIZE - 1;
while c0 < c1 {
let t0 = i - bucket_b[bb(c0, c1)] as isize;
bucket_b[bb(c0, c1)] = i as i32; /* end point */
/* Move all type B* suffixes to the correct position. */
i = t0;
let j = bucket_b[bstar(c0, c1)] as isize;
while j <= k {
sa[i as usize] = sa[k as usize];
i -= 1;
k -= 1;
}
c1 -= 1;
}
bucket_b[bstar(c0, c0 + 1)] = (i - bucket_b[bb(c0, c0)] as isize + 1) as i32; /* start point */
bucket_b[bb(c0, c0)] = i as i32; /* end point */
c0 -= 1;
}
}
m
}
/* Constructs the suffix array by using the sorted order of type B*
* suffixes. */
fn construct_sa(
t: &[u8],
sa: &mut [i32],
bucket_a: &mut [i32],
bucket_b: &mut [i32],
n: isize,
m: isize,
) {
if 0 < m {
/* Construct the sorted order of type B suffixes by using
the sorted order of type B* suffixes. */
let mut c1 = ALPHABET_SIZE - 2;
while 0 <= c1 {
/* Scan the suffix array from right to left. */
let i = bucket_b[bstar(c1, c1 + 1)] as isize;
let mut j = bucket_a[(c1 + 1) as usize] as isize - 1;
let mut k: isize = 0;
let mut c2: i32 = -1;
while i <= j {
let mut s = sa[j as usize];
if 0 < s {
debug_assert_eq!(t[s as usize] as i32, c1);
debug_assert!((s as isize + 1) < n && t[s as usize] <= t[(s + 1) as usize]);
debug_assert!(t[(s - 1) as usize] <= t[s as usize]);
sa[j as usize] = !s;
s -= 1;
let c0 = t[s as usize] as i32;
if 0 < s && (t[(s - 1) as usize] as i32) > c0 {
s = !s;
}
if c0 != c2 {
if 0 <= c2 {
bucket_b[bb(c2, c1)] = k as i32;
}
c2 = c0;
k = bucket_b[bb(c2, c1)] as isize;
}
debug_assert!(k < j);
sa[k as usize] = s;
k -= 1;
} else {
debug_assert!((s == 0 && t[s as usize] as i32 == c1) || s < 0);
sa[j as usize] = !s;
}
j -= 1;
}
c1 -= 1;
}
}
/* Construct the suffix array by using the sorted order of type B
suffixes. */
let mut c2 = t[(n - 1) as usize] as i32;
let mut k = bucket_a[c2 as usize] as isize;
sa[k as usize] = if (t[(n - 2) as usize] as i32) < c2 {
!((n - 1) as i32)
} else {
(n - 1) as i32
};
k += 1;
/* Scan the suffix array from left to right. */
let mut i: isize = 0;
let j = n;
while i < j {
let mut s = sa[i as usize];
if 0 < s {
debug_assert!(t[(s - 1) as usize] >= t[s as usize]);
s -= 1;
let c0 = t[s as usize] as i32;
if s == 0 || (t[(s - 1) as usize] as i32) < c0 {
s = !s;
}
if c0 != c2 {
bucket_a[c2 as usize] = k as i32;
c2 = c0;
k = bucket_a[c2 as usize] as isize;
}
debug_assert!(i < k);
sa[k as usize] = s;
k += 1;
} else {
debug_assert!(s < 0);
sa[i as usize] = !s;
}
i += 1;
}
}
/* --------------------------------------------------------------------- */
/// Rust implementation of the `divsufsort()` entry point used by
/// `ZDICT_trainFromBuffer_legacy()`.
///
/// Integration removes the C function body, so this direct export provides
/// the existing library symbol without a wrapper. The `open_mp` parameter is
/// accepted for signature compatibility only: zstd never defines
/// `LIBBSC_OPENMP`, so the C implementation ignored it as well.
///
/// Returns 0 on success, -1 for invalid arguments, and -2 when the bucket
/// work arrays cannot be allocated, exactly like the C routine.
#[no_mangle]
pub unsafe extern "C" fn divsufsort(
t: *const u8,
sa: *mut c_int,
n: c_int,
open_mp: c_int,
) -> c_int {
let _ = open_mp;
/* Check arguments. */
if t.is_null() || sa.is_null() || n < 0 {
return -1;
}
if n == 0 {
return 0;
}
let text = unsafe { slice::from_raw_parts(t, n as usize) };
let suffix = unsafe { slice::from_raw_parts_mut(sa, n as usize) };
if n == 1 {
suffix[0] = 0;
return 0;
}
if n == 2 {
let m = usize::from(text[0] < text[1]);
suffix[m ^ 1] = 0;
suffix[m] = 1;
return 0;
}
let mut bucket_a: Vec<i32> = Vec::new();
let mut bucket_b: Vec<i32> = Vec::new();
if bucket_a.try_reserve_exact(BUCKET_A_SIZE).is_err()
|| bucket_b.try_reserve_exact(BUCKET_B_SIZE).is_err()
{
/* Match the C implementation's -2 result when malloc fails. */
return -2;
}
bucket_a.resize(BUCKET_A_SIZE, 0);
bucket_b.resize(BUCKET_B_SIZE, 0);
/* Suffixsort. */
let m = sort_type_bstar(text, suffix, &mut bucket_a, &mut bucket_b, n as isize);
construct_sa(text, suffix, &mut bucket_a, &mut bucket_b, n as isize, m);
0
}
#[cfg(test)]
mod tests {
use super::*;
use std::ptr;
fn build_sa(text: &[u8]) -> Vec<i32> {
let mut sa = vec![0i32; text.len()];
let result = unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), text.len() as c_int, 0) };
assert_eq!(result, 0);
sa
}
/// Trivial O(n^2 log n) reference: sort the suffix start positions by the
/// suffixes themselves.
fn reference_sa(text: &[u8]) -> Vec<i32> {
let mut sa: Vec<i32> = (0..text.len() as i32).collect();
sa.sort_by(|&a, &b| text[a as usize..].cmp(&text[b as usize..]));
sa
}
/// Suffix-array invariants: a permutation of `0..n` whose suffixes are in
/// strictly increasing lexicographic order.
fn assert_valid_sa(text: &[u8], sa: &[i32]) {
assert_eq!(sa.len(), text.len());
let mut seen = vec![false; text.len()];
for &p in sa {
let p = usize::try_from(p).expect("suffix index must be non-negative");
assert!(p < text.len(), "suffix index {p} out of range");
assert!(!seen[p], "duplicate suffix index {p}");
seen[p] = true;
}
for pair in sa.windows(2) {
assert!(
text[pair[0] as usize..] < text[pair[1] as usize..],
"suffixes {} and {} are not in sorted order",
pair[0],
pair[1]
);
}
}
/// Fixed-seed numerical-recipes LCG, used to generate reproducible
/// pseudo-random sample buffers.
fn lcg_bytes(len: usize, seed: u32, alphabet: u32) -> Vec<u8> {
let mut state = seed;
(0..len)
.map(|_| {
state = state.wrapping_mul(1_664_525).wrapping_add(1_013_904_223);
((state >> 24) % alphabet) as u8
})
.collect()
}
#[test]
fn rejects_invalid_arguments() {
let text = [0u8; 1];
let mut sa = [0i32; 1];
assert_eq!(
unsafe { divsufsort(ptr::null(), sa.as_mut_ptr(), 1, 0) },
-1
);
assert_eq!(
unsafe { divsufsort(text.as_ptr(), ptr::null_mut(), 1, 0) },
-1
);
assert_eq!(
unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), -1, 0) },
-1
);
}
#[test]
fn sorts_trivial_inputs() {
/* empty */
let text = [0u8; 1];
let mut sa = [i32::MIN; 1];
assert_eq!(
unsafe { divsufsort(text.as_ptr(), sa.as_mut_ptr(), 0, 0) },
0
);
assert_eq!(sa[0], i32::MIN, "n == 0 must not touch the output");
/* single byte */
assert_eq!(build_sa(b"z"), [0]);
/* two bytes: ascending, descending, and equal */
assert_eq!(build_sa(b"ab"), [0, 1]);
assert_eq!(build_sa(b"ba"), [1, 0]);
assert_eq!(build_sa(b"aa"), [1, 0]);
}
#[test]
fn sorts_all_equal_bytes() {
let text = vec![b'q'; 10_000];
let sa = build_sa(&text);
/* For a constant text the shortest suffix sorts first. */
let expected: Vec<i32> = (0..text.len() as i32).rev().collect();
assert_eq!(sa, expected);
}
#[test]
fn sorts_abracadabra_exactly() {
/* Hand-computed: a(10) abra(7) abracadabra(0) acadabra(3) adabra(5)
* bra(8) bracadabra(1) cadabra(4) dabra(6) ra(9) racadabra(2). */
assert_eq!(build_sa(b"abracadabra"), [10, 7, 0, 3, 5, 8, 1, 4, 6, 9, 2]);
}
#[test]
fn matches_reference_on_periodic_text() {
/* Tandem repeats exercise trsort's repeat partitioning. */
let text: Vec<u8> = b"ab".iter().copied().cycle().take(4096).collect();
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
#[test]
fn matches_reference_on_random_bytes() {
let text = lcg_bytes(8192, 0x0BAD_5EED, 256);
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
#[test]
fn matches_reference_on_low_alphabet_text() {
/* A four-symbol alphabet produces the large first-two-character
* buckets that reach sssort's block merging and the deeper trsort
* paths. */
let text = lcg_bytes(16_384, 0xDEAD_BEEF, 4);
let sa = build_sa(&text);
assert_valid_sa(&text, &sa);
assert_eq!(sa, reference_sa(&text));
}
}
-4
View File
@@ -5,8 +5,6 @@ pub mod bitstream;
pub mod common; pub mod common;
pub mod cpu; pub mod cpu;
pub mod debug; pub mod debug;
#[cfg(feature = "dict-builder")]
pub mod divsufsort;
pub mod entropy_common; pub mod entropy_common;
pub mod errors; pub mod errors;
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
@@ -31,8 +29,6 @@ pub mod zstd_compress_frame;
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
pub mod zstd_compress_literals; pub mod zstd_compress_literals;
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
pub mod zstd_compress_params;
#[cfg(feature = "compression")]
pub mod zstd_compress_sequences; pub mod zstd_compress_sequences;
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
pub mod zstd_compress_superblock; pub mod zstd_compress_superblock;
-198
View File
@@ -1,198 +0,0 @@
#![allow(non_camel_case_types)]
#![allow(non_snake_case)]
//! Precise monotonic time measurement for the command-line programs.
//!
//! Port of `programs/timefn.c`. `UTIL_time_t` is a plain nanosecond counter
//! whose absolute value is meaningless; only spans between two measurements
//! are valid. The struct crosses the C ABI by value, so it stays `repr(C)`
//! with the exact `timefn.h` layout.
//!
//! Platform selection mirrors the C preprocessor structure: Windows uses the
//! performance counter, Apple systems use the Mach absolute clock, and other
//! POSIX systems use `clock_gettime(CLOCK_MONOTONIC)`. The C90 `clock()`
//! fallback is never needed on targets Rust supports, so multi-threaded
//! measurements are always supported.
use std::os::raw::c_int;
/// Precise Time (`PTime` in timefn.h): an unsigned 64-bit nanosecond count.
pub type PTime = u64;
/// Nanosecond time counter with the `timefn.h` `UTIL_time_t` layout.
#[repr(C)]
#[derive(Clone, Copy, Debug)]
pub struct UTIL_time_t {
pub t: PTime,
}
const _: () = assert!(std::mem::size_of::<PTime>() == 8);
const _: () = assert!(std::mem::size_of::<UTIL_time_t>() == std::mem::size_of::<PTime>());
#[cfg(windows)]
mod platform {
use super::PTime;
use std::sync::OnceLock;
#[link(name = "kernel32")]
unsafe extern "system" {
/// Takes a `LARGE_INTEGER*`; the union is ABI-identical to `i64*`.
fn QueryPerformanceCounter(count: *mut i64) -> i32;
fn QueryPerformanceFrequency(frequency: *mut i64) -> i32;
}
pub fn monotonic_ns() -> PTime {
static TICKS_PER_SECOND: OnceLock<i64> = OnceLock::new();
let ticks_per_second = *TICKS_PER_SECOND.get_or_init(|| {
let mut frequency = 0i64;
if unsafe { QueryPerformanceFrequency(&mut frequency) } == 0 {
eprintln!(
"timefn::QueryPerformanceFrequency: {}",
std::io::Error::last_os_error()
);
std::process::abort();
}
frequency
});
let mut counter = 0i64;
unsafe { QueryPerformanceCounter(&mut counter) };
(counter as PTime).wrapping_mul(1_000_000_000) / ticks_per_second as PTime
}
}
#[cfg(all(unix, target_vendor = "apple"))]
mod platform {
use super::PTime;
use std::sync::OnceLock;
pub fn monotonic_ns() -> PTime {
static RATE: OnceLock<(PTime, PTime)> = OnceLock::new();
let (numer, denom) = *RATE.get_or_init(|| {
let mut rate = libc::mach_timebase_info { numer: 0, denom: 0 };
unsafe { libc::mach_timebase_info(&mut rate) };
(PTime::from(rate.numer), PTime::from(rate.denom))
});
unsafe { libc::mach_absolute_time() }.wrapping_mul(numer) / denom
}
}
#[cfg(all(unix, not(target_vendor = "apple")))]
mod platform {
use super::PTime;
pub fn monotonic_ns() -> PTime {
// Zero-initialized like the C source, which works around timespec_get
// msan limitations on some targets.
let mut time: libc::timespec = unsafe { std::mem::zeroed() };
if unsafe { libc::clock_gettime(libc::CLOCK_MONOTONIC, &mut time) } != 0 {
eprintln!(
"timefn::clock_gettime(CLOCK_MONOTONIC): {}",
std::io::Error::last_os_error()
);
std::process::abort();
}
(time.tv_sec as PTime)
.wrapping_mul(1_000_000_000)
.wrapping_add(time.tv_nsec as PTime)
}
}
/// Returns the current value of the platform's monotonic nanosecond clock.
#[no_mangle]
pub extern "C" fn UTIL_getTime() -> UTIL_time_t {
UTIL_time_t {
t: platform::monotonic_ns(),
}
}
/// Nanoseconds elapsed between two measurements, with C unsigned wrap-around.
#[no_mangle]
pub extern "C" fn UTIL_getSpanTimeNano(clockStart: UTIL_time_t, clockEnd: UTIL_time_t) -> PTime {
clockEnd.t.wrapping_sub(clockStart.t)
}
/// Microseconds elapsed between two measurements, truncated like C division.
#[no_mangle]
pub extern "C" fn UTIL_getSpanTimeMicro(begin: UTIL_time_t, end: UTIL_time_t) -> PTime {
UTIL_getSpanTimeNano(begin, end) / 1000
}
/// Microseconds elapsed since `clockStart`.
#[no_mangle]
pub extern "C" fn UTIL_clockSpanMicro(clockStart: UTIL_time_t) -> PTime {
UTIL_getSpanTimeMicro(clockStart, UTIL_getTime())
}
/// Nanoseconds elapsed since `clockStart`.
#[no_mangle]
pub extern "C" fn UTIL_clockSpanNano(clockStart: UTIL_time_t) -> PTime {
UTIL_getSpanTimeNano(clockStart, UTIL_getTime())
}
/// Busy-waits until the clock produces a new tick, improving measurement
/// accuracy on platforms with a low timer resolution.
#[no_mangle]
pub extern "C" fn UTIL_waitForNextTick() {
let clockStart = UTIL_getTime();
loop {
let clockEnd = UTIL_getTime();
if UTIL_getSpanTimeNano(clockStart, clockEnd) != 0 {
return;
}
}
}
/// All clock sources used by the Rust port are valid under multi-threaded
/// workloads; only the C90 `clock()` fallback of the C source was not.
#[no_mangle]
pub extern "C" fn UTIL_support_MT_measurements() -> c_int {
1
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn nanosecond_spans_subtract_with_unsigned_wrap_around() {
let start = UTIL_time_t { t: 100 };
let end = UTIL_time_t { t: 350 };
assert_eq!(UTIL_getSpanTimeNano(start, end), 250);
assert_eq!(UTIL_getSpanTimeNano(end, start), u64::MAX - 249);
assert_eq!(UTIL_getSpanTimeNano(start, start), 0);
}
#[test]
fn microsecond_spans_truncate_sub_tick_remainders() {
let start = UTIL_time_t { t: 0 };
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 999 }), 0);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 1_000 }), 1);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 1_999 }), 1);
assert_eq!(UTIL_getSpanTimeMicro(start, UTIL_time_t { t: 2_000 }), 2);
}
#[test]
fn clock_is_monotonic_across_measurements() {
let first = UTIL_getTime();
let second = UTIL_getTime();
assert!(second.t >= first.t);
assert!(UTIL_clockSpanNano(first) >= UTIL_getSpanTimeNano(first, second));
}
#[test]
fn waiting_for_the_next_tick_advances_the_clock() {
let before = UTIL_getTime();
UTIL_waitForNextTick();
let after = UTIL_getTime();
assert!(UTIL_getSpanTimeNano(before, after) > 0);
}
#[test]
fn multi_threaded_measurements_are_supported() {
assert_eq!(UTIL_support_MT_measurements(), 1);
}
}
+9 -194
View File
@@ -10,15 +10,9 @@
//! writes, dictionary loading, streaming, and metadata preservation remain in //! writes, dictionary loading, streaming, and metadata preservation remain in
//! `programs/fileio.c` for this first migration step. //! `programs/fileio.c` for this first migration step.
//! //!
//! Benchmark mode (`-b`) parses here and dispatches through the
//! `ZSTD_rust_cli_bench` bridge in `programs/zstdcli.c`: the run/timing loop
//! (benchfn, timefn) is Rust, while orchestration and result formatting
//! (`benchzstd.c`) remain C behind the preprocessor-gated bridge, so builds
//! with `ZSTD_NOBENCH` never reference benchmark symbols.
//!
//! Remaining C-only CLI boundaries are called out in `unsupported()` below: //! Remaining C-only CLI boundaries are called out in `unsupported()` below:
//! dictionary training, recursive/file-list expansion, tracing, //! benchmark execution, dictionary training, recursive/file-list expansion,
//! alternate-format selection, and the advanced directory modes. //! tracing, alternate-format selection, and the advanced directory modes.
use std::env; use std::env;
use std::ffi::{CStr, CString, OsStr, OsString}; use std::ffi::{CStr, CString, OsStr, OsString};
@@ -36,7 +30,6 @@ use std::os::unix::fs::FileTypeExt;
const DEFAULT_CLEVEL: i32 = 3; const DEFAULT_CLEVEL: i32 = 3;
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
const DEFAULT_MAX_CLEVEL: i32 = 19; const DEFAULT_MAX_CLEVEL: i32 = 19;
const DEFAULT_BENCH_NB_SECONDS: u32 = 3;
const DEFAULT_MEM_LIMIT: u32 = 1 << 27; const DEFAULT_MEM_LIMIT: u32 = 1 << 27;
const DEFAULT_LONG_WINDOW_LOG: u32 = 27; const DEFAULT_LONG_WINDOW_LOG: u32 = 27;
const MAX_FAST_ACCELERATION: i32 = 128 << 10; const MAX_FAST_ACCELERATION: i32 = 128 << 10;
@@ -180,22 +173,6 @@ unsafe extern "C" {
output: *const c_char, output: *const c_char,
dict: *const c_char, dict: *const c_char,
) -> c_int; ) -> c_int;
/// Benchmark bridge implemented by the `programs/zstdcli.c` shim, which
/// owns the `ZSTD_NOBENCH` preprocessor decision. Returns the benchmark
/// result (>= 0), or -1 when benchmarking is compiled out.
fn ZSTD_rust_cli_bench(
file_names: *const *const c_char,
nb_files: c_uint,
dict_file_name: *const c_char,
start_level: c_int,
end_level: c_int,
compression_params: *const ZSTD_compressionParameters,
display_level: c_int,
nb_seconds: c_uint,
block_size: usize,
nb_workers: c_int,
) -> c_int;
} }
#[derive(Clone, Copy, Debug, Eq, PartialEq)] #[derive(Clone, Copy, Debug, Eq, PartialEq)]
@@ -203,7 +180,6 @@ enum Operation {
Compress, Compress,
Decompress, Decompress,
Test, Test,
Bench,
} }
#[derive(Debug)] #[derive(Debug)]
@@ -234,7 +210,6 @@ struct Cli {
mmap_dict: i32, mmap_dict: i32,
progress: i32, progress: i32,
workers: Option<i32>, workers: Option<i32>,
single_thread: bool,
block_size: Option<usize>, block_size: Option<usize>,
mem_limit: Option<u32>, mem_limit: Option<u32>,
ldm: bool, ldm: bool,
@@ -254,8 +229,6 @@ struct Cli {
row_match_finder: i32, row_match_finder: i32,
exclude_compressed: bool, exclude_compressed: bool,
compression_params: ZSTD_compressionParameters, compression_params: ZSTD_compressionParameters,
bench_end_level: Option<i32>,
bench_nb_seconds: Option<u32>,
unsupported_program: Option<String>, unsupported_program: Option<String>,
} }
@@ -281,7 +254,6 @@ impl Cli {
mmap_dict: ZSTD_PS_AUTO, mmap_dict: ZSTD_PS_AUTO,
progress: FIO_PS_AUTO, progress: FIO_PS_AUTO,
workers: None, workers: None,
single_thread: false,
block_size: None, block_size: None,
mem_limit: None, mem_limit: None,
ldm: false, ldm: false,
@@ -301,8 +273,6 @@ impl Cli {
row_match_finder: ZSTD_PS_AUTO, row_match_finder: ZSTD_PS_AUTO,
exclude_compressed: false, exclude_compressed: false,
compression_params: ZSTD_compressionParameters::default(), compression_params: ZSTD_compressionParameters::default(),
bench_end_level: None,
bench_nb_seconds: None,
unsupported_program: None, unsupported_program: None,
}; };
@@ -362,11 +332,8 @@ unsafe fn default_worker_count() -> i32 {
} }
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
unsafe fn resolved_worker_count(workers: Option<i32>, single_thread: bool) -> i32 { unsafe fn resolved_worker_count(workers: Option<i32>) -> i32 {
match workers { match workers {
/* --single-thread pins zero workers; a bare zero (-T0 or the zstdmt
* program name) auto-detects the core count as in the C CLI. */
Some(0) if single_thread => 0,
Some(0) => unsafe { UTIL_countPhysicalCores() }.max(1), Some(0) => unsafe { UTIL_countPhysicalCores() }.max(1),
Some(workers) => workers, Some(workers) => workers,
None => unsafe { default_worker_count() }, None => unsafe { default_worker_count() },
@@ -418,7 +385,7 @@ fn usage(advanced: bool) {
let _ = writeln!(out, "\nImplemented advanced compression controls:"); let _ = writeln!(out, "\nImplemented advanced compression controls:");
let _ = writeln!( let _ = writeln!(
out, out,
" --fast[=#], --ultra, --long[=#], --threads=#, --single-thread, --block-size=#" " --fast[=#], --ultra, --long[=#], --threads=#, --block-size=#"
); );
let _ = writeln!( let _ = writeln!(
out, out,
@@ -432,22 +399,9 @@ fn usage(advanced: bool) {
out, out,
" --adapt[=min=#,max=#], --rsyncable, --[no-]row-match-finder" " --adapt[=min=#,max=#], --rsyncable, --[no-]row-match-finder"
); );
let _ = writeln!(out, "\nBenchmark options:");
let _ = writeln!( let _ = writeln!(
out, out,
" -b# Benchmark file(s) at compression level #" "\nNot yet migrated: benchmark, dictionary training, recursive/file-list expansion,"
);
let _ = writeln!(
out,
" -e# Test all levels from -b# up to # included"
);
let _ = writeln!(
out,
" -i# Set the minimum evaluation time to # seconds"
);
let _ = writeln!(
out,
"\nNot yet migrated: dictionary training, recursive/file-list expansion,"
); );
let _ = writeln!(out, "trace, alternate formats, and output-directory modes."); let _ = writeln!(out, "trace, alternate formats, and output-directory modes.");
} }
@@ -684,7 +638,6 @@ fn parse_long_option(
| "--no-row-match-finder" | "--no-row-match-finder"
| "--row-match-finder" | "--row-match-finder"
| "--rsyncable" | "--rsyncable"
| "--single-thread"
| "--compress-literals" | "--compress-literals"
| "--no-compress-literals" | "--no-compress-literals"
| "--exclude-compressed" | "--exclude-compressed"
@@ -856,14 +809,6 @@ fn parse_long_option(
cli.workers = Some(parse_worker_count(&value)?); cli.workers = Some(parse_worker_count(&value)?);
Ok(None) Ok(None)
} }
"--single-thread" => {
/* As in the C CLI: zero workers plus a latch that suppresses the
* automatic core-count resolution, so fileio runs its
* single-thread streaming mode (slightly different from -T1). */
cli.workers = Some(0);
cli.single_thread = true;
Ok(None)
}
"--memlimit" | "--memory" | "--memlimit-decompress" => { "--memlimit" | "--memory" | "--memlimit-decompress" => {
let value = next_value(attached, args, index, name)?; let value = next_value(attached, args, index, name)?;
cli.mem_limit = Some(parse_u32(&value, "memory limit")?); cli.mem_limit = Some(parse_u32(&value, "memory limit")?);
@@ -909,6 +854,7 @@ fn parse_long_option(
| "--trace" | "--trace"
| "--format" | "--format"
| "--priority" | "--priority"
| "--single-thread"
| "--auto-threads" | "--auto-threads"
| "--fake-stdin-is-console" | "--fake-stdin-is-console"
| "--fake-stdout-is-console" | "--fake-stdout-is-console"
@@ -946,33 +892,6 @@ fn parse_short_options(
'd' => cli.operation = Operation::Decompress, 'd' => cli.operation = Operation::Decompress,
'z' => cli.operation = Operation::Compress, 'z' => cli.operation = Operation::Compress,
't' => cli.operation = Operation::Test, 't' => cli.operation = Operation::Test,
'b' => cli.operation = Operation::Bench,
'e' | 'i' => {
// Benchmark range end (-e#) and duration (-i#): like the C
// parser, digits attach directly and default to 0.
let mut digits_end = offset + 1;
while digits_end < bytes.len() && bytes[digits_end].is_ascii_digit() {
digits_end += 1;
}
let digits = &value[offset + 1..digits_end];
if option == 'e' {
cli.bench_end_level = Some(if digits.is_empty() {
0
} else {
parse_i32(digits, "benchmark end level")?
});
} else {
cli.bench_nb_seconds = Some(if digits.is_empty() {
0
} else {
digits
.parse::<u32>()
.map_err(|_| format!("invalid benchmark duration: {digits:?}"))?
});
}
offset = digits_end;
continue;
}
'c' => { 'c' => {
cli.output = Some(cstring(STDOUT_MARK)?); cli.output = Some(cstring(STDOUT_MARK)?);
cli.force_stdout = true; cli.force_stdout = true;
@@ -1008,7 +927,7 @@ fn parse_short_options(
} }
break; break;
} }
'l' | 'p' | 'P' | 'r' | 's' | 'S' => { 'b' | 'e' | 'i' | 'l' | 'p' | 'P' | 'r' | 's' | 'S' => {
unsupported(&format!("-{option}"))?; unsupported(&format!("-{option}"))?;
} }
_ => return Err(format!("unknown option -{option}")), _ => return Err(format!("unknown option -{option}")),
@@ -1093,7 +1012,7 @@ unsafe fn apply_preferences(cli: &Cli, prefs: *mut FIO_prefs_t, ctx: *mut FIO_ct
}), }),
); );
#[cfg(feature = "compression")] #[cfg(feature = "compression")]
FIO_setNbWorkers(prefs, resolved_worker_count(cli.workers, cli.single_thread)); FIO_setNbWorkers(prefs, resolved_worker_count(cli.workers));
FIO_setLdmFlag(prefs, u32::from(cli.ldm)); FIO_setLdmFlag(prefs, u32::from(cli.ldm));
FIO_setAdaptiveMode(prefs, i32::from(cli.adapt)); FIO_setAdaptiveMode(prefs, i32::from(cli.adapt));
FIO_setRsyncable(prefs, i32::from(cli.rsyncable)); FIO_setRsyncable(prefs, i32::from(cli.rsyncable));
@@ -1299,52 +1218,16 @@ unsafe fn run_decompress(
dictionary, dictionary,
) )
}, },
Operation::Compress | Operation::Bench => { Operation::Compress => unreachable!("compression is dispatched separately"),
unreachable!("compression and benchmark are dispatched separately")
}
} }
} }
/// Runs benchmark mode through the C bridge. Level clamping against
/// `ZSTD_maxCLevel()` happens on the C side, where the symbol is always
/// available when benchmarking is compiled in. No input file means a
/// synthetic-sample benchmark, matching the C CLI.
fn run_bench(cli: &Cli) -> Result<i32, String> {
let inputs: Vec<*const c_char> = cli.inputs.iter().map(|value| value.as_ptr()).collect();
let dictionary = cli
.dictionary
.as_ref()
.map_or(ptr::null(), |value| value.as_ptr());
let result = unsafe {
ZSTD_rust_cli_bench(
inputs.as_ptr(),
inputs.len() as c_uint,
dictionary,
cli.level,
cli.bench_end_level.unwrap_or(cli.level),
&cli.compression_params,
cli.display_level,
cli.bench_nb_seconds.unwrap_or(DEFAULT_BENCH_NB_SECONDS),
cli.block_size.unwrap_or(0),
// The C CLI benchmarks single-threaded unless -T was given.
cli.workers.unwrap_or(1),
)
};
if result < 0 {
return Err("benchmark mode is not available in this build".to_owned());
}
Ok(result)
}
fn run_cli(mut cli: Cli) -> Result<i32, String> { fn run_cli(mut cli: Cli) -> Result<i32, String> {
if let Some(program_name) = &cli.unsupported_program { if let Some(program_name) = &cli.unsupported_program {
return Err(format!( return Err(format!(
"{program_name} compatibility mode is not yet implemented by the Rust CLI frontend" "{program_name} compatibility mode is not yet implemented by the Rust CLI frontend"
)); ));
} }
if cli.operation == Operation::Bench {
return run_bench(&cli);
}
let explicit_input_count = cli.inputs.len(); let explicit_input_count = cli.inputs.len();
filter_symlink_inputs(&mut cli); filter_symlink_inputs(&mut cli);
if explicit_input_count > 0 && cli.inputs.is_empty() { if explicit_input_count > 0 && cli.inputs.is_empty() {
@@ -1451,7 +1334,6 @@ fn run_cli(mut cli: Cli) -> Result<i32, String> {
#[cfg(not(feature = "decompression"))] #[cfg(not(feature = "decompression"))]
unreachable!("unsupported decompression was rejected above") unreachable!("unsupported decompression was rejected above")
} }
Operation::Bench => unreachable!("benchmark mode was dispatched earlier"),
} }
}; };
@@ -1620,36 +1502,6 @@ mod tests {
let cli = parse(&["zstdmt", "input"]); let cli = parse(&["zstdmt", "input"]);
assert_eq!(cli.workers, Some(0)); assert_eq!(cli.workers, Some(0));
assert!(!cli.single_thread);
}
#[test]
fn single_thread_pins_zero_workers() {
let cli = parse(&["zstd", "--single-thread", "input"]);
assert_eq!(cli.workers, Some(0));
assert!(cli.single_thread);
}
#[test]
fn a_later_thread_count_overrides_single_thread_workers() {
/* Mirrors the C CLI: -T after --single-thread wins the worker count,
* while the single-thread latch stays set. */
let cli = parse(&["zstd", "--single-thread", "-T2", "input"]);
assert_eq!(cli.workers, Some(2));
assert!(cli.single_thread);
}
#[test]
fn single_thread_rejects_attached_values() {
let error = parse_args(vec![
OsString::from("zstd"),
OsString::from("--single-thread=1"),
])
.expect_err("an attached value must not activate --single-thread");
assert!(error.contains("does not take an argument"));
} }
#[test] #[test]
@@ -1686,41 +1538,4 @@ mod tests {
assert!(error.contains("not yet implemented")); assert!(error.contains("not yet implemented"));
} }
#[test]
fn bench_mode_parses_level_duration_and_defaults() {
let cli = parse(&["zstd", "-b1", "-i0", "input"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, 1);
assert_eq!(cli.bench_nb_seconds, Some(0));
assert_eq!(cli.bench_end_level, None);
assert_eq!(
cli.inputs
.iter()
.map(|input| input.as_bytes())
.collect::<Vec<_>>(),
vec![&b"input"[..]]
);
}
#[test]
fn bench_range_aggregates_within_a_single_argument() {
let cli = parse(&["zstd", "-b5e6i2", "input"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, 5);
assert_eq!(cli.bench_end_level, Some(6));
assert_eq!(cli.bench_nb_seconds, Some(2));
}
#[test]
fn bench_duration_without_digits_defaults_to_zero() {
let cli = parse(&["zstd", "-b", "-e", "-i"]);
assert_eq!(cli.operation, Operation::Bench);
assert_eq!(cli.level, DEFAULT_CLEVEL);
assert_eq!(cli.bench_end_level, Some(0));
assert_eq!(cli.bench_nb_seconds, Some(0));
}
} }
-1048
View File
@@ -1,1048 +0,0 @@
#![allow(non_camel_case_types)]
#![allow(non_snake_case)]
#![allow(clippy::missing_safety_doc)]
#![allow(clippy::too_many_arguments)]
//! Context-free compression-parameter selection and sizing leaves.
//!
//! This module deliberately does **not** own the public `ZSTD_*` symbols yet.
//! `zstd_compress.c` still owns configuration-sensitive policy: excluded block
//! compressors, private `ZSTD_CCtx_params` layouts, LDM workspace sizing, and
//! ASAN workspace policy. The C integration layer can select a raw table
//! entry here, apply its configured strategy cascade, then use the adjustment
//! and sizing leaves below. That keeps the Rust implementation independent of
//! C preprocessor state while retaining byte-for-byte C policy for reduced
//! builds.
use crate::errors::{ZstdErrorCode, ERROR};
use std::mem::size_of;
use std::os::raw::c_int;
pub const ZSTD_CONTENTSIZE_UNKNOWN: u64 = u64::MAX;
const ZSTD_CLEVEL_DEFAULT: c_int = 3;
const ZSTD_MAX_CLEVEL: c_int = 22;
const ZSTD_TARGETLENGTH_MAX: c_int = 1 << 17;
const ZSTD_BLOCKSIZE_MAX: usize = 1 << 17;
const ZSTD_WINDOWLOG_MIN: c_int = 10;
#[cfg(target_pointer_width = "32")]
const ZSTD_WINDOWLOG_MAX: c_int = 30;
#[cfg(not(target_pointer_width = "32"))]
const ZSTD_WINDOWLOG_MAX: c_int = 31;
const ZSTD_HASHLOG_MIN: c_int = 6;
const ZSTD_HASHLOG_MAX: c_int = 30;
const ZSTD_CHAINLOG_MIN: c_int = ZSTD_HASHLOG_MIN;
#[cfg(target_pointer_width = "32")]
const ZSTD_CHAINLOG_MAX: c_int = 29;
#[cfg(not(target_pointer_width = "32"))]
const ZSTD_CHAINLOG_MAX: c_int = 30;
const ZSTD_SEARCHLOG_MIN: c_int = 1;
const ZSTD_SEARCHLOG_MAX: c_int = ZSTD_WINDOWLOG_MAX - 1;
const ZSTD_MINMATCH_MIN: c_int = 3;
const ZSTD_MINMATCH_MAX: c_int = 7;
const ZSTD_TARGETLENGTH_MIN: c_int = 0;
const ZSTD_FAST: c_int = 1;
const ZSTD_DFAST: c_int = 2;
const ZSTD_GREEDY: c_int = 3;
const ZSTD_LAZY: c_int = 4;
const ZSTD_LAZY2: c_int = 5;
const ZSTD_BTLAZY2: c_int = 6;
const ZSTD_BTOPT: c_int = 7;
const ZSTD_BTULTRA: c_int = 8;
const ZSTD_BTULTRA2: c_int = 9;
const ZSTD_C_COMPRESSION_LEVEL: c_int = 100;
const ZSTD_C_WINDOW_LOG: c_int = 101;
const ZSTD_C_HASH_LOG: c_int = 102;
const ZSTD_C_CHAIN_LOG: c_int = 103;
const ZSTD_C_SEARCH_LOG: c_int = 104;
const ZSTD_C_MIN_MATCH: c_int = 105;
const ZSTD_C_TARGET_LENGTH: c_int = 106;
const ZSTD_C_STRATEGY: c_int = 107;
/// Private compression-parameter modes from `zstd_compress_internal.h`.
///
/// They are ABI-compatible with `ZSTD_CParamMode_e` and intentionally kept as
/// integer constants because that enum remains private C API for now.
pub const ZSTD_RUST_CPM_NO_ATTACH_DICT: c_int = 0;
pub const ZSTD_RUST_CPM_ATTACH_DICT: c_int = 1;
pub const ZSTD_RUST_CPM_CREATE_CDICT: c_int = 2;
pub const ZSTD_RUST_CPM_UNKNOWN: c_int = 3;
/// `ZSTD_ParamSwitch_e` values used by the adjustment and sizing leaves.
pub const ZSTD_RUST_PS_AUTO: c_int = 0;
pub const ZSTD_RUST_PS_ENABLE: c_int = 1;
pub const ZSTD_RUST_PS_DISABLE: c_int = 2;
/// ABI-compatible `ZSTD_compressionParameters` from `zstd.h`.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct ZSTD_compressionParameters {
pub windowLog: u32,
pub chainLog: u32,
pub hashLog: u32,
pub searchLog: u32,
pub minMatch: u32,
pub targetLength: u32,
pub strategy: c_int,
}
/// ABI-compatible `ZSTD_frameParameters` from `zstd.h`.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct ZSTD_frameParameters {
pub contentSizeFlag: c_int,
pub checksumFlag: c_int,
pub noDictIDFlag: c_int,
}
/// ABI-compatible `ZSTD_parameters` from `zstd.h`.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct ZSTD_parameters {
pub cParams: ZSTD_compressionParameters,
pub fParams: ZSTD_frameParameters,
}
/// ABI-compatible `ZSTD_bounds` from `zstd.h`.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
pub struct ZSTD_bounds {
pub error: usize,
pub lowerBound: c_int,
pub upperBound: c_int,
}
/* `clevels.h`, represented as raw fields so the table stays compact and easy
* to compare mechanically against its C source. Field order is W, C, H, S,
* L, TL, strategy. */
const DEFAULT_CPARAMS: [[[u32; 7]; 23]; 4] = [
[
[19, 12, 13, 1, 6, 1, ZSTD_FAST as u32],
[19, 13, 14, 1, 7, 0, ZSTD_FAST as u32],
[20, 15, 16, 1, 6, 0, ZSTD_FAST as u32],
[21, 16, 17, 1, 5, 0, ZSTD_DFAST as u32],
[21, 18, 18, 1, 5, 0, ZSTD_DFAST as u32],
[21, 18, 19, 3, 5, 2, ZSTD_GREEDY as u32],
[21, 18, 19, 3, 5, 4, ZSTD_LAZY as u32],
[21, 19, 20, 4, 5, 8, ZSTD_LAZY as u32],
[21, 19, 20, 4, 5, 16, ZSTD_LAZY2 as u32],
[22, 20, 21, 4, 5, 16, ZSTD_LAZY2 as u32],
[22, 21, 22, 5, 5, 16, ZSTD_LAZY2 as u32],
[22, 21, 22, 6, 5, 16, ZSTD_LAZY2 as u32],
[22, 22, 23, 6, 5, 32, ZSTD_LAZY2 as u32],
[22, 22, 22, 4, 5, 32, ZSTD_BTLAZY2 as u32],
[22, 22, 23, 5, 5, 32, ZSTD_BTLAZY2 as u32],
[22, 23, 23, 6, 5, 32, ZSTD_BTLAZY2 as u32],
[22, 22, 22, 5, 5, 48, ZSTD_BTOPT as u32],
[23, 23, 22, 5, 4, 64, ZSTD_BTOPT as u32],
[23, 23, 22, 6, 3, 64, ZSTD_BTULTRA as u32],
[23, 24, 22, 7, 3, 256, ZSTD_BTULTRA2 as u32],
[25, 25, 23, 7, 3, 256, ZSTD_BTULTRA2 as u32],
[26, 26, 24, 7, 3, 512, ZSTD_BTULTRA2 as u32],
[27, 27, 25, 9, 3, 999, ZSTD_BTULTRA2 as u32],
],
[
[18, 12, 13, 1, 5, 1, ZSTD_FAST as u32],
[18, 13, 14, 1, 6, 0, ZSTD_FAST as u32],
[18, 14, 14, 1, 5, 0, ZSTD_DFAST as u32],
[18, 16, 16, 1, 4, 0, ZSTD_DFAST as u32],
[18, 16, 17, 3, 5, 2, ZSTD_GREEDY as u32],
[18, 17, 18, 5, 5, 2, ZSTD_GREEDY as u32],
[18, 18, 19, 3, 5, 4, ZSTD_LAZY as u32],
[18, 18, 19, 4, 4, 4, ZSTD_LAZY as u32],
[18, 18, 19, 4, 4, 8, ZSTD_LAZY2 as u32],
[18, 18, 19, 5, 4, 8, ZSTD_LAZY2 as u32],
[18, 18, 19, 6, 4, 8, ZSTD_LAZY2 as u32],
[18, 18, 19, 5, 4, 12, ZSTD_BTLAZY2 as u32],
[18, 19, 19, 7, 4, 12, ZSTD_BTLAZY2 as u32],
[18, 18, 19, 4, 4, 16, ZSTD_BTOPT as u32],
[18, 18, 19, 4, 3, 32, ZSTD_BTOPT as u32],
[18, 18, 19, 6, 3, 128, ZSTD_BTOPT as u32],
[18, 19, 19, 6, 3, 128, ZSTD_BTULTRA as u32],
[18, 19, 19, 8, 3, 256, ZSTD_BTULTRA as u32],
[18, 19, 19, 6, 3, 128, ZSTD_BTULTRA2 as u32],
[18, 19, 19, 8, 3, 256, ZSTD_BTULTRA2 as u32],
[18, 19, 19, 10, 3, 512, ZSTD_BTULTRA2 as u32],
[18, 19, 19, 12, 3, 512, ZSTD_BTULTRA2 as u32],
[18, 19, 19, 13, 3, 999, ZSTD_BTULTRA2 as u32],
],
[
[17, 12, 12, 1, 5, 1, ZSTD_FAST as u32],
[17, 12, 13, 1, 6, 0, ZSTD_FAST as u32],
[17, 13, 15, 1, 5, 0, ZSTD_FAST as u32],
[17, 15, 16, 2, 5, 0, ZSTD_DFAST as u32],
[17, 17, 17, 2, 4, 0, ZSTD_DFAST as u32],
[17, 16, 17, 3, 4, 2, ZSTD_GREEDY as u32],
[17, 16, 17, 3, 4, 4, ZSTD_LAZY as u32],
[17, 16, 17, 3, 4, 8, ZSTD_LAZY2 as u32],
[17, 16, 17, 4, 4, 8, ZSTD_LAZY2 as u32],
[17, 16, 17, 5, 4, 8, ZSTD_LAZY2 as u32],
[17, 16, 17, 6, 4, 8, ZSTD_LAZY2 as u32],
[17, 17, 17, 5, 4, 8, ZSTD_BTLAZY2 as u32],
[17, 18, 17, 7, 4, 12, ZSTD_BTLAZY2 as u32],
[17, 18, 17, 3, 4, 12, ZSTD_BTOPT as u32],
[17, 18, 17, 4, 3, 32, ZSTD_BTOPT as u32],
[17, 18, 17, 6, 3, 256, ZSTD_BTOPT as u32],
[17, 18, 17, 6, 3, 128, ZSTD_BTULTRA as u32],
[17, 18, 17, 8, 3, 256, ZSTD_BTULTRA as u32],
[17, 18, 17, 10, 3, 512, ZSTD_BTULTRA as u32],
[17, 18, 17, 5, 3, 256, ZSTD_BTULTRA2 as u32],
[17, 18, 17, 7, 3, 512, ZSTD_BTULTRA2 as u32],
[17, 18, 17, 9, 3, 512, ZSTD_BTULTRA2 as u32],
[17, 18, 17, 11, 3, 999, ZSTD_BTULTRA2 as u32],
],
[
[14, 12, 13, 1, 5, 1, ZSTD_FAST as u32],
[14, 14, 15, 1, 5, 0, ZSTD_FAST as u32],
[14, 14, 15, 1, 4, 0, ZSTD_FAST as u32],
[14, 14, 15, 2, 4, 0, ZSTD_DFAST as u32],
[14, 14, 14, 4, 4, 2, ZSTD_GREEDY as u32],
[14, 14, 14, 3, 4, 4, ZSTD_LAZY as u32],
[14, 14, 14, 4, 4, 8, ZSTD_LAZY2 as u32],
[14, 14, 14, 6, 4, 8, ZSTD_LAZY2 as u32],
[14, 14, 14, 8, 4, 8, ZSTD_LAZY2 as u32],
[14, 15, 14, 5, 4, 8, ZSTD_BTLAZY2 as u32],
[14, 15, 14, 9, 4, 8, ZSTD_BTLAZY2 as u32],
[14, 15, 14, 3, 4, 12, ZSTD_BTOPT as u32],
[14, 15, 14, 4, 3, 24, ZSTD_BTOPT as u32],
[14, 15, 14, 5, 3, 32, ZSTD_BTULTRA as u32],
[14, 15, 15, 6, 3, 64, ZSTD_BTULTRA as u32],
[14, 15, 15, 7, 3, 256, ZSTD_BTULTRA as u32],
[14, 15, 15, 5, 3, 48, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 6, 3, 128, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 7, 3, 256, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 8, 3, 256, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 8, 3, 512, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 9, 3, 512, ZSTD_BTULTRA2 as u32],
[14, 15, 15, 10, 3, 999, ZSTD_BTULTRA2 as u32],
],
];
#[inline]
fn cparams_from_row(row: [u32; 7]) -> ZSTD_compressionParameters {
ZSTD_compressionParameters {
windowLog: row[0],
chainLog: row[1],
hashLog: row[2],
searchLog: row[3],
minMatch: row[4],
targetLength: row[5],
strategy: row[6] as c_int,
}
}
#[inline]
fn bounds(param: c_int) -> ZSTD_bounds {
let (lowerBound, upperBound) = match param {
ZSTD_C_COMPRESSION_LEVEL => (ZSTD_rust_params_minCLevel(), ZSTD_rust_params_maxCLevel()),
ZSTD_C_WINDOW_LOG => (ZSTD_WINDOWLOG_MIN, ZSTD_WINDOWLOG_MAX),
ZSTD_C_HASH_LOG => (ZSTD_HASHLOG_MIN, ZSTD_HASHLOG_MAX),
ZSTD_C_CHAIN_LOG => (ZSTD_CHAINLOG_MIN, ZSTD_CHAINLOG_MAX),
ZSTD_C_SEARCH_LOG => (ZSTD_SEARCHLOG_MIN, ZSTD_SEARCHLOG_MAX),
ZSTD_C_MIN_MATCH => (ZSTD_MINMATCH_MIN, ZSTD_MINMATCH_MAX),
ZSTD_C_TARGET_LENGTH => (ZSTD_TARGETLENGTH_MIN, ZSTD_TARGETLENGTH_MAX),
ZSTD_C_STRATEGY => (ZSTD_FAST, ZSTD_BTULTRA2),
_ => {
return ZSTD_bounds {
error: ERROR(ZstdErrorCode::ParameterUnsupported),
lowerBound: 0,
upperBound: 0,
}
}
};
ZSTD_bounds {
error: 0,
lowerBound,
upperBound,
}
}
#[inline]
fn within_bounds(param: c_int, value: c_int) -> bool {
let bounds = bounds(param);
bounds.error == 0 && value >= bounds.lowerBound && value <= bounds.upperBound
}
#[inline]
fn clamp_unsigned(value: u32, param: c_int) -> u32 {
let bounds = bounds(param);
debug_assert_eq!(bounds.error, 0);
let signed = value as c_int;
if signed < bounds.lowerBound {
bounds.lowerBound as u32
} else if signed > bounds.upperBound {
bounds.upperBound as u32
} else {
value
}
}
#[inline]
fn clamp_cparams(mut cparams: ZSTD_compressionParameters) -> ZSTD_compressionParameters {
cparams.windowLog = clamp_unsigned(cparams.windowLog, ZSTD_C_WINDOW_LOG);
cparams.chainLog = clamp_unsigned(cparams.chainLog, ZSTD_C_CHAIN_LOG);
cparams.hashLog = clamp_unsigned(cparams.hashLog, ZSTD_C_HASH_LOG);
cparams.searchLog = clamp_unsigned(cparams.searchLog, ZSTD_C_SEARCH_LOG);
cparams.minMatch = clamp_unsigned(cparams.minMatch, ZSTD_C_MIN_MATCH);
cparams.targetLength = clamp_unsigned(cparams.targetLength, ZSTD_C_TARGET_LENGTH);
let strategy_bounds = bounds(ZSTD_C_STRATEGY);
if cparams.strategy < strategy_bounds.lowerBound {
cparams.strategy = strategy_bounds.lowerBound;
} else if cparams.strategy > strategy_bounds.upperBound {
cparams.strategy = strategy_bounds.upperBound;
}
cparams
}
#[inline]
fn highbit32(value: u32) -> u32 {
debug_assert_ne!(value, 0);
u32::BITS - 1 - value.leading_zeros()
}
#[inline]
fn cycle_log(hash_log: u32, strategy: c_int) -> u32 {
hash_log.wrapping_sub((strategy >= ZSTD_BTLAZY2) as u32)
}
#[inline]
fn strategy_supports_row_match_finder(strategy: c_int) -> bool {
(ZSTD_GREEDY..=ZSTD_LAZY2).contains(&strategy)
}
#[inline]
fn row_match_finder_used(strategy: c_int, mode: c_int) -> bool {
debug_assert_ne!(mode, ZSTD_RUST_PS_AUTO);
strategy_supports_row_match_finder(strategy) && mode == ZSTD_RUST_PS_ENABLE
}
#[inline]
fn resolve_row_match_finder(mode: c_int, cparams: ZSTD_compressionParameters) -> c_int {
if mode != ZSTD_RUST_PS_AUTO {
return mode;
}
if strategy_supports_row_match_finder(cparams.strategy) && cparams.windowLog > 14 {
ZSTD_RUST_PS_ENABLE
} else {
ZSTD_RUST_PS_DISABLE
}
}
#[inline]
fn dict_and_window_log(window_log: u32, src_size: u64, dict_size: u64) -> u32 {
/* 1ULL << ZSTD_WINDOWLOG_MAX, which is smaller for 32-bit builds. */
const MAX_WINDOW_SIZE: u64 = 1u64 << ZSTD_WINDOWLOG_MAX;
if dict_size == 0 {
return window_log;
}
debug_assert!(window_log <= ZSTD_WINDOWLOG_MAX as u32);
debug_assert_ne!(src_size, ZSTD_CONTENTSIZE_UNKNOWN);
let window_size = 1u64 << window_log;
if window_size >= dict_size.wrapping_add(src_size) {
window_log
} else {
let dict_and_window_size = dict_size.wrapping_add(window_size);
if dict_and_window_size >= MAX_WINDOW_SIZE {
ZSTD_WINDOWLOG_MAX as u32
} else {
highbit32((dict_and_window_size as u32).wrapping_sub(1)) + 1
}
}
}
/// Adjusts a *validated, configuration-resolved* parameter set.
///
/// This is the direct Rust leaf for `ZSTD_adjustCParams_internal()` after the
/// C wrapper has applied its `ZSTD_EXCLUDE_*_BLOCK_COMPRESSOR` cascade. In
/// particular, it does not select a fallback strategy itself. `srcSize == 0`
/// means a known empty input here, just as it does in the C internal helper;
/// the public C wrapper must translate zero to `ZSTD_CONTENTSIZE_UNKNOWN`.
fn adjust_cparams(
mut cparams: ZSTD_compressionParameters,
mut src_size: u64,
mut dict_size: usize,
mode: c_int,
mut use_row_match_finder: c_int,
) -> ZSTD_compressionParameters {
debug_assert_eq!(check_cparams(cparams), 0);
const MIN_SRC_SIZE: u64 = 513;
/* 1ULL << (ZSTD_WINDOWLOG_MAX - 1), which is smaller for 32-bit builds. */
const MAX_WINDOW_RESIZE: u64 = 1u64 << (ZSTD_WINDOWLOG_MAX - 1);
match mode {
ZSTD_RUST_CPM_UNKNOWN | ZSTD_RUST_CPM_NO_ATTACH_DICT => {}
ZSTD_RUST_CPM_CREATE_CDICT => {
if dict_size != 0 && src_size == ZSTD_CONTENTSIZE_UNKNOWN {
src_size = MIN_SRC_SIZE;
}
}
ZSTD_RUST_CPM_ATTACH_DICT => dict_size = 0,
_ => debug_assert!(false, "invalid ZSTD_CParamMode_e"),
}
if src_size <= MAX_WINDOW_RESIZE && (dict_size as u64) <= MAX_WINDOW_RESIZE {
let total_size = src_size.wrapping_add(dict_size as u64) as u32;
let hash_size_min = 1u32 << ZSTD_HASHLOG_MIN;
let src_log = if total_size < hash_size_min {
ZSTD_HASHLOG_MIN as u32
} else {
highbit32(total_size.wrapping_sub(1)) + 1
};
cparams.windowLog = cparams.windowLog.min(src_log);
}
if src_size != ZSTD_CONTENTSIZE_UNKNOWN {
let dict_and_window_log =
dict_and_window_log(cparams.windowLog, src_size, dict_size as u64);
let cycle_log = cycle_log(cparams.chainLog, cparams.strategy);
cparams.hashLog = cparams.hashLog.min(dict_and_window_log + 1);
if cycle_log > dict_and_window_log {
cparams.chainLog = cparams
.chainLog
.wrapping_sub(cycle_log - dict_and_window_log);
}
}
cparams.windowLog = cparams.windowLog.max(ZSTD_WINDOWLOG_MIN as u32);
/* The short-cache tags used for fast and dfast CDicts consume eight bits
* of each 32-bit index. This is a fixed private-header constant today;
* if it becomes configurable, the C shim must pass it as an explicit
* adjustment input rather than silently changing this leaf. */
if mode == ZSTD_RUST_CPM_CREATE_CDICT
&& (cparams.strategy == ZSTD_FAST || cparams.strategy == ZSTD_DFAST)
{
const SHORT_CACHE_TAG_BITS: u32 = 8;
let max_short_cache_hash_log = 32 - SHORT_CACHE_TAG_BITS;
cparams.hashLog = cparams.hashLog.min(max_short_cache_hash_log);
cparams.chainLog = cparams.chainLog.min(max_short_cache_hash_log);
}
if use_row_match_finder == ZSTD_RUST_PS_AUTO {
use_row_match_finder = ZSTD_RUST_PS_ENABLE;
}
if row_match_finder_used(cparams.strategy, use_row_match_finder) {
const ROW_HASH_TAG_BITS: u32 = 8;
let row_log = cparams.searchLog.clamp(4, 6);
let max_hash_log = (32 - ROW_HASH_TAG_BITS) + row_log;
debug_assert!(cparams.hashLog >= row_log);
cparams.hashLog = cparams.hashLog.min(max_hash_log);
}
cparams
}
#[inline]
fn get_cparam_row_size(src_size_hint: u64, dict_size: usize, mode: c_int) -> u64 {
let mut dict_size = dict_size as u64;
match mode {
ZSTD_RUST_CPM_UNKNOWN | ZSTD_RUST_CPM_NO_ATTACH_DICT | ZSTD_RUST_CPM_CREATE_CDICT => {}
ZSTD_RUST_CPM_ATTACH_DICT => dict_size = 0,
_ => debug_assert!(false, "invalid ZSTD_CParamMode_e"),
}
let unknown = src_size_hint == ZSTD_CONTENTSIZE_UNKNOWN;
let added_size = if unknown && dict_size > 0 { 500 } else { 0 };
if unknown && dict_size == 0 {
ZSTD_CONTENTSIZE_UNKNOWN
} else {
src_size_hint
.wrapping_add(dict_size)
.wrapping_add(added_size)
}
}
/// Selects a raw compression-level-table entry, without strategy cascading or
/// source/dictionary adjustment.
///
/// A C wrapper must apply its active excluded-compressor cascade to the
/// returned strategy before calling [`ZSTD_rust_params_adjustCParams`].
fn select_cparams(
compression_level: c_int,
src_size_hint: u64,
dict_size: usize,
mode: c_int,
) -> ZSTD_compressionParameters {
let row_size = get_cparam_row_size(src_size_hint, dict_size, mode);
let table_id = usize::from(row_size <= 256 * 1024)
+ usize::from(row_size <= 128 * 1024)
+ usize::from(row_size <= 16 * 1024);
let row = if compression_level == 0 {
ZSTD_CLEVEL_DEFAULT
} else if compression_level < 0 {
0
} else {
compression_level.min(ZSTD_MAX_CLEVEL)
} as usize;
let mut cparams = cparams_from_row(DEFAULT_CPARAMS[table_id][row]);
if compression_level < 0 {
let clamped = compression_level.max(ZSTD_rust_params_minCLevel());
cparams.targetLength = (-clamped) as u32;
}
cparams
}
#[inline]
fn make_params(cparams: ZSTD_compressionParameters) -> ZSTD_parameters {
ZSTD_parameters {
cParams: cparams,
fParams: ZSTD_frameParameters {
contentSizeFlag: 1,
checksumFlag: 0,
noDictIDFlag: 0,
},
}
}
#[inline]
fn check_cparams(cparams: ZSTD_compressionParameters) -> usize {
if !within_bounds(ZSTD_C_WINDOW_LOG, cparams.windowLog as c_int)
|| !within_bounds(ZSTD_C_CHAIN_LOG, cparams.chainLog as c_int)
|| !within_bounds(ZSTD_C_HASH_LOG, cparams.hashLog as c_int)
|| !within_bounds(ZSTD_C_SEARCH_LOG, cparams.searchLog as c_int)
|| !within_bounds(ZSTD_C_MIN_MATCH, cparams.minMatch as c_int)
|| !within_bounds(ZSTD_C_TARGET_LENGTH, cparams.targetLength as c_int)
|| !within_bounds(ZSTD_C_STRATEGY, cparams.strategy)
{
ERROR(ZstdErrorCode::ParameterOutOfBound)
} else {
0
}
}
/// Returns the highest table-backed compression level (`ZSTD_MAX_CLEVEL`).
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_maxCLevel() -> c_int {
ZSTD_MAX_CLEVEL
}
/// Returns the lowest public fast level (`-ZSTD_TARGETLENGTH_MAX`).
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_minCLevel() -> c_int {
-ZSTD_TARGETLENGTH_MAX
}
/// Returns the default compression level from `zstd.h`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_defaultCLevel() -> c_int {
ZSTD_CLEVEL_DEFAULT
}
/// Returns bounds for the seven core compression parameters and compression
/// level. Other `ZSTD_cParameter` cases remain C-owned.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_getBounds(param: c_int) -> ZSTD_bounds {
bounds(param)
}
/// Checks the seven fields of `ZSTD_compressionParameters`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_checkCParams(cparams: ZSTD_compressionParameters) -> usize {
check_cparams(cparams)
}
/// Clamps the seven fields of `ZSTD_compressionParameters` to public bounds.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_clampCParams(
cparams: ZSTD_compressionParameters,
) -> ZSTD_compressionParameters {
clamp_cparams(cparams)
}
/// C ABI for `ZSTD_cycleLog()`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_cycleLog(hashLog: u32, strategy: c_int) -> u32 {
cycle_log(hashLog, strategy)
}
/// C ABI for private `ZSTD_getCParamRowSize()`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_getCParamRowSize(
srcSizeHint: u64,
dictSize: usize,
mode: c_int,
) -> u64 {
get_cparam_row_size(srcSizeHint, dictSize, mode)
}
/// Returns the unadjusted compression-level-table entry.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_selectCParams(
compressionLevel: c_int,
srcSizeHint: u64,
dictSize: usize,
mode: c_int,
) -> ZSTD_compressionParameters {
select_cparams(compressionLevel, srcSizeHint, dictSize, mode)
}
/// Adjusts a validated, strategy-resolved C parameter set.
///
/// See [`adjust_cparams`] for the required C-side policy step.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_adjustCParams(
cparams: ZSTD_compressionParameters,
srcSize: u64,
dictSize: usize,
mode: c_int,
useRowMatchFinder: c_int,
) -> ZSTD_compressionParameters {
adjust_cparams(cparams, srcSize, dictSize, mode, useRowMatchFinder)
}
/// Builds `ZSTD_parameters` with the public default frame parameters.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_makeParams(
cparams: ZSTD_compressionParameters,
) -> ZSTD_parameters {
make_params(cparams)
}
/// Runtime sizes needed to reproduce `ZSTD_sizeof_matchState()` without
/// exposing private C structures to Rust.
///
/// `asanRedzoneSize` is `ZSTD_CWKSP_ASAN_REDZONE_SIZE` when workspace
/// poisoning is enabled and zero otherwise. `matchTSize` and `optimalTSize`
/// must be `sizeof(ZSTD_match_t)` and `sizeof(ZSTD_optimal_t)` respectively.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default)]
pub struct ZSTD_rustMatchStateSizing {
pub hashLog3Max: u32,
pub matchTSize: usize,
pub optimalTSize: usize,
pub asanRedzoneSize: usize,
}
/// C-only layout inputs for `ZSTD_estimateCDictSize_advanced()`.
#[repr(C)]
#[derive(Clone, Copy, Debug, Default)]
pub struct ZSTD_rustCDictSizing {
pub cdictSize: usize,
pub hufWorkspaceSize: usize,
pub hashLog3Max: u32,
pub matchTSize: usize,
pub optimalTSize: usize,
pub asanRedzoneSize: usize,
}
#[inline]
fn cwksp_alloc_size(size: usize, asan_redzone_size: usize) -> usize {
if size == 0 {
0
} else {
size.wrapping_add(asan_redzone_size.wrapping_mul(2))
}
}
#[inline]
fn cwksp_align(size: usize, alignment: usize) -> usize {
debug_assert!(alignment.is_power_of_two());
size.wrapping_add(alignment - 1) & !(alignment - 1)
}
#[inline]
fn cwksp_aligned64_alloc_size(size: usize, asan_redzone_size: usize) -> usize {
cwksp_alloc_size(cwksp_align(size, 64), asan_redzone_size)
}
#[inline]
fn shift_size(log: u32) -> usize {
debug_assert!(log < usize::BITS);
1usize << log
}
#[inline]
fn allocate_chain_table(strategy: c_int, use_row_match_finder: c_int, for_dds_dict: bool) -> bool {
for_dds_dict
|| (strategy != ZSTD_FAST && !row_match_finder_used(strategy, use_row_match_finder))
}
fn estimate_match_state_size(
cparams: ZSTD_compressionParameters,
use_row_match_finder: c_int,
enable_dedicated_dict_search: bool,
for_cctx: bool,
sizing: ZSTD_rustMatchStateSizing,
) -> usize {
if check_cparams(cparams) != 0 || use_row_match_finder == ZSTD_RUST_PS_AUTO {
return 0;
}
let chain_size = if allocate_chain_table(
cparams.strategy,
use_row_match_finder,
enable_dedicated_dict_search && !for_cctx,
) {
shift_size(cparams.chainLog)
} else {
0
};
let hash_size = shift_size(cparams.hashLog);
let hash_log_3 = if for_cctx && cparams.minMatch == 3 {
cparams.windowLog.min(sizing.hashLog3Max)
} else {
0
};
let hash3_size = if hash_log_3 == 0 {
0
} else {
shift_size(hash_log_3)
};
let table_space = chain_size
.wrapping_mul(size_of::<u32>())
.wrapping_add(hash_size.wrapping_mul(size_of::<u32>()))
.wrapping_add(hash3_size.wrapping_mul(size_of::<u32>()));
let redzone = sizing.asanRedzoneSize;
let opt_potential_space = cwksp_aligned64_alloc_size((52 + 1) * size_of::<u32>(), redzone)
.wrapping_add(cwksp_aligned64_alloc_size(
(35 + 1) * size_of::<u32>(),
redzone,
))
.wrapping_add(cwksp_aligned64_alloc_size(
(31 + 1) * size_of::<u32>(),
redzone,
))
.wrapping_add(cwksp_aligned64_alloc_size(
(1 << 8) * size_of::<u32>(),
redzone,
))
.wrapping_add(cwksp_aligned64_alloc_size(
4099usize.wrapping_mul(sizing.matchTSize),
redzone,
))
.wrapping_add(cwksp_aligned64_alloc_size(
4099usize.wrapping_mul(sizing.optimalTSize),
redzone,
));
let lazy_additional_space = if row_match_finder_used(cparams.strategy, use_row_match_finder) {
cwksp_aligned64_alloc_size(hash_size, redzone)
} else {
0
};
let opt_space = if for_cctx && cparams.strategy >= ZSTD_BTOPT {
opt_potential_space
} else {
0
};
let slack_space = 2 * 64;
table_space
.wrapping_add(opt_space)
.wrapping_add(slack_space)
.wrapping_add(lazy_additional_space)
}
/// Pure leaf for private `ZSTD_sizeof_matchState()`.
///
/// `useRowMatchFinder` must already be resolved to enable or disable. A NULL
/// `sizing` pointer or invalid C parameters returns zero; C never supplies
/// either in a valid estimator call.
#[no_mangle]
pub unsafe extern "C" fn ZSTD_rust_params_estimateMatchStateSize(
cparams: ZSTD_compressionParameters,
useRowMatchFinder: c_int,
enableDedicatedDictSearch: c_int,
forCCtx: u32,
sizing: *const ZSTD_rustMatchStateSizing,
) -> usize {
if sizing.is_null() {
return 0;
}
let sizing = unsafe { *sizing };
estimate_match_state_size(
cparams,
useRowMatchFinder,
enableDedicatedDictSearch != 0,
forCCtx != 0,
sizing,
)
}
#[inline]
fn max_nb_seq(block_size: usize, min_match: u32, use_sequence_producer: bool) -> usize {
let divider = if min_match == 3 || use_sequence_producer {
3
} else {
4
};
block_size / divider
}
/// Pure leaf for private `ZSTD_maxNbSeq()`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_maxNbSeq(
blockSize: usize,
minMatch: u32,
useSequenceProducer: c_int,
) -> usize {
max_nb_seq(blockSize, minMatch, useSequenceProducer != 0)
}
/// Pure leaf for private `ZSTD_resolveMaxBlockSize()`.
#[no_mangle]
pub extern "C" fn ZSTD_rust_params_resolveMaxBlockSize(maxBlockSize: usize) -> usize {
if maxBlockSize == 0 {
ZSTD_BLOCKSIZE_MAX
} else {
maxBlockSize
}
}
fn estimate_cdict_size_from_cparams(
dict_size: usize,
cparams: ZSTD_compressionParameters,
dict_load_method: c_int,
sizing: ZSTD_rustCDictSizing,
) -> usize {
if check_cparams(cparams) != 0 {
return 0;
}
let match_state_sizing = ZSTD_rustMatchStateSizing {
hashLog3Max: sizing.hashLog3Max,
matchTSize: sizing.matchTSize,
optimalTSize: sizing.optimalTSize,
asanRedzoneSize: sizing.asanRedzoneSize,
};
let row_match_finder = resolve_row_match_finder(ZSTD_RUST_PS_AUTO, cparams);
let copied_dict_space = if dict_load_method == 1 {
0
} else {
cwksp_alloc_size(
cwksp_align(dict_size, size_of::<usize>()),
sizing.asanRedzoneSize,
)
};
cwksp_alloc_size(sizing.cdictSize, sizing.asanRedzoneSize)
.wrapping_add(cwksp_alloc_size(
sizing.hufWorkspaceSize,
sizing.asanRedzoneSize,
))
.wrapping_add(estimate_match_state_size(
cparams,
row_match_finder,
true,
false,
match_state_sizing,
))
.wrapping_add(copied_dict_space)
}
/// Parameter-only leaf for `ZSTD_estimateCDictSize_advanced()`.
///
/// C retains ownership of `sizeof(ZSTD_CDict)`, HUF workspace configuration,
/// and sanitizer workspace policy, then passes them in `sizing`. A NULL sizing
/// pointer returns zero.
#[no_mangle]
pub unsafe extern "C" fn ZSTD_rust_params_estimateCDictSizeFromCParams(
dictSize: usize,
cparams: ZSTD_compressionParameters,
dictLoadMethod: c_int,
sizing: *const ZSTD_rustCDictSizing,
) -> usize {
if sizing.is_null() {
return 0;
}
estimate_cdict_size_from_cparams(dictSize, cparams, dictLoadMethod, unsafe { *sizing })
}
#[cfg(test)]
mod tests {
use super::*;
use crate::errors::ERR_isError;
use std::mem::{align_of, size_of};
fn adjust_public(
cparams: ZSTD_compressionParameters,
src_size: u64,
dict_size: usize,
) -> ZSTD_compressionParameters {
let src_size = if src_size == 0 {
ZSTD_CONTENTSIZE_UNKNOWN
} else {
src_size
};
adjust_cparams(
clamp_cparams(cparams),
src_size,
dict_size,
ZSTD_RUST_CPM_UNKNOWN,
ZSTD_RUST_PS_AUTO,
)
}
#[test]
fn abi_parameter_layouts_match_zstd_h() {
assert_eq!(
size_of::<ZSTD_compressionParameters>(),
7 * size_of::<u32>()
);
assert_eq!(align_of::<ZSTD_compressionParameters>(), align_of::<u32>());
assert_eq!(size_of::<ZSTD_frameParameters>(), 3 * size_of::<c_int>());
assert_eq!(size_of::<ZSTD_parameters>(), 10 * size_of::<u32>());
assert_eq!(
size_of::<ZSTD_bounds>(),
size_of::<usize>() + 2 * size_of::<c_int>()
);
}
#[test]
fn level_tables_match_representative_clevels_entries() {
let large = select_cparams(3, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN);
assert_eq!(
large,
ZSTD_compressionParameters {
windowLog: 21,
chainLog: 16,
hashLog: 17,
searchLog: 1,
minMatch: 5,
targetLength: 0,
strategy: ZSTD_DFAST,
}
);
let small = select_cparams(3, 16 * 1024, 0, ZSTD_RUST_CPM_UNKNOWN);
assert_eq!(
small,
ZSTD_compressionParameters {
windowLog: 14,
chainLog: 14,
hashLog: 15,
searchLog: 2,
minMatch: 4,
targetLength: 0,
strategy: ZSTD_DFAST,
}
);
let fast = select_cparams(-5, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN);
assert_eq!(fast.strategy, ZSTD_FAST);
assert_eq!(fast.targetLength, 5);
assert_eq!(
select_cparams(999, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN),
select_cparams(22, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN)
);
}
#[test]
fn adjustment_clamps_and_downsizes_like_the_c_leaf() {
let input = ZSTD_compressionParameters {
windowLog: 31,
chainLog: 30,
hashLog: 30,
searchLog: 1,
minMatch: 3,
targetLength: 0,
strategy: ZSTD_BTOPT,
};
let adjusted = adjust_public(input, 1, 0);
assert_eq!(adjusted.windowLog, 10);
/* C applies WINDOWLOG_ABSOLUTEMIN only after it has used the smaller
* temporary window to downsize the hash and chain logs. */
assert_eq!(adjusted.hashLog, 7);
assert_eq!(adjusted.chainLog, 7);
assert_eq!(ZSTD_rust_params_checkCParams(adjusted), 0);
let clamped = ZSTD_rust_params_clampCParams(ZSTD_compressionParameters {
windowLog: u32::MAX,
chainLog: 0,
hashLog: 0,
searchLog: 0,
minMatch: 0,
targetLength: u32::MAX,
strategy: 99,
});
assert_eq!(clamped.windowLog, ZSTD_WINDOWLOG_MIN as u32);
assert_eq!(clamped.chainLog, ZSTD_CHAINLOG_MIN as u32);
assert_eq!(clamped.hashLog, ZSTD_HASHLOG_MIN as u32);
assert_eq!(clamped.searchLog, ZSTD_SEARCHLOG_MIN as u32);
assert_eq!(clamped.minMatch, ZSTD_MINMATCH_MIN as u32);
assert_eq!(clamped.targetLength, ZSTD_TARGETLENGTH_MIN as u32);
assert_eq!(clamped.strategy, ZSTD_BTULTRA2);
}
#[test]
fn checking_and_bounds_preserve_the_public_error_contract() {
let valid = select_cparams(1, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN);
assert_eq!(check_cparams(valid), 0);
let mut invalid = valid;
invalid.minMatch = 2;
assert!(ERR_isError(check_cparams(invalid)));
let bounds = ZSTD_rust_params_getBounds(ZSTD_C_HASH_LOG);
assert_eq!(bounds.error, 0);
assert_eq!(bounds.lowerBound, 6);
assert_eq!(bounds.upperBound, 30);
assert!(ERR_isError(ZSTD_rust_params_getBounds(-1).error));
}
#[test]
fn params_default_frame_flags_match_c() {
let cparams = select_cparams(5, ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN);
let params = ZSTD_rust_params_makeParams(cparams);
assert_eq!(params.cParams, cparams);
assert_eq!(params.fParams.contentSizeFlag, 1);
assert_eq!(params.fParams.checksumFlag, 0);
assert_eq!(params.fParams.noDictIDFlag, 0);
}
#[test]
fn row_size_preserves_unknown_dictionary_overflow_semantics() {
assert_eq!(
get_cparam_row_size(ZSTD_CONTENTSIZE_UNKNOWN, 0, ZSTD_RUST_CPM_UNKNOWN),
ZSTD_CONTENTSIZE_UNKNOWN
);
assert_eq!(
get_cparam_row_size(ZSTD_CONTENTSIZE_UNKNOWN, 1, ZSTD_RUST_CPM_UNKNOWN),
500
);
assert_eq!(
get_cparam_row_size(128 * 1024, 1, ZSTD_RUST_CPM_ATTACH_DICT),
128 * 1024
);
}
#[test]
fn match_state_sizing_reproduces_table_and_row_rules() {
let sizing = ZSTD_rustMatchStateSizing {
hashLog3Max: 17,
matchTSize: 16,
optimalTSize: 32,
asanRedzoneSize: 0,
};
let fast = ZSTD_compressionParameters {
windowLog: 10,
chainLog: 10,
hashLog: 10,
searchLog: 1,
minMatch: 4,
targetLength: 0,
strategy: ZSTD_FAST,
};
assert_eq!(
estimate_match_state_size(fast, ZSTD_RUST_PS_DISABLE, false, true, sizing),
4096 + 128
);
let row = ZSTD_compressionParameters {
windowLog: 15,
chainLog: 10,
hashLog: 10,
searchLog: 4,
minMatch: 4,
targetLength: 0,
strategy: ZSTD_GREEDY,
};
assert_eq!(
estimate_match_state_size(row, ZSTD_RUST_PS_ENABLE, false, true, sizing),
4096 + 1024 + 128
);
assert_eq!(ZSTD_rust_params_maxNbSeq(100, 3, 0), 33);
assert_eq!(ZSTD_rust_params_maxNbSeq(100, 4, 0), 25);
assert_eq!(ZSTD_rust_params_resolveMaxBlockSize(0), 128 * 1024);
}
}
+4 -39
View File
@@ -87,8 +87,6 @@ RUST_TARGET_DIR := $(RUST_DIR)/target/$(RUST_BUILD_CONFIG)
RUST_STATICLIB := $(RUST_TARGET_DIR)/release/libzstd_rs.a RUST_STATICLIB := $(RUST_TARGET_DIR)/release/libzstd_rs.a
RUST_TARGET_32 ?= i686-unknown-linux-gnu RUST_TARGET_32 ?= i686-unknown-linux-gnu
RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a RUST_STATICLIB_32 := $(RUST_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_rs.a
# Tests build every library module, so they use the crate's default feature
# set (compression, decompression, and dict-builder) plus any forced HUF mode.
RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \ RUST_CARGO_FLAGS := --manifest-path $(RUST_MANIFEST) --release \
--target-dir $(RUST_TARGET_DIR) --target-dir $(RUST_TARGET_DIR)
ifneq ($(RUST_HUF_FEATURE),) ifneq ($(RUST_HUF_FEATURE),)
@@ -102,28 +100,6 @@ $(RUST_STATICLIB): $(RUST_SOURCES)
$(RUST_STATICLIB_32): $(RUST_SOURCES) $(RUST_STATICLIB_32): $(RUST_SOURCES)
$(CARGO) build $(RUST_CARGO_FLAGS) --target $(RUST_TARGET_32) $(CARGO) build $(RUST_CARGO_FLAGS) --target $(RUST_TARGET_32)
# Program-only helpers that were C sources shared with the tests (timefn,
# benchfn) now live in the Rust CLI package. The tests link a helpers-only
# archive, built without the `cli` feature: the parser/dispatch layer needs
# the C fileio backend, which test binaries do not provide.
RUST_CLI_DIR := $(RUST_DIR)/cli
RUST_CLI_MANIFEST := $(RUST_CLI_DIR)/Cargo.toml
RUST_CLI_HELPER_SOURCES := $(RUST_CLI_MANIFEST) $(RUST_CLI_DIR)/Cargo.lock \
$(RUST_CLI_DIR)/src/lib.rs \
$(RUST_DIR)/src/timefn.rs $(RUST_DIR)/src/benchfn.rs
RUST_CLI_HELPERS_TARGET_DIR := $(RUST_DIR)/target/cli-helpers
RUST_CLI_HELPERS_STATICLIB := $(RUST_CLI_HELPERS_TARGET_DIR)/release/libzstd_cli_rs.a
RUST_CLI_HELPERS_STATICLIB_32 := $(RUST_CLI_HELPERS_TARGET_DIR)/$(RUST_TARGET_32)/release/libzstd_cli_rs.a
RUST_CLI_HELPERS_CARGO_FLAGS := --manifest-path $(RUST_CLI_MANIFEST) --release \
--target-dir $(RUST_CLI_HELPERS_TARGET_DIR) \
--no-default-features
$(RUST_CLI_HELPERS_STATICLIB): $(RUST_CLI_HELPER_SOURCES)
$(CARGO) build $(RUST_CLI_HELPERS_CARGO_FLAGS)
$(RUST_CLI_HELPERS_STATICLIB_32): $(RUST_CLI_HELPER_SOURCES)
$(CARGO) build $(RUST_CLI_HELPERS_CARGO_FLAGS) --target $(RUST_TARGET_32)
# These test objects have flat filenames, unlike the configuration-hashed # These test objects have flat filenames, unlike the configuration-hashed
# program objects. Track the HUF mode separately so a C object set compiled # program objects. Track the HUF mode separately so a C object set compiled
# for one decoder is never relinked with a Rust archive for another decoder. # for one decoder is never relinked with a Rust archive for another decoder.
@@ -283,8 +259,8 @@ fuzzer32 : $(ZSTD_FILES)
$(LINK.c) $^ -o $@$(EXT) $(LINK.c) $^ -o $@$(EXT)
# note : broken : requires symbols unavailable from dynamic library # note : broken : requires symbols unavailable from dynamic library
fuzzer-dll : $(LIB_SRCDIR)/common/xxhash.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c $(PRGDIR)/datagen.c fuzzer.c $(RUST_CLI_HELPERS_STATICLIB) fuzzer-dll : $(LIB_SRCDIR)/common/xxhash.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c $(PRGDIR)/datagen.c fuzzer.c
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(RUST_CLI_HELPERS_STATICLIB) $(LDFLAGS) -o $@$(EXT) $(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(LDFLAGS) -o $@$(EXT)
CLEAN += zstreamtest zstreamtest32 CLEAN += zstreamtest zstreamtest32
ZSTREAM_LOCAL_FILES := $(PRGDIR)/datagen.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c seqgen.c zstreamtest.c external_matchfinder.c ZSTREAM_LOCAL_FILES := $(PRGDIR)/datagen.c $(PRGDIR)/util.c $(PRGDIR)/timefn.c seqgen.c zstreamtest.c external_matchfinder.c
@@ -315,8 +291,8 @@ zstreamtest_ubsan : $(ZSTREAMFILES)
# note : broken : requires symbols unavailable from dynamic library # note : broken : requires symbols unavailable from dynamic library
zstreamtest-dll : $(LIB_SRCDIR)/common/xxhash.c # xxh symbols not exposed from dll zstreamtest-dll : $(LIB_SRCDIR)/common/xxhash.c # xxh symbols not exposed from dll
zstreamtest-dll : $(ZSTREAM_LOCAL_FILES) $(RUST_CLI_HELPERS_STATICLIB) zstreamtest-dll : $(ZSTREAM_LOCAL_FILES)
$(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(RUST_CLI_HELPERS_STATICLIB) $(LDFLAGS) -o $@$(EXT) $(CC) $(CPPFLAGS) $(CFLAGS) $(filter %.c,$^) $(LDFLAGS) -o $@$(EXT)
CLEAN += paramgrill CLEAN += paramgrill
paramgrill : DEBUGFLAGS = # turn off debug for speed measurements paramgrill : DEBUGFLAGS = # turn off debug for speed measurements
@@ -375,17 +351,6 @@ $(RUST_LINK_TARGETS_32): $(RUST_STATICLIB_32)
$(RUST_LINK_TARGETS) $(RUST_LINK_TARGETS_32): $(RUST_HUF_C_MODE_STAMP) $(RUST_LINK_TARGETS) $(RUST_LINK_TARGETS_32): $(RUST_HUF_C_MODE_STAMP)
# Tests that compile the timefn/benchfn C shims also link the Rust CLI
# helpers archive, which owns those implementations. The archive is a
# prerequisite so `$^` places it after every C object referencing its symbols.
RUST_CLI_LINK_TARGETS := fullbench fullbench-lib fullbench-dll fuzzer \
zstreamtest zstreamtest_asan zstreamtest_tsan \
zstreamtest_ubsan paramgrill decodecorpus poolTests
$(RUST_CLI_LINK_TARGETS): $(RUST_CLI_HELPERS_STATICLIB)
RUST_CLI_LINK_TARGETS_32 := fullbench32 fuzzer32 zstreamtest32
$(RUST_CLI_LINK_TARGETS_32): $(RUST_CLI_HELPERS_STATICLIB_32)
.PHONY: versionsTest .PHONY: versionsTest
versionsTest: clean versionsTest: clean
$(PYTHON) test-zstd-versions.py $(PYTHON) test-zstd-versions.py