ZstdScope is a pure-Rust parser and structural inspection toolkit for the Zstandard compressed data format.
The project is intentionally focused on inspection, not compression or decompression. It exposes encoded structure and source-byte metadata through a reusable Rust API and a CLI, while leaving compressed block payloads opaque.
The
zstdscopelibrary is published as v0.2.0. The v0.3.0 project release publisheszstdscope-cliv0.3.0 for the first time. The packages use independent versions because the Rust API and CLI/JSON compatibility contracts evolve separately.
Links: library on crates.io · CLI on crates.io · docs.rs · changelog · releases
Add ZstdScope to a Rust project:
cargo add zstdscopeOr add it manually to Cargo.toml:
[dependencies]
zstdscope = "0.2"For optional Serde serialization support:
cargo add zstdscope --features serdeThe CLI package is named zstdscope-cli; it installs the zstdscope binary:
cargo install zstdscope-cli --version 0.3.0 --lockedTo install the current repository revision instead of a crates.io release:
cargo install --git https://github.com/tappe9/zstdscope zstdscope-cli --lockedThe crates.io package is the selected release distribution channel. Prebuilt GitHub Release binaries, signatures, and checksums are not currently provided; they may be added only with an explicit, automated release policy.
Source builds are continuously tested on GitHub-hosted Ubuntu x86_64, Windows x86_64, and macOS arm64 runners. Other Rust-supported targets are best effort. This is a source-build support statement, not a promise of prebuilt binaries for those targets.
The library and CLI use independent package versions. zstdscope-cli 0.3.0 depends on the published zstdscope 0.2.0 API because the parser and public Rust API did not change in the v0.3 project release.
The current structural scope includes:
- Standard and all 16 Skippable Frame magic values;
- concatenated frames with exact frame boundaries;
- Frame Header descriptor fields and derived window size;
- Frame Content Size, including all encoded widths and contradictions detectable from block-level decoded-size bounds;
- Dictionary ID, preserving an explicitly encoded zero separately from an absent field;
- Raw, RLE, and Compressed block headers and encoded content spans;
- the distinction between RLE declared size and its one-byte encoded content;
- structural rejection of Compressed blocks too small to contain their mandatory outer section headers;
- stored content checksum value and span, without claiming checksum verification;
- zero-based source spans for major encoded fields;
- typed, location-aware parse errors for malformed and truncated input.
Parsing literals, sequences, Huffman tables, FSE tables, and other compressed-block internals is intentionally deferred. The parser requires at least one complete frame and consumes the entire input. Empty input, malformed structures that the current structural layer can validate, unknown top-level magic, reserved encodings, impossible structural sizes, and trailing partial frames are errors. Compressed-block internals and content-checksum validity are intentionally not validated.
The primary API is:
pub fn inspect(data: &[u8]) -> Result<ZstdFile, ZstdError>;A simple consumer can read bytes however it chooses and pass the slice to the parser:
use zstdscope::{FrameKind, inspect};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let bytes = std::fs::read("sample.zst")?;
let file = inspect(&bytes)?;
for frame in &file.frames {
match &frame.kind {
FrameKind::Standard(standard) => {
println!(
"frame #{}: standard, {} blocks, offset={}, size={}",
frame.index,
standard.blocks.len(),
frame.span.offset,
frame.span.length
);
}
FrameKind::Skippable(skippable) => {
println!(
"frame #{}: skippable variant {}, payload={} bytes",
frame.index,
skippable.variant,
skippable.declared_payload_size
);
}
}
}
Ok(())
}All public offsets are zero-based byte offsets into the encoded input. Opaque block and Skippable payload bytes are represented by spans rather than copied into the returned model.
inspect() preserves the simple, unlimited frame/block-count behavior. Applications that inspect untrusted or externally supplied inputs can instead apply explicit metadata budgets with inspect_with_limits():
use zstdscope::{InspectionLimits, inspect_with_limits};
let limits = InspectionLimits {
max_frames: 1_024,
max_blocks_per_frame: 2_048,
max_total_blocks: 100_000,
};
let file = inspect_with_limits(&bytes, limits)?;The values above are examples, not universal safe defaults; choose limits for the application's expected workload. A count equal to the configured maximum is accepted. Attempting to parse one more affected frame or block returns the typed ZstdError::ResourceLimitExceeded at the offset where that structure would begin.
These limits bound metadata counts only. They do not cap the size of the caller-owned input slice and do not make the in-memory API streaming. Block and Skippable payloads continue to be skipped without payload-sized copies.
The core crate keeps serialization optional:
[features]
default = []
serde = ["dep:serde"]Enabling the serde feature adds Serialize support to the public inspection model. This library representation is independent from the CLI JSON wire contract. Parsing-only users do not require Serde.
Inspect a file with the human-readable renderer:
zstdscope inspect sample.zst
From a repository checkout:
cargo run -p zstdscope-cli -- inspect sample.zst
The output reports frame type and boundaries, header metadata, block types and sizes, Skippable payload metadata, and stored checksum metadata when present.
The CLI uses the in-memory library API, so an accepted input file is resident in memory while it is inspected. To keep default behavior bounded, zstdscope inspect rejects encoded input larger than 268,435,456 bytes (256 MiB) before parsing.
Raise or lower that boundary explicitly with --max-input-bytes:
zstdscope inspect large.zst --max-input-bytes 1073741824
Raising the limit also raises the maximum memory commitment for the encoded input buffer. The CLI checks file size before the full read when possible and also bounds the actual read, so a file that grows while being read cannot silently bypass the configured limit.
This CLI byte limit is separate from inspect_with_limits(), which only limits frame/block metadata counts in the library. Neither mechanism makes inspection streaming. A future streaming/file-backed library API remains the path for files that should not be held fully in memory.
Use --json for machine-readable output:
zstdscope inspect sample.zst --json
The CLI emits a dedicated JSON DTO rather than serializing the public Rust model directly. The current wire contract has:
- top-level
"schema_version": 1; - explicit
snake_casefield names and enum values; - a tagged
type/dataframe-kind representation; - decimal strings for every Rust
u64value, including offsets, lengths, input size, window size, and Frame Content Size, so JavaScript consumers do not lose integer precision; - JSON numbers for bounded
u8,u32, andusizefields; - preserved distinctions between absent and explicitly encoded zero Dictionary IDs;
- preserved RLE declared-size versus encoded-size semantics.
Before 1.0, backward-compatible additive fields may remain within schema version 1. Removing or renaming fields, changing field types, changing enum/tag representations, or otherwise breaking existing consumers requires a new schema_version and release-note documentation. Human-readable output is a separate contract and is unaffected by JSON DTO refactors.
I/O and parse failures return a non-zero exit status, write diagnostics to stderr, and do not emit partial-success JSON. An input that exceeds the CLI byte limit also returns a non-zero exit status with a structured CLI error. Output write failures are handled without panicking; a downstream process closing a pipe normally is treated as normal CLI termination.
See ADR 0005 for the JSON decision, ADR 0006 for distribution policy, and ADR 0007 for package versioning.
zstdscope/
├── crates/
│ ├── zstdscope/ # Pure Rust parsing library
│ └── zstdscope-cli/ # CLI built on the public library API
├── docs/
└── ARCHITECTURE.md
The accepted API direction is documented in Public API design.
ZstdScope aims to:
- inspect Zstandard structure without decompressing payloads;
- preserve byte-level distinctions useful to inspection and hex-viewer tooling;
- provide precise diagnostics for malformed or unsupported input;
- remain safe on untrusted byte input within the documented in-memory/resource model;
- keep parsing independent from filesystem, terminal, JSON DTO, and CLI concerns;
- keep mandatory core dependencies small;
- remain suitable for
wasm32-unknown-unknowncompilation.
ZstdScope is not intended to be:
- a compressor;
- a decompressor;
- a replacement for the official
zstdCLI orlibzstd; - a decoder for compressed block internals in the current structural scope;
- a content-checksum verifier.
- Requirements
- Architecture
- Zstandard format notes
- Public API design
- Changelog
- Release process
- Supply-chain policy
- Fuzzing guide
- Roadmap
- Architecture decision records
ZstdScope is designed against authoritative Zstandard format documentation:
- RFC 8878 — Zstandard Compression and the
application/zstdMedia Type - Zstandard reference format specification
- Zstandard reference implementation
Where the current reference specification and the RFC differ, the difference must be documented before implementation behavior is chosen.
ZstdScope treats every input byte as untrusted. Parser reads and skips are bounds-checked, offset/size arithmetic is checked, opaque payloads are not copied into the inspection model, and the project forbids authored unsafe Rust in the core crate.
The public parser APIs remain intentionally in-memory. inspect_with_limits() can bound frame/block metadata counts for untrusted inputs, while inspect() retains unlimited count behavior. The CLI adds a separate default 256 MiB encoded-input guard with --max-input-bytes as an explicit override, but any accepted CLI input is still held fully in memory. Streaming/file-backed inspection remains later hardening work.
Parser fuzzing is available through cargo-fuzz; successful fuzz parses are also checked against structural model invariants. See FUZZING.md for setup, execution, and regression-handling instructions. Fuzzing is manual initially and is not part of normal pull-request CI.
CI enforces formatting, Clippy, tests, rustdoc, MSRV, Ubuntu/Windows/macOS tests, package/publish dry-runs, packaged CLI smoke tests, a WebAssembly compile check, and the documented advisory/license/source policy. See Supply-chain policy and SECURITY.md.
The project is developed in public. See CONTRIBUTING.md for the expected workflow.
ZstdScope is dual-licensed under either of:
- Apache License, Version 2.0 (LICENSE-APACHE); or
- MIT license (LICENSE-MIT).
You may choose either license when using or redistributing ZstdScope.