NeuralAmpModeler-rs is a very high-performance and low-latency real-time neural inference DSP engine written in pure Rust. It provides a production-grade DSP library for loading, building, and executing Neural Amp Modeler (NAM) models — WaveNet (A1/A2), LSTM, ConvNet, and Linear FIR/FFT — alongside speaker cabinet impulse response (.wav) convolution, multi-rate sinc resampling, and polyphase anti-aliasing oversampling.
Engineered for seamless embedding into audio plugins (CLAP, VST3, AU), standalone real-time hosts, offline renderers, and embedded audio pipelines, it guarantees strict zero heap allocations, zero mutex locks, and zero blocking system calls on the real-time audio thread.
The crate is host-agnostic, standalone, and general-purpose: public APIs remain decoupled from specific plugin wrappers or audio servers, allowing downstream applications to integrate the engine cleanly into any audio processing graph.
-
Pure Rust & Strict Zero-Allocation Real-Time Safety: Engineered for deterministic audio callbacks under sub-millisecond deadlines. Guarantees zero heap drops, zero locks, and zero blocking syscalls on the audio thread (audited via
CountingAllocator). Parameter changes and model hot-swaps communicate via lock-free SPSC channels, backed by a 3-tier garbage-collection cascade (SPSC queue → 16-slot parking lot → overwrite ring) ensuring safe off-RT resource deallocation without glitches or xruns. -
Extremely Fast SIMD Inference & Zero-Vtable Dispatch: Enforces
x86-64-v3(AVX2/FMA/BMI2) baseline vectorization. Wide use of hand written SIMD code. Thedispatch_simd!engine performs static compile-time monomorphization with zero vtables or indirect function pointer calls. Compute kernels employ tap-major memory layouts and instruction-level parallelism (sum0..sum3) unrolling. AVX2 represents the optimal production sweet spot for NAM channel geometries ($C \le 16$ ); an experimental AVX-512 backend is available for research and cross-ISA validation via the opt-inavx512feature. -
Dual-Oracle Numerical Parity: Formally validated against two co-equal reference oracles: canonical C++ NAMCore (ensuring 100% behavioral compatibility with existing market models) and double-precision f64 reference (measuring numerical ideality). Governed by
docs/quality-contract.json(51 model baselines), typical WaveNet models achieve ESR$< 2.31 \times 10^{-14}$ (SNR$> 136\text{ dB}$ , MR-STFT$< 6.5 \times 10^{-6}$ ), while ConvNet achieves ESR$< 9.33 \times 10^{-16}$ (SNR$> 150\text{ dB}$ ). -
Const-Generic Optimization & Dynamic Topologies: Provides 23 distinct model variants (16 static const-generic profiles + zero-allocation dynamic fallbacks). Channel counts (
$CH=16, 12, 8, 4, 3$ ), kernel sizes, and receptive fields are known at compile time for canonical configurations, enabling aggressive LLVM loop unrolling and register allocation. Non-standard topologies automatically route to dynamic zero-alloc fallbacks (WaveNetModelDyn,LstmModelDyn,WaveNetA2Dyn). -
Complete Native DSP Stack (Zero External Audio Crates): Includes a native Minimum-Phase Polyphase FIR Sinc Resampler (256 phases × 64 taps, Kaiser
$\beta=12$ ,$>105\text{ dB}$ stopband attenuation, zero pre-ringing,$0.7\text{–}1.3\text{ µs}$ execution), Uniform-Partitioned Overlap-Save (UPOLS) FFT cabinet IR convolution ($1.3\text{ µs}$ for 512-sample IRs), and multi-stage Half-Band FIR Anti-Aliasing Oversampling (2×/4×,$>100\text{ dB}$ stopband). -
Adaptive Compute FSM & Pre-Transposed
.nambv2 Container: Dynamic CPU load-monitoring state machine with hysteresis that degrades model complexity under load (Full → Reduced → Minimal) using 32 ms click-free linear crossfades to prevent buffer dropouts. The binary.nambv2 container (Gate-Major LSTM, Interleaved-4 WaveNet) reduces model load time from$\sim 50\text{ ms}$ to$< 1\text{ ms}$ with mandatory IEEE 802.3 CRC32 integrity checks. -
Denormal & Subnormal Armor: Injects symmetric
$-220\text{ dBFS}$ deterministic dither ($1.0 \times 10^{-11}$ ) combined with hardware MXCSR FTZ/DAZ reassertion on every processing block, preventing 10–100× CPU microcode penalties during digital silence with zero net DC drift.
| Feature / Attribute | Technical Implementation | Benefit & Impact |
|---|---|---|
| Inference Topologies | WaveNet (A1/A2), LSTM (1-layer & 2-layer), ConvNet, and Linear FIR/FFT | Complete NAM ecosystem compatibility with native Rust execution speed |
| RT Safety Determinism | Zero heap drop, zero mutex locks, 3-tier lock-free GC cascade | Guaranteed real-time audio stability without dropouts/xruns under sub-millisecond deadlines |
| SIMD Acceleration | Mandatory x86-64-v3 (AVX2/FMA) baseline; static dispatch_simd! dispatch |
WaveNet Standard ≈ 43.3 µs / LSTM 1×16 ≈ 6.8 µs per 64-sample block (AMD Ryzen 7 5700U) |
| Const-Generic Profiles | 23 model variants (16 static const-generic profiles + dynamic fallbacks) | Compile-time LLVM loop unrolling and register allocation for canonical channel counts ( |
| Numerical Parity | Dual-oracle validation: canonical C++ NAMCore |
Bit/float-exact accuracy matching reference models ( |
| Native Polyphase Resampler | 256 phases × 64 taps Kaiser sinc resampler (minimum-phase cepstrum, 0 pre-ringing) | Pristine multi-rate conversion (>105 dB stopband, < 0.05 dB ripple) in 0.7–1.3 µs |
| Cabinet IR Convolution | Uniform-Partitioned Overlap-Save (UPOLS) FFT convolution engine (.wav IRs) | Ultra-low-overhead speaker cabinet simulation (1.3 µs for 512-sample IRs) |
| Oversampling & Anti-Aliasing | Half-band polyphase FIR filters (Off, 2x, 4x, >100 dB stopband) |
Attenuates non-linear high-frequency foldover/aliasing in high-gain amp models |
| Activation Math Modes |
Standard (exact precision Taylor minimax) vs Fast (Padé polynomial minimax) |
User-selectable trade-off between floating-point precision (+89.5 dB SNR) and CPU cycles |
| Adaptive Compute Container | Multi-profile .namb bundle support with runtime fallback switching |
Prevents audio dropouts by dynamically adjusting compute complexity under CPU spikes |
Binary .namb v2 Format |
Pre-transposed memory layout (Gate-Major LSTM, Interleaved-4 WaveNet) with CRC32 | Reduces model loading / hot-swap time from ~50 ms to < 1 ms |
| Denormal Armor | Symmetric −220 dBFS dither injection + hardware MXCSR FTZ/DAZ | Prevents 10–100× CPU microcode stalls on digital silence with zero DC drift |
| Comprehensive QA Suite | 2,000+ unit/integration tests, heap audit, soak, proptest, and Criterion benchmarks | Enterprise-grade software stability and strict regression protection |
| Architecture | Static Profiles | Dynamic Fallback | Notes |
|---|---|---|---|
| WaveNet A1 | Standard (CH=16), Lite (12), Feather (8), Nano (4) | WaveNetModelDyn |
Dilated causal 1D convs, gated tanh/sigmoid |
| WaveNet A2 | Full (CH=8), Lite (CH=3), Cascade | WaveNetA2Dyn |
Headroom-optimized modern NAM architecture |
| LSTM | 10 profiles: 1-layer (hidden 3–40), 2-layer (hidden 8–24) | LstmModelDyn |
Pre-transposed gate-major SIMD GEMV kernels |
| ConvNet | Causal Conv1D + BatchNorm1D + activation | — | Fast feed-forward models (clean / overdrive) |
| Linear | Direct FIR or Partitioned FFT convolution | — | Clean tone equalization & linear filters |
| Slimmable Container | Multi-submodel bundles with runtime quality transitions | — | Adaptive compute bundles with crossfading |
| Dependency | Minimum Requirement | Purpose |
|---|---|---|
| CPU Architecture |
x86_64 with AVX2/FMA (x86-64-v3 baseline) |
SIMD vectorized DSP kernels |
| Rust Toolchain |
|
Public MSRV promise |
| Build Tools |
build-essential, pkg-config, cmake
|
Host build tools & C++ parity oracles |
MSRV Policy:
rust-version = "1.99.0"inCargo.tomlis the guaranteed public MSRV promise. The project builds and validates on stable Rust.
sudo apt update && sudo apt install -y build-essential pkg-config cmakeNeuralAmpModeler-rs requires an x86_64 processor supporting the x86-64-v3 instruction set baseline (avx, avx2, bmi1, bmi2, f16c, fma, lzcnt, movbe). Build-time assertions enforce this target to guarantee SIMD performance across all DSP modules.
Because Cargo does not automatically propagate compiler target flags to downstream dependencies, applications depending on NeuralAmpModeler-rs must instruct rustc to target x86-64-v3:
Option A — In .cargo/config.toml (Recommended for downstream crates):
[build]
rustflags = ["-Ctarget-cpu=x86-64-v3"]Option B — Via environment variable:
RUSTFLAGS="-Ctarget-cpu=x86-64-v3" cargo build --releaseVerify your host CPU and toolchain compatibility directly using the built-in diagnostic probe:
cargo run --bin simd_probeThe probe validates CPU feature flags, operating system vector context support (OSXSAVE), reports the active monomorphized SIMD backend, and executes a synthetic inference smoke test with deterministic checksum verification.
Add NeuralAmpModeler-rs to your Cargo.toml:
[dependencies]
NeuralAmpModeler-rs = "1.0"For test harnesses, audio generators, and perceptual fidelity measurement tools:
[dependencies]
NeuralAmpModeler-rs = { version = "1.0", features = ["testing"] }use std::path::Path;
use neural_amp_modeler_rs::prelude::*;
fn main() -> Result<(), Box<dyn std::error::Error>> {
// 1. Capture system hardware capabilities (SIMD features, CPU topology)
let sys = SystemSnapshot::capture();
// 2. Load a .nam (JSON) or .namb (binary) model
let mut model_pair = load_and_build_model(
Path::new("models/BossWN-standard.nam"),
&sys,
false, // dual_mono: false = mono left-channel only
LoadOptions::default(),
)?;
// 3. Process an audio block on the real-time thread (zero allocations, zero locks)
let input_buffer = [0.0_f32; 64];
let mut output_buffer = [0.0_f32; 64];
if let Some(ref mut model) = model_pair.model_l {
model.process(&input_buffer, &mut output_buffer);
}
Ok(())
}For a complete end-to-end signal processing chain — neural amp model, cabinet impulse response (IR) convolution, and 4× polyphase anti-aliasing oversampling — see the offline_render example:
cargo run --example offline_render -- path/to/model.namThe crate includes runnable examples demonstrating key features:
| Example | Description | Command |
|---|---|---|
load_model |
Off-RT model file loading (.nam/.namb) and SIMD state prewarming |
cargo run --example load_model -- <model.nam> |
inspect_model |
Inspects metadata, architecture, weights, and sample rates (Text/JSON) | cargo run --example inspect_model -- <model.nam> |
offline_render |
Full audio rendering with 4× polyphase oversampling (HQ mode) | cargo run --example offline_render -- <model.nam> |
cabsim |
Standalone cabinet impulse response (IR) convolution and resampling | cargo run --example cabsim -- <ir.wav> |
diagnostics |
Diagnostic bundle (DiagnosticBundle) and log buffer (LogBuffer) export |
cargo run --example diagnostics |
math_activations |
Precision vs speed benchmarks for SIMD activations (Standard vs Fast) |
cargo run --example math_activations |
synthetic_model |
In-memory StaticModel construction without external files |
cargo run --example synthetic_model |
| Module | Purpose |
|---|---|
loader |
Model deserialization and construction (.nam, .namb) |
math |
SIMD math primitives, activation approximations, GEMV/FFT kernels |
models |
Neural network architectures and StaticModel dispatch |
dsp |
Complete DSP engine: polyphase resampler, noise gate, oversampling, cabsim |
common |
System telemetry, diagnostics, lock-free SPSC channels |
testing |
Off-RT testing utilities, perceptual metrics, and test fixtures (feature-gated) |
API Documentation:
- Online: docs.rs/NeuralAmpModeler-rs
- Local:
cargo doc --open
NeuralAmpModeler-rs provides modular Cargo feature flags to tailor capabilities, diagnostic tooling, and benchmarking:
| Feature | Default | Description | Category |
|---|---|---|---|
dual-mono |
Enabled | Independent per-channel neural inference across the DSP pipeline for stereo/dual-mono signals. Disable for single-channel mono. | Production |
heap-audit |
Disabled | Enables real-time allocation tracking via CountingAllocator to enforce zero-heap invariants. |
Diagnostics |
testing |
Disabled | Exposes off-RT test utilities, audio signal generators, synthetic fixtures, and perceptual fidelity measurement oracles. | Tooling / QA |
fft-radix4-planner |
Disabled | Exposes Radix-4 DIT FFT execution planning routines (FftPlannerRadix4) and benchmarks. |
Tooling / Benchmarks |
rt-hardening |
Disabled | Opt-in Linux real-time system hardening (THP disable, mlockall, SCHED_FIFO, MXCSR DAZ/FTZ, CPU affinity). Linux-only. |
Production (Linux RT) |
avx512 |
Disabled | Experimental AVX-512 kernels and instruction dispatch for research and cross-ISA validation. | Research / Experimental |
Note on
avx512: Production builds target thex86-64-v3(AVX2 + FMA) baseline. Benchmarks demonstrate that AVX2 delivers lower latency and higher throughput across canonical NAM channel geometries ($C \le 16$ ). The opt-inavx512feature is maintained for research, hardware benchmarking, and architectural exploration. Seedocs/architecture.mdanddocs/benchmarks.md.
The engine's numerical accuracy is strictly enforced through continuous baseline regression testing (docs/quality-contract.json), tracking 51 model baseline envelopes against both the C++ NAMCore reference and a double-precision
-
Dual-Oracle Parity: WaveNet Standard models maintain an ESR of
$2.31 \times 10^{-14}$ against NAMCore and$9.05 \times 10^{-15}$ against the$f64$ oracle ($136.4\text{ dB}$ SNR). ConvNet ReLU achieves$9.33 \times 10^{-16}$ ESR ($150.3\text{ dB}$ SNR). -
Perceptual Validation: All models verify multi-resolution STFT distance (MR-STFT
$< 6.5 \times 10^{-6}$ ), spectral convergence, and LUFS loudness constancy.
Execution times measured per 64-sample block at 48 kHz (1.33 ms real-time deadline budget) on an AMD Ryzen 7 5700U (AVX2 baseline, rustc 1.98.1):
| Component / Profile | Execution Time | RT Deadline Budget Used |
|---|---|---|
| WaveNet Standard (CH=16) | 43.3 µs | 3.2% |
| WaveNet Lite (CH=12) | 56.4 µs | 4.2% |
| WaveNet Feather (CH=8) | 19.9 µs | 1.5% |
| WaveNet Nano (CH=4) | 17.8 µs | 1.3% |
| WaveNet A2 Full (CH=8) | 25.6 µs | 1.9% |
| WaveNet A2 Lite (CH=3) | 22.6 µs | 1.7% |
| LSTM 1-Layer (Hidden=16) | 6.8 µs | 0.5% |
| LSTM 2-Layer (Hidden=8) | 7.1 µs | 0.5% |
| ConvNet | 8.7 µs | 0.7% |
| Linear FIR (RF=4 Direct) | 0.3 µs | 0.02% |
| DSP Resampler (44.1 kHz → 48 kHz) | 1.2 µs | 0.09% |
| Cabinet IR Simulation (UPOLS 512) | 1.2 µs | 0.09% |
| Full Pipeline Base (Model + CabSim, No OS) | 43.6 µs | 3.3% |
| Full Pipeline HQ (Model + CabSim + 4× OS) | 177.4 µs | 13.3% |
Detailed benchmarking methodologies, throughput curves, and regression envelopes are documented in
docs/benchmarks.md.
Crate consumers only require Cargo and a standard Rust toolchain. Developers contributing to NeuralAmpModeler-rs or validating parity against C++ reference implementations should set up the local mirror environment:
# Clone and configure pinned C++ NAMCore and Plugin mirrors (gitignored):
./utils/setup-third-party.sh
# Optional: Link a local directory of private community test models
NAM_COMMUNITY_MODELS_SRC=/path/to/models ./utils/setup-third-party.shThe ./utils/ suite automates formatting, static analysis, parity verification, and quality gates:
| Script | Purpose & Scope |
|---|---|
utils/lints.sh |
Static Analysis Gate: Runs cargo fmt, strict clippy, compilation checks, doc-tests, and SPDX header verification. |
utils/tests-quick.sh |
Agile QA Suite: Multi-phase test run covering unit tests, C++ parity quick checks, and parser fuzzing. |
utils/quality-dashboard.sh |
Quality & Regression Gate: Verifies benchmark timings and audio fidelity against docs/quality-contract.json. |
utils/check-model.sh |
Model Inspector CLI: Inspects .nam and .namb files, outputting detailed reports, JSON, or manifest arrays. |
utils/simd-probe.sh |
SIMD Diagnostic Probe: Inspects CPU features, OS vector context, and runs deterministic inference smoke tests. |
utils/tests-long.sh |
Pre-Release Audit: Exhaustive suite including soak tests, full proptest/fuzzing, cross-ISA, and heap audits (run by maintainers). |
utils/setup-third-party.sh |
Local Environment Bootstrap: Clones and syncs pinned vendor mirrors into third-party/. |
utils/mod-update.sh |
Dependency Maintenance: Updates the Rust toolchain and dependencies via cargo upgrade. |
Standard verification workflow:
# 1. Static analysis & format checks
./utils/lints.sh
# 2. Agile QA validation
./utils/tests-quick.shComprehensive architectural specifications and engineering guides are available in the docs/ directory:
| Document | Topic |
|---|---|
docs/architecture.md |
Engine architecture, SIMD kernels, memory layout, and DSP pipelines |
docs/audio_fidelity_map.md |
Audio fidelity decisions, frequency response, and quality modes (Live vs HQ) |
docs/fastmath-approximations.md |
Activation approximations (Padé/minimax), polynomial error bounds, and SNR |
docs/namb-spec.md |
Binary .namb v2 container specification, metadata schema, and memory layout |
docs/cpp_parity_map.md |
Bit/float-exact parity verification against canonical C++ NeuralAmpModelerCore |
docs/perceptual_validation.md |
Perceptual audio metrics: ESR, MR-STFT, Spectral Convergence, and LUFS |
docs/benchmarks.md |
Criterion benchmarks, throughput profiles, and performance regression gates |
docs/testing.md |
Test suite organization, oracle hierarchy, and verification policies |
docs/fixtures.md |
Golden vector formats, stress signal generation, and fixture discovery |
docs/functional-tests.md |
Functional test matrix, execution protocols, and certification checklists |
docs/research-references.md |
Scientific literature, DSP bibliography, and neural audio modeling research |
docs/quality-contract.json |
Regression baseline thresholds: audio fidelity and benchmark SLA envelopes |
- Testing, testing, testing! Go on, "cargo add NeuralAmpModeler-rs" into your own project and make NeuralAmpModeler-rs live and evolve!
- Feedback & Issues: Submit detailed bug reports, questions, or feature requests via GitHub Issues.
- Model Compatibility: Test your favorite
.nammodels and impulse responses, and share your benchmarks or findings. - Code Contributions: Help me to make NeuralAmpModeler-rs more useful and correct for the community!
- Steven Atkinson — Creator of Neural Amp Modeler (NAM) for pioneering deep learning amplifier modeling and generously open-sourcing the ecosystem.
- Mike Oliphant — Author of NeuralAudio, whose work provided early insights into WaveNet inference optimization.
The system architecture, DSP design decisions, mathematical verification frameworks, and project orchestration represent the creative and technical direction of the author (Fábio Henrique de Lima Silva). Implementation and iterative development were accelerated through human-directed AI pair programming.
This project is licensed under the Apache License, Version 2.0. See LICENSE.txt for details.