You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
hotkey/button activation
-> microphone ready
-> authoritative audio captured
-> provider request/stream completed
-> final transcript returned
-> final transcript visibly inserted into the intended target
The primary KPI is no longer encoder time or even stop -> requestStarted. It is:
activation_received_to_final_text_observed_ms
For toggle recording, also retain the critical secondary KPI:
stop_requested_to_final_text_observed_ms
The end marker must prove that the exact expected final transcript is visible in the guarded target application. clipboard_set, paste_requested, an API callback, or first_paste alone is not sufficient.
This issue should pursue the lowest total latency for each exact provider, route, model, workload, recording duration, and supported audio format. It should use the fastest verified Rust implementation for the winning format, including pinned Nightly/SIMD builds where they materially improve the complete path.
The detailed provider-format matrix in the following comment is normative input to this issue:
This work depends on the reusable HTTP transport from #16 and should include the OpenRouter STT route from #17 when available.
Product-level optimization rule
Do not choose one universal codec for every provider.
The production objective is:
For every frozen provider + endpoint + model + workload:
choose the fastest verified representation and implementation
that preserves transcript quality, reliability, privacy, and compatibility.
Examples:
Azure MAI / mai-transcribe-1.5
-> benchmark zero-copy WAV, Nightly Rust FLAC, Rust MP3, current FFmpeg MP3
-> no Opus unless the exact endpoint contract changes
Smallest batch / exact Pulse route
-> benchmark passthrough, WAV, MP3, FLAC, Ogg/Opus, WebM/Opus
Deepgram async / exact pre-recorded model
-> benchmark passthrough and documented Opus/MP3/FLAC/WAV choices
OpenAI Transcriptions
-> benchmark accepted MP3/WAV/WebM paths; do not select unsupported Opus/FLAC
Provider-native realtime routes
-> use their exact streaming encoding contract instead of forcing a batch file codec
Unknown models, custom endpoints, and unresolved OpenRouter upstream providers must fail closed to a conservative capability set.
Primary and diagnostic timing contract
Extend the existing HotPathTracer and installed text-target benchmark rather than creating an unrelated timer.
Required start markers
activation_received
hotkey_received # native Tauri QPC marker for global hotkey
button_received # bound UI/native marker for button activation
start_request_dispatched
recording_state_visible
mic_ready
first_audio_frame
first_audible_audio_frame
Required stop/finalization markers
stop_requested
last_audio_frame_captured
last_chunk_sent_to_pipeline
capture_stopped
encoder_tail_started
encoder_tail_completed
request_started
first_request_chunk_sent
last_request_chunk_sent
provider_final_received
transcript_parsed
post_processing_completed # only when enabled
injection_target_validated
clipboard_set
paste_requested
final_text_observed
Because the primary measurement includes the time during which the user is speaking, benchmark comparisons must use identical authoritative audio duration and content. Also report:
For real user sessions, store only privacy-safe aggregate phase timings; never infer or persist spoken content from timing data.
The existing markers such as hotkey_received_to_mic_ready_ms, stop_requested_to_first_paste_ms, stop_requested_to_provider_final_received_ms, and provider_final_received_to_clipboard_set_ms remain useful diagnostics, but the release decision must use the final observed-text KPI.
Critical-path model
Treat the workflow as a directed acyclic graph rather than a sequential list.
A simplified batch/final-only route should become:
activation received
├─ start/lease microphone immediately
├─ freeze provider + endpoint + model + capability entry
├─ acquire/reuse provider HTTP connection
├─ initialize selected encoder/resampler
├─ capture injection-target guard
└─ update overlay/UI
↓
first PCM frame
├─ authoritative PCM spool
├─ VAD/speech diagnostics
└─ selected provider-compatible encoder running during capture
↓
stop requested
├─ stop/drain physical capture
├─ finalize only the small encoder tail
├─ clean audio routing / resume prewarm
├─ validate target and prewarm injection path
└─ prepare multipart headers/request metadata
↓
encoded artifact or stream ready
└─ stream multipart upload over the already warm connection
↓
provider final response
├─ parse/validate final text
├─ finish optional post-processing
└─ revalidate target in parallel where safe
↓
insert text
↓
observe exact final text in target
For provider-native realtime routes, the network stream runs during capture and there is no batch audio-preparation stage after Stop. The relevant optimization is connection/session readiness, provider finalization, and final text insertion.
Parallel execution requirements
1. Idle and application-start prewarm
Perform non-sensitive, non-billable preparation before the first dictation:
keep the existing Rust audio sidecar alive;
prewarm microphone/VAD/SmartTurn resources where already supported;
resolve and cache provider capabilities and selected model metadata;
preallocate bounded PCM and encoder buffers;
initialize reusable resampler plans and codec tables where thread-safe;
warm the native text-injection path without importing legacy fallbacks unnecessarily;
compile/runtime-dispatch SIMD kernels once rather than on the Stop path.
A provider connection may be proactively warmed only through a documented idempotent/non-billable operation. Never issue a transcription POST merely to warm a connection.
2. Work started at activation
At activation_received, launch independent work concurrently:
microphone acquisition and first-frame delivery;
provider route/model freeze;
HTTP pool/connection acquisition;
selected encoder initialization;
resampler initialization if required;
injection-target identity capture;
overlay/UI update.
None of these tasks may delay microphone capture unless it is required for correctness.
3. Encode during capture for buffered/final-only routes
Choose the format before the first audio frame from the frozen capability and benchmark policy. Tee PCM into exactly one production encoder while preserving the authoritative PCM path:
encoder overload fails the candidate locally and preserves a fallback before any provider POST;
use a dedicated CPU pool/thread, not the async/network event loop;
avoid oversubscribing the CPU with VAD, resampling, and many codec workers;
production encodes one selected format, not every candidate simultaneously;
benchmark variants run in separate/interleaved trials so redundant codecs do not distort results.
4. Parallel Stop boundary
On Stop, run safe independent tails concurrently:
capture stop/drain;
final encoder flush/container close;
audio-routing cleanup and microphone-prewarm handoff;
injection-target revalidation;
request metadata/header construction;
UI transition to Transcribing.
Use a structured-concurrency owner so failure or cancellation joins every child task and cleans every handle/artifact.
5. Overlap finalization and upload where the format permits
The selected transport should begin consuming bytes as early as the provider contract safely permits:
WAV: once total PCM bytes are known at Stop, emit a virtual RIFF header followed directly by the authoritative PCM spool; avoid a complete copied WAV file;
MP3: stream already-produced frames and append the final flush frames; do not require nonessential ID3/Xing metadata for STT;
Ogg/Opus or WebM/Opus: stream completed pages/clusters and append the terminal page/cluster;
FLAC: benchmark a streamable/final-header strategy against a seekable encoded spool. Do not delay upload for checksums/metadata that the provider does not require.
Do not start a billable batch transcription POST before the user stops unless that exact provider documents a cancellable streaming-ingest contract and Scriber deliberately implements it. Provider-native realtime sessions are the exception because streaming is their normal contract.
6. Parallel response tail and injection preparation
While awaiting the provider final response, safely prepare:
guarded target revalidation;
clipboard snapshot strategy;
native SendInput/clipboard route readiness;
post-processing connection reuse when post-processing is enabled.
Insert immediately after the final accepted text is ready. Measure actual target observation, not only injection API completion.
Fastest implementation candidates by format
The phrase “fastest Rust implementation” means the fastest implementation on Scriber's reference Windows systems and fixtures, not the fastest upstream marketing number.
WAV / PCM
Implement a zero-copy or near-zero-copy Rust virtual WAV stream:
44-byte RIFF header + existing PCM spool
This is the no-compression lower bound and likely winner for short recordings on fast uploads.
FLAC
Required candidates:
flacenc stable
flacenc pinned-nightly SIMD/default
rezin-flac parallel long-form
FFmpeg FLAC fast control
The current upstream flacenc reports approximately 575x realtime on stable and 1,309x realtime for its Nightly default benchmark, with a compression ratio close to the reference FLAC level-5 result. These figures use a small music corpus and are hypotheses only; Scriber must remeasure 16-kHz mono speech.
Nightly flacenc is the preferred maximum-speed FLAC candidate if it wins the installed Windows benchmark and passes toolchain, packaging, and reproducibility gates. rezin-flac is a long-form parallel challenger; do not assume many workers help short dictation.
Opus
Required candidates for exact provider/model routes that explicitly accept Opus:
ruopus pure Rust SIMD
libopusenc reference control
optional crime/moosicbox façade controls
ruopus reports encode performance at approximately parity with SIMD libopus at matched complexity, between roughly 560x and 1,088x realtime in its published modes. Benchmark speech-oriented bitrate/complexity settings and full Ogg/WebM container finalization.
Prefer ruopus when it is the fastest qualifying implementation and passes conformance, accuracy, packaging, cancellation, and long-session tests. Retain libopusenc as the reference and safe fallback.
MP3
Required candidates:
mp3lame-encoder in-process control
shine-rs / Shine-derived maximum-speed challenger
current FFmpeg MP3 control
Rust Nightly cannot by itself make the native LAME backend a Rust-native codec, but the in-process wrapper can remove FFmpeg process and pipe overhead. The Shine family reports substantially faster encoding than LAME in its own benchmark, but its simple psychoacoustic model can produce lower-quality/larger MP3 output. It is eligible only if actual STT WER/CER and exact-fixture results remain within the quality guardrail.
The production MP3 implementation is whichever candidate minimizes complete latency while passing quality and licensing gates—not automatically the pure-Rust option.
Original-file passthrough
For file, YouTube, and meeting artifacts, unchanged passthrough is the first candidate whenever the exact provider/model accepts the original format. No encoder can outperform zero re-encoding.
Nightly, SIMD, PGO, and build strategy
Maximum-speed experiments must include a pinned and reproducible Nightly toolchain without making the entire application accidentally depend on unpinned Nightly behavior.
pin the exact Nightly date/version in repository configuration;
record compiler, crate, feature, and target fingerprints in benchmark artifacts;
use release builds only;
evaluate fat/Thin LTO, codegen-units=1, panic strategy, and PGO on Scriber speech fixtures;
do not ship one target-cpu=native binary that fails on other users' CPUs;
prefer runtime SIMD dispatch or separately compiled baseline/AVX2 kernels;
validate SSE2/baseline, AVX2, and any optional higher feature path on real supported machines;
include the selected codec/toolchain capabilities in sidecar self-test and runtime attestation;
retain a stable implementation fallback when the Nightly artifact cannot run or fails validation.
Nightly is selected for production only when the complete installed-app path is materially faster and the build is reproducible, supportable, and secure.
Provider-aware selection policy
Use the verified matrix from the linked issue comment and freeze the result per job.
Suggested policy order:
1. Native realtime route:
use the exact provider streaming contract;
avoid batch re-encoding.
2. Direct file/YouTube/meeting route:
pass through the original accepted format unchanged.
3. Buffered/final-only Live Mic route:
choose the measured fastest accepted format for the provider/model,
duration bucket, machine capability, and validated network profile.
4. If the chosen encoder is unavailable or fails locally:
select a pre-approved local fallback before sending any provider request.
5. If no verified representation exists:
fail locally with an explicit unsupported-format error.
Do not use guessed live bandwidth as a routing signal. Network-specific routing may be added only if a bounded, private, stable signal demonstrably improves p95 without oscillation. Initially derive conservative duration/size crossovers from controlled network profiles.
Benchmark design
Workloads
Run each eligible provider/format/implementation against identical authoritative fixtures:
5 s
15 s
30 s
60 s
180 s
600 s where practical
Include German and English clean speech, moderately noisy speech, silence/near-silence lifecycle cases, and the actual 16-kHz/48-kHz capture paths.
Route classes
Report separately:
provider-native realtime
buffered/final-only Live Mic
file direct upload
YouTube direct upload
meeting finalization
local ONNX
Do not mix these routes into one average.
Network profiles
Where reproducible shaping is available:
fast upload / low RTT
normal broadband
constrained upload
higher RTT / packet-loss tail
IPv4-only
IPv6-only
dual-stack warm reuse
Repetitions
For every production candidate:
at least one cold call;
at least 20 warm calls for provider measurements where cost permits;
interleaved/randomized candidate order;
connection-created and connection-reused samples reported separately;
failures/cancellations excluded from successful latency percentiles and reported separately;
exact same model, endpoint, language, vocabulary, post-processing, and injection target.
hotkey -> mic ready
hotkey -> first audible frame
capture stop/drain
time to first encoded byte
encoder tail/finalization
output bytes and compression ratio
request start -> first byte
upload duration
provider processing
provider final -> target validation
provider final -> final text observed
CPU time during capture and after Stop
peak memory
thread/process count
installer size and startup impact
Quality and correctness:
exact transcript fixtures where deterministic;
WER/CER and punctuation comparison where nondeterministic;
no truncation, clipped tail, leading delay, invalid duration, sample-rate mismatch, or channel error;
WAV/FLAC exact PCM round-trip;
MP3/Opus decode through independent decoders;
exact final text observed in the intended guarded target.
in-process MP3 may remove approximately 100–300 ms from common dictations;
encoding during capture may reduce local Stop-to-request work by 80–98%;
complete warm end-to-end gains may commonly be 10–20%;
Opus-compatible routes on constrained uploads may gain 20–35% or several hundred milliseconds;
direct WAV may win short recordings and lose longer recordings on constrained uploads;
Nightly FLAC may make compression cost nearly negligible, but payload size and provider behavior still determine the global winner.
The selected policy must follow measured total latency, not these estimates.
Safety and replay rules
A local encoder or format fallback is allowed only before any billable STT request may have been accepted.
Do not automatically replay audio in another format after:
connect/read/total timeout with ambiguous request state;
cancellation after upload begins;
connection reset after bytes were sent;
provider 5xx or unknown response;
provider/model failure where acceptance is uncertain.
Do not race two paid providers or two billable formats to take the fastest response. Provider-native safe transport recovery before request commitment remains the HTTP stack's responsibility.
Parallelism must not weaken privacy:
no API keys, audio, transcript text, encoded payloads, or personal paths in logs;
bounded request/flow IDs only;
temporary artifacts under allowlisted roots;
deterministic cleanup on success, failure, cancellation, and shutdown.
Implementation ownership
Prefer extending the existing long-lived Rust scriber-audio-sidecar rather than spawning another per-dictation codec process.
Objective
Minimize Scriber's complete user-visible latency:
The primary KPI is no longer encoder time or even
stop -> requestStarted. It is:For toggle recording, also retain the critical secondary KPI:
The end marker must prove that the exact expected final transcript is visible in the guarded target application.
clipboard_set,paste_requested, an API callback, orfirst_pastealone is not sufficient.This issue should pursue the lowest total latency for each exact provider, route, model, workload, recording duration, and supported audio format. It should use the fastest verified Rust implementation for the winning format, including pinned Nightly/SIMD builds where they materially improve the complete path.
The detailed provider-format matrix in the following comment is normative input to this issue:
#18 (comment)
This work depends on the reusable HTTP transport from #16 and should include the OpenRouter STT route from #17 when available.
Product-level optimization rule
Do not choose one universal codec for every provider.
The production objective is:
Examples:
Unknown models, custom endpoints, and unresolved OpenRouter upstream providers must fail closed to a conservative capability set.
Primary and diagnostic timing contract
Extend the existing
HotPathTracerand installed text-target benchmark rather than creating an unrelated timer.Required start markers
Required stop/finalization markers
Required KPIs
Because the primary measurement includes the time during which the user is speaking, benchmark comparisons must use identical authoritative audio duration and content. Also report:
For real user sessions, store only privacy-safe aggregate phase timings; never infer or persist spoken content from timing data.
The existing markers such as
hotkey_received_to_mic_ready_ms,stop_requested_to_first_paste_ms,stop_requested_to_provider_final_received_ms, andprovider_final_received_to_clipboard_set_msremain useful diagnostics, but the release decision must use the final observed-text KPI.Critical-path model
Treat the workflow as a directed acyclic graph rather than a sequential list.
A simplified batch/final-only route should become:
For provider-native realtime routes, the network stream runs during capture and there is no batch audio-preparation stage after Stop. The relevant optimization is connection/session readiness, provider finalization, and final text insertion.
Parallel execution requirements
1. Idle and application-start prewarm
Perform non-sensitive, non-billable preparation before the first dictation:
A provider connection may be proactively warmed only through a documented idempotent/non-billable operation. Never issue a transcription POST merely to warm a connection.
2. Work started at activation
At
activation_received, launch independent work concurrently:None of these tasks may delay microphone capture unless it is required for correctness.
3. Encode during capture for buffered/final-only routes
Choose the format before the first audio frame from the frozen capability and benchmark policy. Tee PCM into exactly one production encoder while preserving the authoritative PCM path:
Requirements:
4. Parallel Stop boundary
On Stop, run safe independent tails concurrently:
Use a structured-concurrency owner so failure or cancellation joins every child task and cleans every handle/artifact.
5. Overlap finalization and upload where the format permits
The selected transport should begin consuming bytes as early as the provider contract safely permits:
Do not start a billable batch transcription POST before the user stops unless that exact provider documents a cancellable streaming-ingest contract and Scriber deliberately implements it. Provider-native realtime sessions are the exception because streaming is their normal contract.
6. Parallel response tail and injection preparation
While awaiting the provider final response, safely prepare:
Insert immediately after the final accepted text is ready. Measure actual target observation, not only injection API completion.
Fastest implementation candidates by format
The phrase “fastest Rust implementation” means the fastest implementation on Scriber's reference Windows systems and fixtures, not the fastest upstream marketing number.
WAV / PCM
Implement a zero-copy or near-zero-copy Rust virtual WAV stream:
This is the no-compression lower bound and likely winner for short recordings on fast uploads.
FLAC
Required candidates:
The current upstream
flacencreports approximately 575x realtime on stable and 1,309x realtime for its Nightly default benchmark, with a compression ratio close to the reference FLAC level-5 result. These figures use a small music corpus and are hypotheses only; Scriber must remeasure 16-kHz mono speech.Nightly
flacencis the preferred maximum-speed FLAC candidate if it wins the installed Windows benchmark and passes toolchain, packaging, and reproducibility gates.rezin-flacis a long-form parallel challenger; do not assume many workers help short dictation.Opus
Required candidates for exact provider/model routes that explicitly accept Opus:
ruopusreports encode performance at approximately parity with SIMDlibopusat matched complexity, between roughly 560x and 1,088x realtime in its published modes. Benchmark speech-oriented bitrate/complexity settings and full Ogg/WebM container finalization.Prefer
ruopuswhen it is the fastest qualifying implementation and passes conformance, accuracy, packaging, cancellation, and long-session tests. Retainlibopusencas the reference and safe fallback.MP3
Required candidates:
Rust Nightly cannot by itself make the native LAME backend a Rust-native codec, but the in-process wrapper can remove FFmpeg process and pipe overhead. The Shine family reports substantially faster encoding than LAME in its own benchmark, but its simple psychoacoustic model can produce lower-quality/larger MP3 output. It is eligible only if actual STT WER/CER and exact-fixture results remain within the quality guardrail.
The production MP3 implementation is whichever candidate minimizes complete latency while passing quality and licensing gates—not automatically the pure-Rust option.
Original-file passthrough
For file, YouTube, and meeting artifacts, unchanged passthrough is the first candidate whenever the exact provider/model accepts the original format. No encoder can outperform zero re-encoding.
Nightly, SIMD, PGO, and build strategy
Maximum-speed experiments must include a pinned and reproducible Nightly toolchain without making the entire application accidentally depend on unpinned Nightly behavior.
Required build variants:
Requirements:
codegen-units=1, panic strategy, and PGO on Scriber speech fixtures;target-cpu=nativebinary that fails on other users' CPUs;Nightly is selected for production only when the complete installed-app path is materially faster and the build is reproducible, supportable, and secure.
Provider-aware selection policy
Use the verified matrix from the linked issue comment and freeze the result per job.
Suggested policy order:
Possible generated choices:
Do not use guessed live bandwidth as a routing signal. Network-specific routing may be added only if a bounded, private, stable signal demonstrably improves p95 without oscillation. Initially derive conservative duration/size crossovers from controlled network profiles.
Benchmark design
Workloads
Run each eligible provider/format/implementation against identical authoritative fixtures:
Include German and English clean speech, moderately noisy speech, silence/near-silence lifecycle cases, and the actual 16-kHz/48-kHz capture paths.
Route classes
Report separately:
Do not mix these routes into one average.
Network profiles
Where reproducible shaping is available:
Repetitions
For every production candidate:
Metrics
Primary:
Supporting:
Quality and correctness:
Baselines and hypotheses
Preserve the current production baseline:
Initial hypotheses to test, not promises:
The selected policy must follow measured total latency, not these estimates.
Safety and replay rules
A local encoder or format fallback is allowed only before any billable STT request may have been accepted.
Do not automatically replay audio in another format after:
Do not race two paid providers or two billable formats to take the fastest response. Provider-native safe transport recovery before request commitment remains the HTTP stack's responsibility.
Parallelism must not weaken privacy:
Implementation ownership
Prefer extending the existing long-lived Rust
scriber-audio-sidecarrather than spawning another per-dictation codec process.Suggested modules:
Suggested Rust boundary:
Suggested backend feature gates:
No broad default feature may silently add every codec to the installer.
Delivery plan
Stage 0 — end-to-end instrumentation
final_text_observedwith exact target verification;Stage 1 — capability registry
gemini_sttcapability coverage.Stage 2 — maximum-speed local codec lab
flacencand stable control;ruopusandlibopusenccontrol;Stage 3 — parallel capture pipeline
Stage 4 — provider end-to-end matrix
Stage 5 — production rollout
Acceptance criteria
Primary outcome
Parallel pipeline
Nightly and codec selection
Reliability and safety
Non-goals
References