Skip to content

Minimize hotkey-to-final-text latency with provider-aware Rust codecs and parallel execution #18

Description

@MyButtermilk

Objective

Minimize Scriber's complete user-visible latency:

hotkey/button activation
    -> microphone ready
    -> authoritative audio captured
    -> provider request/stream completed
    -> final transcript returned
    -> final transcript visibly inserted into the intended target

The primary KPI is no longer encoder time or even stop -> requestStarted. It is:

activation_received_to_final_text_observed_ms

For toggle recording, also retain the critical secondary KPI:

stop_requested_to_final_text_observed_ms

The end marker must prove that the exact expected final transcript is visible in the guarded target application. clipboard_set, paste_requested, an API callback, or first_paste alone is not sufficient.

This issue should pursue the lowest total latency for each exact provider, route, model, workload, recording duration, and supported audio format. It should use the fastest verified Rust implementation for the winning format, including pinned Nightly/SIMD builds where they materially improve the complete path.

The detailed provider-format matrix in the following comment is normative input to this issue:

#18 (comment)

This work depends on the reusable HTTP transport from #16 and should include the OpenRouter STT route from #17 when available.

Product-level optimization rule

Do not choose one universal codec for every provider.

The production objective is:

For every frozen provider + endpoint + model + workload:
    choose the fastest verified representation and implementation
    that preserves transcript quality, reliability, privacy, and compatibility.

Examples:

Azure MAI / mai-transcribe-1.5
    -> benchmark zero-copy WAV, Nightly Rust FLAC, Rust MP3, current FFmpeg MP3
    -> no Opus unless the exact endpoint contract changes

Smallest batch / exact Pulse route
    -> benchmark passthrough, WAV, MP3, FLAC, Ogg/Opus, WebM/Opus

Deepgram async / exact pre-recorded model
    -> benchmark passthrough and documented Opus/MP3/FLAC/WAV choices

OpenAI Transcriptions
    -> benchmark accepted MP3/WAV/WebM paths; do not select unsupported Opus/FLAC

Provider-native realtime routes
    -> use their exact streaming encoding contract instead of forcing a batch file codec

Unknown models, custom endpoints, and unresolved OpenRouter upstream providers must fail closed to a conservative capability set.

Primary and diagnostic timing contract

Extend the existing HotPathTracer and installed text-target benchmark rather than creating an unrelated timer.

Required start markers

activation_received
hotkey_received                 # native Tauri QPC marker for global hotkey
button_received                 # bound UI/native marker for button activation
start_request_dispatched
recording_state_visible
mic_ready
first_audio_frame
first_audible_audio_frame

Required stop/finalization markers

stop_requested
last_audio_frame_captured
last_chunk_sent_to_pipeline
capture_stopped
encoder_tail_started
encoder_tail_completed
request_started
first_request_chunk_sent
last_request_chunk_sent
provider_final_received
transcript_parsed
post_processing_completed       # only when enabled
injection_target_validated
clipboard_set
paste_requested
final_text_observed

Required KPIs

activation_received_to_final_text_observed_ms
hotkey_received_to_final_text_observed_ms
button_received_to_final_text_observed_ms
stop_requested_to_final_text_observed_ms
stop_requested_to_provider_final_received_ms
provider_final_received_to_final_text_observed_ms

Because the primary measurement includes the time during which the user is speaking, benchmark comparisons must use identical authoritative audio duration and content. Also report:

non_speech_overhead_ms = activation_to_final_text - authoritative_fixture_duration

For real user sessions, store only privacy-safe aggregate phase timings; never infer or persist spoken content from timing data.

The existing markers such as hotkey_received_to_mic_ready_ms, stop_requested_to_first_paste_ms, stop_requested_to_provider_final_received_ms, and provider_final_received_to_clipboard_set_ms remain useful diagnostics, but the release decision must use the final observed-text KPI.

Critical-path model

Treat the workflow as a directed acyclic graph rather than a sequential list.

A simplified batch/final-only route should become:

activation received
├─ start/lease microphone immediately
├─ freeze provider + endpoint + model + capability entry
├─ acquire/reuse provider HTTP connection
├─ initialize selected encoder/resampler
├─ capture injection-target guard
└─ update overlay/UI
       ↓
first PCM frame
├─ authoritative PCM spool
├─ VAD/speech diagnostics
└─ selected provider-compatible encoder running during capture
       ↓
stop requested
├─ stop/drain physical capture
├─ finalize only the small encoder tail
├─ clean audio routing / resume prewarm
├─ validate target and prewarm injection path
└─ prepare multipart headers/request metadata
       ↓
encoded artifact or stream ready
└─ stream multipart upload over the already warm connection
       ↓
provider final response
├─ parse/validate final text
├─ finish optional post-processing
└─ revalidate target in parallel where safe
       ↓
insert text
       ↓
observe exact final text in target

For provider-native realtime routes, the network stream runs during capture and there is no batch audio-preparation stage after Stop. The relevant optimization is connection/session readiness, provider finalization, and final text insertion.

Parallel execution requirements

1. Idle and application-start prewarm

Perform non-sensitive, non-billable preparation before the first dictation:

  • keep the existing Rust audio sidecar alive;
  • prewarm microphone/VAD/SmartTurn resources where already supported;
  • create the shared provider HTTP sessions and connection pools from Reuse provider HTTP sessions and add connection-level STT diagnostics #16;
  • resolve and cache provider capabilities and selected model metadata;
  • preallocate bounded PCM and encoder buffers;
  • initialize reusable resampler plans and codec tables where thread-safe;
  • warm the native text-injection path without importing legacy fallbacks unnecessarily;
  • compile/runtime-dispatch SIMD kernels once rather than on the Stop path.

A provider connection may be proactively warmed only through a documented idempotent/non-billable operation. Never issue a transcription POST merely to warm a connection.

2. Work started at activation

At activation_received, launch independent work concurrently:

  • microphone acquisition and first-frame delivery;
  • provider route/model freeze;
  • HTTP pool/connection acquisition;
  • selected encoder initialization;
  • resampler initialization if required;
  • injection-target identity capture;
  • overlay/UI update.

None of these tasks may delay microphone capture unless it is required for correctness.

3. Encode during capture for buffered/final-only routes

Choose the format before the first audio frame from the frozen capability and benchmark policy. Tee PCM into exactly one production encoder while preserving the authoritative PCM path:

capture callback
    -> authoritative PCM/frame-pipe
    -> non-blocking bounded SPSC/ring buffer
    -> dedicated Rust encoder worker
    -> encoded spool/pages/frames

Requirements:

  • capture callback never waits on the encoder;
  • no authoritative PCM is dropped;
  • encoder overload fails the candidate locally and preserves a fallback before any provider POST;
  • use a dedicated CPU pool/thread, not the async/network event loop;
  • avoid oversubscribing the CPU with VAD, resampling, and many codec workers;
  • production encodes one selected format, not every candidate simultaneously;
  • benchmark variants run in separate/interleaved trials so redundant codecs do not distort results.

4. Parallel Stop boundary

On Stop, run safe independent tails concurrently:

  • capture stop/drain;
  • final encoder flush/container close;
  • audio-routing cleanup and microphone-prewarm handoff;
  • injection-target revalidation;
  • request metadata/header construction;
  • UI transition to Transcribing.

Use a structured-concurrency owner so failure or cancellation joins every child task and cleans every handle/artifact.

5. Overlap finalization and upload where the format permits

The selected transport should begin consuming bytes as early as the provider contract safely permits:

  • WAV: once total PCM bytes are known at Stop, emit a virtual RIFF header followed directly by the authoritative PCM spool; avoid a complete copied WAV file;
  • MP3: stream already-produced frames and append the final flush frames; do not require nonessential ID3/Xing metadata for STT;
  • Ogg/Opus or WebM/Opus: stream completed pages/clusters and append the terminal page/cluster;
  • FLAC: benchmark a streamable/final-header strategy against a seekable encoded spool. Do not delay upload for checksums/metadata that the provider does not require.

Do not start a billable batch transcription POST before the user stops unless that exact provider documents a cancellable streaming-ingest contract and Scriber deliberately implements it. Provider-native realtime sessions are the exception because streaming is their normal contract.

6. Parallel response tail and injection preparation

While awaiting the provider final response, safely prepare:

  • guarded target revalidation;
  • clipboard snapshot strategy;
  • native SendInput/clipboard route readiness;
  • post-processing connection reuse when post-processing is enabled.

Insert immediately after the final accepted text is ready. Measure actual target observation, not only injection API completion.

Fastest implementation candidates by format

The phrase “fastest Rust implementation” means the fastest implementation on Scriber's reference Windows systems and fixtures, not the fastest upstream marketing number.

WAV / PCM

Implement a zero-copy or near-zero-copy Rust virtual WAV stream:

44-byte RIFF header + existing PCM spool

This is the no-compression lower bound and likely winner for short recordings on fast uploads.

FLAC

Required candidates:

flacenc stable
flacenc pinned-nightly SIMD/default
rezin-flac parallel long-form
FFmpeg FLAC fast control

The current upstream flacenc reports approximately 575x realtime on stable and 1,309x realtime for its Nightly default benchmark, with a compression ratio close to the reference FLAC level-5 result. These figures use a small music corpus and are hypotheses only; Scriber must remeasure 16-kHz mono speech.

Nightly flacenc is the preferred maximum-speed FLAC candidate if it wins the installed Windows benchmark and passes toolchain, packaging, and reproducibility gates. rezin-flac is a long-form parallel challenger; do not assume many workers help short dictation.

Opus

Required candidates for exact provider/model routes that explicitly accept Opus:

ruopus pure Rust SIMD
libopusenc reference control
optional crime/moosicbox façade controls

ruopus reports encode performance at approximately parity with SIMD libopus at matched complexity, between roughly 560x and 1,088x realtime in its published modes. Benchmark speech-oriented bitrate/complexity settings and full Ogg/WebM container finalization.

Prefer ruopus when it is the fastest qualifying implementation and passes conformance, accuracy, packaging, cancellation, and long-session tests. Retain libopusenc as the reference and safe fallback.

MP3

Required candidates:

mp3lame-encoder in-process control
shine-rs / Shine-derived maximum-speed challenger
current FFmpeg MP3 control

Rust Nightly cannot by itself make the native LAME backend a Rust-native codec, but the in-process wrapper can remove FFmpeg process and pipe overhead. The Shine family reports substantially faster encoding than LAME in its own benchmark, but its simple psychoacoustic model can produce lower-quality/larger MP3 output. It is eligible only if actual STT WER/CER and exact-fixture results remain within the quality guardrail.

The production MP3 implementation is whichever candidate minimizes complete latency while passing quality and licensing gates—not automatically the pure-Rust option.

Original-file passthrough

For file, YouTube, and meeting artifacts, unchanged passthrough is the first candidate whenever the exact provider/model accepts the original format. No encoder can outperform zero re-encoding.

Nightly, SIMD, PGO, and build strategy

Maximum-speed experiments must include a pinned and reproducible Nightly toolchain without making the entire application accidentally depend on unpinned Nightly behavior.

Required build variants:

stable release baseline
pinned Nightly release
pinned Nightly + crate SIMD feature
pinned Nightly + PGO
pinned Nightly + LTO/codegen tuning

Requirements:

  • pin the exact Nightly date/version in repository configuration;
  • record compiler, crate, feature, and target fingerprints in benchmark artifacts;
  • use release builds only;
  • evaluate fat/Thin LTO, codegen-units=1, panic strategy, and PGO on Scriber speech fixtures;
  • do not ship one target-cpu=native binary that fails on other users' CPUs;
  • prefer runtime SIMD dispatch or separately compiled baseline/AVX2 kernels;
  • validate SSE2/baseline, AVX2, and any optional higher feature path on real supported machines;
  • include the selected codec/toolchain capabilities in sidecar self-test and runtime attestation;
  • retain a stable implementation fallback when the Nightly artifact cannot run or fails validation.

Nightly is selected for production only when the complete installed-app path is materially faster and the build is reproducible, supportable, and secure.

Provider-aware selection policy

Use the verified matrix from the linked issue comment and freeze the result per job.

Suggested policy order:

1. Native realtime route:
       use the exact provider streaming contract;
       avoid batch re-encoding.

2. Direct file/YouTube/meeting route:
       pass through the original accepted format unchanged.

3. Buffered/final-only Live Mic route:
       choose the measured fastest accepted format for the provider/model,
       duration bucket, machine capability, and validated network profile.

4. If the chosen encoder is unavailable or fails locally:
       select a pre-approved local fallback before sending any provider request.

5. If no verified representation exists:
       fail locally with an explicit unsupported-format error.

Possible generated choices:

wav_pcm16_virtual
mp3_shine_nightly
mp3_lame_in_process
flac_flacenc_nightly_simd
flac_rezin_parallel
ogg_opus_ruopus
ogg_opus_libopusenc
webm_opus_verified_backend
current_ffmpeg_mp3_fallback

Do not use guessed live bandwidth as a routing signal. Network-specific routing may be added only if a bounded, private, stable signal demonstrably improves p95 without oscillation. Initially derive conservative duration/size crossovers from controlled network profiles.

Benchmark design

Workloads

Run each eligible provider/format/implementation against identical authoritative fixtures:

5 s
15 s
30 s
60 s
180 s
600 s where practical

Include German and English clean speech, moderately noisy speech, silence/near-silence lifecycle cases, and the actual 16-kHz/48-kHz capture paths.

Route classes

Report separately:

provider-native realtime
buffered/final-only Live Mic
file direct upload
YouTube direct upload
meeting finalization
local ONNX

Do not mix these routes into one average.

Network profiles

Where reproducible shaping is available:

fast upload / low RTT
normal broadband
constrained upload
higher RTT / packet-loss tail
IPv4-only
IPv6-only
dual-stack warm reuse

Repetitions

For every production candidate:

  • at least one cold call;
  • at least 20 warm calls for provider measurements where cost permits;
  • interleaved/randomized candidate order;
  • connection-created and connection-reused samples reported separately;
  • failures/cancellations excluded from successful latency percentiles and reported separately;
  • exact same model, endpoint, language, vocabulary, post-processing, and injection target.

Metrics

Primary:

activation_received_to_final_text_observed_ms
stop_requested_to_final_text_observed_ms
non_speech_overhead_ms

Supporting:

hotkey -> mic ready
hotkey -> first audible frame
capture stop/drain
time to first encoded byte
encoder tail/finalization
output bytes and compression ratio
request start -> first byte
upload duration
provider processing
provider final -> target validation
provider final -> final text observed
CPU time during capture and after Stop
peak memory
thread/process count
installer size and startup impact

Quality and correctness:

  • exact transcript fixtures where deterministic;
  • WER/CER and punctuation comparison where nondeterministic;
  • no truncation, clipped tail, leading delay, invalid duration, sample-rate mismatch, or channel error;
  • WAV/FLAC exact PCM round-trip;
  • MP3/Opus decode through independent decoders;
  • exact final text observed in the intended guarded target.

Baselines and hypotheses

Preserve the current production baseline:

PCM spool
    -> post-stop FFmpeg MP3 64 kbit/s
    -> new/current HTTP path
    -> provider final
    -> injection

Initial hypotheses to test, not promises:

  • in-process MP3 may remove approximately 100–300 ms from common dictations;
  • encoding during capture may reduce local Stop-to-request work by 80–98%;
  • complete warm end-to-end gains may commonly be 10–20%;
  • Opus-compatible routes on constrained uploads may gain 20–35% or several hundred milliseconds;
  • direct WAV may win short recordings and lose longer recordings on constrained uploads;
  • Nightly FLAC may make compression cost nearly negligible, but payload size and provider behavior still determine the global winner.

The selected policy must follow measured total latency, not these estimates.

Safety and replay rules

A local encoder or format fallback is allowed only before any billable STT request may have been accepted.

Do not automatically replay audio in another format after:

  • connect/read/total timeout with ambiguous request state;
  • cancellation after upload begins;
  • connection reset after bytes were sent;
  • provider 5xx or unknown response;
  • provider/model failure where acceptance is uncertain.

Do not race two paid providers or two billable formats to take the fastest response. Provider-native safe transport recovery before request commitment remains the HTTP stack's responsibility.

Parallelism must not weaken privacy:

  • no API keys, audio, transcript text, encoded payloads, or personal paths in logs;
  • bounded request/flow IDs only;
  • temporary artifacts under allowlisted roots;
  • deterministic cleanup on success, failure, cancellation, and shutdown.

Implementation ownership

Prefer extending the existing long-lived Rust scriber-audio-sidecar rather than spawning another per-dictation codec process.

Suggested modules:

Frontend/src-tauri/src/audio_codec.rs
Frontend/src-tauri/src/audio_prepare.rs
Frontend/src-tauri/src/provider_audio_capabilities.rs
Frontend/src-tauri/src/latency_markers.rs
src/runtime/provider_http.py              # #16
src/core/provider_audio_formats.py

Suggested Rust boundary:

#[derive(Clone, Copy, Debug, Eq, PartialEq)]
enum EncodedAudioFormat {
    WavPcm16,
    Mp3,
    Flac,
    OggOpus,
    WebmOpus,
}

trait StreamingAudioEncoder: Send {
    fn format(&self) -> EncodedAudioFormat;
    fn write_pcm_i16(&mut self, interleaved: &[i16]) -> Result<(), AudioEncodeError>;
    fn finish(self: Box<Self>) -> Result<AudioPreparationArtifact, AudioEncodeError>;
}

Suggested backend feature gates:

codec-mp3-lame
codec-mp3-shine-nightly
codec-opus-ruopus
codec-opus-libopusenc
codec-flac-flacenc
codec-flac-flacenc-nightly
codec-flac-rezin-parallel

No broad default feature may silently add every codec to the installer.

Delivery plan

Stage 0 — end-to-end instrumentation

  • extend the native hotkey/button marker contract;
  • add final_text_observed with exact target verification;
  • preserve current hot-path metrics;
  • establish cold/warm current-route baselines.

Stage 1 — capability registry

  • encode the documented provider/endpoint/model matrix from the linked comment;
  • separate batch formats from realtime encodings;
  • add evidence and verification dates;
  • fail closed for unknown routes;
  • add missing gemini_stt capability coverage.

Stage 2 — maximum-speed local codec lab

  • implement zero-copy WAV;
  • implement Nightly flacenc and stable control;
  • implement ruopus and libopusenc control;
  • implement in-process LAME and Shine challenger;
  • evaluate PGO/LTO/SIMD and long-form parallel FLAC;
  • publish local latency, quality, resource, and packaging results.

Stage 3 — parallel capture pipeline

  • initialize selected codec at activation;
  • tee PCM through a non-blocking bounded Rust queue during capture;
  • overlap Stop cleanup, encoder tail, target validation, and request preparation;
  • stream multipart output as early as safely possible.

Stage 4 — provider end-to-end matrix

Stage 5 — production rollout

  • enable only the winning implementation per provider/route/model;
  • retain safe stable/local fallbacks;
  • remove or isolate benchmark-only variants;
  • complete installed-app, cancellation, failure-tail, packaging, signing, privacy, and support-bundle gates.

Acceptance criteria

Primary outcome

  • The primary benchmark begins at the authoritative hotkey/button activation marker.
  • It ends only when the exact final transcript is observed in the intended target.
  • Identical audio fixtures make full activation-to-final comparisons valid.
  • Results report p50, p90, p95, maximum, variance, and failure rate.
  • Production chooses the lowest-latency verified route per provider + endpoint + model + workload.

Parallel pipeline

  • Provider/session preparation, encoder initialization, target capture, and mic startup overlap where safe.
  • Buffered routes encode during capture, not entirely after Stop.
  • Capture never blocks on compression and authoritative PCM is never lost.
  • Stop cleanup, encoder tail, request preparation, and target validation overlap where safe.
  • Multipart upload consumes already-produced bytes without unnecessary whole-file copies.
  • Every concurrent child task/thread is owned, cancelled, joined, and cleaned deterministically.

Nightly and codec selection

  • Stable and pinned-Nightly implementations are directly compared.
  • The fastest qualifying implementation is selected per format.
  • Nightly/SIMD/PGO configuration is reproducible and attested.
  • Runtime CPU dispatch prevents incompatible binaries.
  • MP3, Opus, FLAC, WAV, and passthrough candidates are limited to exact supported provider/model routes.
  • Lossy candidates pass STT quality gates; lossless candidates round-trip exactly.

Reliability and safety

  • No new capture stalls, dropped frames, malformed files, deadlocks, orphaned processes/threads, or unclosed transports.
  • No second billable request is caused by codec fallback after an ambiguous failure.
  • No sensitive audio, transcript, credential, filename, or path data is logged.
  • Existing provider, replay, installed workflow, release, and packaging tests continue to pass.

Non-goals

  • A universal codec for all providers.
  • Selecting Nightly because it is newer rather than because total latency is lower.
  • Optimizing only encoder microseconds while worsening upload/provider/injection time.
  • Encoding several production formats simultaneously and racing them.
  • Starting a billable batch request before Stop without an explicit provider streaming contract.
  • Guessing Opus support from generic OGG/WebM support.
  • Choosing thresholds from intuition rather than measured crossover evidence.
  • Replacing provider-native realtime transports with batch-file uploads.

References

  1. Minimize hotkey-to-final-text latency with provider-aware Rust codecs and parallel execution #18 (comment)
  2. Reuse provider HTTP sessions and add connection-level STT diagnostics #16
  3. Add OpenRouter as a native speech-to-text provider #17
  4. https://github.com/MyButtermilk/Scriber/blob/main/src/core/hot_path_tracer.py
  5. https://github.com/MyButtermilk/Scriber/blob/main/scripts/measure_recording_hot_path_baseline.py
  6. https://github.com/MyButtermilk/Scriber/blob/main/src/azure_mai_stt.py
  7. https://github.com/MyButtermilk/Scriber/blob/main/Frontend/src-tauri/src/audio_sidecar.rs
  8. https://github.com/yotarok/flacenc-rs/blob/main/report/report.stable.md
  9. https://github.com/yotarok/flacenc-rs/blob/main/report/report.nightly.md
  10. https://docs.rs/crate/ruopus/latest
  11. https://docs.rs/libopusenc/latest/libopusenc/
  12. https://docs.rs/mp3lame-encoder/latest/mp3lame_encoder/
  13. https://github.com/toots/shine
  14. https://docs.rs/crate/rezin-flac/latest
  15. https://docs.rs/rayon/latest/rayon/
  16. https://doc.rust-lang.org/rustc/profile-guided-optimization.html
  17. https://doc.rust-lang.org/rustc/codegen-options/index.html
  18. Speed up OpenRouter dictation by 22% without forcing IPv4 DevEmperor/DictateKeyboard#210

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions