Skip to content

Spike: evaluate S1-mini as an on-device English transcript normalizer #35

Description

@Pivii

Context

Superwhisper released S1-mini, a 596M-parameter Qwen3 model trained specifically to clean raw ASR transcripts on-device. It is not a speech recognizer. It would run after the existing Whisper or Parakeet provider and replace or extend the current deterministic TextPostProcessor step in DictationService.

This is directly relevant to the Smart Mode Pro roadmap in #6, but it needs a measured spike before any production integration.

Related iOS spike: getdictus/dictus-ios#268

What is verified

Primary sources:

Initial desktop sanity check

The published Q4_K_M file was tested locally through a freshly built llama.cpp using the exact documented prompt.

This was a Linux x86_64 functional check, not an Android benchmark.

  • 3/3 English fixtures were acceptable, including filler removal, self-correction, punctuation, product names, and a spoken date.
  • French cannot be considered supported:
    • It failed to resolve vendredi non plutôt jeudi to Thursday.
    • It changed support arobase getdictus point com into the wrong address, support@arobase.getdictus.com.
    • A simpler French dictation happened to produce acceptable text, which makes silent language fallback especially risky.
  • Warm-server RSS was about 997 MiB on the host with a 2,048-token context. The GGUF file size is therefore not a sufficient proxy for runtime memory.
  • Mean request time for three English fixtures was 0.34 s on the host, around 92.6 generated tokens/s. Do not extrapolate these numbers to Android.

Scope: decision spike, not production integration

Answer: Can S1-mini improve English Dictus transcripts on representative Android devices without causing unacceptable latency, memory pressure, thermal load, or battery drain?

Runtime prototype

  • Build a throwaway llama.cpp Android/JNI runner for arm64-v8a.
  • Pin a model revision and verify the downloaded file hash/size.
  • Apply the documented prompt exactly, with greedy decoding and thinking disabled.
  • Keep the experiment behind a developer-only switch. Do not add it to the user-facing model catalogue yet.

Quality evaluation

  • Capture or curate real raw outputs from both Whisper and Parakeet.
  • Score at least these categories: filler removal, punctuation, false starts, self-corrections, numbers/dates/currency, email addresses, proper nouns, empty/noise input, and hallucinated content.
  • Compare S1-mini output against the current TextPostProcessor baseline.
  • Include French rejection fixtures and prove the feature cannot silently run when the transcript language is not English.
  • Treat semantic changes to names, addresses, dates, amounts, or the speaker's final correction as hard failures.

Device measurements

Measure on at least one constrained device and one modern device, with the real STT pipeline:

  • Model load time and first-token latency.
  • Total post-processing latency for realistic short and long dictations.
  • Peak RSS/PSS while the active Whisper or Parakeet model is still resident.
  • Peak RSS/PSS and total latency if the STT provider is released before loading S1-mini.
  • Thermal and battery behavior over a burst of repeated dictations.
  • Behavior while DictationService is running as the foreground service.
  • Recovery after cancellation, timeout, low-memory pressure, and model-load failure.

The existing strict single-STT-provider policy explicitly releases one engine before loading another to prevent OOM. This spike must make the same cohabitation decision for STT plus LLM rather than assuming both can stay resident.

Distribution and product decision

  • Review the custom naming clause and define where the exact attribution would appear.
  • Decide whether the 484 MB model is optional-download only and how storage is communicated.
  • End with one explicit recommendation: ship English-only, gate by device tier, wait for a multilingual release, or reject.

Non-goals

  • Shipping Smart Mode UI.
  • Replacing Whisper or Parakeet.
  • Claiming French support from incidental plausible outputs.
  • Bundling the model in the APK.

Acceptance criteria

  • A reproducible benchmark report includes quality fixtures, device models, Android versions, runtime revision, model revision, memory, latency, and thermal observations.
  • The report compares keeping STT resident against unloading it before LLM inference.
  • No production dependency or user-facing download is added unless the report recommends shipping.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions