Feature request: Audio Time Compression for STT #271
Replies: 3 comments
|
That's a really great feature idea! I already added a silence cutter, which will reduce the audio length and API cost by cutting all the silences, but your idea makes even more sense. I will add an issue for that asap and try to fit it in the next update. Thanks a lot! :) |
|
This is in now — thanks again for the idea, @Ochauke202608. Settings → Dictate → Recording → Speed up audio, a slider from 1.0× to 2.0×, off by default. Pitch-preserving (WSOLA, not a resample), it applies after the existing pause trimmer, and it reaches every provider at once because it happens where the upload file is decided — including the individual segments of a long-form dictation. Since the whole point is that it should not cost accuracy, it was measured before shipping rather than assumed: 13 clips, 5 rates, 3 systems (Groq whisper-large-v3-turbo plus both on-device Parakeet models), 195 transcriptions. Up to 1.5× nothing measurable is lost; the first system degrades at 1.75×, and at 2× every one of them does — German harder than English. So 2× stays available, but from 1.75× on the setting says what it costs. Full tables and the reasoning are in #272. Your point about the ratio being model-dependent held up, by the way: at 2× Groq's English stayed flat at 4.5 % word error rate while its German went from 5.8 % to 32.6 %. Same rate, same encoder, very different outcome — which is exactly why this is a slider and not a fixed setting. |
|
@DevEmperor Thank you so much for taking the idea seriously and actually measuring the accuracy. The result that there was no measurable accuracy loss at 1.5× is especially exciting for me. I expected that range to be a particularly good balance between cost savings and recognition accuracy. The difference between English and German at 2× is also fascinating. I expected that the optimal speed might depend on the model and language, so it's really interesting to see that reflected in the actual results. I'm a Japanese speaker, so I'm also very interested to see how Japanese will behave. I'm really looking forward to seeing this released on Google Play. This is exactly the kind of feature I was hoping for. Thanks again! |
Uh oh!
There was an error while loading. Please reload this page.
I would like to see an option in Dictate Keyboard to compress the duration of audio before sending it to an STT API.
Many modern STT APIs charge based on audio duration, so compressing the audio duration before upload could potentially reduce API costs without significantly affecting transcription accuracy.
For example, with a duration-based STT service such as GPT Transcribe, compressing the audio to 1.5x its original speed would theoretically reduce the billable audio duration to about two-thirds while preserving the same spoken content.
Ideally, the feature would support:
I think this could be a useful general-purpose feature for any duration-based STT API, rather than something specific to a single provider.
All reactions