Repository navigation
Transcribe while recording, and run a hook when a recording starts - #25
Merged
Merged
Conversation
`live_transcript` (off by default, a toggle under Recordings in Settings) makes a session write `live.jsonl` as it records: every buffer the two track files get is also resampled and accumulated, and a timer that exists only while a session is live cuts a chunk when the speaker pauses or every ten seconds, transcribes it on the daemon's already-loaded model, and appends one JSON line. The post-stop transcript is untouched and remains the record. `on_start` is the counterpart of `on_stop`: a shell command given the session folder the moment both tracks are recording. Both hooks share one launcher now and both are in Settings. Cost, measured on the default model (M-series, p50 of 7): a 5 s chunk transcribes in 43 ms, a 20 s chunk in 65 ms, so a two-track call spends about 1% of the time on the ANE while live is on. Digital silence is skipped before it reaches the model. Memory is unchanged: same model, same actor as dictation. Claude-Session: https://claude.ai/code/session_01TG4DEV5UyCPRFegrkNK3oE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
live_transcript(off by default; toggle under Recordings in Settings): a session also writeslive.jsonlas it records. Every buffer the track files get is resampled to 16 kHz mono and accumulated per track; a timer that only exists while a session is live cuts a chunk when the speaker pauses (0.6 s under a quiet threshold) or, failing that, at the quietest 300 ms in the last 3 s once 10 s have built up. Each chunk goes through the daemon's already-loaded model and lands as one JSON line (speaker,start_ms,end_ms,text). The post-stop transcript is unchanged and remains the record.on_start: counterpart ofon_stop, given the session folder the moment both tracks are recording. Both hooks share one launcher (Hook.swift) and both have a field in Settings.Rules this keeps
Transcriber, same as the post-stop pass. No streaming checkpoint, no second resident model.Measured
yap bench --audioonparakeet-tdt-ctc-110m, 7 timed runs:A chunk is 3 to 10 s of one speaker, so two tracks cost roughly 100 ms of ANE per 10 s of call, about 1%. Digital silence (the far side while you talk) is skipped before the model. A dictation press during a live session queues behind at most one chunk.
Tests
LiveChunkerTests: the cut rules on synthetic audio (wait under 3 s, pause takes all, 10 s cuts after the quietest lull, final takes the tail, silence detection).LiveTranscriptTests: the whole path on real speech fromsay, through the recorder sink, the timer, the model and the file. Asserts a chunk is on disk while audio is still arriving and the final flush catches the tail. Skips when the model is not downloaded, so CI builds without it.https://claude.ai/code/session_01TG4DEV5UyCPRFegrkNK3oE