Skip to content

Transcribe while recording, and run a hook when a recording starts - #25

Merged
TerrifiedBug merged 1 commit into
mainfrom
feat/live-transcript
Oct 5, 2026
Merged

TerrifiedBug merged 1 commit into
mainfrom
feat/live-transcript

Conversation

@TerrifiedBug

Copy link
Copy Markdown
Owner

What

  • live_transcript (off by default; toggle under Recordings in Settings): a session also writes live.jsonl as it records. Every buffer the track files get is resampled to 16 kHz mono and accumulated per track; a timer that only exists while a session is live cuts a chunk when the speaker pauses (0.6 s under a quiet threshold) or, failing that, at the quietest 300 ms in the last 3 s once 10 s have built up. Each chunk goes through the daemon's already-loaded model and lands as one JSON line (speaker, start_ms, end_ms, text). The post-stop transcript is unchanged and remains the record.
  • on_start: counterpart of on_stop, given the session folder the moment both tracks are recording. Both hooks share one launcher (Hook.swift) and both have a field in Settings.

Rules this keeps

  • One loaded model: the live path borrows the daemon's Transcriber, same as the post-stop pass. No streaming checkpoint, no second resident model.
  • Nothing runs while idle: the timer is created at session start and invalidated at stop.
  • No new dependencies, no new processes.

Measured

yap bench --audio on parakeet-tdt-ctc-110m, 7 timed runs:

clip p50 max
5 s 43 ms 48 ms
20 s 65 ms 71 ms

A chunk is 3 to 10 s of one speaker, so two tracks cost roughly 100 ms of ANE per 10 s of call, about 1%. Digital silence (the far side while you talk) is skipped before the model. A dictation press during a live session queues behind at most one chunk.

Tests

  • LiveChunkerTests: the cut rules on synthetic audio (wait under 3 s, pause takes all, 10 s cuts after the quietest lull, final takes the tail, silence detection).
  • LiveTranscriptTests: the whole path on real speech from say, through the recorder sink, the timer, the model and the file. Asserts a chunk is on disk while audio is still arriving and the final flush catches the tail. Skips when the model is not downloaded, so CI builds without it.
  • Serializer order test updated for the new template key.

https://claude.ai/code/session_01TG4DEV5UyCPRFegrkNK3oE

`live_transcript` (off by default, a toggle under Recordings in Settings)
makes a session write `live.jsonl` as it records: every buffer the two
track files get is also resampled and accumulated, and a timer that
exists only while a session is live cuts a chunk when the speaker pauses
or every ten seconds, transcribes it on the daemon's already-loaded
model, and appends one JSON line. The post-stop transcript is untouched
and remains the record.

`on_start` is the counterpart of `on_stop`: a shell command given the
session folder the moment both tracks are recording. Both hooks share
one launcher now and both are in Settings.

Cost, measured on the default model (M-series, p50 of 7): a 5 s chunk
transcribes in 43 ms, a 20 s chunk in 65 ms, so a two-track call spends
about 1% of the time on the ANE while live is on. Digital silence is
skipped before it reaches the model. Memory is unchanged: same model,
same actor as dictation.

Claude-Session: https://claude.ai/code/session_01TG4DEV5UyCPRFegrkNK3oE
@TerrifiedBug
TerrifiedBug merged commit 4741d20 into main Oct 5, 2026
1 check passed
@TerrifiedBug
TerrifiedBug deleted the feat/live-transcript branch October 5, 2026 15:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant