Skip to content

Latest commit

 

History

History
281 lines (227 loc) · 15 KB

File metadata and controls

281 lines (227 loc) · 15 KB

On-device AI

Ledgeur transcribes, diarizes and reasons entirely on the user's device.

There are two engines. The native one (this document's main subject) is behind a Cargo feature so the base app stays fast to build. The browser one runs everywhere else — the web app, and any desktop build without the native feature — and since the 2026-08 overhaul it does speaker separation too, so an unbuilt native engine no longer means a transcript with one speaker on every line.

The browser engine

Capability Model Where
Transcription Whisper (Xenova/whisper-*, onnx-community/whisper-*) packages/asr/transcribe.worker.js
Speaker segmentation pyannote 3.0 (onnx-community/pyannote-segmentation-3.0) packages/asr/diarize.worker.js
Speaker embeddings WeSpeaker ResNet34 (onnx-community/wespeaker-voxceleb-resnet34-LM) same worker
Clustering + identification pure TypeScript packages/core/src/diarize/

All of it runs through transformers.js, pinned to one version so a page never loads two onnxruntime majors at once. The models stream from the Hugging Face CDN on first use and are then cached by the browser; after that it works offline.

A live meeting analyses each drained slice as it arrives and keeps only turns and vectors, never the audio — an hour at 16 kHz is ~230 MB of Float32. Clustering runs once at the end, over everything, because "which of these voices is the same person" is not answerable twenty seconds at a time.

Verifying it

The unit tests cover the logic, but they cannot tell you that the models return the shapes this code reads. packages/asr/verify/diarize.mjs runs the real models over a real recording and prints what comes out. It is how the clustering threshold was set — see the comments in packages/core/src/diarize/cluster.ts, which record the measurements rather than the reasoning that produced the first, wrong guess.

The native engine

What runs where

Capability Engine Where
Real-time transcription whisper.cpp (whisper-rs) Rust command transcribe_chunk — the model is loaded once per process, not per chunk
Live speaker labels sherpa-onnx speaker embeddings + running centroids Inside transcribe_chunk; reset_live_speakers clears them between takes
Speaker diarization + confidence sherpa-onnx (sherpa-rs) Rust command diarize_meeting — speaker turns only; it does not re-transcribe
Speaker identification (named voices) sherpa-onnx speaker embeddings + cosine match enroll_voice / list_voice_profiles / delete_voice_profile; applied inside diarize_meeting
Copilot chat · coaching suggestions · post-meeting notes llama.cpp in-process (llama-cpp-2) Rust commands llm_chat / llm_status / download_llm — no server, no third-party app
RAG embeddings (Ask semantic search) OpenAI-compatible HTTP endpoint VITE_LOCAL_LLM_URL (BYO key or external llama.cpp) — optional; Ask falls back to keyword search

Speaker identification

Enrol a voice in Settings → On-device AI → Voice profiles (~10 s of clear speech). The embedding is stored in voices.json in the app data dir — voice prints never leave the device. On stop, diarize_meeting embeds each diarized speaker's audio (up to 12 s) and cosine-matches against enrolled profiles; matches at ≥ 0.5 similarity label the transcript with the real name and a confidence figure (speaker_confidence). Unmatched speakers stay anonymous "Speaker N" — identity is never guessed.

Folding away phantom speakers

sherpa-onnx's fast clustering (DIARIZE_DISTANCE_THRESHOLD in engine.rs) is a plain threshold over cosine distance, with no notion of cluster size — every turn that lands outside the threshold of every voice heard so far becomes its own permanent "speaker." pyannote's segmentation model routinely cuts far more turns than there are people in a room (an interjection, a cough, a word caught mid-hand-over), so left alone this over-splits badly on real recordings — a real 4-person meeting once came back as 126 "speakers."

fold_tiny_clusters (ai/mod.rs) runs immediately after sherpa's pass: any cluster whose total speaking time stays under MIN_SPEAKER_MS (2s) is folded into whichever surviving cluster its own audio — embedded via engine::cluster_embeddings, one CAM++ pass per cluster, reusing the same clip-collection identify_speakers already does for voice-profile matching — sounds most similar to, however weakly. Pure and unit-tested independently of the model. This is the native-Rust equivalent of the MIN_SPEAKER_SECONDS correction described below for the browser path; the two are separate implementations over separate embedding spaces and were fixed separately — see the 2026-09-22 changelog entry.

Live speakers vs. the pass on stop

Two different problems, solved two different ways.

The pass on stop sees the whole recording and can cluster globally. Live, the only thing in hand is the utterance that just finished, so each one is embedded once and matched against the running centroid of every voice heard so far in the meeting (src-tauri/src/ai/live_speakers.rs). An index, once handed out, belongs to that voice for the rest of the take.

Both thresholds are measured, not chosen, and they do not transfer between these two uses even though both are cosine distances over the same model — see DIARIZE_DISTANCE_THRESHOLD and LIVE_SPEAKER_DISTANCE in engine.rs for the data and the sweeps that produced them. In particular, a clip shorter than about three seconds carries no usable speaker identity at all, so utterances below that are left unattributed and render with no speaker chip. The pass on stop re-labels the transcript from a global view regardless.

Measuring it

The transcription path has a set of #[ignore]d measurements next to it, run by hand against the real models and a real speech clip:

curl -sL https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/ted_60.wav -o /tmp/ted.wav
cd apps/desktop/src-tauri
LEDGEUR_TEST_WAV=/tmp/ted.wav cargo test --release --features native-ai \
  measures_the_live_loop -- --ignored --nocapture

measures_transcription_speed, measures_the_live_loop, measures_speaker_separability, assignment_threshold_sweep and sweeps_the_clustering_threshold are all there. Run them before and after touching anything in this path — every number quoted in the comments came from one of them, and the ones that were guessed instead were wrong.

The webview fallback (transformers.js Whisper) is used automatically when the native engine isn't compiled/available, so the app always works.

Build with the native engine

Requires a C/C++ toolchain + CMake (both present on macOS with Xcode + Homebrew).

# Desktop, native engine on:
pnpm --filter @ledgeur/desktop tauri:dev:ai      # dev
pnpm --filter @ledgeur/desktop tauri:build:ai    # release

# Or check just the Rust crate:
cd apps/desktop/src-tauri && cargo check --features native-ai

whisper-rs and llama-cpp-2 compile whisper.cpp / llama.cpp via CMake (picking up Metal on macOS, CUDA/Vulkan where the toolchain provides it); sherpa-rs downloads prebuilt sherpa-onnx libs from GitHub releases (set UNSAFE_DISABLE_CHECKSUM_VALIDATION=1 only if a release checksum lags a new version).

Offline / CI build (no GitHub release download)

If the build machine can't fetch GitHub release assets, download the sherpa-onnx lib bundle once, set sherpa-rs's default-features = false in src-tauri/Cargo.toml, and point SHERPA_LIB_PATH at it:

# one-off, from anywhere with network:
curl -L -o sherpa.tar.bz2 \
  https://github.com/k2-fsa/sherpa-onnx/releases/download/v1.12.9/sherpa-onnx-v1.12.9-osx-universal2-shared.tar.bz2
tar xf sherpa.tar.bz2
export SHERPA_LIB_PATH="$PWD/sherpa-onnx-v1.12.9-osx-universal2-shared"
cargo check --features native-ai   # verified compiling on macOS arm64

Models (downloaded on first use)

Integrations → On-device AI → Download models (or the download_models command) fetches into the app data dir:

Transcription/diarization models (download_models):

File Source
ggml-base.en.bin huggingface.co/ggerganov/whisper.cpp
pyannote-segmentation-3.0.onnx sherpa-onnx pyannote segmentation
speaker-embedding-en-campplus.onnx WeSpeaker CAM++, VoxCeleb (English)

The embedding model is English-specific on purpose. It was previously 3D-Speaker's ..._sv_zh-cn_..., which is trained on Mandarin. That model separated English speakers poorly: voices that sound obviously different landed close enough together for the clusterer to merge them into one person. It was also several times more expensive to run, and this model is evaluated once per diarization window. Superseded files are deleted on the next download, so an install that has been through an upgrade does not keep carrying old weights.

Both diarization models run with num_threads set from the machine's core count. sherpa_rs::diarize::Diarize hardcodes 1 and exposes no way to change it, so src/ai/engine.rs builds the sherpa-onnx config against the C API directly. That is the only reason it uses the raw bindings.

Copilot LLM (download_llm, one tap in Settings → On-device AI → Download assistant, or the inline prompt the first time you use the copilot):

File Source Size
qwen2.5-1.5b-instruct-q4_k_m.gguf huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF ~1.1 GB

Copilot LLM (in-process, no server)

Chat, coaching suggestions and post-meeting notes run inside the app via llama-cpp-2 (src-tauri/src/ai/llm.rs) — there is no separate process and nothing for the user to install. The weights are downloaded once (streamed with progress) into the app data dir and cached; the model is loaded once and reused.

  • Frontend entry point: src/lib/llm.ts (chatComplete) — native first, then an OpenAI-compatible HTTP fallback (VITE_LOCAL_LLM_URL, for a BYO cloud key or an external llama.cpp), then an explicit error. Never fabricates.
  • Prompt format: Qwen ChatML (render_chatml, unit-tested). Sampler: top-k/top-p
    • temperature. Context window 8192; over-long prompts keep the most recent tokens (the question is at the tail).
  • Post-meeting notes: src/lib/notes.ts asks the model for structured JSON (summary / action items / decisions / questions) and falls back to the local heuristic extractor (packages/core summarizeTranscript) when no model is available — so notes are always real, never blank, never invented.

RAG embeddings still use the HTTP endpoint (point it at nomic-embed-text, 768-dim, matching the embeddings.embedding vector(768) column). Native embeddings are a follow-up; Ask degrades to keyword search without them.

Native system audio (Core Audio Process Tap)

A separate feature from everything above — it captures audio, not AI. The webview's only way to hear the other side of a call is getDisplayMedia (video permission + a screen-share picker, just to get audio). On macOS 14.2+, apps/desktop/src-tauri/src/audio/ uses Core Audio's Process Tap API (AudioHardwareCreateProcessTap) instead: no picker, no video, no menu-bar recording indicator — a private aggregate device combines the tap with the system's default output device, and PCM is streamed to the frontend over a Tauri event (system-audio:chunk). Modelled on Apple's own reference sample (insidegui/AudioCap), using the objc2-core-audio crate's bindings rather than hand-rolled FFI.

On by default. system-audio-tap is in the crate's default features. It shipped opt-in at first, which meant every real build had it off and still fell back to getDisplayMedia — the exact prompt the module exists to remove. Its dependencies are declared under the macOS target only, so the feature is inert on other platforms, and the module's own cfg keeps the command stubs in place there (and on macOS older than 14.2, where the API doesn't exist).

cd apps/desktop/src-tauri && cargo check                       # tap included
cargo check --no-default-features                              # stubs only
cargo check --features native-ai                               # with the AI engine

useRecorder.start() (desktop only) checks system_audio_tap_available and uses the tap when it reports true, falling back to getDisplayMedia automatically everywhere else (Windows, older macOS, the plain website) — so this can't regress anyone it doesn't apply to. Where the tap is available, the "System audio" toggle defaults on: there is no picker and no recording indicator to spring on someone, only a one-time OS permission. Where it isn't, the toggle stays off and says it needs screen-recording permission.

For comparison, Granola requires macOS's "Screen & System Audio recording" permission to do the same job, so this is a real difference rather than parity.

Still not verified capturing real audio here — it compiles cleanly (by default, with native-ai, and with no default features), but actually tapping system audio, the macOS permission prompt (NSAudioCaptureUsageDescription), and signed-build / entitlement behaviour all need a real Mac with a signed build. See the manual checklist below.

Manual test checklist (can't be verified headless)

  • tauri:dev:ai launches; ai_status reports compiled: true, llm_status reports compiled: true.
  • Download models; record a short meeting → live transcript appears as chat bubbles; on stop, multiple Speaker N labels + per-segment confidence.
  • First copilot message shows the one-tap Download assistant (~1 GB) prompt; after download, llm_status.model_ready is true.
  • In a meeting, type in the bottom input → copilot answers as a gold bubble; quoting a transcript line prepends it to the question. Proactive suggestions appear when enabled (Settings → Meeting copilot).
  • Stop the meeting → summary + action items are written by the model (or the heuristic fallback if the model isn't downloaded). With "Save copilot chat" off (default), the saved meeting holds only the transcript.
  • tauri:dev (the tap is now a default feature) on macOS 14.2+: "System audio" is already on and shows the no-picker copy, Start recording triggers the OS's one-time audio-recording permission prompt (not a screen-share picker), and playing audio from another app during the meeting shows up in the transcript. No menu-bar recording indicator should appear.
  • The same in a signed, notarised build, which is where entitlements differ from a dev build — an unsigned binary can be allowed to tap where a signed one is refused, or the reverse.
  • Contextely ingest end-to-end: add Ledgeur as a source in Contextely with the endpoint from Integrations, and confirm meetings arrive with a title and real content. The record shape is pinned by tests in packages/mcp/test/run.mts, but nothing here can exercise a live Contextely workspace.