Ledgeur transcribes, diarizes and reasons entirely on the user's device.
There are two engines. The native one (this document's main subject) is behind a Cargo feature so the base app stays fast to build. The browser one runs everywhere else — the web app, and any desktop build without the native feature — and since the 2026-08 overhaul it does speaker separation too, so an unbuilt native engine no longer means a transcript with one speaker on every line.
| Capability | Model | Where |
|---|---|---|
| Transcription | Whisper (Xenova/whisper-*, onnx-community/whisper-*) |
packages/asr/transcribe.worker.js |
| Speaker segmentation | pyannote 3.0 (onnx-community/pyannote-segmentation-3.0) |
packages/asr/diarize.worker.js |
| Speaker embeddings | WeSpeaker ResNet34 (onnx-community/wespeaker-voxceleb-resnet34-LM) |
same worker |
| Clustering + identification | pure TypeScript | packages/core/src/diarize/ |
All of it runs through transformers.js, pinned to one version so a page never loads two onnxruntime majors at once. The models stream from the Hugging Face CDN on first use and are then cached by the browser; after that it works offline.
A live meeting analyses each drained slice as it arrives and keeps only turns and vectors, never the audio — an hour at 16 kHz is ~230 MB of Float32. Clustering runs once at the end, over everything, because "which of these voices is the same person" is not answerable twenty seconds at a time.
The unit tests cover the logic, but they cannot tell you that the models return
the shapes this code reads. packages/asr/verify/diarize.mjs runs the real
models over a real recording and prints what comes out. It is how the clustering
threshold was set — see the comments in
packages/core/src/diarize/cluster.ts, which record the measurements rather than
the reasoning that produced the first, wrong guess.
| Capability | Engine | Where |
|---|---|---|
| Real-time transcription | whisper.cpp (whisper-rs) |
Rust command transcribe_chunk — the model is loaded once per process, not per chunk |
| Live speaker labels | sherpa-onnx speaker embeddings + running centroids | Inside transcribe_chunk; reset_live_speakers clears them between takes |
| Speaker diarization + confidence | sherpa-onnx (sherpa-rs) |
Rust command diarize_meeting — speaker turns only; it does not re-transcribe |
| Speaker identification (named voices) | sherpa-onnx speaker embeddings + cosine match | enroll_voice / list_voice_profiles / delete_voice_profile; applied inside diarize_meeting |
| Copilot chat · coaching suggestions · post-meeting notes | llama.cpp in-process (llama-cpp-2) |
Rust commands llm_chat / llm_status / download_llm — no server, no third-party app |
| RAG embeddings (Ask semantic search) | OpenAI-compatible HTTP endpoint | VITE_LOCAL_LLM_URL (BYO key or external llama.cpp) — optional; Ask falls back to keyword search |
Enrol a voice in Settings → On-device AI → Voice profiles (~10 s of clear
speech). The embedding is stored in voices.json in the app data dir — voice
prints never leave the device. On stop, diarize_meeting embeds each
diarized speaker's audio (up to 12 s) and cosine-matches against enrolled
profiles; matches at ≥ 0.5 similarity label the transcript with the real name
and a confidence figure (speaker_confidence). Unmatched speakers stay
anonymous "Speaker N" — identity is never guessed.
sherpa-onnx's fast clustering (DIARIZE_DISTANCE_THRESHOLD in engine.rs) is
a plain threshold over cosine distance, with no notion of cluster size — every
turn that lands outside the threshold of every voice heard so far becomes its
own permanent "speaker." pyannote's segmentation model routinely cuts far more
turns than there are people in a room (an interjection, a cough, a word caught
mid-hand-over), so left alone this over-splits badly on real recordings — a
real 4-person meeting once came back as 126 "speakers."
fold_tiny_clusters (ai/mod.rs) runs immediately after sherpa's pass: any
cluster whose total speaking time stays under MIN_SPEAKER_MS (2s) is folded
into whichever surviving cluster its own audio — embedded via
engine::cluster_embeddings, one CAM++ pass per cluster, reusing the same
clip-collection identify_speakers already does for voice-profile matching —
sounds most similar to, however weakly. Pure and unit-tested independently of
the model. This is the native-Rust equivalent of the MIN_SPEAKER_SECONDS
correction described below for the browser path; the two are separate
implementations over separate embedding spaces and were fixed separately —
see the 2026-09-22 changelog entry.
Two different problems, solved two different ways.
The pass on stop sees the whole recording and can cluster globally. Live, the
only thing in hand is the utterance that just finished, so each one is embedded
once and matched against the running centroid of every voice heard so far in the
meeting (src-tauri/src/ai/live_speakers.rs). An index, once handed out,
belongs to that voice for the rest of the take.
Both thresholds are measured, not chosen, and they do not transfer between
these two uses even though both are cosine distances over the same model — see
DIARIZE_DISTANCE_THRESHOLD and LIVE_SPEAKER_DISTANCE in engine.rs for the
data and the sweeps that produced them. In particular, a clip shorter than
about three seconds carries no usable speaker identity at all, so utterances
below that are left unattributed and render with no speaker chip. The pass on
stop re-labels the transcript from a global view regardless.
The transcription path has a set of #[ignore]d measurements next to it, run by
hand against the real models and a real speech clip:
curl -sL https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/ted_60.wav -o /tmp/ted.wav
cd apps/desktop/src-tauri
LEDGEUR_TEST_WAV=/tmp/ted.wav cargo test --release --features native-ai \
measures_the_live_loop -- --ignored --nocapturemeasures_transcription_speed, measures_the_live_loop,
measures_speaker_separability, assignment_threshold_sweep and
sweeps_the_clustering_threshold are all there. Run them before and after
touching anything in this path — every number quoted in the comments came from
one of them, and the ones that were guessed instead were wrong.
The webview fallback (transformers.js Whisper) is used automatically when the native engine isn't compiled/available, so the app always works.
Requires a C/C++ toolchain + CMake (both present on macOS with Xcode + Homebrew).
# Desktop, native engine on:
pnpm --filter @ledgeur/desktop tauri:dev:ai # dev
pnpm --filter @ledgeur/desktop tauri:build:ai # release
# Or check just the Rust crate:
cd apps/desktop/src-tauri && cargo check --features native-aiwhisper-rs and llama-cpp-2 compile whisper.cpp / llama.cpp via CMake (picking
up Metal on macOS, CUDA/Vulkan where the toolchain provides it); sherpa-rs
downloads prebuilt sherpa-onnx libs from GitHub releases (set
UNSAFE_DISABLE_CHECKSUM_VALIDATION=1 only if a release checksum lags a new
version).
If the build machine can't fetch GitHub release assets, download the sherpa-onnx
lib bundle once, set sherpa-rs's default-features = false in
src-tauri/Cargo.toml, and point SHERPA_LIB_PATH at it:
# one-off, from anywhere with network:
curl -L -o sherpa.tar.bz2 \
https://github.com/k2-fsa/sherpa-onnx/releases/download/v1.12.9/sherpa-onnx-v1.12.9-osx-universal2-shared.tar.bz2
tar xf sherpa.tar.bz2
export SHERPA_LIB_PATH="$PWD/sherpa-onnx-v1.12.9-osx-universal2-shared"
cargo check --features native-ai # verified compiling on macOS arm64Integrations → On-device AI → Download models (or the download_models
command) fetches into the app data dir:
Transcription/diarization models (download_models):
| File | Source |
|---|---|
ggml-base.en.bin |
huggingface.co/ggerganov/whisper.cpp |
pyannote-segmentation-3.0.onnx |
sherpa-onnx pyannote segmentation |
speaker-embedding-en-campplus.onnx |
WeSpeaker CAM++, VoxCeleb (English) |
The embedding model is English-specific on purpose. It was previously
3D-Speaker's ..._sv_zh-cn_..., which is trained on Mandarin. That model
separated English speakers poorly: voices that sound obviously different landed
close enough together for the clusterer to merge them into one person. It was
also several times more expensive to run, and this model is evaluated once per
diarization window. Superseded files are deleted on the next download, so an
install that has been through an upgrade does not keep carrying old weights.
Both diarization models run with num_threads set from the machine's core
count. sherpa_rs::diarize::Diarize hardcodes 1 and exposes no way to change
it, so src/ai/engine.rs builds the sherpa-onnx config against the C API
directly. That is the only reason it uses the raw bindings.
Copilot LLM (download_llm, one tap in Settings → On-device AI → Download
assistant, or the inline prompt the first time you use the copilot):
| File | Source | Size |
|---|---|---|
qwen2.5-1.5b-instruct-q4_k_m.gguf |
huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF | ~1.1 GB |
Chat, coaching suggestions and post-meeting notes run inside the app via
llama-cpp-2 (src-tauri/src/ai/llm.rs) — there is no separate process and
nothing for the user to install. The weights are downloaded once (streamed with
progress) into the app data dir and cached; the model is loaded once and reused.
- Frontend entry point:
src/lib/llm.ts(chatComplete) — native first, then an OpenAI-compatible HTTP fallback (VITE_LOCAL_LLM_URL, for a BYO cloud key or an external llama.cpp), then an explicit error. Never fabricates. - Prompt format: Qwen ChatML (
render_chatml, unit-tested). Sampler: top-k/top-p- temperature. Context window 8192; over-long prompts keep the most recent tokens (the question is at the tail).
- Post-meeting notes:
src/lib/notes.tsasks the model for structured JSON (summary / action items / decisions / questions) and falls back to the local heuristic extractor (packages/coresummarizeTranscript) when no model is available — so notes are always real, never blank, never invented.
RAG embeddings still use the HTTP endpoint (point it at nomic-embed-text,
768-dim, matching the embeddings.embedding vector(768) column). Native
embeddings are a follow-up; Ask degrades to keyword search without them.
A separate feature from everything above — it captures audio, not AI. The
webview's only way to hear the other side of a call is getDisplayMedia
(video permission + a screen-share picker, just to get audio). On macOS
14.2+, apps/desktop/src-tauri/src/audio/ uses Core Audio's Process Tap API
(AudioHardwareCreateProcessTap) instead: no picker, no video, no menu-bar
recording indicator — a private aggregate device combines the tap with the
system's default output device, and PCM is streamed to the frontend over a
Tauri event (system-audio:chunk). Modelled on Apple's own reference sample
(insidegui/AudioCap), using the objc2-core-audio crate's bindings rather
than hand-rolled FFI.
On by default. system-audio-tap is in the crate's default features. It
shipped opt-in at first, which meant every real build had it off and still
fell back to getDisplayMedia — the exact prompt the module exists to remove.
Its dependencies are declared under the macOS target only, so the feature is
inert on other platforms, and the module's own cfg keeps the command stubs in
place there (and on macOS older than 14.2, where the API doesn't exist).
cd apps/desktop/src-tauri && cargo check # tap included
cargo check --no-default-features # stubs only
cargo check --features native-ai # with the AI engineuseRecorder.start() (desktop only) checks system_audio_tap_available and
uses the tap when it reports true, falling back to getDisplayMedia
automatically everywhere else (Windows, older macOS, the plain website) — so
this can't regress anyone it doesn't apply to. Where the tap is available, the
"System audio" toggle defaults on: there is no picker and no recording
indicator to spring on someone, only a one-time OS permission. Where it isn't,
the toggle stays off and says it needs screen-recording permission.
For comparison, Granola requires macOS's "Screen & System Audio recording" permission to do the same job, so this is a real difference rather than parity.
Still not verified capturing real audio here — it compiles cleanly (by
default, with native-ai, and with no default features), but actually tapping
system audio, the macOS permission prompt
(NSAudioCaptureUsageDescription), and signed-build / entitlement behaviour all
need a real Mac with a signed build. See the manual checklist below.
-
tauri:dev:ailaunches;ai_statusreportscompiled: true,llm_statusreportscompiled: true. - Download models; record a short meeting → live transcript appears as chat
bubbles; on stop, multiple
Speaker Nlabels + per-segment confidence. - First copilot message shows the one-tap Download assistant (~1 GB)
prompt; after download,
llm_status.model_readyis true. - In a meeting, type in the bottom input → copilot answers as a gold bubble; quoting a transcript line prepends it to the question. Proactive suggestions appear when enabled (Settings → Meeting copilot).
- Stop the meeting → summary + action items are written by the model (or the heuristic fallback if the model isn't downloaded). With "Save copilot chat" off (default), the saved meeting holds only the transcript.
-
tauri:dev(the tap is now a default feature) on macOS 14.2+: "System audio" is already on and shows the no-picker copy,Start recordingtriggers the OS's one-time audio-recording permission prompt (not a screen-share picker), and playing audio from another app during the meeting shows up in the transcript. No menu-bar recording indicator should appear. - The same in a signed, notarised build, which is where entitlements differ from a dev build — an unsigned binary can be allowed to tap where a signed one is refused, or the reverse.
- Contextely ingest end-to-end: add Ledgeur as a source in Contextely with
the endpoint from Integrations, and confirm meetings arrive with a title
and real content. The record shape is pinned by tests in
packages/mcp/test/run.mts, but nothing here can exercise a live Contextely workspace.