Skip to content

Repository files navigation

Ledgeur

The centre of your company's brain. Ledgeur records meetings, transcribes them, and works out who said what — all on your own device. Name a voice once and it is recognised in every meeting after that. Drag in recordings you already have and they are treated exactly like live ones.

The app is free, permanently, for one person: unlimited recording, transcription, speaker separation, notes, search and export. The paid tier is what happens when the record has to leave your machine — sync across your devices, a shared team library, and a Model Context Protocol endpoint that opens your meetings to Claude, ChatGPT or Cursor.

Nothing goes in the price list unless it ships. See apps/marketing/lib/plans.ts, which carries that rule and a test that enforces it.

One app for macOS, Windows, iOS and Android, built on Tauri 2 + Supabase, on one design system (packages/ui: Plus Jakarta Sans, a neutral canvas, six pastel families, light and dark — every colour pairing measured). The phone app is the same code as the desktop app and syncs with it under the same meeting ids. See docs/ARCHITECTURE.md, docs/MOBILE.md, docs/REDESIGN.md and docs/ROADMAP.md.

The Vercel marketing app sends errors, low-volume performance traces, source maps, and user feedback to Sentry; no telemetry runs if NEXT_PUBLIC_SENTRY_DSN is absent.

Monorepo

Path What
apps/desktop The app — Tauri 2 + Vite + React (all platforms)
apps/marketing Next.js 16 marketing/SEO site (Vercel)
packages/core Shared domain model, diarization logic, the meeting library, browser controllers, auth wording, notes/audio logic, Supabase client
packages/asr Browser speech-to-text and speaker-diarization workers + their load plans (synced into each app's public/)
packages/ui Design tokens, the shared theme.css, and the React primitives both apps render
supabase/ Database schema (migrations) — source of truth for data

Develop

pnpm install                 # install the whole workspace

# App (browser preview :1420) + marketing (:3000) together
pnpm dev

# Just the app (browser preview — fast UI iteration, on-device transcription)
pnpm desktop:dev             # http://localhost:1420

# The app (native window with the Rust core)
pnpm --filter @ledgeur/desktop tauri:dev

# iOS (requires Xcode) and Android (requires the SDK, NDK and JDK 17) — see docs/MOBILE.md
pnpm --filter @ledgeur/desktop ios:dev
pnpm --filter @ledgeur/desktop android:dev

# Regenerate the design tokens' CSS after editing packages/ui/src/tokens.ts
pnpm --filter @ledgeur/ui build:theme

# The marketing site
pnpm marketing:dev

# The paid MCP server (needs Supabase creds — run it on its own, not part of dev)
LEDGEUR_SUPABASE_URL=… LEDGEUR_SUPABASE_ANON_KEY=… LEDGEUR_REFRESH_TOKEN=… \
  pnpm --filter @ledgeur/mcp-server start

# Everything
pnpm build      # turbo build across all packages
pnpm test       # turbo test
pnpm lint

Use pnpm, not npm — this is a pnpm workspace. npm run dev also works but prints harmless Unknown project config warnings for pnpm-only .npmrc keys.

Configure the backend (optional for local UI work)

The app runs in local-only mode with no configuration (recordings are cached in IndexedDB). To enable accounts, sync and the hive mind, create apps/desktop/.env from apps/desktop/.env.example and apply the schema in supabase/migrations to your Supabase project.

How the microphone is captured

Ledgeur opens the mic passively, with no echo cancellation, no noise suppression and no automatic gain, everywhere it records: meetings, dictation and voice enrolment. This is deliberate and it is not a missing feature.

Asking for echo cancellation is not a request for a software filter. On macOS, WebKit maps it straight onto the audio unit subtype (m_shouldUseVPIO = enableEchoCancellation()), so echoCancellation: true opens the mic through VoiceProcessingIO, the same "I am a call client" path Zoom, Meet and FaceTime use, instead of a passive one. That reconfigures the shared input device. Measured on a built-in MacBook mic, it fires 32 Core Audio property changes rewriting the device's physical and virtual stream formats, against five benign lifecycle events for passive capture. When somebody is already on a call, that renegotiation lands underneath the call app while it holds the mic, and the far end hears them go quiet. Recording a call must never degrade the call.

The constraints live in one table, MIC_PROCESSING in packages/core/src/browser/capture.ts, with every flag stated explicitly. An omitted flag is not "off": browsers default all three to true, which is how this shipped as a bug once already.

This fixed Meet. It didn't fix Zoom, because Zoom's quiet-mic symptom has a different cause: Zoom ships its own microphone auto-gain-control ("Automatically adjust microphone volume," on by default, separate from its noise-suppression toggles) that is independently documented to reset input volume when triggered. Ledgeur has no API into another app's settings, so apps/desktop/src-tauri/src/callapps.rs detects when Zoom is running and the Record screen tells people which setting to turn off, instead of requiring them to already know.

How speaker separation works

Two models, both in the browser, both free:

Stage Model What it answers
Segmentation onnx-community/pyannote-segmentation-3.0 Where does the voice change? Handles up to three people talking at once.
Embedding onnx-community/wespeaker-voxceleb-resnet34-LM What does this stretch of speech sound like, as a vector?

The deciding — clustering those vectors into people, and matching them against voices you have already named — is pure TypeScript in packages/core/src/diarize, so it is unit-tested without a browser and shared by the live and imported paths.

A live meeting analyses each drained slice as it arrives and keeps only the turns and their vectors, never the audio: an hour at 16 kHz is ~230 MB of Float32, and holding that in a tab to diarize at the end is not reasonable. Clustering still runs once over everything at the end, because "which of these voices is the same person" cannot be answered twenty seconds at a time.

Clustering weighs each turn by its duration, not by count: a short, noisy turn (an interjection, a word caught mid-hand-over) gives the embedding model its least reliable signal, so a handful of confident seconds from one voice outweighs several brief ones from another when deciding who two clusters belong to. Any cluster whose total speaking time stays under two seconds after the main pass is folded into its nearest neighbour regardless of similarity — a phantom split from noise is far likelier than a real participant who barely spoke.

Voice prints live in IndexedDB and are never synced, not even on the paid plan — a voice print identifies a person after the transcript is deleted.

Putting names to the voices

Separation gives you "Speaker 1" and "Speaker 2", which is useful once. Meetings usually say who is present, though — someone introduces themselves, or answers to their name — so when a recording finishes, the on-device model reads the transcript and names the voices it can prove.

Nothing here guesses from patterns. There is deliberately no regex pulling "I'm X" out of a transcript: it cannot tell "I'm Max" from "I'm afraid not". The model proposes, and packages/core/src/diarize/names.ts then throws out anything it cannot check —

  • the name must actually be spoken in the transcript;
  • the model's quoted evidence must be a real line, and must be the line that says the name;
  • it must clear a belief threshold (0.75 to label, 0.85 to teach the voice);
  • one name per voice, one voice per name.

Every name that survives is shown as a guess — a mark on the chip, the belief, and the words it came from — everywhere it appears, and is one click to accept, change, or reject. Correcting a guess also un-teaches whatever it taught the voice store, so a wrong name cannot quietly propagate into later meetings. A single misattributed line can be moved on its own, without touching the voice. If clustering splits one person into two speakers, "Merge into…" on either chip folds one entirely into the other, relabelling every line and combining their speaking time in one step.

To recognise somebody next time, a few seconds of their speech is kept with the meeting — chosen for blandness rather than convenience. Every candidate window is scored for card numbers, credentials, salaries, health and the like (snippet.ts), the least sensitive one wins, and if a person's every stretch looks sensitive, nothing is kept at all. Like voice prints, these samples never leave the device — asserted by a test, not just by this paragraph.

Status

Phases 0–5 are code-complete. The 2026-07 "Library of Record" redesign shipped, followed by the seamless on-device copilot update:

  • The copilot, coaching suggestions and post-meeting notes run in-process (llama.cpp via llama-cpp-2) — nothing to install; the model is auto-downloaded once (one tap) and cached. Notes fall back to a local heuristic when offline.
  • The live meeting is one continuous chat thread — transcript, copilot and your questions as bubbles you can quote; the right rail is just your notes.
  • The whole app is a chat surface with an ever-present bottom input; each screen renders as an embedded window card.

The 2026-08 overhaul pass added on-device speaker separation everywhere, voice prints that persist between meetings, drag-and-drop import, a real web app with a searchable library, accounts and billing on the web, and a price list that describes only what exists. It also fixed a checkout that took money without activating anything and an access-token scheme in which every token issued was unusable. See docs/OVERHAUL.md and the changelog.

The 2026-09 grounding pass rebuilt what a question is allowed to see. Asking something mid-meeting now reaches the live transcript with speakers and timestamps, who has spoken, your own typed notes, and — in the same prompt — Contextely company memory, Notion, the org's indexed meetings and your past recordings, on a five-second deadline that names whatever did not arrive in time. Every answer shows the sources it was grounded in. The same pass added per-line provenance from notes back to the transcript, follow-up email drafts, user-written note recipes, 32 spoken languages, spaces, a derived people directory, signed outbound webhooks, calendar auto-start and SAML SSO.

The 2026-09-06 redesign, phone and sync pass replaced the design system outright — one sans family, a neutral canvas with six pastel families, light and dark, a generated token sheet the two apps cannot drift from — and rebuilt every screen of the app and every page of the site on it. The same pass generated the iOS and Android projects (the phone is the same app with a phone shell) and replaced one-shot, one-way sync with an engine that pushes every edit, pulls every change, honours deletions, and listens for the other device over Realtime. See docs/REDESIGN.md and docs/MOBILE.md.

The 2026-09-08 capture pass gave the product a second way in. Recording a meeting was the only one, so the thought you have on the way out of one went nowhere. There is now a box for it (⌘⇧K, a phone tab, a home-screen and lock-screen widget) that takes a thought typed or spoken and keeps it in one action. The on-device model then decides whether it was a task or a note and which space it belongs to, under the same rules speaker naming works by: it may only choose a space you already have, it has to quote the words it decided on, and it has to clear a belief threshold or leave the thought in the inbox. The thought is on disk before any of that runs, so a missing or wrong model can change where it lands but never whether it survived. Spaces became projects in the same pass: a space now has its own page holding its meetings, its tasks and its notes, and a finished meeting files itself into one.

The 2026-09-09 strategy pass made the product and the site agree with each other. Two of its findings were bugs rather than opinions: /security and /pricing published contradictory claims about SAML, and sync, the headline paid feature, was working on free accounts. The first is fixed by giving the list of what is missing one home (apps/marketing/lib/gaps.ts) and its own page at /what-we-dont-have; the second by migration 0009_sync_is_paid.sql, which gates writes on a paid plan in the database while deliberately leaving reads and deletes open, so cancelling never strands a library. The same pass added one-click task push to Linear, Todoist and Asana, moved the Team tier to $12, and built the pages the site was missing: /templates, /guides, /speaker-identification and /company-memory. See docs/seo_geo_content_plan.md.

Earlier: editorial design system, ⌘K palette, mobile tab bar, recordings that survive navigation. See docs/ROADMAP.md for what's next, docs/NATIVE_AI.md for the on-device engine, and docs/MANUAL_TESTING.md for flows that need live services or a device, and docs/DEPLOYMENT.md for what has to be configured — including SUPABASE_JWT_SECRET, which is new and which the hosted agent endpoint cannot work without.

Licence

MIT © Ledgeur

About

Open-source, private AI meeting notes that run 100% in your browser. On-device Whisper transcription — no bot joins your call, no cloud upload, free for individuals. The open alternative to Granola & Otter.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages