Local, speaker-attributed transcription for Apple Silicon — who said what, on your Mac, nothing leaves the machine.
whosaid pairs Whisper speech-to-text with speaker diarization to turn a meeting,
interview, call, or podcast recording into a transcript that says who spoke when — running fully
offline and on-device on an Apple Silicon Mac, with no cloud service, no API keys, and no Hugging
Face token. Think of it as whisper + speaker labels + voice-based speaker recognition, in one
command.
Example output (meeting.speakers.txt):
[00:00:04] Alice: Thanks for jumping on, I know it's late for you.
[00:00:11] SPEAKER_01: No problem at all, happy to make it work.
[00:00:19] Alice: Let's start with the roadmap for next quarter.
[00:00:27] SPEAKER_01: Sounds good, I've got a few updates on that.
- Meetings, interviews, calls, podcasts — get a transcript where every turn is attributed to a person, not just a wall of text.
- Privacy by construction — audio, text, and voice embeddings never leave your Mac. The only
cloud step is one you opt into:
[summarizer] engine = "claude"sends each transcript's text to Anthropic to draft action items. Leave it unset and nothing is sent. - Your name on your own lines — a one-time ~45s voice enrollment teaches whosaid your voice, so
your turns read as your name instead of
SPEAKER_00. - Tells you how many people spoke — the number of distinct speakers is auto-detected and reported up front, with per-speaker turn counts and talk time.
- Speaker cards to identify who's who — for every speaker, whosaid writes a card of that voice's most representative snippets, so you can read a few lines and know who was talking.
- Remembers people you name — identify a speaker once with
whosaid relabel, and their voiceprint is saved to a private local registry so they're auto-named in every future transcript. - Fast on long recordings — recordings over ~15 min are diarized in parallel windows and stitched back into consistent speakers by voiceprint, so an 80-minute meeting is minutes, not tens of minutes, recovering the same speakers as a single-pass run.
- No accounts, no API keys, no Hugging Face token — every model comes from an open, ungated source.
- Search everything you have recorded.
whosaid indexbuilds a local full-text (and, with Ollama, meaning) index over every transcript in a meeting workspace;whosaid searchfinds the turn,whosaid contextshows the minute around it, andwhosaid graphanswers who committed to what, when. See Search your meetings. - Hands-free from Voice Memos.
whosaid watch installputs a launchd agent on the Voice Memos folder (or any folder): stop recording, and a few minutes later the meeting is transcribed, summarized, and searchable. See Hands-free ingest. - Usable from an AI agent, too.
whosaid mcpexposes transcribe, relabel, doctor, and the read-only workspace search tools as MCP tools for Claude Code, Claude Desktop, Kiro, and other MCP clients, with the same local-only guarantee as the CLI.
Looking for a MacWhisper, whisperX, or aTrain alternative? If you want speaker-attributed
transcription that runs fully offline on Apple Silicon and can put real names on voices — not just
SPEAKER_00 labels — that's the gap whosaid fills. Here's how it stacks up:
| whosaid | whisperX | plain mlx-whisper | cloud transcription APIs | |
|---|---|---|---|---|
| Speaker labels | Yes | Yes | No | Varies by provider |
| Names speakers by voice | Yes (enrollment) | No | No | No |
| Remembers speakers across meetings | Yes — persistent local voiceprint registry | No | No | No |
| Runs fully offline | Yes | Partial — needs a gated model download | Yes | No |
| Parallel diarization on long audio | Yes | No | — (no diarization) | Varies by provider |
| Needs an account / token | No | Yes — Hugging Face token for gated pyannote models | No | Yes — API key |
| Install weight | ffmpeg + uv, ephemeral environments |
torch + pyannote + the full HF stack |
mlx-whisper only |
None (network client only) |
git clone https://github.com/sblattj/whosaid && cd whosaid # get the code
./bootstrap.sh # check deps, download models, install ~/.local/bin/whosaid
whosaid enroll # ~45s reading a printed passage — teaches whosaid your voice
whosaid path/to/meeting.m4a # transcribe + diarize + label -> meeting.speakers.txt (and friends)If ~/.local/bin is not on your shell's PATH, add it or invoke the installed command by its
absolute path. WHOSAID_INSTALL_DIR=/another/bin ./whosaid install selects another install
directory. The installed command is a symlink to the checkout, so updating the checkout updates the
command without copying or duplicating the implementation.
whosaid's helper scripts need Python 3.11 or newer, because whosaid.toml is read with tomllib.
The launcher picks one interpreter: $WHOSAID_PYTHON if it is set, otherwise the first of python3,
python3.13, python3.12 and python3.11 on your PATH that is 3.11+, otherwise
uv python find '>=3.11'. It skips the stock macOS python3 (3.9) automatically. If no interpreter
qualifies, commands that run Python stop with an error; uv python install 3.13 fixes that.
whosaid doctor shows which interpreter was picked.
Homebrew is optional when uv, ffmpeg, and ffprobe are already installed. The launcher
and watcher search ~/.local/bin, ~/bin, and ~/homebrew/bin as well as the usual system
prefixes. Setup can install a pinned, checksum-verified Apple Silicon uv archive into
~/.local/bin. Homebrew FFmpeg installation requires a bottle; a nonstandard prefix will not
silently trigger a source build.
For a Mac without Homebrew, or one where no bottle is available, use:
./bootstrap.sh --build-ffmpegThis builds pinned FFmpeg source into user-local ffmpeg and ffprobe executables. It needs
Apple's Command Line Tools (xcode-select --install if missing), compilation time, and free
disk space, but no administrator-owned install prefix or GPG installation. The source archive
is checked with macOS shasum against a checksum verified from FFmpeg's signed release.
Existing usable tools are retained. Setup then warms the Python environments and downloads
the models as usual.
| Command | What it does |
|---|---|
./bootstrap.sh [--yes] [--build-ffmpeg] (also whosaid setup) |
Capability check, dependency install, model pre-download, and command installation. --build-ffmpeg selects the verified user-local source build when FFmpeg tools are missing. Idempotent — safe to re-run. |
whosaid install [--force] |
Install/update the command symlink in ~/.local/bin (or WHOSAID_INSTALL_DIR). Refuses to replace an unrelated command, and refuses (unless --force) to install from a checkout under /tmp, /private/tmp, /var/tmp, or $TMPDIR — that symlink would dangle once the OS cleans the temp directory up. |
whosaid enroll [Name] |
Records ~45s from the mic reading a printed passage, saves voices/<Name>.wav. See Speaker identity: enrollment clips vs. the registry. |
whosaid enroll <Name> --from FILE [--ss T] [--t D|--to T] [--force] |
Extracts a clip from an existing recording instead of the mic (extract → verify → save, no interaction). --ss/--t/--to accept seconds or M:SS/H:MM:SS; same ≥15s/non-silent bar as mic enrollment. |
whosaid record [--label L] |
Foreground mic capture to recordings/<timestamp>[-label].m4a, then transcribes automatically. |
whosaid <audio>… [flags] |
The default command: transcribe + diarize + label one or more audio files. whosaid transcribe <audio>… is the same command written explicitly (matching the whosaid_transcribe MCP tool name). |
whosaid relabel <base> SPEAKER_02=Jane … |
Put real names on clusters after reading the speaker cards. Rewrites the transcript + cards and saves each named voiceprint to the local registry for future transcripts — refusing (unless --force) to overwrite an existing registry entry the new cluster doesn't match (similarity below the match threshold, default 0.50). No re-transcription. |
whosaid relabel <base> SPEAKER_04=Alice --no-save [--note TEXT] |
Transcript-only label: renames the cluster in the transcript + cards and never writes to the registry. For a speaker the conversation makes obvious but whose cluster is a poor voiceprint — a long mixed cluster saved as that person would degrade their enrolled print. The label is kept in the sidecar (local_labels), survives relabel --auto and whosaid samples, and each card is marked Alice_Example (transcript-only label; registry untouched). See Speaker identity: enrollment clips vs. the registry. |
whosaid relabel <base> --auto [--fold-unknown] |
Re-apply naming to an existing transcript with no assignments: re-runs registry matching + the absorb pass over the cached sidecar and rewrites the transcript + cards. Picks up voices enrolled after the transcript was made. Add --fold-unknown to repair anonymous phantom clusters in cached auto-diarization; it requires --auto and does not assign an identity. No re-transcription, no re-diarization. In a meeting workspace the base is transcript. See Speaker identity: enrollment clips vs. the registry. |
whosaid backfill <ws> [--dry-run] |
Apply current registry names and roles to historical meetings, regenerate commitments, and refresh roll-up, worklists, search, graph, and wiki. Reports conflicts before changing meeting files. |
whosaid speakers export --out FILE |
Export the private speaker registry to a new backup file, including voiceprints, roles, and metadata. |
whosaid speakers import FILE [--merge|--overwrite] |
Restore a registry backup. An existing destination requires an explicit conflict policy; --overwrite replaces the entire registry. |
whosaid samples <base> [-o DIR] [--audio FILE] [--per-speaker N] [--seconds S] [--json] |
Export one short representative WAV per speaker cluster — the longest diarized segment, clamped to --seconds (default 8) — so you can listen and confirm an identity before trusting an auto-label or enrolling. Cuts from the sidecar's own source.path, or an explicit --audio FILE for a sidecar written before that metadata existed. |
whosaid ingest <audio>… --into DIR [--action-items] [--engine E] [--index] |
Transcribe a batch into dated meeting folders (idempotent by content hash). --engine picks the action-item summarizer, --index runs roll-up + index afterwards. See Meeting workspaces. |
whosaid roll-up <ws> [--action-items] [--index] [--owner NAME|me] [--all-owners] |
Rebuild the workspace index (_INDEX.md), audit, the deduplicated action-item and dev-commitments corpora, and the owner's ranked _WORKLIST-<Owner>.md; --index then rebuilds the search index too. |
whosaid commitments <ws> [--owner NAME|me] [--all-owners] [--json] [-o FILE] |
Print the ranked personal worklist (P1/P2/P3) on demand from the corpora, without a roll-up. See Dev-commitments. |
whosaid index <ws> [--no-embed] [--rebuild] |
Build <ws>/_search.db (full text, optional embeddings, entity graph) and <ws>/_WIKI.md. See Search your meetings. |
whosaid search <ws> "<query>" [--mode exact|meaning|hybrid] [--speaker S] [--meeting M] [-k N] [--json] |
Search every speaker-labeled transcript; each hit is a meeting, a timestamp, a speaker, and a snippet. |
whosaid context <ws> <meeting> <HH:MM:SS> [--before S] [--after S] [--json] |
The verbatim turns around a moment, for reading a hit in place. |
whosaid graph <ws> items|item AI-NNN|person [Name]|prs|meetings|speakers|commitments [--json] |
Entity views: action items with filters, one item's history, a person's commitments and asks, PR mentions, meeting coverage, talk share. commitments lists the unified CM table (--owner, --source transcript|action-items, --status filters) — one model behind the worklist, whosaid commitments, and the wiki since #23: the roll-up folds self-owned action-items bullets into the same CM-NNN corpus as the spoken clauses (one item, occurrences from both sources, the bullet's timestamp included), and graph build loads that corpus instead of re-parsing the bullets. |
whosaid wiki <ws> [--stdout] [-o FILE] |
Regenerate <ws>/_WIKI.md from the graph. |
whosaid watch run|install|uninstall|status --into <ws> |
Hands-free ingest of new recordings via a launchd agent. See Hands-free ingest. |
whosaid memos list|pull|delete "<title>"|shortcut-recipe |
Voice Memos helpers that never touch the app's files or database. See Voice Memos helpers. |
whosaid doctor |
Read-only environment report, including Ollama reachability and the index status of $WHOSAID_WORKSPACE. |
whosaid version | whosaid --version | whosaid -V |
Print the installed version (plus a git describe suffix when run from a git checkout). |
Already have a clip of the person speaking? Skip the mic and cut a reference straight from it:
whosaid enroll Alice --from meeting.m4a --ss 3:20 --t 20 # or --to 3:40One-off transcriptions are files; a recurring meeting series is a corpus. The meeting-workspace layer gives that corpus a home: every recording is transcribed into a dated folder, each meeting can carry generated action items and dev commitments, and one roll-up produces the index, the audit, a living action-item list, and the dev-commitments corpus across all of them — still entirely offline.
whosaid ingest weekly/*.m4a --into ./meetings --folder-by created \
--tz America/Los_Angeles --action-itemsEvery file is transcribed with the full set of whosaid <audio>… flags passed through, into
meetings/YYYY-MM-DD-HHMM/. The timestamp comes from the recording's own container
creation_time rendered in --tz (default UTC), falling back to the file's mtime — so folder
order reflects when meetings actually happened, not when you got around to copying the files.
Ingest is idempotent by source sha256: re-running a batch never re-transcribes or duplicates a
recording.
With --action-items, each meeting also gets an action-items.md. Generation is pluggable: the
hook command receives the speaker-labeled transcript on stdin, plus WHOSAID_SPEAKERS and
WHOSAID_TRANSCRIPT_PATH in its environment, and whatever it writes to stdout becomes the
markdown. Pass it per-run with --hook CMD, or set WHOSAID_ACTION_ITEMS_HOOK once. With no
hook, a skeleton is written instead and everything stays offline. Example hook — illustrative
only; any command that turns stdin into markdown works:
#!/bin/sh
# Drafts action items with a local Ollama model.
ollama run llama3.2 "List this meeting's action items as markdown bullets (Owner: task):"You no longer need a hook to get a real draft: --engine ollama (or the default --engine auto, which uses Ollama when it is running) turns on the built-in summarizer described in
Action items drafted by a local model; --engine claude
is the opt-in cloud alternative. --engine
implies --action-items. Add --index and, once every file is in, ingest runs
whosaid roll-up <ws> --action-items and whosaid index <ws> for you, so the new meetings are
searchable in the same command:
whosaid ingest weekly/*.m4a --into ./meetings --engine ollama --indexWith --commitments, each meeting also gets a commitments.md — the first-person commitments
you made in it (see Dev-commitments). Roles set once via
whosaid relabel --role decide whose cues count: tag your own voice self (and your manager
boss), and the extractor tracks what you promised, ranking boss-requested items higher.
--min-words N tunes the fragment filter for the run (fragments such as "I'll do that" are
dropped; default [commitments] min_words, else 2; 0 disables).
whosaid roll-up ./meetings --action-items- Coverage index —
_INDEX.mdholds one row per meeting (created date, duration, and whether it is transcribed, diarized, and has action items), plus a nothing-missing audit that flags orphan directories and stale manifest entries, and a recurring-topics section that surfaces themes appearing across meetings. Once any dev commitments exist, a one-line open/total count section is appended too. Both output paths are overridable with-oand--action-items-out. - Action-item corpus — with
--action-items,_ACTION-ITEMS.mddeduplicates items across meetings (by text similarity; threshold--similarity-threshold, 0.5–1.0, default 0.82) and groups them by owner, then status. Ids are stable (AI-001… and never renumber), each item carriesfirst_seen/last_seendates and its occurrence list, and open/ongoing/resolved statuses survive re-runs — so the corpus reads as living history across the series, not a per-meeting snapshot. Items can carry a free-form type, rendered in parens after the status, and pairs scoring just under the threshold are surfaced in a Possible duplicates (review) section at the end. - Dev-commitments corpus — each meeting's
commitments.json(fromingest --commitments) folds into_COMMITMENTS.mdautomatically whenever any exists: stableCM-NNNids that never renumber, the same 0.82 text-similarity dedupe and 0.10 near-miss review band, grouped by status then speaker, with**[boss]**marking boss-requested items. Hand edits in_COMMITMENTS.md(mark onedone, retitle, merge via(merged CM-NNN)) fold back on the next roll-up, exactly like the action-item corpus. The fold applies the same[commitments] min_wordsfragment rule as extraction, so per-meeting files written before the filter (or by a hook) never put "I'll do that" into the corpus: skipped entries are listed under Dropped fragments (review) and kept in_commitments.jsonasdropped;CMitems already in the corpus are never touched. - Personal worklist
_WORKLIST-<Owner>.md: whenever a commitments corpus or owner-attributed action items exist, roll-up also writes the owner's open items ranked into P1/P2/P3 (see Dev-commitments). The owner is--owner NAME, else theself-roled speaker, else[workspace] owner;--all-ownerswrites one file per participant. It is a regenerated view, never reconciled: edit the corpora, not the worklist.
Dedupe is text similarity (difflib, the --similarity-threshold above) and, when a local
Ollama answers on 127.0.0.1 with [search] embed = true, also embedding cosine: two texts
match when either the difflib ratio or the [search] embed_model cosine clears its threshold
([commitments] embed_threshold, default 0.90), so "I'll write the rollout runbook for the
platform team" and "for the platform team write the rollout runbook" fold into one item. No
Ollama, embed = false, or a non-loopback URL means difflib only, with one log line saying so.
Roll-up is incremental and append-only by default: re-running with nothing new writes nothing.
--rebuild is the escape hatch — it resets the manifest and corpus and regenerates both from the
folders on disk.
State is plain JSON in the workspace directory — _workspace.json (the manifest; it points at
the dev-commitments corpus via commitments_corpus) and _action-items.json (the corpus; it
also records the similarity_threshold in effect), plus the parallel _commitments.json.
All are safe to read and hand-edit — marking an item resolved by hand is the intended way to
close one the extractor phrased wrong. Hand edits made directly in _ACTION-ITEMS.md (and,
same contract, _COMMITMENTS.md) are folded back on the next roll-up and survive re-runs:
statuses, types, retitles, and (merged AI-NNN) merge annotations (the merged item stays at
its id, rendered collapsed as [merged → AI-NNN]). Only --rebuild discards them. --index
runs whosaid index <ws> right after the roll-up.
Many decisions and commitments are made in Teams chat, not in recorded calls. whosaid teams ingest adds a chat export to the same workspace. There is no diarization and no Whisper, because
every message already names its author.
whosaid teams ingest teams-export.json --into ~/meetings --tz America/Los_Angeles --self Alice_Example
whosaid roll-up ~/meetings --indexThe input is the output of the browser scraper in
contrib/teams-scraper/. That can be its raw_by_chat map,
a flat list or NDJSON of {chat, author, ts, epoch, text} records, or the same records with
timestamp_iso/epoch_ms. The scraper's README also covers how to run it and the Teams web DOM
contract it depends on. Ingest writes one dated folder per chat per local
day (YYYY-MM-DD-HHMM, where the time is that day's first message). Each folder holds:
teams.speakers.txt:[HH:MM:SS] Speaker: textturns, stamped with the local time of day.teams.diarization.json: the provenance sidecar.source.kindis"teams-chat", and each segment keeps the message's exactepoch_ms.
Search, the graph, commitments and the worklist all read teams.speakers.txt the same way they
read a recording's transcript. A search hit's file is teams.speakers.txt, and whosaid graph <ws> meetings --kind teams-chat (or the kind argument of the whosaid_meetings MCP tool)
tells chat days apart from recordings. A folder with no audio takes its created time from the
sidecar's source.creation_time. The roll-up audit treats it as text-only rather than flagging a
missing transcript.
Map Teams display names to the names you use elsewhere, and give people roles, in whosaid.toml:
[teams.names]
"Doe, Jane (Vendor, consultant)" = "Jane_Example"
[teams.roles]
Jane_Example = "boss"Unmapped authors keep their display name. When a role is not set in [teams.roles], it comes
from the voice registry's role for the same name. --self NAME overrides both. Re-ingesting an
overlapping export is safe: messages are deduplicated on epoch_ms and merged into the existing
folder, and --dry-run shows what would be written.
A workspace with a dozen meetings is a corpus you cannot reread. whosaid index turns it into
something you can query: an exact full-text index over every turn, optional meaning search
through a local embedding model, an entity graph (people, meetings, action items, timestamped
commitments, PR mentions), and a generated wiki. Everything lives in the workspace, everything is
rebuildable, and nothing leaves the machine.
whosaid index ./meetings # build _search.db + _WIKI.md (embeds if Ollama is up)
whosaid index ./meetings --no-embed # exact search only, no Ollama needed
whosaid search ./meetings "budget" # every turn that says budget
whosaid search ./meetings "who is blocked" --mode meaning
whosaid search ./meetings '"token exchange"' --speaker Bob_Example --meeting 2026-09-16 -k 5
whosaid context ./meetings 2026-09-16-0900 00:00:06 # the verbatim minute around a hit
whosaid graph ./meetings items --owner Alice_Example --status open
whosaid graph ./meetings item AI-003 # one item: occurrences + timestamped commitments
whosaid graph ./meetings person Bob_Example # talk share, commitments made, asks received
whosaid graph ./meetings prs | whosaid graph ./meetings meetings | whosaid graph ./meetings speakers
whosaid wiki ./meetings --stdout # the generated wiki, regenerated from the graph<ws> can be left out everywhere: the commands fall back to $WHOSAID_WORKSPACE, then to the
current directory when it holds _workspace.json or whosaid.toml. Any subfolder with a
*.speakers.txt counts as a meeting, dated or hand-named, so transcripts you made by hand are
indexed too.
Three search modes. --mode exact (the default) is SQLite FTS5: words, "quoted phrases",
AND/OR/NOT, and prefix*, and it needs nothing but Python. --mode meaning embeds the
query with nomic-embed-text through Ollama on 127.0.0.1:11434 and returns the nearest turns,
so "who is blocked" finds "I'm stuck until the review lands" with no shared words. --mode hybrid
runs both and fuses the rankings, which is usually what you want once the index has embeddings.
Every hit prints the same way, and --json gives you one object per hit with the same fields:
$ whosaid search ./meetings "budget"
whosaid: engine: exact
[2026-09-16-0900 @ 00:00:06] Bob_Example: I will review the »budget« spreadsheet before Thursday.
1 hit(s).
$ whosaid context ./meetings 2026-09-16-0900 00:00:06 --before 10 --after 10
whosaid: 2026-09-16-0900: 3 turn(s) between 0s and 16s
[00:00:01] Alice_Example: Let's start with the roadmap for next quarter.
[00:00:06] Bob_Example: I will review the budget spreadsheet before Thursday.
[00:00:14] Alice_Example: Great, and I will send the draft out to the team today.
Set Ollama up once if you want meaning search: brew install ollama && brew services start ollama && ollama pull nomic-embed-text. Without it, index skips the embedding pass with a note and
search runs exact-only; nothing fails.
Graph views. The entity tables are built from reliable signals only: the manifest, the
deduplicated action-item corpus, each meeting's action-items.md, and the transcripts
themselves. items filters by --owner, --requester, --status, and --type; item AI-NNN
shows one item's occurrences and every timestamped commitment behind it; person [Name] is the
per-owner view (talk share, what they committed to, what they asked others for); prs lists
pull-request numbers mentioned in speech or in the notes; meetings is coverage per folder; and
speakers is talk share across the workspace. Every view takes --json.
The wiki. _WIKI.md is a generated rollup: people, meetings, open items, and PR mentions,
each action item citing the [meeting @ time] it was committed at. It is regenerated by index
and by whosaid wiki, so it never drifts from the data. Edit the corpus, not the wiki.
Per-workspace settings: whosaid.toml. Optional, next to the transcripts. Every key has a
default; WHOSAID_OWNER, WHOSAID_OLLAMA, and WHOSAID_SUMMARIZER_MODEL override the matching
keys per run.
[workspace]
owner = "Alice_Example" # whose action items this workspace tracks
aliases = ["Ali", "Alicia"] # how the transcript may misspell the owner (default: first name)
[groups] # optional, ordered; drives the action-item sections
leadership = ["Bob_Example"]
team = ["Carol_Example", "Dan_Example"]
[summarizer]
engine = "auto" # auto | ollama | claude | hook | none (claude: opt-in, cloud)
model = "qwen2.5:14b"
timeout = 900 # seconds per model call
num_predict = 2048 # max tokens per model reply; caps a runaway generation
think = false # thinking models (qwen3, qwen3.x) skip reasoning; much faster
# per model: think = { "qwen3:14b" = true, default = false }
claude_model = "opus" # engine = "claude": any `claude --model` value
claude_timeout = 900 # seconds for the one claude call
claude_bin = "" # path to `claude` (default: WHOSAID_CLAUDE_BIN, PATH, ~/.local/bin)
fallback = "ollama" # engine = "claude" failed: "ollama" (if it is up) or "none" (skeleton)
[search]
ollama = "http://127.0.0.1:11434"
embed_model = "nomic-embed-text"
embed = true # false: exact search only, never contact Ollama
[commitments] # optional; ranks _WORKLIST-<Owner>.md. A list you set replaces
boss = [] # the default list, so omit a key to keep the built-in cues.
deadline_cues = ["today", "tonight", "tomorrow", "eod", "end of day", "this week", "next week"]
blocking_cues = ["blocking", "blocked", "urgent", "asap", "critical", "hotfix", "outage", "prod issue", "release blocker", "customer escalation"]
negators = ["not", "no", "non", "never", "isn't", "won't", "don't", "without", "hardly"]
embed_threshold = 0.90 # cosine at or above this is a duplicate (with [search] embed)
min_words = 2 # content words a commitment clause needs; 0 keeps every fragment
[commitments.weights] # score = sum of the signals that fired
boss = 5
blocking = 4
deadline = 4
overdue = 1 # replaces deadline once a relative deadline has passed
repeat = 2 # per extra meeting
recent = 1
strong = 1
requested = 1
negative = -3
[watch]
source = "" # folder to watch (default: the macOS Voice Memos store)
# speakers = 2 # exact diarization speaker count (a 1:1-call workspace)
# min_speakers = 2 # lower bound on auto-detected speaker count
# max_speakers = 4 # upper bound (a ceiling for a mixed-size meeting workspace)
# expected_speakers = ["Alice_Example", "Bob_Example"] # roster, or "Alice_Example,Bob_Example"What _search.db is. One SQLite file in the workspace: an FTS5 table of every turn (meeting,
timestamp, speaker, text), a table of stored embeddings keyed by turn, and the graph tables.
It is a derived artifact. Delete it whenever you like and run whosaid index again; --rebuild
does the same and also re-embeds every turn.
whosaid ingest --action-items can draft each meeting's action-items.md with a local model
instead of a hook. --engine (or [summarizer] engine in whosaid.toml) picks how:
| Engine | What happens |
|---|---|
auto (default) |
The hook if one is configured, else Ollama if it answers on localhost, else the skeleton. |
ollama |
The built-in summarizer below. If Ollama is down or the model is missing, it warns and writes the skeleton (exit 0). |
claude |
Opt-in, cloud. The transcript text goes to Anthropic. See The claude engine. auto never picks it. |
hook |
The --hook / WHOSAID_ACTION_ITEMS_HOOK command, exactly as before. |
none |
The skeleton only (speakers listed, no items). |
The model never gets to invent timestamps or evidence. The script picks the candidate turns
deterministically (every substantive turn by the workspace owner, every turn that names the
owner, plus one model pass over the rest for unnamed asks), then the model reads one turn at a
time (long turns in sentence groups) and answers either SKIP or bullets that each carry a short
verbatim quote. Every quote is checked against its turn before it is kept; a mismatch is flagged
rather than dropped silently, and the evidence turns are appended verbatim at the bottom of the
file.
Sections come from your config: one section per [groups] entry (asks from leadership, asks
from the team, and so on), team directives from leadership, the owner's own commitments, and an
inferred section for the rest. A team directive is a priority, deadline, process or expectation
that leadership sets for the whole team, such as "from now on every PR needs a linked ticket". It
counts as the owner's item even when the owner is never named and never speaks. Leadership is the
[groups] entry named leadership plus anyone the transcript role-tags boss; without either,
the section is left out. The
file opens with a DRAFT banner so a reader knows it has not been reviewed, and ends with an
evidence block of the exact turns each bullet came from. Read the evidence before folding
anything into _ACTION-ITEMS.md.
Model choice: the default is qwen2.5:14b (ollama pull qwen2.5:14b; about a minute for an
hour-long meeting on an M-series Mac). WHOSAID_SUMMARIZER_MODEL=qwen2.5:7b is roughly twice as
fast and roughly twice as noisy. Ollama is only ever contacted on 127.0.0.1.
engine = "claude" (or --engine claude) drafts the same file with Claude through the
Claude Code CLI you are already signed in to: one claude -p
call per meeting reads the whole transcript and returns the items as JSON. This sends the
transcript's text to Anthropic. Audio and voiceprints stay local. Only turn it on for
recordings you may send to a cloud model; auto never picks it.
The rest is the local engine's contract: the same sections (the plan decides them, not the model), the same guards (a team directive from someone outside leadership, or on a sentence that opens "Sam, ...", is dropped), and the same verification. The speaker and time of every item come from the transcript turn that holds its quote, a quote no turn holds is flagged ⚠, and the evidence block lists the verified turns verbatim. The note under the banner names the model and says the transcript went to Anthropic.
The call is isolated: no tools, no MCP servers, no settings, hooks, skills, or slash commands, no
saved session, an empty temp directory as its cwd, and ANTHROPIC_API_KEY /
ANTHROPIC_AUTH_TOKEN scrubbed so it runs on your Claude login, not an API key. Model-alias
overrides (ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_OPUS_MODEL, ...) and a parent Claude Code
session's CLAUDECODE / CLAUDE_CODE_* variables are dropped too, so claude_model alone picks
the model and a run from inside Claude Code matches the watcher's. The binary is
[summarizer] claude_bin, else WHOSAID_CLAUDE_BIN, else claude on PATH, else
~/.local/bin/claude, so the whosaid watch LaunchAgent finds it too. If the call fails, the
engine warns and falls back to the local Ollama engine when fallback = "ollama" and Ollama is
up, else writes the skeleton (exit 0). On the six eval fixtures the claude engine scores F1 0.960
in 6 calls, against 0.899 for the same Opus model run through the per-turn pipeline (84 calls) and
0.895 for qwen3:14b (docs/eval.md).
Record a meeting in Voice Memos, walk away, and have it transcribed, summarized, and searchable a
few minutes after you press stop. whosaid watch is a launchd LaunchAgent that watches a folder
(the macOS Voice Memos store by default) and pushes each new recording through
whosaid ingest --action-items --commitments, then roll-up (which also refreshes the
worklist), then index.
whosaid watch install --into ~/meetings --seed # install; mark the memos already there as done
whosaid watch install --into ~/meetings --dry-run # print the plist and every step; change nothing
whosaid watch status --into ~/meetings # installed? loaded? last run? (--json too)
whosaid watch run --into ~/meetings --dry-run # what WOULD be ingested right now
whosaid watch uninstall --into ~/meetings --purge # stop the agent, remove state + staging + interpreterHow a pass works: every audio file in the source folder that is not yet in
<ws>/.watch_state.json and whose mtime has been stable for [watch] stable_seconds (default
120) is copied into <ws>/.watch_staging/ and ingested from there; a file whose mtime is still
moving is treated as still recording or still syncing, and the pass waits (bounded by
max_wait_seconds) and rescans. A .watch.lock keeps two passes from overlapping. The agent
fires on folder changes (WatchPaths) and every --interval seconds (default 900) as a safety
net. --seed marks the recordings already present as done so an existing library is never
reprocessed. --source DIR watches any folder, not just Voice Memos; --offline bakes the
Hugging Face offline variables into the agent so a machine with no network access never tries to
fetch; --env K=V adds any other environment the agent should carry; --engine E on run picks
the summarizer. The recordings source and the meeting workspace (transcripts, corpora,
_search.db) must be separate folders — never the same directory, never nested inside each
other — or install/run refuse to start.
Diarization speaker hints. The watcher can never ask how many people were actually in the
room, and blind speaker auto-detect can over-split a long call into a pile of phantom speakers
when it should have found two. Set [watch] speakers = 2 for a workspace that only records 1:1
calls, or max_speakers as a ceiling for a workspace with mixed-size meetings; min_speakers
and expected_speakers (see the --expected-speakers flag reference below) are read from the same
block. These are workspace-wide, not per-recording — every ingested recording gets the same
hint — and take effect on the very next pass with no reinstall, since run re-reads
whosaid.toml each time it runs; install's --dry-run summary echoes the hints it would use.
An invalid value (a non-whole-number, min_speakers above max_speakers, an empty name) makes
both run and install refuse, naming the bad key.
Reinstall preserves saved environment settings, including custom tool paths. Current exported
certificate settings and explicit --env values can override saved values; --env takes final
precedence. On a managed network, the watcher can use the organization's installed macOS trust
roots or its CA bundle without turning off certificate verification:
whosaid watch install --into ~/meetings --no-fda --env UV_SYSTEM_CERTS=1
# If your organization supplies a PEM bundle instead:
whosaid watch install --into ~/meetings --no-fda \
--env SSL_CERT_FILE=/path/to/organization-ca.pem --env UV_SYSTEM_CERTS=1The installer also carries exported SSL_CERT_DIR, REQUESTS_CA_BUNDLE, CURL_CA_BUNDLE,
and the legacy UV_NATIVE_TLS setting into launchd. Certificate failures retain the original
error and provide this setup guidance. To discard saved environment settings completely,
uninstall the watcher, then install it with the desired settings. Uninstall leaves shared
user-local tools and CA files intact.
Full Disk Access, scoped to one binary. The Voice Memos store is TCC-protected: only an
executable you have granted Full Disk Access can read it, and macOS grants that per executable
path. Granting it to your terminal or your everyday Python would privilege everything they run.
So install provisions one dedicated, ad-hoc-signed interpreter at
$HOME/.local/opt/whosaid-watch/bin/whosaid-watch (named so that is what the FDA list shows),
points the agent at it, and opens the right System Settings pane. That interpreter runs nothing
but the watcher; it copies each recording out of the store and hands the copy to whosaid, so
whosaid, uv, ffmpeg, and the models read an ordinary file and need no grant at all. Adding
that one path is the single manual step. Already granted a binary? --interpreter PATH reuses
it instead of provisioning a new one. Several workspaces can coexist: the label defaults to
com.whosaid.watch.<8 hex of the workspace path>, or pass --label.
Logs land in <ws>/.watch.log (launchd captures the watcher's stdout and stderr there).
uninstall unloads and removes the agent; --purge also removes the dedicated interpreter and
the workspace's .watch_state.json, .watch.lock, and .watch_staging/, never the log.
The Full Disk Access toggle asks for an administrator's password — even though install scopes
the grant to one dedicated binary — and on a company-managed Mac the Privacy & Security pane may
be hidden entirely, so a standard user often cannot grant it at all. The escape hatch needs no
admin time: watch already reads any folder outside ~/Library with no FDA, and --no-fda
makes that the default (it watches ~/Recordings, creating it if needed, when no --source and
no [watch] source is set):
whosaid watch install --into ~/meetings --no-fda # watches ~/Recordings; no FDA, no adminThis mode reuses the launcher's Python and does not create a dedicated copied interpreter.
For Voice Memos FDA installs, the installer selects a Python that supports copied virtual
environments. If only Apple's Command Line Tools Python is available, install a compatible
user-local interpreter with uv python install 3.12, then rerun; --interpreter PATH still
selects an existing interpreter explicitly.
Three ways to get recordings into that plain folder:
- iPhone Shortcut "Record Audio -> Save File" -> Dropbox (the one fully automatic
option): in the Shortcuts app, chain the Record Audio and Save File actions with a
Dropbox folder (say
whosaid) as the destination, then bind the shortcut to the Action Button (Settings > Action Button > Shortcut) or Back Tap (Settings > Accessibility > Touch > Back Tap). Recordings sync to~/Dropbox/whosaidon the Mac. - QuickTime Player (built in, nothing to install): File > New Audio Recording, record,
stop, and save into
~/Recordings. - Drag and drop: drag memos out of the Voice Memos app into
~/Recordings, or from iOS use Share > Save to Files > the Dropbox folder. The watcher ingests them on the next pass like any other file.
Dropbox beats iCloud Drive here: ~/Dropbox is a plain folder outside ~/Library, while iCloud
Drive's on-disk root (~/Library/Mobile Documents) is TCC-adjacent and not verifiably readable
without FDA. The same applies to capture apps that sync via iCloud Drive (e.g. Just Press
Record): their files land under ~/Library, so they cannot feed the no-FDA watcher — an App
Store recorder is only an option if it writes into a folder you choose outside ~/Library.
The menu bar glyph works on a no-admin setup too, with one expected gap. Every glyph the
SwiftBar plugin (whosaid watch menubar) shows except 🔴 REC comes from watch status --json
and the workspace's .watch.log — no grant needed — so the idle/waiting/ingesting/done/
warn/not-loaded glyphs, the dropdown links and actions, and the transition notifications all
work untouched against a --no-fda watcher. 🔴 REC is the exception for two independent
reasons:
- REC is a Voice-Memos-store signal by design. "Recording right now" means a store row
with no synced file and 0 seconds; a plain folder has no in-progress signal, so a
--no-fdawatcher never shows REC even with every grant in place — absence of the dot is expected, not a failure. - SwiftBar itself holds a separate Full Disk Access grant, admin-gated exactly like the watcher's. Without it the plugin's store probe fails silently — the REC dot never appears, and the dropdown says so with a red "SwiftBar cannot read the Voice Memos store" line, a link to the grant pane, and a relaunch action. Everything else in the menu keeps working.
See Menu bar glyph for the full story, including why a fresh grant needs a relaunch.
whosaid watch menubar install [--workspace <ws>] [--interval 10s] symlinks a stdlib-only
SwiftBar plugin (contrib/swiftbar/whosaid.10s.py) into SwiftBar's plugin directory, so the
watcher's state lives in your menu bar at a glance; menubar uninstall removes it and
menubar status reports whether SwiftBar is installed and running and whether the configured
recording source is readable. whosaid doctor prints the same report. A watcher using an
ordinary folder, such as ~/Recordings, does not require a separate Voice Memos access grant.
If SwiftBar has no plugin directory preference, the installer configures ~/.config/swiftbar
and verifies the preference before linking the plugin. An existing directory preference is
preserved. Refresh or relaunch SwiftBar after installation to load the plugin; a successful
preference write and symlink do not by themselves prove the app has loaded it.
| glyph | meaning |
|---|---|
| 🎙 | idle, watching the store |
| 🔴 REC | a memo is being recorded right now |
| ⏳ 90s | new memo, waiting for the file to settle |
| ⚙️ 3m | transcribing / diarizing / action items (elapsed) |
| ✅ | ingested (held 30 min, then back to idle) |
| needs a look | |
| 🎙✗ | launch agent not loaded |
The dropdown carries the stage log tail, the agent line, links to the workspace, the latest
dated meeting folder, every _WORKLIST-<Owner>.md, and _WIKI.md, plus actions to follow the
log, kick the watcher now (launchctl kickstart -k), and refresh or relaunch SwiftBar.
Two platform notes the plugin cannot fix for you: when Voice Memos is the configured source,
reading its store requires
SwiftBar itself to hold Full Disk Access (the watcher's grant is separate — the plugin
shows the pane link and a relaunch action when the probe fails, because a fresh grant only
applies after a relaunch); and on a notched display a new status item can land behind the
notch — defaults write com.ameba.SwiftBar "NSStatusItem Preferred Position <plugin path>" -float <n> plus a relaunch fixes it.
whosaid memos list # titles and sync state (reads a COPY of the database)
whosaid memos pull --latest -o ./inbox # copy the newest recording out of the store
whosaid memos pull --title "Standup" -o . # or one by exact title
whosaid memos delete "Standup" --yes # delete through the app's own action
whosaid memos shortcut-recipe # build (or print how to build) the Shortcut delete usesdelete never touches CloudRecordings.db or the .m4a files. It runs the Voice Memos
"Delete Recordings" App Intent through a one-action Shortcut named "Delete Voice Memo", which is
exactly what tapping Delete in the app does: iCloud stays in sync across your devices, the memo
lands in Recently Deleted (restorable for 30 days), and the app's database stays consistent.
Editing the store behind the app's back desyncs iCloud and can corrupt the library, which is why
there is no --force path that does. shortcut-recipe builds and signs that Shortcut for you
when signing is available; on a Mac with no iCloud account it prints the one-action recipe to
build by hand in the Shortcuts app (--no-sign prints the recipe only).
A dev-commitment is a first-person promise the self-roled speaker made to someone else
("I'll send the migration plan by Friday"). Each meeting's transcript yields a
commitments.md + commitments.json; roll-up folds those into one living corpus,
_COMMITMENTS.md / _commitments.json — so "what did I promise across this whole series?"
has a single answer.
Extraction is a stdlib heuristic, no LLM and no network: a clause starting with a first-person
cue (i'll, i will, i plan to, let me, i owe, …) records the clause; negations
(i won't, i can't) are kept but flagged negative, and question clauses are skipped. Roles
drive it: with roles set only the self speaker's cues count, and when the turn before a
commitment was a different speaker asking or directing ("can you…", "please…"), the item
records requested_by — priority high when that speaker's role is boss. A pluggable hook
(--hook CMD on the commitments subcommand, or WHOSAID_COMMITMENTS_HOOK) can replace the
heuristic; it receives the transcript on stdin plus WHOSAID_SPEAKERS and WHOSAID_ROLES.
What gets dropped. A cue alone is not a commitment: "I'll do that", "I'll check", "I'll
bring that up" carry nothing to act on (the thing promised lives in the previous turn), so
they would only clutter the worklist as rows with no object that never dedupe against the real
item. The extractor therefore requires at least [commitments] min_words content words per
clause (default 2), where content words are what is left after removing a stoplist of pronouns
(including the cue itself), determiners, particles, prepositions, auxiliaries and fillers: "I'll
bump the version" has two (bump, version); "I'll see what I can do" has one; "I'll do that"
has none. One content word is enough when the item has a requested_by or names a deadline or
urgency (today, tomorrow, this week, by Friday, urgent, blocking), since the request
supplies the object: after "Can you own the rollout?", "I'll own it" is kept. Clause hygiene
runs first: a stuttered cue collapses ("I'll I'll start investigating that while I" becomes
"I'll start investigating that"), a clause ends at a subordinator (while, because, if,
when, unless, until, which, ...: "I can create an Epic if needed" becomes "I can create
an Epic"), and a dangling trailing pronoun or preposition is trimmed. Set min_words = 0 to
keep every clause, or pass --min-words N on whosaid ingest for one run.
Seeing the near misses. Dropped clauses are reported the way low-similarity merges are:
each meeting's commitments.md ends with a Dropped fragments (review) section (- (Alice_Example) I'll check @ 00:00:22 (fragment: 1 content word, min 2)) mirrored as dropped in
commitments.json, and the extractor logs dropped N fragment(s) (min_words=2). Roll-up
applies the same rule when folding, so older commitments.json files never seed fragments into
the corpus; its skips land under Dropped fragments (review) in _COMMITMENTS.md and persist
in _commitments.json. If a real commitment shows up there, lower min_words in
whosaid.toml and re-run (roll-up --rebuild re-folds every meeting).
The corpus works exactly like the action-item corpus: stable CM-NNN ids that never renumber,
the same 0.82 text-similarity dedupe with a 0.10 near-miss review band, and hand edits in
_COMMITMENTS.md folded back on the next roll-up. An item's rendered shape:
- **CM-002** [open] (Stephen) 2026-09-14-1802 → 2026-09-21-1802 (2×): **[boss]** update the exec deck
The corpus answers "what did I promise"; the worklist answers "what should I do first". Every
roll-up (and whosaid commitments <ws> on demand) takes the union of the owner's commitments
(CM-NNN, speaker = owner) and the action items assigned to them (AI-NNN, owner matched
case-insensitively with _/space interchangeable, plus [workspace] aliases), folds the two
sources together (an action item that restates a commitment appears once, as (also AI-012)),
and ranks what is open. Deterministic, no LLM:
| Tier | Rule |
|---|---|
| P1 | boss-requested (priority high, requester role boss, [commitments] boss, or a [groups] leadership name when no role is recorded), a blocking/urgency cue (blocking, urgent, asap, hotfix, outage, prod issue, release blocker, customer escalation, …), a deadline cue (tomorrow, eod, by Friday, next week, an ISO date, …) that has not already passed, or seen in 3+ meetings |
| P2 | seen in 2 meetings, requested by anyone, or a strong cue (i'll own, i will, i promise, i owe, …) in the latest meeting |
| P3 | the rest (weak cues such as i can, let me, i plan to earn nothing) |
A negated commitment (i won't) is never P1 and carries a penalty. Within a tier the order is
score (sum of the signal weights), then most recent, then id, and every line says why:
- **CM-002** [open] 2026-09-14-1802 → 2026-09-21-1802 (2×) P1 · boss · due=tomorrow · 2 meetings: update the exec deck (also AI-012)
Three rules keep the cues honest (GitHub issue #21):
- Negation. A cue with a negator in the three words before it, inside the same clause, does
not fire: "send non-urgent questions to the channel", "not blocking anyone", "no rush", "isn't
really critical", "won't block us" earn nothing, while "not done yet, this is urgent" still
does (the comma starts a new clause). The negator list is
[commitments] negators(defaultnot,no,non,never,isn't,aren't,wasn't,won't,don't,doesn't,didn't,without,nothing,hardly, plus the apostrophe-free spellings a transcript may produce). - Bare nouns are not emergencies.
prod,production,release,ship,customer(s)are not blocking cues by default: "post it in the release chat" and "shipping version 2.5" are plain work. The defaults carry phrases instead (prod issue,production is down,release blocker,blocking the release,before we ship,customer escalation,customer is waiting, …). A workspace where the bare noun really does mean pressure adds it back through[commitments] blocking_cues(the list you set replaces the default, so include what you keep). - Relative deadlines expire.
today,tonight,eod,tomorrow,this week,eow,next week,by Friday,before the 14th,on sept 3,by 9/20and ISO dates resolve against the date of the meeting the item was last seen in (this weekis that week's Friday, or the next Friday from a weekend;next weekthe Friday after that;the 14throlls to the next month once it has passed). While the resolved date is still ahead the line saysdue=<cue>and the deadline counts toward P1; once it is behind today the line saysoverdue=YYYY-MM-DD, the smalloverdueweight (default 1) replacesdeadline, and the item is no longer P1 on the deadline alone, so a "today" from two weeks ago sits below this morning's real work. Cues with no calendar meaning (this sprint,before the demo,before the release) never expire.WHOSAID_TODAY=YYYY-MM-DDpins "today" for tests and replays.
Resolved and merged items sit under ## Done / history. The file is a regenerated view (hand
edits are overwritten on the next run; ids never renumber; the JSON corpora stay the source of
truth), so close or retitle items in _COMMITMENTS.md / _ACTION-ITEMS.md. The owner is
--owner NAME, or --owner me (the default): the speaker with role self in the newest
meeting's commitments.json or # Role: header, else [workspace] owner / WHOSAID_OWNER.
--all-owners writes one file per participant. whosaid commitments <ws> --json emits
{owner, generated_from, items: [{id, source, text, status, tier, score, why, first_seen, last_seen, occurrences, requested_by, negative, …}]} for scripts and the MCP tool. Cue lists,
boss names, weights, and the embedding threshold are overridable in whosaid.toml
[commitments] (see the example under Search your meetings).
whosaid's local, private, GPU transcription and speaker diarization are also exposed as MCP tools,
so any MCP client — Claude Code, Claude Desktop, and others — can call them directly instead of
shelling out to the CLI. Launch is whosaid mcp, a stdio server that needs only uv, which whosaid
already requires; audio never leaves the machine, exactly as with the CLI.
Add it to your MCP client config:
{
"mcpServers": {
"whosaid": { "command": "/ABSOLUTE/PATH/TO/whosaid", "args": ["mcp"] }
}
}command is the path to the whosaid script itself — the checkout's ./whosaid, or the installed
~/.local/bin/whosaid symlink.
| Tool | What it does |
|---|---|
whosaid_transcribe |
Transcribes + diarizes an audio file and writes the labeled transcript, speaker cards, and a sidecar for relabeling. |
whosaid_relabel |
Names SPEAKER_NN clusters and remembers them — saved to the local registry and auto-applied to every future transcript. Optional roles ({Name: role}) tags speakers (self/boss/peer/report/external). |
whosaid_list_speakers |
Read-only: lists enrolled voices and registry names (with role tags) already known. |
whosaid_doctor |
Read-only readiness check — models cached, deps present, mic/audio devices — run this first when a transcribe fails. |
whosaid_enroll_from_file |
Enrolls a named voice from an existing audio clip, no mic needed. |
whosaid_samples |
Exports one short representative WAV per speaker cluster (longest segment, clamped to seconds) so you can listen and confirm an identity before trusting a label. |
enroll and record (microphone capture) stay CLI-only — they need an interactive terminal and
Microphone permission. The first whosaid_transcribe call downloads ~1.5 GB of models; call
whosaid_doctor first to check readiness.
The workspace search layer is exposed too, so an agent can answer "what
did we decide about the budget" from the transcripts instead of reading them whole. Every tool
takes an optional workspace argument and otherwise uses WHOSAID_WORKSPACE from the server's
environment; there is no current-directory fallback. None of them rebuilds the index: whosaid index (or the watcher) owns writes, and the tools only read _search.db.
| Tool | What it does |
|---|---|
whosaid_search |
Turn-level hits (meeting folder, timestamp, speaker, snippet) for a query in exact, meaning, or hybrid mode, with speaker/meeting filters. |
whosaid_context |
The verbatim turns around one hit, so an agent reads a minute instead of a transcript. |
whosaid_items |
The action-item corpus, filterable by owner, requester, status, and type. |
whosaid_item |
One item with its occurrences and timestamped commitments. |
whosaid_person |
A person's talk share, the commitments they made, and the asks they received. |
whosaid_meetings |
Every meeting folder with coverage figures. |
whosaid_prs |
Pull-request numbers mentioned in speech or in the notes. |
whosaid_speakers |
Talk share per speaker across the workspace. |
whosaid_workspace_status |
Whether the index exists, how many turns it holds, and whether Ollama is reachable. |
whosaid_worklist |
The owner's ranked worklist (P1/P2/P3 with score and why) from the commitments and action-item corpora, the whosaid commitments --json payload; owner defaults to me. tier (P1/P2/P3), limit, and offset filter and page it (tier first, then offset, then limit); compact=true keeps only id, source, text, tier, score and why per item; with neither limit nor offset the top 50 items come back plus total/omitted and a paging hint. |
Resources, all reading WHOSAID_WORKSPACE: whosaid://workspace/wiki (_WIKI.md),
whosaid://workspace/action-items (_ACTION-ITEMS.md), whosaid://workspace/index
(_INDEX.md), and per meeting whosaid://workspace/meeting/{folder}/transcript and
whosaid://workspace/meeting/{folder}/action-items.
Register it with the workspace set. Claude Code:
claude mcp add --scope user whosaid -e WHOSAID_WORKSPACE=$HOME/meetings -- whosaid mcpKiro (~/.kiro/settings/mcp.json) and any client that takes the same JSON shape:
{
"mcpServers": {
"whosaid": {
"command": "whosaid",
"args": ["mcp"],
"env": { "WHOSAID_WORKSPACE": "/Users/<you>/meetings" }
}
}
}Write the workspace as an absolute path (most clients do not expand $HOME inside env), and use
the absolute path to the whosaid script for command if it is not on the client's PATH.
| Flag | Meaning |
|---|---|
-o, --outdir DIR |
Output directory (default: alongside the input file). |
-m, --model NAME |
Whisper model to use. |
--accurate |
Use the full large-v3 model instead of the default large-v3-turbo. |
-l, --lang LANG |
Force the transcription language. |
-f, --format FMT |
Output format: txt, srt, vtt, tsv, json, or all. |
-n, --name NAME |
Override the output base name (single input only; default: derived from the input filename). |
--speakers N |
Exact number of speakers. Overrides auto-detect (and any --min/--max-speakers). |
--min-speakers N |
Lower bound on the auto-detected speaker count. |
--max-speakers N |
Upper bound on the auto-detected speaker count; also lowers the hard cap of 20. Use a range when you know roughly who was in the room but not exactly — an estimate that lands on the bound is reported as unreliable rather than shipped as a result. |
--expected-speakers A,B,C |
Roster of people you expect in this recording (comma-separated, repeatable). Each name must already be a known voice — an enrolled clip or a registry entry. Clustering is anchored to their voiceprints: a turn within --anchor-threshold of one of them is pinned to that person, everyone else is clustered into new speakers, and a listed person who never speaks is dropped. This is the flag for a recurring team with varying attendance, where --speakers N would merge distinct people. |
--anchor-threshold F |
Per-turn cosine required before --expected-speakers pins a turn to a known voice. Default 0.70; turns below it fall through to ordinary clustering rather than taking a low-confidence name. |
-j, --jobs N |
Parallel diarization workers for long audio (default: auto, ~cores−2, capped at 8). |
--chunk-seconds S |
Window length for parallel diarization (default: auto — about --jobs windows, min 300s). |
--no-chunk |
Diarize the whole file in a single pass (disable parallel chunking). |
--match-threshold F |
Cosine similarity a known voice must reach before it may claim a cluster (alias --ref-threshold). Default 0.50; a cluster whose best candidate scores below F keeps its anonymous SPEAKER_NN label rather than taking a low-confidence name. Raise it (e.g. 0.6) if you see wrong names, lower it to catch more. |
--absorb-threshold F |
Cosine similarity at which a still-unnamed cluster is folded into a known voice, merging phantom splits of one person. Default 0.85. |
--no-diarize |
Skip diarization; write the plain transcript only. |
| Variable | Purpose |
|---|---|
WHOSAID_PYTHON |
Python 3.11+ interpreter for the helper scripts (default: auto-detected, see Quickstart). |
WHOSAID_MODEL |
Default Whisper model, overridden by -m. |
WHOSAID_LANG |
Default transcription language, overridden by -l. |
WHOSAID_VOICE_REFS |
Override the directory of enrollment voice clips (default: voices/). |
WHOSAID_SPEAKER_DB |
Local speaker registry of named voiceprints (default: ~/.config/whosaid/speakers.json). Private, never pushed. |
WHOSAID_MATCH_THRESHOLD |
Default registry/reference match threshold, overridden by --match-threshold (default: 0.50). |
WHOSAID_ABSORB_THRESHOLD |
Default absorb-pass threshold, overridden by --absorb-threshold (default: 0.85). |
WHOSAID_ANCHOR_THRESHOLD |
Default per-turn anchoring threshold for --expected-speakers, overridden by --anchor-threshold (default: 0.70). |
DIARIZE_EMB_NAME |
Speaker-embedding model. Default is NeMo nemo_en_titanet_small.onnx (English-native, ~2.5× faster than ERes2Net in sherpa's benchmark). Alternatives from the same release: 3dspeaker_speech_eres2net_sv_en_voxceleb_16k.onnx (English ERes2Net) or …_zh-cn_… for Mandarin. Registry voiceprints are keyed by model, so switching re-enrolls speakers. |
WHOSAID_REC_DEVICE |
avfoundation input device used by record and enroll. |
WHOSAID_INSTALL_DIR |
Command install directory used by whosaid install (default: ~/.local/bin). |
WHOSAID_ACTION_ITEMS_HOOK |
Default action-items hook for ingest --action-items (transcript on stdin, markdown on stdout), overridden by --hook. |
WHOSAID_WORKSPACE |
Default meeting workspace for index, search, context, graph, wiki, watch, for the index-status line in whosaid doctor, and for the read-only workspace tools and resources of whosaid mcp. |
WHOSAID_OWNER |
Whose action items a workspace tracks; overrides [workspace] owner in whosaid.toml. |
WHOSAID_TODAY |
YYYY-MM-DD the worklist treats as today when deciding whether a relative deadline (today, tomorrow, by Friday, …) is overdue (default: the clock). |
WHOSAID_OLLAMA |
Ollama base URL for embeddings and the built-in summarizer (default: http://127.0.0.1:11434). Localhost is the only supported destination. |
WHOSAID_SUMMARIZER_MODEL |
Ollama model for --engine ollama (default: qwen2.5:14b); overrides [summarizer] model. |
WHOSAID_CLAUDE_BIN |
The claude binary for --engine claude when [summarizer] claude_bin is unset (default: claude on PATH, else ~/.local/bin/claude). |
WHOSAID_BIN |
The whosaid command the watcher runs (default: the script that launched whosaid watch). |
WHOSAID_COMMITMENTS_HOOK |
Default hook for dev-commitments extraction (ingest --commitments), overridden per-run by --hook. Receives the transcript on stdin plus WHOSAID_SPEAKERS/WHOSAID_ROLES. |
HF_HOME |
Hugging Face cache location (where the Whisper model lands). |
SHERPA_DIARIZE_CACHE |
Diarization model cache location (default: ~/.cache/sherpa-diarization). |
whosaid has two independent ways to put a name on a SPEAKER_NN cluster, and knowing which one
is driving a given label matters — this was GitHub issue #1's other point of confusion.
1. Enrollment clips (voices/, WHOSAID_VOICE_REFS). Every *.wav/*.m4a file in the
directory is embedded fresh on every transcribe and passed to the diarizer as --ref Name=path (the --ref "$refname=$ref" loop in the whosaid script, run before each diarize
call). Written by whosaid enroll [Name] (mic, or --from FILE to cut a clip from an existing
recording). Because a clip is re-embedded every run, it survives a DIARIZE_EMB_NAME change
with no extra step: the embedding is always current for whichever model is active.
2. The registry (~/.config/whosaid/speakers.json, WHOSAID_SPEAKER_DB). A JSON file of
{name, model, embedding, added} records, written once by whosaid relabel SPEAKER_XX=Name
(or --save-speaker on a transcribe) and reused forever with no re-embedding. Each entry is
keyed by the embedding model it was saved under, so switching DIARIZE_EMB_NAME silently
orphans every registry entry for the old model — they don't error, they just stop matching —
until you relabel again under the new model.
Role tags. A registry entry may carry an optional "role" string — the conventional set is
self, boss, peer, report, external, and free-form lowercase tags are allowed. Set one
with whosaid relabel <base> --role Karen=boss (repeatable) or the roles map on the MCP
whosaid_relabel tool; it is stored on the registry entry, preserved when the voiceprint is
re-saved, and rendered as # Role: Karen = boss header lines in <base>.speakers.txt, a
Karen [boss] label on the speaker card, and a top-level "roles" key in the sidecar.
Roles carry semantics downstream: self marks your own voice, and a boss-roled speaker's
requests rank higher in action items and the dev-commitments corpus.
Apply registry changes to historical meetings. Relabeling a meeting updates that meeting; use backfill to bring an existing workspace into line with the current registry:
whosaid backfill ~/meetings --dry-run # inspect per-meeting changes and conflicts
whosaid backfill ~/meetings # apply, then refresh all derived viewsBackfill works from cached diarization sidecars without transcribing audio again. It preserves transcript prose and transcript-only labels, updates speaker role headers and cards, and regenerates heuristic commitments. Conflicting edits to a meeting's commitments are reported before meeting files are written; reconcile those edits before retrying. Corpus curation, including edited commitment text and statuses, survives source refresh. Roll-up fingerprints commitment inputs and roles so previously folded meetings are reconsidered when those inputs change. Run backfill after changing registry roles to update the per-meeting inputs as well.
Back up or move the registry. Both commands respect WHOSAID_SPEAKER_DB and work without
loading a model:
whosaid speakers export --out ./speakers-backup.json
whosaid speakers import ./speakers-backup.json # destination must not exist
whosaid speakers import ./speakers-backup.json --merge # add non-conflicting entriesExports preserve names, model identifiers, embeddings, roles, notes, and other registry metadata
inside a versioned format. Import validates the complete file before writing and refuses merge
conflicts. To deliberately replace every destination entry, use --overwrite; this also removes
voices absent from the backup. Export refuses an existing output path. Registry writes are
atomic and backup/import files have owner-only permissions. These files contain private
voiceprints: keep them out of source control and use a private, encrypted transfer when moving
them between machines. Enrollment audio in voices/ is separate from the registry backup.
Which one actually names a cluster. name_clusters() (lib/diarize_sherpa.py) runs three
passes in order, and a later pass only touches a cluster the earlier ones left unnamed: (1)
registry one-best, (2) --ref enrollment clips, (3) absorb (folds a still-unnamed cluster into
a known voice — registry or clip — above --absorb-threshold). Passes 1 and 2 both gate on
--match-threshold (default 0.50, env WHOSAID_MATCH_THRESHOLD): below it a cluster keeps
its anonymous label rather than take a low-confidence name.
Inspecting each. ls voices/ lists enrollment clips; whosaid doctor prints the registry's
voiceprint count and path; the MCP whosaid_list_speakers tool lists both at once.
<base>.diarization.json's registry_matches records which pass named (or nearly named) every
cluster, so you can tell exactly which mechanism fired — and whosaid samples <base> lets you
listen to a clip per cluster to confirm a label before trusting either one.
Transcript-only labels (relabel --no-save). Naming a cluster doesn't have to mean rewriting the
registry. whosaid relabel <base> SPEAKER_04=Alice_Example --no-save [--note TEXT] renames the
cluster in the transcript and cards only and never writes to the registry — what you want when the
conversation makes the speaker obvious but the cluster doesn't represent one clean voice. That is
GitHub issue #19's case: in a 7-person meeting, SPEAKER_04 was identifiable as Alice from what she
said, but her 176-turn cluster was a mixture of voices, and saving it would have replaced her
enrolled print with a degraded one. The label is stored in the sidecar under local_labels (with the
optional --note alongside it), so it survives relabel --auto and whosaid samples, and each
speaker card reads Alice_Example (transcript-only label; registry untouched) so the label's
provenance is never in doubt. Related guard, on the other side of the same decision: the
registry-writing paths (relabel with an assignment, transcribe --save-speaker) now refuse to
replace an existing print whose similarity to the new cluster is below the match threshold (default
0.50) — pass --force to overwrite deliberately, or --no-save to label this transcript only.
audio file
|
v
MLX Whisper (Metal GPU) -------------> <base>.txt / .srt / .vtt / .tsv / .json
|
v
sherpa-onnx diarization (CPU) --------> <base>.rttm
|
v
cosine-match vs voices/*.wav ---------> <base>.speakers.txt
Transcription and diarization run as two independent local stages that get merged at the end.
Transcription uses MLX Whisper (mlx-community/whisper-large-v3-turbo by default, or the full
large-v3 with --accurate) on the Mac's GPU via Metal, through a hallucination-hardened decode
path: the temperature-fallback ladder stays enabled, condition_on_previous_text is turned off, and
a hallucination-silence threshold keeps dead air from turning into repeated-token filler. Diarization
runs on the CPU via sherpa-onnx offline diarization: pyannote's segmentation-3.0 ONNX model finds
who's speaking when, a speaker-embedding model (NeMo TitaNet-small by default, configurable via
DIARIZE_EMB_NAME) embeds each turn, and clustering (optionally hinted by
--speakers N) groups turns into speakers. Both diarization models are small (~30 MB total), ungated
GitHub releases — no Hugging Face token required — cached locally in ~/.cache/sherpa-diarization/.
Finally, every clip in voices/ — plus every voiceprint in your local registry — is embedded the
same way and matched to a cluster by cosine similarity (a match at or above --match-threshold,
default 0.50, names the cluster);
unmatched clusters keep a SPEAKER_00-style label. Aside from the one-time model downloads,
everything runs in ephemeral uv environments, so there's no persistent Python install left behind
on your machine.
Tell it who to expect. When you already know the roster, --expected-speakers Alice,Bob,Carol
turns diarization from a guess into a lookup. Every listed name must already be a known voice (an
enrolled clip in voices/, or a registry entry), and clustering is then anchored: each turn's
voiceprint is compared to each expected person's, a turn at cosine ≥ --anchor-threshold (default
0.70) is pinned directly to that person, and only the residual turns — guests, strangers, and
the occasional turn that scored low — go through the ordinary count estimate and clustering. Anyone
on the list who never speaks is silently dropped, so one roster works for a standing meeting whose
attendance changes week to week. That is the difference from --speakers N: an exact count merges
two distinct people the moment fewer than N show up, and auto-detect over-segments and then leaves a
major speaker's averaged cluster unmatched. The threshold sits deliberately high — TitaNet-small
scores same-speaker turns at ~0.90 (p10 ~0.72) against ~0.25 for strangers (p90 ~0.45), and a false
anchor puts a real name on the wrong words, while a missed one merely falls through to clustering
where the absorb pass can still recover it. Each anchor's result (turns, mean_cosine) lands in
the sidecar under anchors, and each anchored cluster in registry_matches with pass: "anchor".
Because anchoring needs per-turn voiceprints, this flag forces the chunked diarization path at any
recording length.
Long recordings run in parallel. For audio over ~15 minutes, diarization splits into
non-overlapping time windows that are segmented and embedded concurrently across CPU workers, then a
single global clustering pass over every turn's voiceprint recovers speakers that stay consistent
across window boundaries. Because the clustering sees all turns at once, it recovers the same
speakers as a single-pass run while finishing several times faster. Pass --no-chunk to force a single pass, or
--jobs/--chunk-seconds to tune it.
| File | Contents |
|---|---|
<base>.txt |
Plain transcript. |
<base>.srt |
SubRip subtitles. |
<base>.vtt |
WebVTT subtitles. |
<base>.tsv |
Tab-separated segments with timestamps. |
<base>.json |
Full Whisper segment output. |
<base>.rttm |
Raw diarization turns, standard RTTM format. |
<base>.speakers.txt |
Speaker-labeled transcript: Whisper text merged with diarization turns and enrollment names; carries # Role: NAME = ROLE header lines when speakers have role tags. |
<base>.speaker-cards.txt |
One card per speaker with turn count, talk time, and representative snippets — read it to identify who each SPEAKER_NN is, then name them with whosaid relabel. |
<base>.diarization.json |
Cached segments + per-cluster voiceprints, so whosaid relabel can rename and persist speakers without re-diarizing. Also carries registry_matches and source (below). |
<base>.samples/ |
Created on demand by whosaid samples <base>: one short representative WAV per speaker cluster (SPEAKER_NN[-Name].wav), for a quick human listen before trusting an auto-label. |
A meeting workspace (whosaid ingest --into <ws>) adds these at the workspace root. Underscored
files are generated and safe to delete; the dotfiles are the watcher's working state.
| File | Contents |
|---|---|
<ws>/YYYY-MM-DD-HHMM/ |
One meeting: the source audio, transcript.* outputs as above, and action-items.md / action-items.json when requested. |
_workspace.json |
The manifest: one entry per ingested meeting with its source sha256 (what makes ingest idempotent). |
_INDEX.md |
Coverage index and nothing-missing audit, from whosaid roll-up. |
_ACTION-ITEMS.md, _action-items.json |
The deduplicated action-item corpus and its state (stable AI-NNN ids, statuses, occurrences). Hand-editable. |
_search.db |
SQLite: FTS5 over every turn, stored embeddings, and the entity graph tables. Rebuilt by whosaid index; delete freely. |
_WIKI.md |
The generated wiki, regenerated by whosaid index and whosaid wiki. |
whosaid.toml |
Optional per-workspace settings (owner, groups, summarizer, search, watch). The one file here you write by hand. |
.watch_state.json, .watch.lock, .watch_staging/, .watch.log |
whosaid watch state: recordings already handled, the overlap guard, copies staged out of the Voice Memos store, and the agent's log. |
The sidecar's machine-readable extras, so a consumer never has to scrape stderr or shell out to
ffprobe:
| Sidecar key | Contents |
|---|---|
registry_matches |
One record per naming decision, including near-misses: {"cluster": "SPEAKER_03", "name": "Alice", "similarity": 0.919, "threshold": 0.5, "matched": true, "pass": "registry"}. pass is registry, ref, absorb, or anchor; matched: false means that decision did not identify the voice. After an anonymous fold, cluster points to the retained cluster and optional original_cluster records which original cluster supplied the similarity evidence. Refreshed by whosaid relabel --auto, and also printed in the transcribe JSON line. |
source |
Recording provenance: {"path": "/abs/path.m4a", "duration_seconds": 1834.2, "creation_time": "2026-09-14T18:02:11.000000Z"}. creation_time is the container tag, or null when the file carries none. |
roles |
Optional — present only when at least one speaker carries a registry role: {"Karen": "boss"}. Read by downstream consumers such as the commitments extractor. |
count_before_fold, count_after_fold, fold_note |
Physical speaker-cluster counts before and after the anonymous-phantom check, plus a note when it ran. Equal counts and fold_note: null mean no fold was applied. These counts do not claim that the remaining speakers were identified. Older sidecars are read as their observed num_speakers for both counts. |
In a meeting workspace, ingest --commitments additionally writes per-meeting
commitments.md / commitments.json, which roll-up folds into the workspace-level
_COMMITMENTS.md / _commitments.json corpus — see Dev-commitments.
Match confidence and the threshold. Auto-naming only asserts a name when the cluster's cosine
similarity to a known voiceprint reaches --match-threshold (default 0.50); below it the cluster
keeps its SPEAKER_NN label — see the >= ref_threshold guards in name_clusters()
(lib/diarize_sherpa.py). The default was raised from 0.40 to 0.50 because on real meeting
audio TitaNet-small produced wrong assertions in the 0.40–0.53 band, while genuine same-speaker
matches score far higher — in the end-to-end test the enrolled reference matches its cluster at
0.986, against 0.194 for the nearest stranger, so 0.50 sits in a wide empty gap. Use
registry_matches to see exactly how close every near-miss came, then lower the threshold
deliberately if a real speaker is being missed.
whosaid: command not foundafter it used to work.whosaid installmakes~/.local/bin/whosaida symlink to the checkout — it doesn't copy the implementation. If that checkout lived in a temp directory (/tmp,/private/tmp,$TMPDIR, e.g. a clone under/private/tmp/whosaid-latest/), the OS eventually purges it and the symlink starts pointing at nothing, so the shell reports a plain "command not found" with no hint why.whosaid install(and./bootstrap.sh) now refuse to install from a temp-directory checkout in the first place unless you pass--force. To fix an already-dangling install: clone/move the checkout to a stable location (e.g.~/code/whosaid) and runwhosaid installagain from there. Because the installed symlink itself is what's broken,whosaid doctorcan't be run by name to confirm this — run./whosaid doctorfrom the fresh checkout instead; it reports aDANGLINGinstall line when it finds this.- Enroll/record produces silence. This is almost always a macOS microphone permission problem: go to System Settings → Privacy & Security → Microphone and grant access to your terminal app. macOS feeds an unauthorized app silent zeros instead of an error, so whosaid detects this by checking the captured volume rather than trusting a clean exit code.
- Doesn't run on my Intel Mac. MLX is Apple-Silicon-only, so whosaid requires an
arm64Mac. - First run is slow. The first
./bootstrap.sh(or first transcribe, if you skip it) downloads the Whisper model (~1.5 GB) and the diarization models (~30 MB). Every run after that uses the local cache. - Long recordings degrade into repeated text. This is Whisper's well-known
hallucination/repetition-collapse failure mode, most likely on long or low-signal audio. It's why
whosaid calls the
mlx-whisperlibrary directly (lib/transcribe_mlx.py) instead of the bare CLI: the bare CLI's single-temperature default is exactly the configuration that lets this happen, whereas the library call keeps the temperature-fallback ladder, disables conditioning on previous text, and applies a hallucination-silence threshold. - Speakers show up as
SPEAKER_00/SPEAKER_01instead of a name. No enrolled voice or registry entry matched closely enough. Read<base>.speaker-cards.txtto tell who each cluster is, then runwhosaid relabel <base> SPEAKER_01=Name— this labels them and remembers them for next time. (Enrollment viawhosaid enroll <Name>still works too.) Naming uses a cosine-similarity threshold (--match-threshold, default 0.50), so a short or noisy sample can fall just short of it. Checkregistry_matchesin<base>.diarization.jsonfor the exact similarity of every near-miss, then lower the threshold deliberately if a real speaker is being missed. - The same person shows up as two speakers (phantom split), or a known voice stays
UNIDENTIFIED. On long recordings the diarizer can split one voice across several clusters. A registry/enrolled voice names its single closest cluster, so the extra clusters used to stay unnamed. whosaid now runs an absorb pass: any still-unnamed cluster whose voiceprint is within--absorb-threshold(default 0.85) of a known voice is folded into that person, and the speaker cards merge those clusters into one card. To apply this to a transcript you already have, runwhosaid relabel <base> --auto— it re-names from the registry + absorb pass with no re-transcription. - Distinct people get merged into one speaker (or the count is too low). The speaker-embedding
model must match the spoken language. whosaid defaults to an English-native model (NeMo
TitaNet-small); on English audio the Mandarin-trained model cannot tell similar voices apart and
collapses them. For
predominantly Mandarin audio, set
DIARIZE_EMB_NAMEto the…zh-cn…model from the same release. - The speaker count looks one too high, with a cluster that has ~1 second of speech. Forcing
--speakers Ntoo high can carve a phantom cluster out of crosstalk. Omit--speakersto auto-detect, which is usually more accurate. - Auto-detect reports far too many speakers on a long recording (it used to always say 20).
Fixed in the current version. The speaker count is now estimated by average-linkage
agglomerative clustering of the per-turn voiceprints, cut at cosine
0.58(envWHOSAID_COUNT_THRESHOLD; calibrated against real TitaNet-small embeddings, where same-speaker turns measure ~0.90 and different speakers ~0.25). The previous estimator compared each turn to a single earlier turn rather than to a cluster, so on any recording with real channel variation it kept opening new clusters until it pinned at the hard cap of 20. The cap is now a bound, not a target: if the estimate lands on it, the count is reported as untrustworthy — aWARNING:line is written into<base>.speaker-cards.txtright under the count. Suspicious estimates first try stable clustering cuts using turns of at least one second, then assign the shorter turns. A short turn whose voiceprint is too far from those substantive voices prevents that fallback.count_estimatepreserves the primary count and recovery evidence;count_warningexplains the action taken or remaining uncertainty. On a problematic fresh auto run, whosaid also checks anonymous phantom clusters for repair after clustering and recordscount_before_fold,count_after_fold, andfold_notein the sidecar and MCP result. For an existing cached auto sidecar, runwhosaid relabel <base> --auto --fold-unknown; it only merges confident anonymous clusters and does not give them names. Cached repairs preserve their original similarity evidence infold_evidence, so repeating the command cannot create a chain of merges through averaged voices. Named speakers and transcript-only labels are protected. If you know roughly how many people were present, pass--min-speakers/--max-speakers; if you know exactly, pass--speakers N. Note this estimator runs on the chunked path (recordings over 15 minutes, or any--chunk-seconds); shorter recordings use sherpa's own clustering, which takes an exact count only, so a--min-speakers/--max-speakersrange there is reported as unenforced unless the two are equal. whosaid search --mode meaningsays Ollama is unreachable, orindexskipped embeddings. Meaning and hybrid search need Ollama withnomic-embed-texton127.0.0.1:11434(brew install ollama && brew services start ollama && ollama pull nomic-embed-text, or pointWHOSAID_OLLAMAat where it listens). Nothing else breaks:indexstill builds the full-text tables andsearchfalls back to exact mode, andwhosaid doctorprints an Ollama line so you can see which state you are in. Once Ollama is up, runwhosaid index <ws>again to embed the turns that were skipped.no search index at <ws>/_search.db; run: whosaid index <ws>. Search, context, the graph views, and the MCP workspace tools only read the index; none of them builds it. Runwhosaid index <ws>once (and again after adding meetings, or useingest --index/roll-up --index/ the watcher so it stays current). The same message fromwhosaid doctormeansWHOSAID_WORKSPACEpoints at a workspace that has not been indexed yet.whosaid watchnever sees new memos, orwatch runexits 3 with "cannot read the source". The dedicated interpreter does not have Full Disk Access yet. Open System Settings, Privacy & Security, Full Disk Access, and add the exact pathwhosaid watch installprinted (default$HOME/.local/opt/whosaid-watch/bin/whosaid-watch), thenlaunchctl kickstart -k gui/$(id -u)/<label>or just wait for the next interval.whosaid watch status --into <ws>shows the label and whether the agent is loaded;<ws>/.watch.loghas the pass-by-pass detail. Granting your terminal FDA instead would work but privileges everything you run from it, which is exactly what the dedicated binary avoids.
./test/e2e.sh is a fully offline smoke test: it synthesizes a two-speaker dialog with two macOS
say voices, builds a one-clip voice enrollment for one of them, and runs the real transcribe +
diarize + name pipeline against it end to end — then asserts the speaker-labeled transcript names
the enrolled speaker and labels the other speaker distinctly. Run ./bootstrap.sh once first so the
models are cached locally; the test itself makes no network calls.
./test/version_test.sh is a fast, offline unit test that checks whosaid version, --version,
and -V all print a matching whosaid X.Y.Z line and exit 0.
./test/install_guard_test.sh covers the temp-directory install guard: refusal without
--force, success with it, install --check-only, and whosaid doctor detecting a dangling
install symlink.
./test/enroll_from_file_test.sh covers whosaid enroll --from in isolation (time-window parsing,
the 15s/silence floor, and the --force overwrite guard) with synthesized say audio — no model
download required.
./test/samples_test.sh covers whosaid samples (longest-segment picking, the --seconds clamp,
--per-speaker, naming, output format, and the source/--audio fallback) against a hand-written
sidecar and synthesized say audio — no model download required.
The workspace-search layer has its own offline tests, all against synthetic workspaces with
placeholder speakers and --no-embed so Ollama is never contacted:
./test/cli_test.shchecks the launcher: every new command inwhosaid help, usage hints and--helppaths, then an end-to-endroll-up,index,search,context,graph,wiki,roll-up --index,doctor, and a stubingest --indexover a two-meeting workspace../test/search_test.shcoverslib/search.py: the FTS5 build, exact queries with filters,context,speakers,status, and the no-Ollama fallback.python3 test/graph_test.pycoverslib/graph.py: the entity tables, each view, and the wiki.python3 test/teams_chat_test.pycoverswhosaid teams ingest: chat-day folders, turn parsing, name and role mapping, idempotent re-ingest, and a CLI ingest → index → search run.python3 test/teams_manifest_test.pycovers audio-less folders in the roll-up and graph:createdfrom the sidecar, thekindfield, andmeetings --kind.python3 test/action_items_test.pycovers the built-in summarizer with a fake model: candidate selection, quote verification, section assignment, and the evidence block.python3 test/eval_test.pycovers the action-item eval harness with fake models: the scorer, record and replay, and the reference backend's isolation../test/watch_test.shcoverslib/watch.pyagainst a temp source folder and a stubwhosaid: the stable-mtime wait, staging, state,--seed,--dry-run, and the plist.python3 test/mcp_descriptions_test.pyalso checks the new read-only workspace tools.
test/eval/run_eval.py measures how good the built-in summarizer's drafts are. It runs the
pipeline over synthetic meetings with hand-labeled action items, reports precision, recall and
F1, and records each run so it can be replayed offline. It can also draft the same meetings with a
cloud model for comparison. That backend is opt-in developer tooling that sends only the committed
synthetic fixtures; whosaid itself never calls a cloud model. The recorded baseline is F1 0.717
for the default qwen2.5:14b against 0.921 for the Claude opus reference, and most of the gap is
precision. See docs/eval.md.
./test/roles_test.sh covers speaker role tags offline: --role/--save-role validation, role
preservation across registry re-saves, the # Role: header lines in .speakers.txt, the
sidecar's roles key, and the [role] card label — no model download required.
./test/commitments_test.sh covers the dev-commitments extractor and corpus offline: cue,
negation, and question handling, self-role gating, boss-requested priority, the CM-NNN
roll-up dedupe and hand-edit reconcile, the semantic dedupe under WHOSAID_EMBED_FAKE=1
versus the difflib fallback when the embed server is unreachable, and the fragment filter
(clause hygiene, min_words at extraction and at fold, --min-words, the [commitments] min_words key, and the dropped-fragments review sections); no model download required.
./test/worklist_test.sh covers the ranked worklist offline: every tier rule, ordering, why
strings, owner resolution (--owner, me, the toml owner, aliases), the CM + AI union and its
cross-source dedupe, _WORKLIST-<Owner>.md as a regenerated view, --all-owners, --json, and
the whosaid commitments launcher command.
The client setup and history regressions have focused suites as well:
./test/bootstrap_test.shexercises the no-Homebrew, required-tool, bottle-failure, and checksum-failure paths with isolated tool mocks.python3 test/menubar_test.pyexercises the real plugin and watcher CLI, including active versus stopped PIDs, configured-source permissions, and first-run preferences.uv run --with numpy python test/backfill_test.pyexercises historical backfill, corpus freshness and curation, dry-run behavior, malformed inputs, and rollback.python3 test/speaker_registry_migration_test.pyexercises private registry round trips, validation, conflict policies, permissions, and concurrent creation.
Each suite is run directly; the runner depends on what it imports:
*.shsuites:./test/<name>.sh.- Stdlib-only Python suites (
graph_test.py,teams_chat_test.py, ...):python3 test/<name>.py. - numpy suites (
anchor,diarize_absorb,diarize_recovery,estimate_k,registry_matchesimport numpy;local_labelsandbackfillneed it for the CLI they drive):uv run --quiet --with numpy python test/<name>_test.py. - MCP suites (
mcp_descriptions_test.py,mcp_workspace_tools_test.pyimportmcp):uv run --with "mcp[cli]" python test/<name>.py.
Couplings to know before editing code:
test/graph_test.pypins themeetingtable as row tuples (SELECT * FROM meeting), themeetings --jsonoutput, and the| folder | kind | created | ...wiki table row, so adding ameetingscolumn means editing those snapshots.- A speaker-registry fixture's
modelmust be an ONNX basename ("<name>.onnx");_validate_modelinlib/speaker_registry.pymakesread_jsonreject anything else. - macOS has no GNU
timeout. In-script deadlines useperl -e 'alarm shift; exec @ARGV' <secs> <cmd>(seetest/enroll_from_file_test.sh).
MIT — see LICENSE.
