Runs offline
Capture, transcription, and speaker separation all run locally through sherpa-onnx. The speech models live on your disk. No audio is streamed anywhere to be understood.
Local-first · Windows desktop
Live transcription, speaker-aware, on your machine. Acousma turns the conversation into notes that write themselves — summary, decisions, action items — and knows which voice is you. Offline except the one AI call you choose.
Free · No account or sign-up Windows 10/11 · x64 · ~350 MB one-time model download
“…let’s lock the launch for the 14th — Priya owns rollout.”
Meeting notes
Summary
Team aligned on a June 14 launch; scope is frozen after Thursday’s review.
Decisions
Action items
Local-first
Granola, Fathom, and Otter are cloud products — your conversations live on their servers. Acousma is a desktop app that keeps the whole pipeline on your machine.
Capture, transcription, and speaker separation all run locally through sherpa-onnx. The speech models live on your disk. No audio is streamed anywhere to be understood.
Use Claude, OpenAI, or Gemini for the assistant and your API key is encrypted per-user with Windows DPAPI — never stored in plaintext, never logged. Or read it from an environment variable.
Point the assistant at Ollama, LM Studio, or any OpenAI-compatible server on your own machine and the loop closes: nothing at all leaves the computer, not even the AI call.
Four shapes
Switch modes mid-meeting for free — the speaker clustering is identical underneath. Only the workflow document out front and how much the transcript claims about speakers change.
Speaker-colored bubbles, a live question queue, and a one-press draft of your answer to the latest thing someone else asked. Prepared-story matching keeps you on message.
A living meeting document — summary, decisions, action items — that fills itself in as the meeting runs. Type your own bullets and the model hangs the specifics it heard underneath, never rewriting your words.
Decisions
Point it at a talk — in the room, or playing on your device — for a rolling summary, plus an “Explain that term” reply the moment a concept flies past you.
Explain: “eventual consistency”
Replicas converge over time, not instantly.
Single-speaker dictation with punctuation on. Speak a draft, then polish it — the raw transcript is yours and the model only ever writes a cleaned copy beside it.
Draft the follow-up email. Confirm the timeline, then ask about budget.
Speaker-aware
Acousma separates speakers from a single microphone using per-utterance voice embeddings — no second channel, no system audio tap. Mark one bubble as “me” and your turns line up on the right for the rest of the meeting, and forever after in saved sessions.
We’re candid about the limit: through one mic in a room, you-vs-everyone-else is the distinction that holds. Telling two remote voices apart through a shared loudspeaker is genuinely hard, so the app leans on the call it can make and never fakes the one it can’t.
You’ve spoken 68% of the last few minutes — leave them some room.
Living notes
Type bullets while the meeting runs. They appear at the top of the document exactly as you typed them, and the assistant hangs the specifics it heard — names, numbers, dates, commitments — underneath each one.
My notes
Meeting notes
Ask about the Q3 budget
$1.2M approved; the hiring freeze lifts in August.
Who owns the migration?
You took it end-to-end; plan due Thursday.
The document keeps three separate authors — you, the summary, and the supporting detail — and the boundary is structural, not a polite request to the model. Reword or delete a bullet and its detail follows. Cost tracks the document’s size, not the meeting’s length, so a two-hour call updates as cheaply as a ten-minute one.
How it works
Audio never leaves the room to be transcribed or sorted by speaker. Only the assistant reaches out — and only to the provider you pick, which can be a model running on the same computer.
Hosted models for quality, or a server on your own machine for total privacy — one dropdown in Settings, switch anytime.
One self-contained ~75 MB executable — no installer, no .NET runtime to hunt down. Run it, press Download models once, and you’re live.
Free · No account or sign-up Windows 10/11 · x64 · ~350 MB one-time model download
Good afternoon
Start something
Set up