Skip to content
Acousma

Local-first · Windows desktop

The meeting
hears itself.

Live transcription, speaker-aware, on your machine. Acousma turns the conversation into notes that write themselves — summary, decisions, action items — and knows which voice is you. Offline except the one AI call you choose.

Free · No account or sign-up Windows 10/11 · x64 · ~350 MB one-time model download

Local-first

Your meetings never leave the room.

Granola, Fathom, and Otter are cloud products — your conversations live on their servers. Acousma is a desktop app that keeps the whole pipeline on your machine.

Runs offline

Capture, transcription, and speaker separation all run locally through sherpa-onnx. The speech models live on your disk. No audio is streamed anywhere to be understood.

Your keys stay yours

Use Claude, OpenAI, or Gemini for the assistant and your API key is encrypted per-user with Windows DPAPI — never stored in plaintext, never logged. Or read it from an environment variable.

Or fully local

Point the assistant at Ollama, LM Studio, or any OpenAI-compatible server on your own machine and the loop closes: nothing at all leaves the computer, not even the AI call.

Four shapes

One app, four ways to listen.

Switch modes mid-meeting for free — the speaker clustering is identical underneath. Only the workflow document out front and how much the transcript claims about speakers change.

Interview

Speaker-colored bubbles, a live question queue, and a one-press draft of your answer to the latest thing someone else asked. Prepared-story matching keeps you on message.

How would you scale this?
Shard by tenant, then
Answer ↵ Shorter

Notes

A living meeting document — summary, decisions, action items — that fills itself in as the meeting runs. Type your own bullets and the model hangs the specifics it heard underneath, never rewriting your words.

Decisions

Ship the beta Friday
Freeze scope after review

Lecture

Point it at a talk — in the room, or playing on your device — for a rolling summary, plus an “Explain that term” reply the moment a concept flies past you.

Explain: “eventual consistency”

Replicas converge over time, not instantly.

Memo

Single-speaker dictation with punctuation on. Speak a draft, then polish it — the raw transcript is yours and the model only ever writes a cleaned copy beside it.

Draft the follow-up email. Confirm the timeline, then ask about budget.

Polish draft ↵

Speaker-aware

It knows which voice is you.

Acousma separates speakers from a single microphone using per-utterance voice embeddings — no second channel, no system audio tap. Mark one bubble as “me” and your turns line up on the right for the rest of the meeting, and forever after in saved sessions.

We’re candid about the limit: through one mic in a room, you-vs-everyone-else is the distinction that holds. Telling two remote voices apart through a shared loudspeaker is genuinely hard, so the app leans on the call it can make and never fakes the one it can’t.

  • Right-aligned, accent-striped bubbles for your turns
  • A gentle talk-time nudge when you’re dominating the room
  • Claiming “me” costs more than any other label — a false “you” is the one error that matters
Can you own the migration timeline?
Yes — I’ll have a plan by Thursday.
★ This is me voiceprint saved

You’ve spoken 68% of the last few minutes — leave them some room.

Living notes

Your words are never rewritten.

Type bullets while the meeting runs. They appear at the top of the document exactly as you typed them, and the assistant hangs the specifics it heard — names, numbers, dates, commitments — underneath each one.

My notes

  • Ask about the Q3 budget
  • Who owns the migration?

Meeting notes

  • Ask about the Q3 budget

    $1.2M approved; the hiring freeze lifts in August.

  • Who owns the migration?

    You took it end-to-end; plan due Thursday.

The document keeps three separate authors — you, the summary, and the supporting detail — and the boundary is structural, not a polite request to the model. Reword or delete a bullet and its detail follows. Cost tracks the document’s size, not the meeting’s length, so a two-hour call updates as cheaply as a ten-minute one.

How it works

Six steps on your machine. One you choose.

Audio never leaves the room to be transcribed or sorted by speaker. Only the assistant reaches out — and only to the provider you pick, which can be a model running on the same computer.

On your machine · offline
  1. 1 Microphone WASAPI · 16 kHz mono
  2. 2 Streaming ASR sherpa-onnx zipformer
  3. 3 Endpoint utterance boundary
  4. 4 Voice embedding per-utterance
  5. 5 Clustering you vs. others
  6. 6 Transcript live, on disk
Assistant Claude · GPT · Gemini · or fully local

Bring any brain. Or none of ours.

Hosted models for quality, or a server on your own machine for total privacy — one dropdown in Settings, switch anytime.

  • Anthropic Claude (default)
  • OpenAI GPT
  • Google Gemini
  • Ollama local
  • LM Studio local
  • Any OpenAI-compatible llama.cpp · vLLM · OpenRouter · Groq

Download Acousma.

One self-contained ~75 MB executable — no installer, no .NET runtime to hunt down. Run it, press Download models once, and you’re live.

Free · No account or sign-up Windows 10/11 · x64 · ~350 MB one-time model download

Good afternoon

Start something

Interview
Notes
Lecture
Memo

Set up

  • Speech models installed
  • Assistant provider chosen
  • Train my voice