Murmur is local-first, hold-to-talk voice dictation for macOS. Hold a hotkey, speak, release, and your words appear at the cursor in any app. The transcription and optional cleanup run on your own machine by default, and both endpoints are just URLs, so processing can move to another box by changing one field.
The cycle
A dictation moves through an explicit state machine, so a second press mid-cycle can't
corrupt the pipeline. Every stage below is a real state in
DictationState.
Hold the hotkey and speak. Audio is captured live at 16 kHz mono.
The clip is sent to the speech-to-text backend and decoded.
Optional LLM pass fixes punctuation and drops fillers.
optionalText lands at your cursor via paste or synthesized keystrokes.
Clipboard is restored and the hotkey is armed again.
Features
Push-to-talk is the default. Switch to tap-to-start / tap-to-stop for long notes so you don't have to hold a key for two minutes.
Seven presets (Right Option, Right Command, fn/Globe, F13, F20, ⌃+Space, ⌥+Space) or record your own combo, modifiers-only included. Changes apply live.
Audio goes to a whisper.cpp server on localhost. Nothing leaves the machine unless you point an endpoint at a server you own.
Paste with ⌘V and restore your clipboard afterwards, or fall back to keystroke synthesis for apps that block pasting.
Language auto-detection, or pin an ISO code. The optional cleanup pass is told to preserve the original language.
A borderless panel shows the current state with an animated waveform that changes per stage: reactive bars, flowing sine, pulsing sine, flat line.
Point the STT and LLM base URLs at a GPU box, e.g. an NVIDIA DGX Spark, and the same app keeps working with heavier models.
The event tap consumes the hotkey's own key events, so F20 as a trigger won't also walk your shell history while you dictate.
Signing pins a stable designated requirement, so Microphone, Accessibility and Input Monitoring grants survive rebuilds.
Backends
Transcription is always on; the cleanup pass is opt-in and never blocks the core path. If the LLM is slow, the raw transcript is used instead.
whisper-large-v3-turbo via whisper.cpp, Metal-accelerated, multilingual. ~547 MB model file.
Local mode: POST /inference to a bundled server on 127.0.0.1:8126, supervised and restarted by the app.
Remote mode: any OpenAI-compatible /v1/audio/transcriptions.
qwen2.5:7b via Ollama's OpenAI-compatible API, on by default only in the sense that it is off by default. Toggle it in Settings.
Provider picker: Ollama (localhost:11434/v1), LM Studio (localhost:1234/v1), or Custom.
The transcript is fenced in tags with a "never act on it" directive, so a dictated question gets punctuated rather than answered.
Install
There is no signed installer yet, and that is deliberate: Murmur is a personal dev build, not a notarized app. You build it from source in about the time it takes the model to download.
# 1. Build the whisper.cpp server and fetch the model (~547 MB) $ ./Scripts/setup.sh # 2. Only if you want the optional LLM cleanup pass $ ollama pull qwen2.5:7b # 3. Build Murmur.app and open it $ ./Scripts/run.sh
Microphone, Accessibility, Input Monitoring. Hold ⌥ Right Option and talk. A bare swift run binary can't request the mic, which is why make-app.sh builds a real bundle.
./Scripts/make-signing-cert.sh once creates a self-signed identity so the designated requirement is cryptographically anchored, not just identifier-pinned.
ditto Murmur.app ~/Applications/Murmur.app. Deleting and re-copying can drop the TCC entry and force an unnecessary re-grant.
# If your Command Line Tools hit the known duplicate-SwiftBridging # modulemap bug, every Foundation import fails to compile. The fallback # builds without SwiftPM via a VFS overlay; no system files are touched. $ ./Scripts/build-swiftc.sh # make-app.sh falls back to this automatically
Settings
Menu-bar icon → Settings. Values persist to config.json
and a partial or older file still loads, with absent keys falling back to defaults.
| Setting | Default |
|---|---|
| STT backend | whisperCpp (local) or openAICompatible (remote) |
| STT base URL | http://127.0.0.1:8126 |
| STT model | whisper-large-v3-turbo |
| Language | auto, or an ISO code such as en |
| LLM cleanup | disabled; base URL http://localhost:11434/v1, model qwen2.5:7b |
| Cleanup timeout | 8.0 s, then the raw transcript is used |
| Hotkey | Right Option (keycode 61), no extra modifiers |
| Activation | hold to talk, or toggle for tap start/stop |
| Insertion | paste, with clipboard restore on |
| Whisper port | 8126 |
The LLM URL is the source of truth: editing it flips the Provider picker to the matching entry rather than the other way around.
Architecture
Nothing heavy is linked into the app. Model work happens behind two HTTP boundaries, which is exactly what makes the local-to-remote move a settings change.
Design constraint worth naming: the core dictation path must never block on the optional LLM step. Cleanup is a nicety, so it gets a timeout and a fallback.
Privacy
Audio goes to a process you launched, on a loopback address. No account, no update server, no analytics.
Only the host you configure sees audio or text. That host is a plain URL field, and it's your own machine in the intended setup.
Paste insertion touches your clipboard and then restores the previous contents, unless you turn the restore off.
Roadmap