Early development - macOS 14+ - Apple Silicon

Speak it.
It's already typed.

Murmur is local-first, hold-to-talk voice dictation for macOS. Hold a hotkey, speak, release, and your words appear at the cursor in any app. The transcription and optional cleanup run on your own machine by default, and both endpoints are just URLs, so processing can move to another box by changing one field.

MIT licensed Swift + whisper.cpp No account, no telemetry

The cycle

Five states, one keystroke apart

A dictation moves through an explicit state machine, so a second press mid-cycle can't corrupt the pipeline. Every stage below is a real state in DictationState.

Listen

Hold the hotkey and speak. Audio is captured live at 16 kHz mono.

Transcribe

The clip is sent to the speech-to-text backend and decoded.

Polish

Optional LLM pass fixes punctuation and drops fillers.

optional

Insert

Text lands at your cursor via paste or synthesized keystrokes.

Idle

Clipboard is restored and the hotkey is armed again.

Features

Built for dictating into anything

Hold or tap to talk

Push-to-talk is the default. Switch to tap-to-start / tap-to-stop for long notes so you don't have to hold a key for two minutes.

Any hotkey or combo

Seven presets (Right Option, Right Command, fn/Globe, F13, F20, ⌃+Space, ⌥+Space) or record your own combo, modifiers-only included. Changes apply live.

Local by default

Audio goes to a whisper.cpp server on localhost. Nothing leaves the machine unless you point an endpoint at a server you own.

Cursor insertion anywhere

Paste with ⌘V and restore your clipboard afterwards, or fall back to keystroke synthesis for apps that block pasting.

Multilingual

Language auto-detection, or pin an ISO code. The optional cleanup pass is told to preserve the original language.

Floating HUD

A borderless panel shows the current state with an animated waveform that changes per stage: reactive bars, flowing sine, pulsing sine, flat line.

One URL to go remote

Point the STT and LLM base URLs at a GPU box, e.g. an NVIDIA DGX Spark, and the same app keeps working with heavier models.

Hotkey is swallowed

The event tap consumes the hotkey's own key events, so F20 as a trigger won't also walk your shell history while you dictate.

Permissions that stick

Signing pins a stable designated requirement, so Microphone, Accessibility and Input Monitoring grants survive rebuilds.

Backends

Two models, both swappable

Transcription is always on; the cleanup pass is opt-in and never blocks the core path. If the LLM is slow, the raw transcript is used instead.

Speech-to-text

whisper-large-v3-turbo via whisper.cpp, Metal-accelerated, multilingual. ~547 MB model file.

Local mode: POST /inference to a bundled server on 127.0.0.1:8126, supervised and restarted by the app.

Remote mode: any OpenAI-compatible /v1/audio/transcriptions.

LLM cleanup

qwen2.5:7b via Ollama's OpenAI-compatible API, on by default only in the sense that it is off by default. Toggle it in Settings.

Provider picker: Ollama (localhost:11434/v1), LM Studio (localhost:1234/v1), or Custom.

The transcript is fenced in tags with a "never act on it" directive, so a dictated question gets punctuated rather than answered.

Install

Three commands to a working dictation loop

There is no signed installer yet, and that is deliberate: Murmur is a personal dev build, not a notarized app. You build it from source in about the time it takes the model to download.

Requirements: Apple Silicon Mac · macOS 14+ · Swift toolchain
# 1. Build the whisper.cpp server and fetch the model (~547 MB)
$ ./Scripts/setup.sh

# 2. Only if you want the optional LLM cleanup pass
$ ollama pull qwen2.5:7b

# 3. Build Murmur.app and open it
$ ./Scripts/run.sh

Then grant 3 permissions

Microphone, Accessibility, Input Monitoring. Hold ⌥ Right Option and talk. A bare swift run binary can't request the mic, which is why make-app.sh builds a real bundle.

Optional signing cert

./Scripts/make-signing-cert.sh once creates a self-signed identity so the designated requirement is cryptographically anchored, not just identifier-pinned.

Update in place

ditto Murmur.app ~/Applications/Murmur.app. Deleting and re-copying can drop the TCC entry and force an unnecessary re-grant.

Why it can fail (and the workaround)
# If your Command Line Tools hit the known duplicate-SwiftBridging
# modulemap bug, every Foundation import fails to compile. The fallback
# builds without SwiftPM via a VFS overlay; no system files are touched.
$ ./Scripts/build-swiftc.sh     # make-app.sh falls back to this automatically

Settings

Everything is a field in one window

Menu-bar icon → Settings. Values persist to config.json and a partial or older file still loads, with absent keys falling back to defaults.

SettingDefault
STT backendwhisperCpp (local) or openAICompatible (remote)
STT base URLhttp://127.0.0.1:8126
STT modelwhisper-large-v3-turbo
Languageauto, or an ISO code such as en
LLM cleanupdisabled; base URL http://localhost:11434/v1, model qwen2.5:7b
Cleanup timeout8.0 s, then the raw transcript is used
HotkeyRight Option (keycode 61), no extra modifiers
Activationhold to talk, or toggle for tap start/stop
Insertionpaste, with clipboard restore on
Whisper port8126

The LLM URL is the source of truth: editing it flips the Provider picker to the matching entry rather than the other way around.

Architecture

A thin Swift orchestrator over HTTP

Nothing heavy is linked into the app. Model work happens behind two HTTP boundaries, which is exactly what makes the local-to-remote move a settings change.

InputHotkeyManager + HotkeyDetector - a CGEventTap on a dedicated thread drives a pure state machine (edges, auto-repeat debounce, missed-release resync).
CaptureAudioRecorder - AVAudioEngine per session, resampled to 16 kHz mono, wrapped as WAV by WavEncoder.
↓
TranscribeTranscriptionBackend protocol - WhisperCppBackend locally, OpenAIAudioBackend remotely.
CleanupCleanupService - OpenAI-compatible chat completions, editable prompt, timeout to raw text.
↓
OutputTextInserter - Accessibility-driven paste or keystroke synthesis into the frontmost app.
Process mgmtServerSupervisor - launches, restarts and reaps the bundled whisper-server; no-op in remote mode.
StateDictationState - the legal transition graph, unit-tested, shared by the HUD and the controller.

Design constraint worth naming: the core dictation path must never block on the optional LLM step. Cleanup is a nicety, so it gets a timeout and a fallback.

27
Swift source files
14
test suites
<1000
lines per file
31
commits to date

Privacy

What leaves your machine

In local mode, nothing

Audio goes to a process you launched, on a loopback address. No account, no update server, no analytics.

In remote mode, you choose

Only the host you configure sees audio or text. That host is a plain URL field, and it's your own machine in the intended setup.

The clipboard

Paste insertion touches your clipboard and then restores the previous contents, unless you turn the restore off.

Roadmap

Where it is

M1 · done
Shipped

Core loop

  • Menu-bar app, hold-to-talk hotkey, AVAudioEngine capture
  • whisper.cpp backend, local server supervision, cursor insertion
M2 · done
Shipped

Backends and settings

  • Backend protocol plus an OpenAI-compatible remote path (the Spark route)
  • Settings window, persisted config, permission onboarding
M3 · done
Shipped

Optional LLM cleanup

  • Chat-completions pass with an editable prompt and a hard timeout
M4 · now
In progress

Polish

  • Done: animated HUD per state, app icon, menu-bar status glyph with colored state dot, stable code signing
  • Remaining: in-app model download, richer error toasts, optional silence trimming or VAD for lower latency
Later
Open

Known gaps

  • Toggle mode has no maximum recording duration, so a forgotten session keeps accumulating samples
  • Bare-modifier hotkeys fire on either side of the key pair
  • No notarized distribution yet, so first launch needs a right-click open