A free, fully-offline push-to-talk dictation app for macOS. Hold a hotkey anywhere on the system, speak, release, and the transcribed text is typed at the cursor in whatever app is focused. Everything runs locally: no cloud calls, no API keys, no accounts. Your own "Wispr Flow."
- A menu bar app (built with
rumps) shows the current state as an icon: idle (mic), recording (red circle), transcribing (hourglass), or downloading a model (down arrow) the first time you use it. - A global keyboard listener (
pynput) watches for a configurable hotkey (default: hold Right Command). While held,sounddevicerecords the microphone (16kHz mono) into memory. - On release, if the key was held at least 300ms, the audio is transcribed
locally with
faster-whisper(CPU, int8) using the configured model (tiny / base / small). - The transcript is cleaned up (trimmed, vocabulary corrections applied, trailing space added) and injected at the cursor: the current clipboard is saved, the transcript is copied, Cmd+V is simulated, and the original clipboard is restored a moment later. A character-by-character typing fallback is available for apps where paste misbehaves.
Nothing leaves your machine. The only network access is the one-time model download from Hugging Face on first use of a given model size (cached locally afterward).
- macOS (Apple Silicon or Intel), Python 3.9+.
ffmpegonly needed for the optional manual verification script (not for runtime).py2app(inrequirements.txt) is only needed to run./build_app.sh; it's a build-only tool, never imported at runtime, and irrelevant to the./run.shdevelopment path.
| Daily use | Development | |
|---|---|---|
| Command | ./build_app.sh once, then launch Dictation.app |
./run.sh |
| What actually runs | A real macOS app bundle (Contents/MacOS/Dictation) |
A bare CPython interpreter from Xcode Command Line Tools |
| Microphone permission prompt | Shows up in System Settings, works | Never appears at all — see below |
| Accessibility / Input Monitoring grants | Tied to a stable bundle identity (com.ayodeji.dictation) |
Tied to a raw CLT binary path; unstable across changes |
| Use when | You just want to dictate | You're editing src/dictation/*.py and want fast iteration |
Why this distinction matters: on current macOS, an app must declare
NSMicrophoneUsageDescription in a real Info.plist before the OS will even
list the Microphone permission in System Settings — a bare interpreter
(what ./run.sh execs) has no Info.plist, so the request silently fails
with no entry ever appearing to grant. Dictation.app (built via
build_app.sh, see below) is a genuine bundle and doesn't have this problem.
Accessibility/Input Monitoring grants also become far more stable once
they're attached to the app's bundle identity instead of a CLT python binary
whose path can shift.
cd Dictation
./build_app.shThis creates/updates .venv, installs dependencies (including the
build-only py2app), and produces a standalone dist/Dictation.app —
ad-hoc code-signed so macOS will run it (see Gatekeeper note below). The
build does not need .venv or this repo checkout to exist afterward; move
the app wherever you like:
cp -R dist/Dictation.app /Applications/First launch — Gatekeeper: since the app is only ad-hoc signed (no Apple notarization), the first time you open it macOS will refuse with "cannot be opened because the developer cannot be verified." Either:
- Right-click (or Control-click)
Dictation.app> Open > Open again in the dialog, or - Try to open it normally once, then go to System Settings > Privacy & Security, scroll down, and click Open Anyway next to the Dictation mention.
You only need to do this once per build.
The app needs three grants under System Settings > Privacy & Security,
now against Dictation.app itself rather than Terminal/python:
- Microphone — required to record audio.
- Accessibility — required to simulate Cmd+V (paste injection) and to type text (fallback injection mode).
- Input Monitoring — required for the global hotkey listener to receive key events system-wide.
macOS will usually prompt automatically the first time each capability is
used (first recording, first paste simulation, first hotkey press). If a
grant is missing, the relevant feature will silently fail (no recording, no
paste, no hotkey events) — check the three Settings panes above if the app
doesn't respond, add Dictation.app with the + button if it isn't
listed, and relaunch after granting.
On startup the app prints a best-effort permission report
([dictation] Permission check ...) — check Console.app (search "dictation")
or, if launched via Terminal (open dist/Dictation.app), your shell. This is
heuristic — macOS doesn't expose a clean, promptless way to check all three
grants from a script — so treat "looks ungranted" as "go check Settings,"
not as ground truth.
If you previously tried running this via python directly: grants made
while ./run.sh was running are attached to the raw CLT python3 binary
(/Library/Developer/CommandLineTools/.../python3.9), not to Dictation.app
— they're now clutter tied to a binary that's no longer what's running.
Clean up:
- Open System Settings > Privacy & Security > Accessibility (and
separately Input Monitoring). Find the entry that looks like a raw
python/python3.9binary (not an app icon) and remove it with the – button. - Add
Dictation.appfresh in each of the three panes (Microphone, Accessibility, Input Monitoring) — either let macOS prompt you on first use, or add it yourself with + (Cmd+Shift+G in the file picker lets you type a path like/Applications/Dictation.appdirectly).
Hold your configured hotkey (Right Command by default), say something, release. After a moment the transcript should appear wherever your cursor is focused.
cd Dictation
./run.shrun.sh creates a .venv if missing, installs dependencies, and execs the
bare interpreter directly against main.py — no bundle, no Info.plist. Use
this while iterating on src/dictation/*.py: it's faster to restart than
rebuilding the app, and pytest targets the same source tree. It has the
Microphone-prompt limitation described above, so it's not a substitute for
Dictation.app for actual daily dictation use — only for development.
The first time you pick a model (default: small), it will be downloaded from
Hugging Face to ~/.cache/huggingface — this can take a minute depending on
your connection. The menu bar icon shows a down-arrow while this happens.
This applies the same way whether launched via run.sh or the built app.
- Enable/Disable — menu bar toggle; when disabled the hotkey is ignored.
- Model — tiny / base / small, checkmark shows the active one. Changing
it updates
config.jsonimmediately; the new model downloads on first use if not already cached. - Hotkey — shown in the menu (edit
config.jsonto change it; see below). - Open config — opens
config.jsonin your default text editor. - Quit — stops the hotkey listener and exits.
Config lives at ~/.config/dictation/config.json, seeded on first run from
config.example.json in this repo. Fields:
| Key | Default | Description |
|---|---|---|
hotkey |
"cmd_r" |
Key to hold. One of: cmd_r, cmd_l, cmd, alt_r, alt_l, ctrl_r, ctrl_l, shift_r, shift_l, f13, f14, f15. |
model |
"small" |
tiny, base, or small. |
injection_mode |
"paste" |
"paste" (clipboard + Cmd+V, default) or "type" (character-by-character, for apps where paste misbehaves). |
recording_mode |
"toggle" |
"toggle" (double-tap the hotkey to start recording, tap once — or click the on-screen badge — to stop and transcribe; starting needs a double-tap so a single press of the key still works in normal keyboard shortcuts) or "hold" (push-to-talk: record only while the key is held). |
double_tap_window_ms |
400 |
Toggle mode only: max gap in ms between the two taps of a starting double-tap. |
min_hold_ms |
300 |
Hold mode only: presses shorter than this are ignored (avoids accidental triggers). |
sample_rate |
16000 |
Microphone capture rate in Hz; 16000 matches what Whisper expects. |
corrections |
{} |
Vocabulary correction map, e.g. {"acme corp": "Acme Corp"}. Matching is case-insensitive on whole words; the replacement is inserted exactly as written. |
restore_clipboard_delay_s |
0.6 |
Delay before restoring your previous clipboard contents after a paste injection. |
language |
null |
Force a transcription language (e.g. "en"); null = auto-detect. |
Edit the file directly (or via "Open config") and restart the app to pick up hotkey changes; model and other menu-driven fields update live.
Measured on this machine (Apple Silicon, CPU int8) transcribing a synthetic
test phrase ("hello world this is a dictation test") generated with macOS
say -o test.aiff, resampled to 16kHz mono with ffmpeg, then run through
transcribe_file() after each model was already cached locally (warm
latency — first use of a given model also pays a one-time download cost):
| Model | Warm latency | Transcript produced |
|---|---|---|
| tiny | 2.22s | "Hello, well this is a dictation test." |
| base | (see note) | "Hello world this is a dictation test." |
| small | 2.88s | "Hello well this is a dictation test." |
Note: base was measured on first download (79.6s, dominated by the model
fetch); warm latency for base is expected to be similar to tiny/small
(~2-3s) once cached. All three models recognized the phrase correctly except
for "world" vs. "well" in tiny/small — a minor artifact more likely with
robotic TTS-synthesized audio than natural speech. base happened to get
this particular synthetic clip exactly right; in general small has the
best accuracy of the three on real speech and is the recommended default.
Rule of thumb: start with small; drop to tiny/base on older hardware or
if you want lower latency and can tolerate more transcription errors.
Now that Dictation is a real .app, this is the simplest method and needs
no scripts or plists:
- Build it (
./build_app.sh) and move it somewhere permanent, e.g./Applications/Dictation.app. - System Settings > General > Login Items, click
+, and addDictation.app.
That's it — macOS launches the bundle directly at login, using the same bundle identity you already granted permissions to.
A plist is included (com.user.dictation.plist) for anyone who wants
launchd semantics instead (custom throttling, KeepAlive, scripted
install/uninstall, etc.). It points at the built app's actual executable —
Contents/MacOS/Dictation inside the bundle — not run.sh; run.sh
execs the bare CLT interpreter, which is exactly the setup this whole
packaging change moves away from.
./build_app.sh
cp -R dist/Dictation.app /Applications/
mkdir -p ~/Library/LaunchAgents
cp com.user.dictation.plist ~/Library/LaunchAgents/
sed -i '' "s#REPLACE_WITH_APP_PATH#/Applications/Dictation.app#g" \
~/Library/LaunchAgents/com.user.dictation.plist
launchctl load ~/Library/LaunchAgents/com.user.dictation.plistTo uninstall:
launchctl unload ~/Library/LaunchAgents/com.user.dictation.plist
rm ~/Library/LaunchAgents/com.user.dictation.plistLogs go to /tmp/dictation.log / /tmp/dictation.err.log (edit the plist's
StandardOutPath/StandardErrorPath if you want them elsewhere — they're no
longer inside the project directory since the app doesn't need the repo
checkout to exist once built).
- Menu bar icon never leaves the down-arrow / model never finishes
downloading: check your network connection; the model downloads from
Hugging Face on first use of that size and is cached under
~/.cache/huggingfaceafterward. - Nothing happens when I hold the hotkey: almost always an Input
Monitoring permission issue. Check
System Settings > Privacy & Security > Input Monitoring. - Recording seems to happen but nothing gets pasted: check Accessibility
permission (needed to simulate Cmd+V). As a workaround, switch
injection_modeto"type"inconfig.json. - Paste replaces the wrong thing / lands in the wrong field: make sure the target app's cursor is actually focused before releasing the hotkey; the app pastes into whatever has focus at release time.
- Transcription is slow: switch to
tinyorbasein the model menu. - Wrong words for jargon/names: add entries to
correctionsinconfig.json, e.g.{"acme corp": "Acme Corp"}. - App won't start /
ModuleNotFoundError(dev path): delete.venvand re-run./run.shto reinstall dependencies from scratch. Dictation.appwon't open / "developer cannot be verified": it's only ad-hoc signed (no Apple notarization). Right-click > Open, or System Settings > Privacy & Security > "Open Anyway." See "Daily use" above.- Permissions granted but nothing works after switching from
./run.shtoDictation.app: old grants are attached to the CLT python binary, not the new app — see "If you previously tried running this via python directly" above. ./build_app.shfails or the built app crashes on launch: rerun withrm -rf build dist && ./build_app.shfor a clean build. py2app's dependency scanner occasionally misses a package used by faster-whisper's dependency tree (ctranslate2/tokenizers/onnxruntime/av all ship compiled extensions + vendored dylibs); if you seeModuleNotFoundErrorin Console.app after launching, add the missing top-level package name to thepackageslist insetup.pyand rebuild.
Run the test suite (pure-logic parts: config, corrections, audio buffering,
hotkey press/release logic, and clipboard injection where pbcopy/pbpaste
are available):
source .venv/bin/activate
python3 -m pytest -qIterate on the source with ./run.sh (see "Two ways to run this" above);
build/verify packaging changes with ./build_app.sh.
Project layout:
main.py entrypoint
setup.py py2app build spec (Info.plist, bundled packages)
build_app.sh builds + ad-hoc signs dist/Dictation.app
src/dictation/config.py config load/save
src/dictation/corrections.py vocabulary + whitespace post-processing
src/dictation/audio.py mic capture buffer + WAV writer
src/dictation/hotkey.py global push-to-talk key listener
src/dictation/transcribe.py faster-whisper wrapper
src/dictation/inject.py clipboard save/paste/restore + typing fallback
src/dictation/permissions.py best-effort permission checks
src/dictation/app.py rumps menu bar app wiring it together
tests/ pytest suite
run.sh venv + launch (development path, bare interpreter)
com.user.dictation.plist LaunchAgent for start-at-login (advanced/optional)