Skip to content

About

A free, fully-offline push-to-talk dictation app for macOS

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Dictation

A free, fully-offline push-to-talk dictation app for macOS. Hold a hotkey anywhere on the system, speak, release, and the transcribed text is typed at the cursor in whatever app is focused. Everything runs locally: no cloud calls, no API keys, no accounts. Your own "Wispr Flow."

How it works

  1. A menu bar app (built with rumps) shows the current state as an icon: idle (mic), recording (red circle), transcribing (hourglass), or downloading a model (down arrow) the first time you use it.
  2. A global keyboard listener (pynput) watches for a configurable hotkey (default: hold Right Command). While held, sounddevice records the microphone (16kHz mono) into memory.
  3. On release, if the key was held at least 300ms, the audio is transcribed locally with faster-whisper (CPU, int8) using the configured model (tiny / base / small).
  4. The transcript is cleaned up (trimmed, vocabulary corrections applied, trailing space added) and injected at the cursor: the current clipboard is saved, the transcript is copied, Cmd+V is simulated, and the original clipboard is restored a moment later. A character-by-character typing fallback is available for apps where paste misbehaves.

Nothing leaves your machine. The only network access is the one-time model download from Hugging Face on first use of a given model size (cached locally afterward).

Requirements

  • macOS (Apple Silicon or Intel), Python 3.9+.
  • ffmpeg only needed for the optional manual verification script (not for runtime).
  • py2app (in requirements.txt) is only needed to run ./build_app.sh; it's a build-only tool, never imported at runtime, and irrelevant to the ./run.sh development path.

Two ways to run this

Daily use Development
Command ./build_app.sh once, then launch Dictation.app ./run.sh
What actually runs A real macOS app bundle (Contents/MacOS/Dictation) A bare CPython interpreter from Xcode Command Line Tools
Microphone permission prompt Shows up in System Settings, works Never appears at all — see below
Accessibility / Input Monitoring grants Tied to a stable bundle identity (com.ayodeji.dictation) Tied to a raw CLT binary path; unstable across changes
Use when You just want to dictate You're editing src/dictation/*.py and want fast iteration

Why this distinction matters: on current macOS, an app must declare NSMicrophoneUsageDescription in a real Info.plist before the OS will even list the Microphone permission in System Settings — a bare interpreter (what ./run.sh execs) has no Info.plist, so the request silently fails with no entry ever appearing to grant. Dictation.app (built via build_app.sh, see below) is a genuine bundle and doesn't have this problem. Accessibility/Input Monitoring grants also become far more stable once they're attached to the app's bundle identity instead of a CLT python binary whose path can shift.

Daily use: build and run Dictation.app

cd Dictation
./build_app.sh

This creates/updates .venv, installs dependencies (including the build-only py2app), and produces a standalone dist/Dictation.app — ad-hoc code-signed so macOS will run it (see Gatekeeper note below). The build does not need .venv or this repo checkout to exist afterward; move the app wherever you like:

cp -R dist/Dictation.app /Applications/

First launch — Gatekeeper: since the app is only ad-hoc signed (no Apple notarization), the first time you open it macOS will refuse with "cannot be opened because the developer cannot be verified." Either:

  • Right-click (or Control-click) Dictation.app > Open > Open again in the dialog, or
  • Try to open it normally once, then go to System Settings > Privacy & Security, scroll down, and click Open Anyway next to the Dictation mention.

You only need to do this once per build.

Grant macOS permissions

The app needs three grants under System Settings > Privacy & Security, now against Dictation.app itself rather than Terminal/python:

  • Microphone — required to record audio.
  • Accessibility — required to simulate Cmd+V (paste injection) and to type text (fallback injection mode).
  • Input Monitoring — required for the global hotkey listener to receive key events system-wide.

macOS will usually prompt automatically the first time each capability is used (first recording, first paste simulation, first hotkey press). If a grant is missing, the relevant feature will silently fail (no recording, no paste, no hotkey events) — check the three Settings panes above if the app doesn't respond, add Dictation.app with the + button if it isn't listed, and relaunch after granting.

On startup the app prints a best-effort permission report ([dictation] Permission check ...) — check Console.app (search "dictation") or, if launched via Terminal (open dist/Dictation.app), your shell. This is heuristic — macOS doesn't expose a clean, promptless way to check all three grants from a script — so treat "looks ungranted" as "go check Settings," not as ground truth.

If you previously tried running this via python directly: grants made while ./run.sh was running are attached to the raw CLT python3 binary (/Library/Developer/CommandLineTools/.../python3.9), not to Dictation.app — they're now clutter tied to a binary that's no longer what's running. Clean up:

  1. Open System Settings > Privacy & Security > Accessibility (and separately Input Monitoring). Find the entry that looks like a raw python / python3.9 binary (not an app icon) and remove it with the – button.
  2. Add Dictation.app fresh in each of the three panes (Microphone, Accessibility, Input Monitoring) — either let macOS prompt you on first use, or add it yourself with + (Cmd+Shift+G in the file picker lets you type a path like /Applications/Dictation.app directly).

Try it

Hold your configured hotkey (Right Command by default), say something, release. After a moment the transcript should appear wherever your cursor is focused.

Development: ./run.sh

cd Dictation
./run.sh

run.sh creates a .venv if missing, installs dependencies, and execs the bare interpreter directly against main.py — no bundle, no Info.plist. Use this while iterating on src/dictation/*.py: it's faster to restart than rebuilding the app, and pytest targets the same source tree. It has the Microphone-prompt limitation described above, so it's not a substitute for Dictation.app for actual daily dictation use — only for development.

The first time you pick a model (default: small), it will be downloaded from Hugging Face to ~/.cache/huggingface — this can take a minute depending on your connection. The menu bar icon shows a down-arrow while this happens. This applies the same way whether launched via run.sh or the built app.

Usage

  • Enable/Disable — menu bar toggle; when disabled the hotkey is ignored.
  • Model — tiny / base / small, checkmark shows the active one. Changing it updates config.json immediately; the new model downloads on first use if not already cached.
  • Hotkey — shown in the menu (edit config.json to change it; see below).
  • Open config — opens config.json in your default text editor.
  • Quit — stops the hotkey listener and exits.

Config reference

Config lives at ~/.config/dictation/config.json, seeded on first run from config.example.json in this repo. Fields:

Key Default Description
hotkey "cmd_r" Key to hold. One of: cmd_r, cmd_l, cmd, alt_r, alt_l, ctrl_r, ctrl_l, shift_r, shift_l, f13, f14, f15.
model "small" tiny, base, or small.
injection_mode "paste" "paste" (clipboard + Cmd+V, default) or "type" (character-by-character, for apps where paste misbehaves).
recording_mode "toggle" "toggle" (double-tap the hotkey to start recording, tap once — or click the on-screen badge — to stop and transcribe; starting needs a double-tap so a single press of the key still works in normal keyboard shortcuts) or "hold" (push-to-talk: record only while the key is held).
double_tap_window_ms 400 Toggle mode only: max gap in ms between the two taps of a starting double-tap.
min_hold_ms 300 Hold mode only: presses shorter than this are ignored (avoids accidental triggers).
sample_rate 16000 Microphone capture rate in Hz; 16000 matches what Whisper expects.
corrections {} Vocabulary correction map, e.g. {"acme corp": "Acme Corp"}. Matching is case-insensitive on whole words; the replacement is inserted exactly as written.
restore_clipboard_delay_s 0.6 Delay before restoring your previous clipboard contents after a paste injection.
language null Force a transcription language (e.g. "en"); null = auto-detect.

Edit the file directly (or via "Open config") and restart the app to pick up hotkey changes; model and other menu-driven fields update live.

Model tradeoffs

Measured on this machine (Apple Silicon, CPU int8) transcribing a synthetic test phrase ("hello world this is a dictation test") generated with macOS say -o test.aiff, resampled to 16kHz mono with ffmpeg, then run through transcribe_file() after each model was already cached locally (warm latency — first use of a given model also pays a one-time download cost):

Model Warm latency Transcript produced
tiny 2.22s "Hello, well this is a dictation test."
base (see note) "Hello world this is a dictation test."
small 2.88s "Hello well this is a dictation test."

Note: base was measured on first download (79.6s, dominated by the model fetch); warm latency for base is expected to be similar to tiny/small (~2-3s) once cached. All three models recognized the phrase correctly except for "world" vs. "well" in tiny/small — a minor artifact more likely with robotic TTS-synthesized audio than natural speech. base happened to get this particular synthetic clip exactly right; in general small has the best accuracy of the three on real speech and is the recommended default. Rule of thumb: start with small; drop to tiny/base on older hardware or if you want lower latency and can tolerate more transcription errors.

Start at login (optional)

Recommended: Login Items

Now that Dictation is a real .app, this is the simplest method and needs no scripts or plists:

  1. Build it (./build_app.sh) and move it somewhere permanent, e.g. /Applications/Dictation.app.
  2. System Settings > General > Login Items, click +, and add Dictation.app.

That's it — macOS launches the bundle directly at login, using the same bundle identity you already granted permissions to.

Advanced: LaunchAgent

A plist is included (com.user.dictation.plist) for anyone who wants launchd semantics instead (custom throttling, KeepAlive, scripted install/uninstall, etc.). It points at the built app's actual executable — Contents/MacOS/Dictation inside the bundle — not run.sh; run.sh execs the bare CLT interpreter, which is exactly the setup this whole packaging change moves away from.

./build_app.sh
cp -R dist/Dictation.app /Applications/
mkdir -p ~/Library/LaunchAgents
cp com.user.dictation.plist ~/Library/LaunchAgents/
sed -i '' "s#REPLACE_WITH_APP_PATH#/Applications/Dictation.app#g" \
  ~/Library/LaunchAgents/com.user.dictation.plist
launchctl load ~/Library/LaunchAgents/com.user.dictation.plist

To uninstall:

launchctl unload ~/Library/LaunchAgents/com.user.dictation.plist
rm ~/Library/LaunchAgents/com.user.dictation.plist

Logs go to /tmp/dictation.log / /tmp/dictation.err.log (edit the plist's StandardOutPath/StandardErrorPath if you want them elsewhere — they're no longer inside the project directory since the app doesn't need the repo checkout to exist once built).

Troubleshooting

  • Menu bar icon never leaves the down-arrow / model never finishes downloading: check your network connection; the model downloads from Hugging Face on first use of that size and is cached under ~/.cache/huggingface afterward.
  • Nothing happens when I hold the hotkey: almost always an Input Monitoring permission issue. Check System Settings > Privacy & Security > Input Monitoring.
  • Recording seems to happen but nothing gets pasted: check Accessibility permission (needed to simulate Cmd+V). As a workaround, switch injection_mode to "type" in config.json.
  • Paste replaces the wrong thing / lands in the wrong field: make sure the target app's cursor is actually focused before releasing the hotkey; the app pastes into whatever has focus at release time.
  • Transcription is slow: switch to tiny or base in the model menu.
  • Wrong words for jargon/names: add entries to corrections in config.json, e.g. {"acme corp": "Acme Corp"}.
  • App won't start / ModuleNotFoundError (dev path): delete .venv and re-run ./run.sh to reinstall dependencies from scratch.
  • Dictation.app won't open / "developer cannot be verified": it's only ad-hoc signed (no Apple notarization). Right-click > Open, or System Settings > Privacy & Security > "Open Anyway." See "Daily use" above.
  • Permissions granted but nothing works after switching from ./run.sh to Dictation.app: old grants are attached to the CLT python binary, not the new app — see "If you previously tried running this via python directly" above.
  • ./build_app.sh fails or the built app crashes on launch: rerun with rm -rf build dist && ./build_app.sh for a clean build. py2app's dependency scanner occasionally misses a package used by faster-whisper's dependency tree (ctranslate2/tokenizers/onnxruntime/av all ship compiled extensions + vendored dylibs); if you see ModuleNotFoundError in Console.app after launching, add the missing top-level package name to the packages list in setup.py and rebuild.

Development

Run the test suite (pure-logic parts: config, corrections, audio buffering, hotkey press/release logic, and clipboard injection where pbcopy/pbpaste are available):

source .venv/bin/activate
python3 -m pytest -q

Iterate on the source with ./run.sh (see "Two ways to run this" above); build/verify packaging changes with ./build_app.sh.

Project layout:

main.py                    entrypoint
setup.py                   py2app build spec (Info.plist, bundled packages)
build_app.sh               builds + ad-hoc signs dist/Dictation.app
src/dictation/config.py    config load/save
src/dictation/corrections.py  vocabulary + whitespace post-processing
src/dictation/audio.py     mic capture buffer + WAV writer
src/dictation/hotkey.py    global push-to-talk key listener
src/dictation/transcribe.py  faster-whisper wrapper
src/dictation/inject.py    clipboard save/paste/restore + typing fallback
src/dictation/permissions.py  best-effort permission checks
src/dictation/app.py       rumps menu bar app wiring it together
tests/                     pytest suite
run.sh                     venv + launch (development path, bare interpreter)
com.user.dictation.plist   LaunchAgent for start-at-login (advanced/optional)

About

A free, fully-offline push-to-talk dictation app for macOS

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages