Entune guide · Local speech
Set up free dictation with local speech models.
Entune can turn your recordings into text with local speech recognition. Choose Whisper.cpp on macOS, Windows or Linux, or Parakeet on an Apple Silicon Mac.
Free and open source · No speech API key for local models.
Choose a local speech model
Entune is free and MIT licensed. Local speech recognition needs no Entune subscription, speech API key or per-minute service payment. Cloud speech, cloud decision models and dictionary generation use your own accounts and can cost money. The setup below starts with local speech and optional processing off.
Whisper.cpp: open Models, find a local Whisper model and download it. Entune offers large-v3-turbo, its compact build, small.en and base.en. The .en models are English-only. Check the download size shown in the app and choose a model your machine can comfortably run.
Parakeet on Apple Silicon: install its separate engine once in a terminal, then return to Models:
uv tool install parakeet-mlx
Download parakeet-tdt-0.6b-v3 in Entune and select it. This path requires an Apple Silicon Mac; it is not the Windows, Linux or Intel Mac option.
Both choices need an initial download. Local models use your computer’s memory and processing resources. Start with a short recording and judge the text and wait time on your own hardware.
Make your first local dictation
- Install Entune. Use the command below with uv installed. uv handles Python for you; the pip alternative needs Python 3.12 or newer.
- Download and select your local model. Confirm it is the selected speech model in Models. Installing Entune alone does not choose local speech for every setup.
- Allow the requested permissions. macOS needs Microphone, Accessibility and Input Monitoring. Windows needs microphone access. On Linux, follow Entune’s keyboard-access instruction, then log out and back in. See the Linux setup reference for system libraries and platform details.
- Start with optional processing off. In Settings → Corrections & formatting, keep dictionary correction, filler removal and formatting off for a simple local speech setup. You can configure these separately later.
- Choose a shortcut and speak. Focus a text field, hold the shortcut, say a short sentence and release. With fast mode off, transcription happens after recording, then Entune pastes the result. You can also use Record in Entune’s window.
For longer dictations, optional fast mode transcribes pieces at natural pauses while you speak. It works with local models too; short dictations are still transcribed whole.
Check History if you need to play the audio, copy the transcript or retry with another model. On Windows, pasting into an elevated application may require copying the result yourself.
Keep speech and optional processing separate
- Local speech: Whisper.cpp and Parakeet transcribe on your computer without sending your recording to a speech service.
- Contextual corrections, filler removal and formatting: enabled processing uses the decision model you choose. Cloud choices send the relevant text to their provider and bill your own API key, even with local speech. Laya is a separately installed local option, tested on macOS; follow the decision-model setup and its engine requirements.
- Dictionary suggestions: building a dictionary uses a cloud language model, even when a local model transcribes the source audio. Suggestions are reviewed before you apply them.
- Downloads: models and their engines need network access to install. Downloading weights does not upload your dictation.
Local speech with the optional processing switches off is a useful starting point when you want audio-to-text processing to stay on your machine. Selecting a local speech model alone does not make every feature offline.
Entune has no account requirement, telemetry or hosted history. Recordings, transcripts and settings stay in its local data folder and do not automatically expire. Optional Langfuse tracing sends dictionary-generation requests and replies to a configured host when its keys are saved. Read the data and privacy reference before using sensitive recordings.
If setup does not work
- Parakeet is unavailable: check that the Mac uses Apple Silicon and the separate engine installation completed. Whisper.cpp is the local alternative on other supported systems.
- A model download stops: downloads resume if interrupted. Check disk space and finish the download in Models before selecting it.
- Imported audio will not transcribe: Entune records WAV. Local transcription of imported MP3, M4A, FLAC, Ogg or WebM needs ffmpeg installed to convert it.
- Names or jargon come out wrong: a local speech model can still mishear words. Compare models using your own recordings, or set up a personal dictionary with the processing and privacy choices above.
Linux support is newer. macOS has been tested by hand; Windows and Linux are covered by automated tests.
Try it with your own dictation
Install Entune, set up a speech model, allow the permissions it requests, and choose a shortcut. Hold the shortcut to record; release it to transcribe and paste.
uv tool install entune && entune
Needs uv, which installs Python for you, or pip with Python 3.12+. Free and open source.
Puts Entune in Applications and opens it. The first start can take up to a minute.