Entune guide · The basics
What is dictation software?
Dictation software turns what you say into text where you are typing. Here is how it works, what your computer already includes, and what to look for when you need more.
Speech-to-text, explained · Mac, Windows and Linux
Updated
How dictation works
- You speak while holding a key, or after tapping one, and your microphone records it.
- A speech model turns the audio into text. This is speech recognition, also called speech-to-text. The model runs either on your computer or on a provider’s servers.
- Optional processing tidies the text: punctuation, paragraphs and lists, removing hesitation sounds, or correcting names the model misheard.
- The text lands where you were typing, in an email, a document, a chat or a terminal.
Dictation puts your words into whatever app you are using as you go. Transcription is the related job of turning a recording you already have into text; some dictation apps do both.
What your computer already includes
- Mac: turn on Dictation in System Settings › Keyboard › Dictation, then start it with the Microphone key, your Dictation shortcut, or Edit › Start Dictation. It stops after 30 seconds without speech, and its settings show whether text dictation is processed on your Mac. See Apple’s guide.
- Windows: press Windows key+H to start voice typing in a text box. It needs an internet connection, because it uses Microsoft’s online speech recognition. See Microsoft’s guide.
- Linux: most desktops include no system-wide dictation, so it comes from an app. See dictation on Linux.
The built-in options are a good start. People look further when they need a choice of speech model, dictation that stays on their computer, the same tool on every system they use, or names and jargon written correctly.
Local or cloud speech models
- Local models, such as Whisper.cpp or Parakeet on Apple Silicon, run on your computer. Your audio stays there and there is no per-minute charge, but they need a download and use your computer’s memory and processor.
- Cloud models run on a provider’s servers. You send them your audio and usually pay per minute, often through your own account with that provider.
Neither is best for every voice. The surest test is to try two models on your own recordings and compare.
What makes dictation accurate
- Your microphone and room: a close microphone and a quiet room help every model.
- Your language and accent: models differ in how well they handle each, which is another reason to compare.
- Your vocabulary: names, product terms and jargon are the hardest, because a model can only write words it expects. A personal dictionary fixes them, but replacing every match would also change ordinary words that sound the same, so the right spelling depends on the sentence. Correct names and technical terms explains how Entune checks each match in context.
Where Entune fits
Entune is free, open-source dictation for macOS, Windows and Linux. Hold a shortcut in any app, speak, and it pastes the text where you are. You choose a local speech model or a cloud one on your own key, and a personal dictionary corrects your names and terms, choosing by the sentence when a word could be either. There is no Entune account and no telemetry.
uv tool install entune && entune
Needs uv, which installs Python for you, or pip with Python 3.12+. Free and open source.
Puts Entune in Applications and opens it. The first start can take up to a minute.Puts Entune in the Start menu and opens it, in Windows PowerShell.Puts Entune in your applications menu and opens it. Linux needs a few system libraries; see the Linux guide.