Entune guide · Coding prompts
Keep your technical terms right when you dictate to Claude Code.
Speaking a prompt is quick, until a library name, a flag or the agent’s own name comes out wrong. Entune learns the words your speech model gets wrong and fixes them, choosing by the sentence when a word could mean either.
Free and open source · Local speech models · Works in any app you type in
Updated
Try it with one prompt See what it does →
Try it with one prompt
You need no account or speech API key to start. Entune itself sets no recording limit; how long a recording each speech model handles is up to that model. With AssemblyAI, recordings of over 20 minutes have been transcribed.
- Install Entune with the command below.
- Download a local speech model in Models: a Whisper.cpp model on macOS, Windows or Linux, or Parakeet on an Apple Silicon Mac after its one-time engine install.
- Allow the permissions Entune asks for, and set a shortcut.
- Click into Claude Code’s prompt, hold the shortcut, say your prompt and release. Entune pastes into the text input you are in. On a Mac this has been tested in the Ghostty terminal; wherever Entune finds no input to paste into, the text stays on the clipboard and its pill says so. Every transcript is also kept in History.
uv tool install entune && entune
Needs uv, which installs Python for you, or pip with Python 3.12+. Free and open source.
Puts Entune in Applications and opens it. The first start can take up to a minute.Puts Entune in the Start menu and opens it, in Windows PowerShell.Puts Entune in your applications menu and opens it. Linux needs a few system libraries; see the Linux guide.
Claude Code’s built-in /voice is the quickest start when you sign in with a Claude.ai account: it recognizes common development terms and your project and branch names. Entune fits when Claude Code uses an Anthropic API key, Amazon Bedrock, Google Cloud’s Agent Platform or Microsoft Foundry, when you want your speech kept on your computer, or want a vocabulary of your own that also works in every other app you write in.
Your words, learned from your own recordings
In Dictionary, choose Learn from audio to use recordings you already have, in Entune, in another dictation app on your Mac, or in a folder. Or choose Get suggestions to read your recent transcripts.
Entune transcribes the audio with the speech model you dictate with, so that model’s own mistakes show up. A language model you choose proposes entries, and you review each one before it is applied.
Heard by Parakeet in our October test
“cloud” for Claude, “in tune” for Entune, “open A” for OpenAI, “Jeff” for Jev
In that test, the dictionary learned from earlier recordings covered 39 of the 94 names and terms Parakeet got wrong; it never learned the ones the recordings rarely mentioned. Entries belong to the speech model they were learned on, because every model mishears differently. Pin the spellings you care about: a pinned entry applies with every model and is not learned again.
The sentence decides between meanings
A name can sound like an ordinary word. After transcription, a decision model looks at each match and the words around it, and applies only the entries that fit. It chooses from your entries and never rewrites the rest of your text.
One sound, two meanings, from the same test
“code”: code or quote. “top”: top or tab. “start”: start or star. “from”: from or Chrome.
Most of the dictionary’s risky entries pair a name or term with an everyday word like these. In the test, the everyday word was right far more often, which is what the decision step is for.
In our October test, the dictionary offered a different word at 121 places, and 47 of them should change:
- Replace every match: 121 changes, 72 of them wrong.
- Jev: 39 changes, 1 wrong, and 38 of the 47 needed changes caught.
- Perplexity: 47 changes, 3 wrong, and 43 of the 47 caught.
Two matches could not be judged and are excluded; Perplexity changed one of them. The 121 places were in 51 of 160 test recordings, 2.1 hours from the maintainer’s own use over six days: one speaker, one run per model. OpenAI’s Decisions API, Laya and the method are in the full results.
Corrections are off until you turn them on, and they need a decision model: Jev, OpenAI’s Decisions API or Perplexity on your own key, or Laya on your computer, in English.
Let Claude Code add the words you confirm
Entune’s local API takes corrections that you have confirmed with an agent. Nothing is added unless you say yes. Add an instruction like this to your project’s CLAUDE.md:
When a word in my prompt looks mistranscribed, ask one short question to
confirm what I meant. Once I confirm, send it to Entune with a short
description, and drop the request silently after two seconds if Entune is
not running:
curl -s -m 2 -X POST localhost:4187/api/dictionary/corrections \
-H 'content-type: application/json' \
-d '{"entries": [{"spelling": "OpenAI", "description": "the AI company", "heard": ["open A"]}], "source": "claude-code"}'
Send only what I confirmed, whole words or phrases, never guesses.
A confirmed word is pinned, so it applies with every speech model, and with corrections on, Entune uses it wherever it hears that sound. That suits a sound you never mean literally, like “open A” for OpenAI in our test. When a sound has two meanings you use, such as “cloud” and Claude, give it both meanings in Dictionary: the decision model then reads their descriptions and picks one for each sentence. Read the agents’ API for the full shape. The local API has no password, so any program on your computer can call it: use Entune on a computer only you use.
What leaves your computer
- Speech: a local model keeps your audio on your computer. A cloud model sends it to the provider you chose, on your key.
- Corrections: a cloud decision model receives up to 160 characters of your transcript on each side of a match, with each matching entry’s spelling, definition and personal context. Laya keeps all of it on your computer.
- Filler removal and formatting: when you turn them on, they send the whole transcript to your decision model.
- Building the dictionary: runs only when you ask, and sends the transcripts it reads, with the entries that occur in them, to the language model you chose.
More detail: correct names and technical terms, set up local dictation, and the data and privacy reference.