Free, open-source dictation · macOS, Windows, Linux

Learns your words. Knows when not to use them.

Dictate into any app with free local speech models. Hold a shortcut, speak, and Entune pastes the text where you are.

  • 01 · Dictionary Suggests your names and terms from your own dictation, for you to review. A decision model then uses each one only where the sentence fits.
  • 02 · Your model Local speech with no API key, or cloud speech with your own keys and provider charges. Corrections, formatting and filler removal work with either.
  • 03 · Your data No Entune account, no telemetry, no hosted history. Recordings and transcripts stay in Entune's folder on your computer.
uv tool install entune && entune

Needs uv, which installs Python for you, or pip with Python 3.12+. Free and open source.

42-second tour · Entune on a Mac

What it does for you

You already dictate. This is what changes.

  1. 01 / 09

    Corrections that fit

    A dictionary of your names and terms. A decision model checks each match against the sentence and applies only the ones that fit. It never rewrites you.

  2. 02 / 09

    A dictionary from day one

    Pick your speech model, then Learn from audio. Entune transcribes recordings you already have, from a folder or another dictation app on your Mac, and a language model you choose proposes entries for you to review.

  3. 03 / 09

    Cloud or local, one page

    Paste a key (each provider has a Get a key link), or download a Whisper.cpp model in one click. On Apple Silicon, Parakeet runs locally too, after a one-time engine install from the terminal.

  4. 04 / 09

    Formatting, words untouched

    Paragraph breaks and bullets where your dictation clearly has them. A decision model chooses where, and every word is kept. Off until you turn it on.

  5. 05 / 09

    Hesitation sounds, removed

    Remove um, uh, er, erm, ah and hmm when the decision model identifies hesitation. Repeated like becomes one. Quotes, code and uncertain cases are kept. English only, and off until you turn it on.

  6. 06 / 09

    Drop any recording

    Drop audio files onto the window. Each one is transcribed with your default model and lands in History, like a dictation.

  7. 07 / 09

    Anonymous mode

    The eye switch blurs every transcript in History, for screen recordings and sharing. The newest stays readable until you copy it, then it blurs too.

  8. 08 / 09

    Retry with another model

    Every recording is kept. Play it, see what each step changed, or transcribe it again with another model. If a provider fails, the recording waits for a retry.

  9. 09 / 09

    A pill that says what happened

    A small pill in the corner shows level bars while you record, then each stage, then what happened: pasted, or copied and why. An error stays on it with Retry, and it never takes the focus.

Corrections that consider the context

A match is not always a replacement.

After transcription, a decision model looks at each dictionary match and the words around it, and applies only the ones that fit. It chooses from your own entries. It never generates or rewrites your text.

Four decision models: Jev, which made the fewest wrong changes in our test, sends the matched context to TypeSafe on your own key. With OpenAI's Decisions API or Perplexity, it goes to that provider on your own API key. Laya runs on your computer, in English, after a one-time install.

Decision modelEntry: cloud · Claude

Ask cloudClaude to review this.

cloudClaudeReplaced: the name fits here

Move the files to cloudClaude storage.

cloudClaudeKept: the everyday word is right

Same entry, two sentences. Find-and-replace would break one of them.

A test on real dictation

Which model knows when to fix a misheard word?

The dictionary offered a different word in 121 places. 47 should change. 72 should stay.

  1. Perplexity Decisions API

    Wrong swaps: 3
    Correct fixes: 43
  2. Jev

    Wrong swaps: 1
    Correct fixes: 38
  3. OpenAI Decisions API

    Wrong swaps: 8
    Correct fixes: 42
  4. Laya Running locally on a Mac

    Wrong swaps: 8
    Correct fixes: 28
  5. Replace every match Hypothetical baseline

    Wrong swaps: 72
    Correct fixes: 47

In this test, Perplexity caught the most fixes.
Jev made the fewest wrong swaps.

The data

  • Dictionary learned from 7.5 hours of the maintainer's dictation, transcribed by Parakeet, with entries suggested by GPT-6 Astra on a ChatGPT plan.
  • Tested on the next 2.1 hours it had never seen: 160 recordings.
  • Expected answers written down before any model ran. Jev, OpenAI and Laya were tested on October 7; Perplexity on October 8 with the same transcripts, dictionary and expected answers.
  • One speaker, one reviewer judging from text, one run per model.

Two matches could not be judged and are excluded from the correct and wrong counts; Perplexity changed one of them. Both sides use the same count scale. These results do not establish a statistical ranking, including between 1 and 3 wrong swaps, or measure overall transcription accuracy. Data, method and limits

In the separate formatting test, of 92 places judged to need a paragraph or bullet, Jev placed 35 with 12 unwanted, Perplexity 32 with 5, OpenAI 8 with 1, and Laya 5 with 37.

Your models

Three jobs. You choose the model for each.

Entune is free and MIT licensed, with no Entune subscription. Local speech models run on your computer without a speech API key or per-minute payment. Cloud speech, cloud decision models and dictionary generation use your own accounts and can cost money. Dictionary generation always uses a cloud language model. Starting points are marked.

  1. 01 · Speech

    Turn audio into text

    After every recording, or at natural pauses with optional fast mode

    CloudAssemblyAIGroqSonioxElevenLabsxAI
    LocalParakeet Apple SiliconWhisper.cpp
  2. 02 · Decisions

    Apply your dictionary and formatting

    After transcription, with enabled processing steps running together

    CloudJev fewest wrong in our testOpenAI's Decisions APIPerplexity
    LocalLaya English
  3. 03 · Dictionary

    Build your dictionary

    When you choose Get suggestions

    ChatGPTSign in with ChatGPT
    KeyOpenAIAnthropicGoogle GeminiGroqMistral

In History, transcribe any recording again with another model to compare them on your own voice. Choosing models

A dictionary that learns

Teach your speech model your words.

Local open-source models such as Parakeet and Whisper.cpp are fast and keep your audio on your computer, but they can mishear names and jargon that a large cloud model gets right. Entune helps close that gap with recordings you already have: it transcribes them with the model you dictate with and builds a dictionary from the words that model gets wrong. It works the same with any speech model.

Built on your ChatGPT plan

Sign in with ChatGPT. No API key.

Choose Sign in with ChatGPT, approve Entune on OpenAI's page, and you are back in Entune. There is nothing to set up in ChatGPT first. Your plan's models and usage limits apply, and in ChatGPT's Settings › Usage you can set a weekly limit for Entune or disconnect it. The sign-in builds your dictionary only; speech and decision models use their own keys or run on your computer. Prefer a key? OpenAI, Anthropic, Google Gemini, Groq and Mistral work too.

  1. 01

    Bring your recordings

    In Dictionary, learn from recordings already in Entune, another dictation app on your Mac, or a folder. Only the audio is copied; the other app's transcripts are never read.

  2. 02

    Pick the model you dictate with

    Entune transcribes the audio with it, so its own mistakes show up: a local model one recording at a time, a cloud model four at a time.

  3. 03

    Find what it gets wrong

    A language model reads the transcripts in parts and suggests entries for that speech model. It runs in the background while you keep dictating. A large import can take hours; Stop keeps finished parts for Continue.

  4. 04

    Review, apply, dictate

    Edit any suggestion, leave out what you don't want, and apply the rest. Turn on Apply your dictionary, and your decision model uses each entry only where the sentence agrees.

Entries belong to the speech model they were learned on, because every model mishears differently. Pin an entry to use it with every model; suggestions can't remove pinned entries. The same audio can build a dictionary for another model. Already dictating? Get suggestions reads your newest transcripts. Later runs can revise or remove learned entries, and you review each change.

History you keep

Every recording, on your disk, with what happened to it.

Replay the audio. Copy the text. See what each processing step changed. Transcribe a recording again with another model and compare. Provider failures stay visible. Nothing is uploaded to Entune, because there is no Entune server.

Optional processing, each with its own switch
  • Apply your dictionaryPicks the meaning that fits each sentence
  • Remove fillersHesitation sounds removed; repeated like becomes one
  • Paragraphs and bulletsAdded where the dictation clearly has them. Every word is kept

All three start off, in Settings › Corrections & formatting, and each needs a decision model. The original transcript is always kept in History.

Privacy, said plainly

No account. No telemetry. No hosted history.

Cloud speech sends audio to the provider you chose. Local speech keeps it on your computer.

Corrections, filler removal and formatting are off until you turn them on. Then your decision model reads the text: the words around each match with the entries that could apply, or the whole transcript for fillers and formatting. Jev sends it to TypeSafe, OpenAI's Decisions API to OpenAI, and Perplexity to Perplexity, even when speech is local. Laya keeps it on your computer.

Building your dictionary runs only when you ask. It sends the transcripts it reads, and the entries in them, to the language model you chose.

Audio, transcripts, your dictionary, settings and keys live in one folder on your computer. Settings › Data & privacy exports your audio and transcripts, or deletes everything.

Use Entune on a computer only you use: its local API has no password, so other programs on it can read your recordings. Data and privacy, in full

What leaves your computer
MicrophoneHold the shortcut, speak
Stays here
Speech model
Stays here
Your dictionaryEntries that match the transcript
Stays here
Decision model
Stays here
Pasted where you were typingSaved to Entune's folder
Stays here

Local speech and Laya: dictation stays on your computer.

With corrections, filler removal and formatting all off, nothing goes to a decision model. Building the dictionary is separate and runs only when you ask.

Install

One command. Then three steps.

uv tool install entune && entune

Needs uv. Without it, on Python 3.12 or newer: python -m pip install entune && entune

Puts Entune in Applications and opens it. The first start can take up to a minute.

From then on, open it like any other app. To upgrade: uv tool upgrade entune && entune

Free and open source. Setting up on Linux? See the Linux guide.

  1. 01

    Set up a speech model

    In Models, paste a provider's API key or download a local model. The first one becomes your default. Parakeet, on Apple Silicon, first needs its engine: uv tool install parakeet-mlx.

  2. 02

    Allow permissions

    On macOS: Microphone, Accessibility and Input Monitoring, each from its own button.

  3. 03

    Set a shortcut

    Hold it in any app and speak, or toggle hands-free recording. The text is pasted where you are, or copied if no text field is active. You can also click Record in Entune's window.

You can start without a dictionary or decision model and add them later.