DOCUMENTATION / 05

Choose the right models

Compare WhisperInk’s speech and writing models by language, speed, quality, privacy, memory, and image support.

Speech models vs. writing models

Speech models turn audio into text for both Dictation and Command. This always happens locally. Writing models are used only by Command to transform an instruction and its context into a finished result.

Speech models

ModelDownloadLanguagesBest for
Pocket31 MBEnglishFastest notes and commands; bundled fallback
Velocity356 MB25 European languagesVery fast everyday multilingual dictation
Swift322 MBEnglishFast, high-quality English
Balanced181 MB99 languagesCompact multilingual use
Studio547 MB99 languagesBest speech quality when speed matters less

Local writing models

Downloaded local writing models work offline. The first request after choosing a model can take longer while it loads into memory. WhisperInk uses your Mac’s memory to suggest suitable choices.

ModelDownloadImagesNotes
Apple On-DeviceBuilt inNoPrivate default; requires Apple Intelligence and macOS 26
Qwen Writer2.5 GBNoMultilingual, strong instruction following
Qwen Vision2.8 GBYesText and screenshot context
DeepSeek Compact1.1 GBNoSmall, deliberate rewriting model
Gemma 4 E2B3.4 GBYesResponsive local vision
Gemma 4 E4B5.2 GBYesHigher-quality local vision; 16 GB memory recommended

Optional OpenAI models

OpenAI models require your own API key and can create API charges on your OpenAI account. The key is stored in macOS Keychain. Dictation and speech transcription remain local, and audio is not sent to OpenAI.

ModelImagesProfile
LunaYesFast and economical
TerraYesBalanced speed and quality
SolYesQuality-first

Create an OpenAI API key and control costs

OpenAI API usage is billed separately by OpenAI. WhisperInk cannot set or enforce a spending limit for you, so configure the limit in your OpenAI Platform account before selecting an OpenAI model.

  1. 1

    Sign in to the OpenAI Platform and create a dedicated project for WhisperInk if you do not already have one.

  2. 2

    Open API keys, select the project, create a new secret key, and copy it. Paste it into WhisperInk under Models → Command. Never share the key or save it in ordinary notes.

  3. 3

    Open the project’s Limits page. Under Spend, choose Edit spend limit and enter a monthly amount you are comfortable with.

  4. 4

    Turn on Enforce a hard limit, then save. A spend alert alone only sends a notification and does not stop API requests.

  5. 5

    Add one or more spend alerts below the hard limit and check the Usage dashboard occasionally for unexpected activity.

Import a custom GGUF model

WhisperInk can import a valid GGUF writing model and an optional multimodal projector. It checks the file header and runs a smoke test before importing. If the projector fails validation, the model can still be imported as text-only with a warning.

Model downloads and managed imports use integrity checks. Remove models from inside WhisperInk when you no longer need them to reclaim storage.