Choose the right models
Compare WhisperInk’s speech and writing models by language, speed, quality, privacy, memory, and image support.
Speech models vs. writing models
Speech models turn audio into text for both Dictation and Command. This always happens locally. Writing models are used only by Command to transform an instruction and its context into a finished result.
Speech models
| Model | Download | Languages | Best for |
|---|---|---|---|
| 31 MB | English | Fastest notes and commands; bundled fallback | |
| Velocity | 356 MB | 25 European languages | Very fast everyday multilingual dictation |
| Swift | 322 MB | English | Fast, high-quality English |
| Balanced | 181 MB | 99 languages | Compact multilingual use |
| Studio | 547 MB | 99 languages | Best speech quality when speed matters less |
Local writing models
Downloaded local writing models work offline. The first request after choosing a model can take longer while it loads into memory. WhisperInk uses your Mac’s memory to suggest suitable choices.
| Model | Download | Images | Notes |
|---|---|---|---|
| Apple On-Device | Built in | No | Private default; requires Apple Intelligence and macOS 26 |
| Qwen Writer | 2.5 GB | No | Multilingual, strong instruction following |
| Qwen Vision | 2.8 GB | Yes | Text and screenshot context |
| DeepSeek Compact | 1.1 GB | No | Small, deliberate rewriting model |
| Gemma 4 E2B | 3.4 GB | Yes | Responsive local vision |
| Gemma 4 E4B | 5.2 GB | Yes | Higher-quality local vision; 16 GB memory recommended |
Optional OpenAI models
OpenAI models require your own API key and can create API charges on your OpenAI account. The key is stored in macOS Keychain. Dictation and speech transcription remain local, and audio is not sent to OpenAI.
| Model | Images | Profile |
|---|---|---|
| Luna | Yes | Fast and economical |
| Terra | Yes | Balanced speed and quality |
| Sol | Yes | Quality-first |
Create an OpenAI API key and control costs
OpenAI API usage is billed separately by OpenAI. WhisperInk cannot set or enforce a spending limit for you, so configure the limit in your OpenAI Platform account before selecting an OpenAI model.
- 1
Sign in to the OpenAI Platform and create a dedicated project for WhisperInk if you do not already have one.
- 2
Open API keys, select the project, create a new secret key, and copy it. Paste it into WhisperInk under Models → Command. Never share the key or save it in ordinary notes.
- 3
Open the project’s Limits page. Under Spend, choose Edit spend limit and enter a monthly amount you are comfortable with.
- 4
Turn on Enforce a hard limit, then save. A spend alert alone only sends a notification and does not stop API requests.
- 5
Add one or more spend alerts below the hard limit and check the Usage dashboard occasionally for unexpected activity.
Import a custom GGUF model
WhisperInk can import a valid GGUF writing model and an optional multimodal projector. It checks the file header and runs a smoke test before importing. If the projector fails validation, the model can still be imported as text-only with a warning.
Model downloads and managed imports use integrity checks. Remove models from inside WhisperInk when you no longer need them to reclaim storage.