AI Model Rankings
We measured every AI model on cleaning up real dictation — quality, speed, and which ones to actually pick.
VivaDicta lets you point AI processing at a lot of different models, and a model name on its own tells you very little about whether it is any good at cleaning up dictation. So we measured them: every model below was sent the same real transcripts through the app's own prompt, and graded on what came back.
The short version: almost any current model is good enough. The top twenty score between 9.4 and 9.8, which is close enough to be a tie. What actually separates them is speed, and whether they quietly do something wrong.
Our picks
gpt-5.5 - the best results of anything measured. 9.8 quality, and at 1.0s it is still quick enough that you are not sitting waiting for it.
gpt-5.4-nano - the best all-round choice. 9.7 quality at 0.7s, and one of the cheapest models OpenAI sells.
grok-4.20-non-reasoning - very fast, and the same 9.7 at 0.7s. With reasoning off there is no thinking pause before the text comes back, so it feels immediate even on long dictation.
gemma-4-31b on Cerebras - the best free option, and the quickest model here: 9.6 at 0.29s. Cerebras asks for an account with a payment method before it hands out an API key, but you can delete the card again straight after signing up, so nothing can ever be charged - and the $5 signup credit it gives you lasts a very long time at dictation-sized requests.
mistral-medium-latest - free with nothing but an account: no card, no credit to spend down. 9.4 at 0.4s, which makes it the easiest of the free picks to get running.
gemma4:31b on Ollama Cloud - another free option. 9.4 at 0.6s.
claude-haiku-4-5 or gemini-3.1-flash-lite - if you already have credit with Anthropic or Google, both are sub-second and score 9.4-9.6.
How this was measured
Each model was given five raw dictation transcripts through the exact prompt VivaDicta sends, and graded on four things: whether it returned only cleaned text (rather than answering a question it was asked to clean up), whether every name and number survived, how well it removed fillers and stutters, and formatting. Obedience is weighted hardest - a model that replies to your dictation instead of tidying it is useless here, however good it sounds.
The cases deliberately target the hard parts: a speaker correcting themselves mid-sentence, a transcript that is entirely a question, technical terms and numbers, a spoken list, and non-English text.
Speed is the median of three runs. Value combines quality and speed as (2 × quality + speed) / 3, with speed on a curve - under a second is indistinguishable while you wait, and past about eight seconds you have already given up. Cost is not part of the score.
Full ranking
48 models, sorted by value - tap Quality, Speed or Value to re-sort. Graded by claude-opus-5 over 5 cases, measured August 29, 2026.
| # | Provider | Model | |||
|---|---|---|---|---|---|
| 1 | OpenAI | gpt-5.5 | 9.8 | 1.0s | 9.9 |
| 2 | OpenAI | gpt-5.4-nano | 9.7 | 0.7s | 9.8 |
| 3 | Grok (xAI) | grok-4.20-non-reasoning | 9.7 | 0.7s | 9.8 |
| 4 | Cerebras | gemma-4-31bFree tier | 9.6 | 0.3s | 9.7 |
| 5 | OpenAI | gpt-5.4-mini | 9.6 | 0.8s | 9.7 |
| 6 | OpenAI | gpt-5.6-luna | 9.5 | 0.9s | 9.7 |
| 7 | Anthropic | claude-haiku-4-5 | 9.6 | 0.9s | 9.7 |
| 8 | Mistral | mistral-medium-latestFree tier | 9.4 | 0.4s | 9.6 |
| 9 | Ollama Cloud | gemma4:31bFree tier | 9.4 | 0.6s | 9.6 |
| 10 | Gemini | gemini-3.1-flash-lite | 9.4 | 0.7s | 9.6 |
| 11 | Gemini | gemini-3.5-flash-lite | 9.4 | 0.8s | 9.6 |
| 12 | Anthropic | claude-opus-4-7 | 9.7 | 1.1s | 9.6 |
| 13 | OpenAI | gpt-4o-mini | 9.2 | 0.7s | 9.5 |
| 14 | Ollama Cloud | gpt-oss:120bFree tier | 9.2 | 0.7s | 9.5 |
| 15 | Groq | qwen/qwen3.6-27bFree tier | 9.1 | 0.3s | 9.4 |
| 16 | Gemini | gemini-3.7-flash | 9.4 | 1.1s | 9.4 |
| 17 | Gemini | gemini-3.5-flash | 9.5 | 1.2s | 9.4 |
| 18 | Groq | openai/gpt-oss-20bFree tier | 9.0 | 0.4s | 9.3 |
| 19 | Ollama Cloud | kimi-k2.7-code | 9.6 | 1.4s | 9.3 |
| 20 | Gemini | gemini-3.6-flash | 9.6 | 1.4s | 9.3 |
| 21 | Anthropic | claude-sonnet-4-6 | 9.6 | 1.4s | 9.3 |
| 22 | Cerebras | gpt-oss-120bFree tier | 8.8 | 0.3s | 9.2 |
| 23 | Ollama Cloud | glm-5.3-flash | 9.2 | 1.2s | 9.2 |
| 24 | Groq | openai/gpt-oss-120bFree tier | 8.6 | 0.3s | 9.1 |
| 25 | Mistral | mistral-small-latestFree tier | 8.4 | 0.5s | 8.9 |
| 26 | Ollama Cloud | deepseek-v4-flash:0731 | 9.7 | 2.3s | 8.9 |
| 27 | HuggingFace | Qwen/Qwen3-235B-A22B-Instruct-2507Free tier | 8.2 | 0.7s | 8.8 |
| 28 | Ollama Cloud | kimi-k2.6 | 9.6 | 2.4s | 8.8 |
| 29 | Anthropic | claude-sonnet-5 | 9.6 | 2.4s | 8.8 |
| 30 | OpenCode Go | minimax-m3 | 9.7 | 3.1s | 8.6 |
| 31 | Ollama Cloud | glm-5.2 | 9.6 | 3.2s | 8.5 |
| 32 | OpenCode Go | glm-5.3-flash | 9.7 | 3.3s | 8.5 |
| 33 | Grok (xAI) | grok-4.3 | 9.7 | 3.4s | 8.5 |
| 34 | OpenCode Go | kimi-k2.7-code | 9.7 | 3.4s | 8.5 |
| 35 | Ollama Cloud | nemotron-3-nano:30bFree tier | 8.3 | 1.5s | 8.3 |
| 36 | Ollama Cloud | nemotron-3-superFree tier | 8.7 | 2.2s | 8.3 |
| 37 | OpenCode Go | glm-5.3 | 9.7 | 4.2s | 8.2 |
| 38 | Ollama Cloud | glm-5.1 | 9.7 | 4.4s | 8.2 |
| 39 | OpenCode Go | qwen3.8-max | 9.8 | 4.5s | 8.2 |
| 40 | Ollama Cloud | kimi-k3 | 9.6 | 4.5s | 8.1 |
| 41 | Ollama Cloud | minimax-m2.7 | 8.9 | 3.3s | 8.0 |
| 42 | Ollama Cloud | minimax-m3Free tier | 9.0 | 3.5s | 8.0 |
| 43 | Gemini | gemini-2.5-pro | 9.6 | 5.3s | 7.8 |
| 44 | Ollama Cloud | deepseek-v4-pro:0813 | 9.1 | 4.7s | 7.7 |
| 45 | Mistral | mistral-large-latestFree tier | 7.4 | 1.8s | 7.6 |
| 46 | Ollama Cloud | mistral-large-3:675b | 5.9 | 1.9s | 6.5 |
| 47 | Gemini | gemini-2.5-flash-lite | 4.0 | 0.7s | 6.0 |
| 48 | Ollama Cloud | nemotron-3-ultraFree tier | 9.0 | 134.6s | 6.0 |
Tap a column heading to re-sort. Free tier marks models you can run without paying, though the terms differ: Groq is free forever with rate limits, Mistral and HuggingFace give you a monthly allowance, Cerebras gives $5 of credit but needs a card on file first, and on Ollama Cloud only some models are free — those we verified against a free key rather than trusting the vendor's tier list.
What the numbers do not tell you
These are 48 of the models VivaDicta offers, not all of them. A model missing from this table was not measured - that is not a verdict on it.
Scores within about 0.3 of each other are ties. Five cases graded by a single model is enough to separate good from broken, not enough to rank 9.6 above 9.5 with a straight face.
Speed depends on where you are. These were measured from one machine on one network. The ordering held up across runs; the exact seconds will not match what you see.
Two things models get wrong regardless of score. Several ignore the instruction to avoid the Russian letter "ё", and several use em-dashes despite being told not to. Both show up in models that otherwise score well.
Where this runs in the app
Pick a model under AI Providers - on macOS the top ten are listed there directly. Once a provider is configured you can set a different model per Viva Mode, so a fast cheap model can handle everyday dictation while something stronger handles the work that needs it.