AI Model Rankings

We measured every AI model on cleaning up real dictation — quality, speed, and which ones to actually pick.

VivaDicta lets you point AI processing at a lot of different models, and a model name on its own tells you very little about whether it is any good at cleaning up dictation. So we measured them: every model below was sent the same real transcripts through the app's own prompt, and graded on what came back.

The short version: almost any current model is good enough. The top twenty score between 9.4 and 9.8, which is close enough to be a tie. What actually separates them is speed, and whether they quietly do something wrong.

Our picks

gpt-5.5 - the best results of anything measured. 9.8 quality, and at 1.0s it is still quick enough that you are not sitting waiting for it.

gpt-5.4-nano - the best all-round choice. 9.7 quality at 0.7s, and one of the cheapest models OpenAI sells.

grok-4.20-non-reasoning - very fast, and the same 9.7 at 0.7s. With reasoning off there is no thinking pause before the text comes back, so it feels immediate even on long dictation.

qwen-3.8-27b on Cerebras - the quickest free option, 9.1 at 0.32s, and what VivaDicta now selects by default on Cerebras. It replaces gemma-4-31b, which scored higher (9.6 at 0.29s, and is still in the table below) but which Cerebras retired from its public endpoints in September 2026. Cerebras asks for an account with a payment method before it hands out an API key, but you can delete the card again straight after signing up, so nothing can ever be charged - and the $5 signup credit it gives you lasts a very long time at dictation-sized requests.

mistral-medium-latest - free with nothing but an account: no card, no credit to spend down. 9.4 at 0.4s, which makes it the easiest of the free picks to get running.

gemma4:31b on Ollama Cloud - another free option. 9.4 at 0.6s.

claude-haiku-4-5 or gemini-3.1-flash-lite - if you already have credit with Anthropic or Google, both are sub-second and score 9.4-9.6.

How this was measured

Each model was given five raw dictation transcripts through the exact prompt VivaDicta sends, and graded on four things: whether it returned only cleaned text (rather than answering a question it was asked to clean up), whether every name and number survived, how well it removed fillers and stutters, and formatting. Obedience is weighted hardest - a model that replies to your dictation instead of tidying it is useless here, however good it sounds.

The cases deliberately target the hard parts: a speaker correcting themselves mid-sentence, a transcript that is entirely a question, technical terms and numbers, a spoken list, and non-English text.

Speed is the median of three runs. Value combines quality and speed as (2 × quality + speed) / 3, with speed on a curve - under a second is indistinguishable while you wait, and past about eight seconds you have already given up. Cost is not part of the score.

Full ranking

49 models, sorted by value - tap Quality, Speed or Value to re-sort. Graded by claude-opus-5 over 5 cases, last measured September 8, 2026.

#ProviderModel
1OpenAIgpt-5.59.81.0s9.9
2OpenAIgpt-5.4-nano9.70.7s9.8
3Grok (xAI)grok-4.20-non-reasoning9.70.7s9.8
4Cerebrasgemma-4-31bFree tier9.60.3s9.7
5OpenAIgpt-5.4-mini9.60.8s9.7
6OpenAIgpt-5.6-luna9.50.9s9.7
7Anthropicclaude-haiku-4-59.60.9s9.7
8Mistralmistral-medium-latestFree tier9.40.4s9.6
9Ollama Cloudgemma4:31bFree tier9.40.6s9.6
10Geminigemini-3.1-flash-lite9.40.7s9.6
11Geminigemini-3.5-flash-lite9.40.8s9.6
12Anthropicclaude-opus-4-79.71.1s9.6
13OpenAIgpt-4o-mini9.20.7s9.5
14Ollama Cloudgpt-oss:120bFree tier9.20.7s9.5
15Groqqwen/qwen3.6-27bFree tier9.10.3s9.4
16Cerebrasqwen-3.8-27bFree tier9.10.3s9.4
17Geminigemini-3.7-flash9.41.1s9.4
18Geminigemini-3.5-flash9.51.2s9.4
19Groqopenai/gpt-oss-20bFree tier9.00.4s9.3
20Ollama Cloudkimi-k2.7-code9.61.4s9.3
21Geminigemini-3.6-flash9.61.4s9.3
22Anthropicclaude-sonnet-4-69.61.4s9.3
23Cerebrasgpt-oss-120bFree tier8.80.3s9.2
24Ollama Cloudglm-5.3-flash9.21.2s9.2
25Groqopenai/gpt-oss-120bFree tier8.60.3s9.1
26Mistralmistral-small-latestFree tier8.40.5s8.9
27Ollama Clouddeepseek-v4-flash:07319.72.3s8.9
28HuggingFaceQwen/Qwen3-235B-A22B-Instruct-2507Free tier8.20.7s8.8
29Ollama Cloudkimi-k2.69.62.4s8.8
30Anthropicclaude-sonnet-59.62.4s8.8
31OpenCode Gominimax-m39.73.1s8.6
32Ollama Cloudglm-5.29.63.2s8.5
33OpenCode Goglm-5.3-flash9.73.3s8.5
34Grok (xAI)grok-4.39.73.4s8.5
35OpenCode Gokimi-k2.7-code9.73.4s8.5
36Ollama Cloudnemotron-3-nano:30bFree tier8.31.5s8.3
37Ollama Cloudnemotron-3-superFree tier8.72.2s8.3
38OpenCode Goglm-5.39.74.2s8.2
39Ollama Cloudglm-5.19.74.4s8.2
40OpenCode Goqwen3.8-max9.84.5s8.2
41Ollama Cloudkimi-k39.64.5s8.1
42Ollama Cloudminimax-m2.78.93.3s8.0
43Ollama Cloudminimax-m3Free tier9.03.5s8.0
44Geminigemini-2.5-pro9.65.3s7.8
45Ollama Clouddeepseek-v4-pro:08139.14.7s7.7
46Mistralmistral-large-latestFree tier7.41.8s7.6
47Ollama Cloudmistral-large-3:675b5.91.9s6.5
48Geminigemini-2.5-flash-lite4.00.7s6.0
49Ollama Cloudnemotron-3-ultraFree tier9.0134.6s6.0

Tap a column heading to re-sort. Free tier marks models you can run without paying, though the terms differ: Groq is free forever with rate limits, Mistral and HuggingFace give you a monthly allowance, Cerebras gives $5 of credit but needs a card on file first, and on Ollama Cloud only some models are free — those we verified against a free key rather than trusting the vendor's tier list.

What the numbers do not tell you

These are 49 of the models VivaDicta offers, not all of them. A model missing from this table was not measured - that is not a verdict on it.

Scores within about 0.3 of each other are ties. Five cases graded by a single model is enough to separate good from broken, not enough to rank 9.6 above 9.5 with a straight face.

Speed depends on where you are. These were measured from one machine on one network. The ordering held up across runs; the exact seconds will not match what you see.

Two things models get wrong regardless of score. Several ignore the instruction to avoid the Russian letter "ё", and several use em-dashes despite being told not to. Both show up in models that otherwise score well.

Where this runs in the app

Pick a model under Settings → AI Providers, where the top ten are listed directly. Once a provider is configured you can set a different model per Viva Mode, so a fast cheap model can handle everyday dictation while something stronger handles the work that needs it.