Speaker Labels

Prefix each turn with Speaker A, Speaker B - configured per Viva Mode, on the models that support it.

Speaker labels prefix each turn in a transcript with Speaker A, Speaker B, and so on — so a meeting or interview reads as a conversation rather than one block of text.

It Belongs to the Mode

Speaker labels are a property of the Viva Mode, not a global switch. A meetings mode can label speakers while your dictation mode stays clean, and switching apps switches the behaviour with it.

The toggle only appears when the mode's selected transcription model can actually produce labels.

Supported Models

Speaker labels are available on:

  • Deepgram, ElevenLabs, Mistral, Soniox
  • Gladia, Speechmatics, AssemblyAI, xAI
  • OpenAI, through its dedicated gpt-4o-transcribe-diarize model
  • Gemini, through gemini-3.5-transcribe

Local models — Whisper, Parakeet, Apple Speech — do not produce speaker labels. Cloud model cards show which ones do.

Two Things to Know

Streaming stands down. Realtime providers emit plain partial text with no speaker information, so with labels on VivaDicta uses the upload path instead. You trade a little latency for the labels.

Gemini transcribes verbatim. A mode with speaker labels on switches Gemini 3.5 Transcribe out of its "smart" mode, because the API rejects diarization inside it. That means hesitation sounds and self-corrections are no longer removed for you.

Related

  • Viva Modes — where the toggle lives
  • Transcription Models — which models support labels
  • Translation — combining labels with native translation