Voxtral Transcribe 2 Review (2026): The Fast Half Is Free, the Smart Half Costs Money

By ICON Team · Aug 10, 2026 · 7 min read
Voxtral Transcribe 2 Review (2026): The Fast Half Is Free, the Smart Half Costs Money

Most speech to text launches ask you to pick a vendor and pay by the minute. Mistral did something stranger with Voxtral Transcribe 2. It gave away the fast streaming model under an open licence, kept the more capable batch model behind its API, and called both by the same name. For developers that is a good deal. It is also why so many people still ask whether Voxtral is free: the honest answer is "half of it".

Eight months after the February launch, and with Mistral's text to speech model now sitting beside it, here is where Voxtral Transcribe 2 stands and who should use it.

DetailVoxtral Transcribe 2
Made byMistral AI, Paris
ReleasedFebruary 4, 2026
What it isA speech to text (ASR) model family
The two modelsVoxtral Mini Transcribe V2 (batch) and Voxtral Realtime (streaming)
Languages13: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, Dutch
API price$0.003 per minute (batch), $0.006 per minute (realtime)
Open weightsRealtime model only, Apache 2.0 licence
Realtime model size4 billion parameters, about an 8.87 GB download
Stated latencyUnder 200 milliseconds (realtime), configurable
Headline featuresSpeaker diarization, context biasing, word level timestamps, on device deployment
Where to get itMistral API, Hugging Face, community projects on GitHub
Icon Polls rating3.2 out of 5

Two Models Wearing One Badge

Think of Voxtral Transcribe 2 as a pair of tools for different jobs. Voxtral Mini Transcribe V2 is the batch model. You hand it a finished recording, such as a meeting, an interview or a podcast episode, and it returns a transcript with speaker labels (diarization) and word level timestamps.

Voxtral Realtime is the streaming model. It transcribes audio as it arrives, and Mistral says the delay is under 200 milliseconds. That is quick enough for live captions or a voice agent that has to answer without an awkward pause.

The catch is that the features don't line up. The batch model can tell you who spoke. The realtime model can't. If you want speaker labels during a live call, this family won't give you them.

What You Pay, and What You Don't

The realtime model is free in the sense that counts. It ships as open weights under Apache 2.0, so you can use it in personal projects or commercial products without paying licence fees. You do pay in hardware. A 4 billion parameter model typically needs a GPU with at least 16 GB of VRAM.

The batch model is not free. It runs only through Mistral's API, at $0.003 per minute of audio, which works out to roughly $0.18 per hour. Realtime access through the API costs $0.006 per minute. The original documentation mentions no free tier, but the audio playground in Mistral Studio lets you upload files and see transcripts with diarization and timestamps before you commit.

Even on the paid side, that pricing sits at the cheap end of the market. For anyone moving thousands of hours of recordings, the cost argument is hard to ignore.

Running It Yourself

The official weights live on Hugging Face as mistralai/Voxtral-Mini-4B-Realtime-2602, published in BF16. The model card includes Python examples using the Transformers library (version 5.2.0 and above), deployment notes for vLLM, and compatibility notes for ExecuTorch if you are aiming at on device use. For most developers, that page is the best place to start, and the discussion threads are useful when something breaks.

Mistral doesn't maintain a dedicated GitHub repository for the model, but the community has filled the gap. Projects include local transcription apps for macOS, Rust builds that run in the browser, a C implementation called voxtral.c for squeezing performance out of embedded hardware, and a fully local setup that pairs a macOS app with a GPU server. Nothing in that last one touches the cloud, which matters if your recordings are sensitive.

Where It Earns Its Keep

  • Recorded meetings and interviews. Diarization plus timestamps is exactly what minute takers and editors need, and the batch price is low.
  • Live captions and voice agents. Sub 200 millisecond latency is the realtime model's whole reason to exist.
  • Privacy first projects. Because the realtime model runs locally, audio never has to leave your machine.
  • Accuracy on paper. Benchmarks cited for the release put its word error rate at about 5.9% on the FLEURS multilingual test, against 7.4% for Whisper large v3. Mistral also claims it beats GPT 4o mini Transcribe and Gemini 2.5 Flash. Treat vendor benchmarks as a starting point, not a guarantee for your own audio.

Where It Falls Short

Thirteen languages is a respectable list, but some rivals cover 50 or more. If your users speak Swahili, Polish or Vietnamese, look elsewhere. Context biasing, which lets you feed in names and jargon so they're spelled correctly, is most dependable in English and still experimental in other languages.

Then there's the audience problem. This is a developer product. There's no polished desktop app, no simple consumer tool, nothing for someone who just wants to drop in an MP3 and get a Word document back. The playground is a demo, not a workflow. Google and OpenAI also bundle extras like sentiment analysis and topic detection that Voxtral doesn't offer.

And because the batch model can't be self hosted, you can never fully own the stack. The best features stay on Mistral's servers.

The Sibling That Talks Back

On March 23, 2026, Mistral released Voxtral TTS, which goes the other direction and turns text into speech. It is also a 4 billion parameter model, covers 9 languages, offers 20 preset voices, and can clone a voice from about 3 seconds of sample audio. API pricing is $16 per million characters, and the weights are on Hugging Face as mistralai/Voxtral-4B-TTS-2603 under a CC BY-NC licence, which means no commercial use without a separate arrangement. Community quantized versions cut the download from about 8 GB to roughly 2.67 GB.

Put the two together and Mistral has a full listen and speak pipeline for voice agents and support bots. That's the bigger story here: Voxtral is starting to look like a platform, not a one off model.

Our Verdict: 3.2 Out of 5

Icon Polls rates Voxtral Transcribe 2 at 3.2 out of 5. The speed is real, the pricing is aggressive and the open realtime model is a generous gift to developers. But the split between a free model with fewer features and a paid model you can't host, the narrow language list, and the lack of anything a non technical user could pick up keep it out of the top tier. If you build software, try it. If you just need transcripts, wait for someone to wrap it in an app.

Language coverage is the weak spot here, and it raises a fair question about which languages tech companies should prioritise next. Have your say in our poll on the most powerful languages in the world.

Can I use Voxtral Transcribe 2 offline?

The realtime model, yes. Once downloaded from Hugging Face and set up on suitable hardware, it runs without an internet connection. The batch model always needs Mistral's cloud API.

Is it good for meeting notes?

For recorded meetings, yes. The batch model's speaker labels and timestamps suit that job well, and the price is low. For live meetings you get fast text but no speaker labels.

What's the difference between Voxtral Transcribe 2 and Voxtral TTS?

Transcribe 2 turns speech into text. TTS turns text into speech. They're separate releases, about six weeks apart, with different licences.