xAI has released Grok Voice Transcribe 2.0, its latest speech-to-text model for developers and businesses. The model is available through xAI’s Speech-to-Text API.
According to xAI, Transcribe 2.0 delivers twice the accuracy of its predecessor in the company’s real-world evaluations. Meanwhile, pricing remains unchanged at $0.10 per hour for batch transcription and $0.20 per hour for streaming.
Improved Accuracy for Real-World Audio
Grok Voice Transcribe 2.0 builds on the audio foundation model behind Grok Voice. As a result, xAI has focused the model on audio conditions that commonly challenge speech recognition systems.
These conditions include noisy telephone calls, conversations, spoken credentials, and multilingual voice commands. Furthermore, xAI evaluated the model using four internal datasets based on production traffic.
According to xAI, Transcribe 2.0 improved on Grok Voice Transcribe 1.0 across all four evaluation sets. It also ranked first for accuracy among 32 streaming models on the Artificial Analysis leaderboard, the company says.
In addition, xAI says the model performs strongly with compressed 8 kHz telephony audio. This capability could support customer-service applications where recordings often contain background noise and limited audio bandwidth.
The company has also published comparisons against models including Gemini 3.5 Transcribe, MAI-Transcribe-2, ElevenLabs Scribe v2, Deepgram Nova-3, and Whisper Large v3. However, these comparisons come from xAI’s own evaluation methodology and should be viewed in that context.
Batch and Real-Time Transcription
Developers can use Grok Voice Transcribe 2.0 for both recorded and real-time audio. The API supports batch transcription through REST and low-latency streaming through WebSockets.
The API supports several common audio formats, including WAV, MP3, WebM, OGG, and M4A. It also supports multiple languages, speaker diarization, timestamps, key-term prompting, and real-time interim results.
The model also includes Smart Turn for streaming applications. This feature uses machine learning to estimate whether a speaker has finished a thought, helping voice applications manage conversational turns.
xAI has also retained the same pricing structure as the previous model. Batch transcription costs $0.10 per hour, while streaming transcription costs $0.20 per hour.
As a result, businesses already using Transcribe 1.0 can access the newer model without a higher published transcription rate. xAI says Transcribe 2.0 will become the default model in its Speech-to-Text API, while version 1.0 will be deprecated in the coming weeks.
Expanding Voice AI Applications
The release extends xAI’s broader push into developer-focused voice technology. Earlier this year, the company introduced standalone Speech-to-Text and Text-to-Speech APIs based on the technology behind Grok Voice.
Since then, xAI has expanded its speech capabilities for voice agents, transcription services, and other audio applications. The company says Grok Voice technology already supports customer-support calls, video narration, and voice interactions in physical products.
For businesses, the new model could support applications such as customer-service transcription, meeting records, voice assistants, accessibility tools, and media processing. Furthermore, its streaming API gives developers a route to build applications that require near real-time transcription.
At the same time, xAI is positioning accuracy and pricing as key elements of the release. The company says Transcribe 2.0 improves substantially over its previous model without increasing the published API rates.
The model is currently available through xAI’s Speech-to-Text API. Therefore, developers can integrate it into existing applications through batch or streaming transcription workflows.








