Artificial intelligence

Google Launches New Gemini Models for Building Real-Time Voice Applications

Google has announced the availability of Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.5 Transcribe for developers through the Gemini API and Google AI Studio. The models are designed to build voice agents capable of maintaining conversations and completing tasks, alongside low-error speech-to-text conversion across more than 85 languages.

2026-09-15
4 min read
3 views
فريق تحرير certi.news
Google Launches New Gemini Models for Building Real-Time Voice Applications

Google has announced the availability of a new set of audio models for developers through the Gemini API and Google AI Studio, including Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for real-time voice conversations, and Gemini 3.5 Transcribe for converting streaming speech into text. The company says the new models are designed to build voice experiences that do more than respond: they can reason and complete tasks while maintaining the flow of conversation.

Voice agents that complete tasks during conversations

Gemini 3.8 Live provides direct speech-to-speech voice capabilities, allowing the agent to handle requests and carry out actions without pausing the conversation. Its asynchronous function-calling feature enables API and tool calls to run in the background while the user's audio response continues streaming.

The model can also use live visual inputs to provide context for what the user is saying and seeing, and accurately handle alphanumeric data such as confirmation codes, claim numbers, and technical data. The models support more than 97 languages while maintaining accent consistency, in addition to combining real-time audio with structured data updates within the same response.

Gemini 3.8 Live Extended Thinking adds a configurable option for background reasoning when handling complex, multistep problems. According to Google, this model ranks first on Artificial Analysis's Speech-to-Speech leaderboard. The announcement did not clarify the evaluation methodology or the full comparison, so this result remains tied to the measurement cited in the source.

Speech-to-text conversion across 85 languages

Google also launched Gemini 3.5 Transcribe, a model designed for low-latency speech recognition. The company says the model achieved an average word error rate of 4.0% for streaming transcription and 2.6% for non-streaming transcription. It supports more than 85 languages and can automatically handle switching between languages within a sentence or between sentences.

Developers can provide a custom vocabulary list containing up to 1,000 terms to guide recognition toward specialized terminology, uncommon names, and company names. Smart Transcription mode provides text that is more ready to read through structured formatting, correction of some phrasing, and removal of filler and hesitation words.

What changes practically for developers?

This availability combines listening, reasoning, and tool execution in real-time voice interfaces, reducing the need to build a separate chain of speech-to-text models, a conversational model, and a text-to-speech engine. However, the announcement does not eliminate the infrastructure requirements for audio streaming; therefore, Google is also providing access to the models through integration partners such as Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents.

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available through the Live API at an announced price of $0.005 per minute of audio input and $0.018 per minute of audio output. The source notes in its footnote that the cost estimate is based on $3 per million input tokens and $12 per million output tokens, requiring developers to review the actual billing mechanism before scaling.

The models can be tried through ai.studio/live, or developers can use sample applications from GitHub or take advantage of the Live API skill. Google's audio lineup also includes Gemini 3.5 Live Translate for speech-to-speech translation across more than 70 languages, Gemini 3.1 Flash TTS for speech generation, and Lyria 3.5 for music generation.

News source
Google Official News
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news