securityonline.info 16 Sept 2026, 04:13 UTC

Google Unveils Voice AI That Speaks 97 Languages and Acts in Background

Google Unveils Voice AI That Speaks 97 Languages and Acts in Background
CyberSIXT Evidence Panel Source marked as original reporting

GOOGLE has unveiled two voice-focused artificial intelligence models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The first is designed for large-scale, cost-efficient deployment and can switch between 97 languages during a conversation. It also supports “visual grounding”, allowing users to stream images or other visual data while asking questions in real time.

The Extended Thinking version is intended for complex, multi-step tasks and can make API calls in the background while verbally updating users on its progress. Google says it achieved a score of 82.6 on the Artificial Analysis Speech-to-Speech Quality Index.

Both models are available to developers through the Gemini API and Google AI Studio. Google lists pricing at $0.005 per minute for audio input and $0.018 per minute for audio output. The company has also added its Gemini 3.5 Transcribe model to its audio tools; it supports more than 85 languages, has a reported streaming Word Error Rate of 4.0%, and offers custom vocabulary and intelligent transcription features. Integrations with LiveKit, LangChain and Vercel are intended to simplify deployment of real-time voice applications.

Google says generated audio includes invisible SynthID watermarks to support traceability and help address deceptive content. Gemini 3.8 Live is being deployed in Google Search Live, while the Extended Thinking model is planned for Gemini Live users and Google AI Pro and Ultra subscribers.

The article presents the models as commercially available native speech-to-speech systems that can converse while running background software actions, although it provides no independent evidence beyond Google’s stated specifications.

View full article

Article by CyberSIXT