Google pushed a new pair of voice models into the world on September 15, 2026, and the timing says almost as much as the release notes. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking arrived just two weeks after Gemini 3.8 Flash and 3.8 Flash Cyber, according to Google’s own announcement. That cadence, roughly one major Gemini release every three weeks since midsummer, is now the story as much as the models themselves.

The pitch is simple: a voice assistant that can see what you’re looking at, keep talking while it works a problem in the background, and switch between 97 languages mid-conversation without you noticing the seam. Google frames Gemini 3.8 Live as the everyday, cost-efficient option and Gemini 3.8 Live Extended Thinking as the version built for jobs that need real multi-step reasoning, not just fast replies. Coverage from Unite.AI and Google’s developer documentation lay out a launch that’s less about a single flashy demo and more about closing gaps competitors have been probing for a year.

What Gemini 3.8 Live and 3.8 Live Extended Thinking Actually Do

Both models are audio-to-audio systems, meaning they take in speech, video frames, images, and text, and hand back speech or text without the old round trip through a separate transcription step. Google’s blog post and technical documentation describe near real-time reasoning during a live conversation, automatic language detection, and the ability to call outside tools and APIs while still talking. That last point is the headline feature: instead of going silent to “think,” the model gives an early verbal acknowledgment, keeps the conversation moving, and reports back once a background task finishes.

Gemini 3.8 Live Extended Thinking pushes that idea further with what Google calls an asynchronous reasoning protocol. The model narrates progress out loud while it works through a harder problem in the background, rather than freezing the conversation until it has a final answer. Google’s own framing, drawn from its developer documentation, describes this as thinking in the background while streaming speech, aimed squarely at multi-step tasks like debugging a script, walking through a tax scenario, or planning a multi-leg trip.

On the technical side, Gemini 3.8 Live accepts 16-bit PCM audio at 16kHz and returns audio at 24kHz, plus JPEG visual input sampled at up to one frame per second. Google has not published token limits or per-minute pricing for either model, which puts Gemini 3.8 Live in the same disclosure gap as most of its rivals: developers can test the capability today, but budgeting a production voice agent still means waiting on a rate card.

The People Behind the Announcement

The launch post was written on behalf of Google’s Gemini Audio team by Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, according to the byline on Google’s blog. Their framing leans on a specific example: a model that can watch you attempt a chess move through your camera, keep the game going conversationally, and quietly plan three moves ahead without breaking the back-and-forth. It’s a deliberately mundane demo, chosen to show fluency under normal use rather than a scripted best case.

Gemini 3.8 Live vs. Gemini 3.8 Live Extended Thinking

Google is positioning these as two ends of the same product line rather than a flagship-and-budget split. The table below breaks down how Google itself differentiates the two, based on the specifications published alongside the launch.

AttributeGemini 3.8 LiveGemini 3.8 Live Extended Thinking
Primary use caseGeneral conversation, cost-efficient scaleHigh-complexity, multi-step reasoning
Reasoning styleFast, near real-time responsesBackground reasoning while narrating progress
Language support97 languages, automatic detection97 languages, automatic detection
Visual inputJPEG frames, up to 1 fpsJPEG frames, up to 1 fps
Audio input/output16-bit PCM, 16kHz in / 24kHz out16-bit PCM, 16kHz in / 24kHz out
Consumer accessSearch Live, Gemini LiveGemini Live
Enterprise accessGemini Enterprise (private preview)Gemini Enterprise (private preview)
WatermarkingSynthID on all audio outputSynthID on all audio output

The Benchmark Numbers Google Is Leading With

Google backed the launch with a specific set of third-party and internal benchmark scores rather than vague marketing language, which is a shift from how some earlier Gemini Live updates were pitched. According to figures published in Google’s announcement and echoed by Unite.AI’s coverage, Gemini 3.8 Live Extended Thinking scored 82.6 on the Artificial Analysis Speech to Speech Quality Index and 97.7% on Big Bench Audio, a benchmark that tests audio reasoning rather than plain transcription accuracy.

On agentic benchmarks, the model completed 68.6% of tasks on the τ-Voice agentic task completion test and 35.1% on Sierra’s τ-Voice-banking benchmark, which simulates a customer calling a bank and needing multi-step account help. Gemini 3.8 Live placed second on the Artificial Analysis Speech Agent Arena leaderboard. Google also says the model pushes the Pareto frontier on ServiceNow’s EVA-Bench, a benchmark built around complex enterprise workflows rather than single-turn chat.

BenchmarkScoreWhat it measures
Artificial Analysis Speech-to-Speech Quality Index82.6Overall voice conversation quality
Big Bench Audio97.7%Audio-based reasoning accuracy
τ-Voice agentic task completion68.6%Multi-step task completion via voice
Sierra τ-Voice-banking35.1%Complex banking-scenario task completion
Artificial Analysis Speech Agent Arena2nd placeHead-to-head ranking against rival voice agents

The banking benchmark score is worth sitting with. A 35.1% completion rate on a task built to mimic a real customer service call means these models still fail more often than they succeed on genuinely complex, multi-turn financial scenarios. That gap is exactly why Google is pitching Extended Thinking as a distinct product rather than folding background reasoning into the base model by default.

Where the Models Show Up First

Rollout is staggered by audience, which is typical for a Gemini release but worth tracking closely if you’re deciding when to build on top of it. Developers get the fastest path: both models are available now through the Gemini API and Google AI Studio. Consumers get a split experience, with Gemini 3.8 Live reaching Search Live and the Gemini Live app, while Extended Thinking is currently scoped to Gemini Live itself. Google AI Pro, Ultra, and other AI-subscription tiers also gain access inside Google Workspace apps, including Docs, Gmail, and Keep.

Enterprise customers are behind a gate for now. Gemini Enterprise access is listed as private preview, with Gemini Enterprise for Customer Experience described as coming soon rather than live. That staging matters for anyone evaluating this for a call-center or support-automation project: the version aimed at handling real customer interactions at scale isn’t broadly available yet, even though the underlying models are.

The Infrastructure Partners Building on Top

Google named a specific set of partners integrating the new models into their own products: Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents on the infrastructure side, with Salesforce, Genspark, and Lumeris listed as implementation partners building customer-facing products on top. That’s a notably broader partner list than Google published for the 3.8 Flash launch three weeks earlier, and it signals Google wants Gemini Live treated as a platform other companies build voice products on, not just a feature inside Google’s own apps.

How This Fits the Last Two Months of Gemini Releases

Context matters here. Google shipped Gemini 3.7 Flash roughly three weeks before Gemini 3.8 Flash and 3.8 Flash Cyber landed on September 2, 2026, according to reporting from The Register and 9to5Google. Gemini 3.8 Live and 3.8 Live Extended Thinking now arrive roughly two weeks after that, which puts three distinct Gemini releases inside a single month. That’s a materially faster cadence than the quarterly-ish rhythm Google kept through most of 2025.

The acceleration isn’t happening in a vacuum. It follows a stretch in which Google’s cloud and AI businesses have been fighting for headlines against Nvidia’s chip dominance and Microsoft’s Copilot distribution advantage, and it comes as rival labs have been shipping their own updates on a similarly tight schedule. Shipping fast, publicly, with benchmark numbers attached, reads as a deliberate signal that Google intends to compete on release velocity, not just headline capability.

Competitive Landscape: Gemini Live Against the Rest of the Voice AI Field

Google’s own launch materials don’t mention competitors by name, but the comparison is unavoidable given how crowded real-time voice AI has become. OpenAI’s Realtime API, built on its GPT-4o family, remains the most widely deployed alternative for developers building voice agents. Amazon has pushed hard on price and latency with Nova Sonic, which TechCrunch reported Amazon positioned as dramatically cheaper than OpenAI’s offering, at roughly $0.015 per minute versus a GPT-4o Realtime rate several times higher, with Amazon also claiming lower perceived latency in its own published comparisons.

ModelMakerNotable claimPricing disclosed
Gemini 3.8 Live Extended ThinkingGoogle82.6 speech-to-speech quality index, background reasoningNot disclosed
GPT-4o Realtime APIOpenAIBroad developer adoption, mature WebRTC toolingPer-minute, published
Nova SonicAmazonVendor-claimed ~$0.015/minute, lower latency vs. GPT-4oPublished, low-cost positioning

The gap that stands out is disclosure. Amazon and OpenAI both publish per-minute pricing for their real-time voice APIs, which lets developers model costs before committing engineering time. Google has not done that for Gemini 3.8 Live or Extended Thinking, which leaves cost as the biggest open question for any team deciding between the three right now. On raw capability, Google’s benchmark scores suggest it’s caught up to, and in some cases ahead of, where OpenAI and Amazon’s real-time models stood as of their last public benchmarking round, but independent, apples-to-apples testing across all three hasn’t been published yet.

The Trust Layer: SynthID Watermarking

Every piece of audio these models generate carries a SynthID watermark, Google’s audio watermarking system designed to let detection tools flag AI-generated speech even after compression or editing. That’s not a new technology for Google, which has used SynthID across image and video outputs for several years, but extending it by default to a live conversational voice product is a meaningful commitment given how often voice cloning shows up in fraud and disinformation cases. It also puts pressure on Amazon, OpenAI, and other voice AI vendors to either match that default or explain why they haven’t.

Why the Enterprise Angle Matters More Than the Consumer Demo

The consumer-facing chess demo is memorable, but the money is in the enterprise rollout. Gemini Enterprise for Customer Experience, still listed as coming soon, is aimed directly at the call-center and support-automation market that companies like Salesforce, Genesys, and NICE have spent years building specialized products around. If Google can get Extended Thinking’s background-reasoning trick working reliably on real customer calls, not just controlled chess demos, it has a credible pitch against purpose-built contact-center AI vendors, most of which don’t have a foundation model of this scale to draw on.

The 35.1% score on Sierra’s banking benchmark is the reason that pitch isn’t finished yet. Enterprises buying automation for regulated, high-stakes conversations like banking or healthcare intake need failure rates far lower than a coin flip on complex tasks, and Google’s own published number shows Extended Thinking still failing on the majority of a benchmark specifically designed to look like a real banking call.

Historical Context: From Gemini 1.0 to Live Voice in Two Years

Google’s Gemini line has moved from a text-first chatbot competitor into a multimodal, real-time voice platform in a relatively short window. The jump from static, turn-based chat to a model that can watch a video feed, listen, reason in the background, and speak back without an obvious pause is the kind of capability that used to require stitching together three or four separate services: speech-to-text, a language model, a tool-calling layer, and text-to-speech. Collapsing that pipeline into a single audio-to-audio model is what made products like OpenAI’s Realtime API and now Gemini Live technically possible, and it’s also what’s driving the compressed release schedule industry-wide. Once the architecture exists, shipping incremental versions is faster than the original multi-service approach ever was.

Developer Documentation Signals a Platform Push

Google’s own Gemini API model documentation lists both new models alongside the rest of the Gemini family, treating Live and Extended Thinking as standard, general-availability options for developers rather than a limited experiment. That framing, plus the named list of infrastructure partners, points toward Google trying to make Gemini Live a default choice for anyone building a voice product from scratch in 2026, the same way Gemini’s text models became a default choice for chatbot builders in 2024 and 2025.

Market Impact: What This Means for Google’s AI Strategy

Voice is the interface Google arguably has the most existing distribution to win. Search Live puts Gemini 3.8 Live in front of anyone already using Google Search, Workspace access puts it in front of every AI Pro and Ultra subscriber typing into Docs or Gmail, and the developer partner list gives it a foothold in products Google doesn’t own outright. That’s a wider distribution surface than either OpenAI or Amazon can claim for their real-time voice products today, since neither has anything close to Google Search’s daily reach.

The absence of published pricing is the clearest sign this is still an early-stage rollout rather than a finished product launch. Google typically publishes rate cards once it’s confident in unit economics at scale, and holding that back for both Gemini 3.8 Live and Extended Thinking suggests the compute cost per conversation, especially for the background-reasoning Extended Thinking variant, is still being tuned before Google commits to a number publicly.

What to Watch Next

  • Pricing disclosure: expect Google to publish per-minute or per-token rates for both models once enterprise private preview moves toward general availability, likely within the next one to two Gemini release cycles.
  • Wider Extended Thinking access: given Gemini 3.8 Live already reached Search Live at launch, Extended Thinking reaching Search Live and broader consumer surfaces is a reasonable next step rather than a long-term holdout.
  • Independent benchmarking: third-party researchers are likely to run Gemini 3.8 Live head-to-head against GPT-4o Realtime and Nova Sonic on shared benchmarks, since Google, OpenAI, and Amazon have each only published favorable internal or vendor-adjacent comparisons so far.
  • Contact-center pilots: watch for named enterprise customers moving from Gemini Enterprise private preview into public case studies, which would signal the banking-benchmark failure rate is closing fast enough for regulated use cases.
  • Faster release cadence continuing: three Gemini releases in roughly one month suggests Google is unlikely to slow back down to a quarterly cycle, and a Gemini 3.9 announcement inside Q4 2026 would be consistent with the pace set since August.

The Bottom Line for Developers and Businesses

For developers, the practical takeaway is that Gemini 3.8 Live and Extended Thinking are usable today through the Gemini API and AI Studio, with real benchmark numbers to evaluate against rather than marketing adjectives. For businesses evaluating voice AI for customer support or internal tools, the enterprise version is still gated behind private preview, and the banking-style benchmark score is a clear signal to pilot carefully before betting a regulated workflow on it. The takeaway for the wider market: Google has closed the technical gap with OpenAI and Amazon on real-time voice AI, but pricing transparency and independent, cross-vendor benchmarking are the two things still missing before anyone can call a clear winner.

Frequently Asked Questions

What is Gemini 3.8 Live?
Gemini 3.8 Live is Google’s general-purpose, real-time voice AI model, built for fluid conversation with visual grounding and support for 97 languages. It’s designed for cost-efficient, everyday use cases rather than the most complex reasoning tasks.

What is Gemini 3.8 Live Extended Thinking?
It’s the higher-reasoning variant in the same family, built for multi-step tasks. It uses an asynchronous reasoning protocol that lets it keep talking or narrating progress while it works through a harder problem in the background, instead of pausing the conversation.

How do I access Gemini 3.8 Live as a developer?
Both models are available now through the Gemini API and Google AI Studio, according to Google’s official announcement. Enterprise access through Gemini Enterprise is currently in private preview.

Is Gemini 3.8 Live available to regular consumers?
Yes. Gemini 3.8 Live is available in Search Live and the Gemini Live app. Gemini 3.8 Live Extended Thinking is currently available within Gemini Live, and both are accessible inside Google Workspace apps like Docs, Gmail, and Keep for AI Pro, Ultra, and other AI-subscription tiers.

How much does Gemini 3.8 Live cost?
Google has not published per-minute or per-token pricing for either model as of this announcement. That’s a notable gap compared to rivals like OpenAI and Amazon, which both publish real-time voice API pricing.

How does Gemini 3.8 Live compare to OpenAI’s Realtime API and Amazon Nova Sonic?
Google’s published benchmark scores, including 82.6 on the Artificial Analysis Speech-to-Speech Quality Index and second place on the Speech Agent Arena, suggest it’s competitive with both rivals on quality. Amazon has publicly claimed a pricing and latency advantage with Nova Sonic, but no independent lab has yet run all three models through the same shared benchmark suite.

What is SynthID and why does it matter here?
SynthID is Google’s watermarking system, applied to all audio Gemini 3.8 Live and Extended Thinking generate. It’s designed to let detection tools identify AI-generated speech even after the audio has been edited or compressed, which matters as voice cloning becomes a bigger fraud and disinformation risk.

Can Gemini 3.8 Live handle complex customer service calls today?
Not reliably yet. Google’s own published Sierra τ-Voice-banking benchmark score of 35.1% shows the model still fails on the majority of complex, multi-step banking-style scenarios, which is why Gemini Enterprise for Customer Experience remains in a “coming soon” state rather than general availability.