Google's Gemini 3.8 Live Goes Native Speech-to-Speech

Google's Gemini 3.8 Live Goes Native Speech-to-Speech
Google answered on September 15 with Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, native speech-to-speech models now in the Gemini API and AI Studio [3][4]. Both handle real-time dialogue, visual inputs, asynchronous function calls, and background tool use without breaking conversational flow — and the standard model tops the Speech-to-Speech Quality Index at 82.6 [5].
Pricing is aggressive for developers ($0.005/min input, $0.018/min output audio), with an enterprise preview rolling into Gemini Enterprise. Notably, it ships with live transcription and translation across 70+ languages out of the box — a detail that matters more than it might first appear for anyone building multilingual meeting tools.
X reaction focused on the Workspace and Search integration angle, with several voices flagging the low latency and background tool-use as the real unlock for business applications, not just consumer voice assistants.
MLCommons Standardizes Benchmarks for RAG and Agentic Systems
MLCommons dropped MLPerf Inference v6.1 on September 16, and the headline addition is an End-to-End RAG benchmark that finally tests the full embed-retrieve-rerank-reason pipeline, plus an Edge Agentic Inference benchmark for multi-turn tasks with growing context [6][7]. Participation hit a record, signaling the industry now treats retrieval-augmented systems as core infrastructure worth measuring rigorously, not a side quest.
NVIDIA's submissions showed the payoff — 6.4x faster performance on Jetson AGX Thor for agentic workloads [8]. For anyone building or buying knowledge retrieval systems, this is the first time there's a standardized, comparable way to ask "how good is this RAG pipeline, actually?" rather than trusting vendor marketing.
AI researchers on X called this overdue, noting that enterprise buyers evaluating knowledge management and retrieval tools have lacked apples-to-apples benchmarks — until now.
Salesforce and NVIDIA Ship Koa, a CRM-Native Reasoning Model
At Dreamforce on September 15, Salesforce and NVIDIA unveiled Koa, Salesforce's first CRM-specific reasoning model, post-trained on NVIDIA Nemotron 3 Super using synthetic workflows drawn from 27 years of CRM data across 14+ industries — notably, no customer data was used in training [9][10]. Jensen Huang framed it at Dreamforce as a turning point: "now we can know everything and do anything" [11].
Folded into Agentforce, Koa delivers an 11% precision bump, 2.1x better context recall, and 3x fewer errors on CRM Bench tasks like opportunity updates and case routing [10]. Pilot customers include Formula 1 and UChicago Medicine, with general availability slated for Winter 2026.
Sales and enterprise voices on X zeroed in on the token-efficiency angle — a domain-specific reasoning model that wastes less context and makes fewer errors is a direct signal that vertical, workflow-native AI is starting to beat generalist models at their own game.
What This Means For Your Meetings
Today's news is really one story told four ways: voice, reasoning, retrieval, and domain-specific memory are converging into a single expectation — that AI systems should hear what happens, understand it, and recall it accurately months later. GPT-Live-1 and Gemini 3.8 Live both prove that full-duplex, low-latency voice is now table stakes, not a differentiator. The differentiator is what happens after the conversation — how well the system retrieves and reasons over what was said.
That's exactly where MLCommons' new RAG benchmark and Salesforce's Koa are instructive. Koa's gains came not from a bigger model, but from training on realistic workflow data and rewarding precise recall over generic fluency — the same principle that separates a meeting tool that transcribes words from one that actually builds usable institutional memory. A knowledge graph of your meetings is only as good as its retrieval quality, and now there's finally a standardized way to measure that instead of taking a vendor's word for it.
For teams relying on meeting intelligence platforms, the message is clear: the era of "just transcribe it" is over. The bar is now full-duplex conversational understanding, accurate speaker attribution, and retrieval that holds up to rigorous benchmarking — across your entire meeting history, not just the last call.
Key takeaway: Voice AI is commoditizing fast — the real competitive edge is moving to retrieval quality and long-term memory, which is precisely the ground meeting intelligence tools like Proudfrog are built to own.
Sources
- https://openai.com/index/gpt-6-astra-next-generation-work/
- https://openai.com/index/introducing-gpt-live-1-in-the-api/
- https://www.wired.com/story/openai-says-gpt-6-can-use-a-computer-better-than-a-human/
- https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
- https://ai.google.dev/gemini-api/docs/models/gemini-3-8-live
- https://mlcommons.org/2026/09/mlperf-inference-v6-1-results/
- https://techstrong.ai/articles/mlperf-inference-v6-1-updated-for-rag-agents-and-api-based-serving/
- https://developer.nvidia.com/blog/tensorrt-edge-llm-completes-the-mlperf-edge-agentic-benchmark-6-4x-faster-on-jetson-agx-thor/
- https://www.salesforce.com/agentforce/koa/
- https://blogs.nvidia.com/blog/jensen-huang-dreamforce/
- https://investor.salesforce.com/news/news-details/2026/Announcing-Koa-Salesforces-First-CRM-Reasoning-Model-Built-on-NVIDIA-Nemotron/default.aspx
Get the daily briefing
AI, knowledge graphs, and the future of work — in your inbox every morning.
No spam. Unsubscribe anytime.