快讯
Google Gemini 3.8 Live 端到端语音模型亮相
DeepLearningAI 介绍,Google 的 Gemini 3.8 Live 模型采用语音到语音架构,在一个系统中完成聆听、推理和回应。推文称,Extended Thinking 版本在 Artificial Analysis 的 Speech to Speech Index 中排名第一;标准版在真人盲评现场对话中排名第二,输入音频价格为每小时 0.84 美元,为该指数最低。两个版本还支持图像和视频输入。
查看原文
Most voice assistants work like a relay race: speech becomes text, text goes to a model, the answer becomes speech again. Every handoff adds a pause. ⏱️ Speech-to-speech models skip the relay. Google's new Gemini 3.8 Live models listen, reason, and respond in one system. 🎯 The Extended Thinking version ranks first on Artificial Analysis' Speech to Speech Index 🗣️ The standard version ranks second in blind live conversations judged by people 💰 The standard version costs $0.84 per hour of input audio, the lowest in the index Both models also take in image and video input. Picture asking an assistant for help with whatever is on your screen, and it simply answers. 📱 Read the full story in The Batch 👉 https://hubs.la/Q04ztmlL0 #DeepLearningAI #VoiceAgents #AI