跳到正文
原文
Google DeepMind·· 2026-08-27精选AI 评分64

Google DeepMind 发布 Gemini 3.5 Transcribe 语音转写模型

Intelligent transcription with Gemini 3.5 Transcribe

AI 导读

Google DeepMind 发布 Gemini 3.5 Transcribe,定位为更精准、支持智能交互的实时语音转写模型。模型可将原始音频直接转为格式化文本,自动清除口头语和自我纠正,支持自定义词表、超过 85 种语言、最多三名说话人区分,并可通过 function calling 调用其他 Gemini 模型。

推荐理由

官方公布了模型在噪声环境、自定义词表和多说话人场景的能力细节与 API 入口,读者可据此评估语音转写工作流能否迁移。

来源:Google DeepMind · deepmind.google