- 开始
- 2026年9月23日可信度 90%来自来源
Gemini 3.8 Flash TTS 与 3.8 Flash-Lite TTS 发布
由原文自动翻译
原标题: Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS Released
- 谷歌于 2026 年 9 月 23 日推出 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS,称其为“我们迄今表现力最强的音频生成模型”:Flash TTS 面向创意指导和角色设计,Flash-Lite TTS 面向高量、低成本的配音和语音智能体[1]
- Flash TTS 可根据 100 多种语言和方言的自然语言提示创建新声音,提供 2,000 多种现成声音,并可在原声主人录制的口头同意与参考发音人匹配后,通过 30 秒样本复刻其声音[1]
- 两款模型都能逐行接受关于节奏、情感和口音的指导,在数小时的音频中保持声音稳定,可根据单一脚本呈现双人对话场景,并能呈现<笑声>、<叹气>等提示,以及“嗯”这类附和声[1]
- 两款模型已在 Gemini API 和 Google AI Studio 中推出,Flash TTS 登陆 Gemini Notebook,Flash-Lite TTS 登陆 Google Vids;Gemini Enterprise 的访问权限即将开放[1][5]
- Gemini API 更新日志将此次发布列在 9 月 22 日,并新增了一个 Voices 接口;Flash-Lite TTS 取代了此前的 Gemini 3.1 Flash TTS 预览版[2]
- AI Studio 的声音复刻功能在英国、欧洲经济区、印度、得克萨斯州和伊利诺伊州受到限制[5]
值得关注的功能
- 在 Hume AI 的语音设计基准测试中综合得分第一(71.4),口音建模项目也排名第一(60.8);在 Hume 的综合质量指数上,Flash 和 Flash-Lite 分别位列第一和第二,不过在 Hume 的各项单独语音质量类别中,ElevenLabs 的得分更高[1][5]
- 截至 2026 年 12 月 31 日,Flash TTS 的价格为每百万文本输入 token 0.50 美元、每百万音频输出 token 9.00 美元(约合每 10 秒 0.00225 美元);2027 年 1 月 1 日起价格翻倍,分别升至 1.00 美元和 18.00 美元;Flash-Lite TTS 的价格为 0.50 美元和 6.00 美元,届时将升至 1.00 美元和 12.00 美元[3]
- 在 Gemini API 上,文本输入最多支持 8,192 个 token,输出最多 16,384 个 token[4]
- 每段音频都带有 SynthID 水印,复刻的声音还带有 C2PA 凭证[1][5]
- 集成这些模型的合作伙伴包括 Figma、HeyGen、Wondercraft、99.co 和 Ollang[1]
参考资料 6可信度 87%总体可信度: 87%该图钉的来源和参考资料对其日期的支持程度。包括来源在内的 6 份资料对图钉开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半显示所有可信度不低于 75% 的图钉
第一项始终是图钉的来源。总体可信度是各份资料对上方所用开始和结束时间支持程度的加权平均;资料每比最新的一份旧 180 天,权重减半。
- [1]90%blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speechblog.google· 发布于 2026年9月28日· 开始 2026年9月23日 ✓· 占评分 17%
谷歌博文日期为“2026 年 9 月 23 日”,文中称两款模型均“从今天起开始推出”;The Next Web[6] 报道称“谷歌于周三发布了 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS”。Gemini API 更新日志则将正式发布记录在早一天,即 9 月 22 日。
- [2]85%Release notes | Gemini APIai.google.dev· 添加于 2026年9月28日· 占评分 17%
The Gemini API changelog under September 22, 2026 lists both models as generally available with the new Voices endpoint, voice design, voice replication and an extended library of 150+ voices, and says Flash-Lite TTS replaces gemini-3.1-flash-tts-preview.[1]
- [3]85%Gemini Developer API pricingai.google.dev· 添加于 2026年9月28日· 占评分 17%
Flash TTS costs $0.50 per million text input tokens and $9.00 per million audio output tokens through 31 December 2026, then $1.00 and $18.00; Flash-Lite TTS $0.50 and $6.00, then $1.00 and $12.00.[1]
- [4]85%Gemini 3.8 Flash TTS | Gemini APIai.google.dev· 添加于 2026年9月28日· 占评分 17%
The API model page for gemini-3.8-flash-tts: text in, audio out, an 8,192-token input limit and 16,384 output tokens (Gemini API serving limit), latest update September 2026.[1]
- [5]85%Gemini can now clone your voice and perform scripts like an actorandroidauthority.com· 发表于 2026年9月24日· 占评分 16%
Android Authority on the rollout to Gemini Notebook and Google[2][3][4] Vids; it notes ElevenLabs still scored higher in Hume's individual voice-quality and human-like-variation categories, and that AI Studio voice replication is blocked in the UK, the EEA, India, Texas and Illinois.
建议更正
有遗漏或错误吗?用你自己的话说明:能佐证此图钉的链接、不同的开始或结束日期及理由,或缺失、有误的信息。AI 会对照此图钉的来源进行核实,搜索更好的来源,并添加任何支持你说法的页面。图钉自身的来源仍然最重要。AI 也会查看图片:显示的是别的东西或显示效果差的图片会被移到后面或替换。