Gemini 2.5: Enhancing Our Most Advanced Models Even Further
New Gemini 2.5 Capabilities
Native Audio Output and Improvements to Live API
Today, the Live API is introducing a preview version of audio-visual input and native audio out dialogue, enabling you to directly build conversational experiences with a more natural and expressive Gemini.
It also allows users to steer the model’s tone, accent, and speaking style. For example, you can instruct the model to adopt a dramatic voice when narrating a story. Additionally, it supports tool use to assist in searches on your behalf.
You can experiment with a set of early features, including:
- Affective Dialogue: The model detects emotion in the user’s voice and responds appropriately.
- Proactive Audio: The model can ignore background conversations and know when to engage.
- Thinking in the Live API: This feature leverages Gemini’s thinking capabilities to support more complex tasks.
We’re also releasing new previews for text-to-speech in 2.5 Pro and 2.5 Flash. These features offer first-of-their-kind support for multiple speakers, enabling text-to-speech with two voices via native audio out.
Like Native Audio dialogue, text-to-speech is expressive and can capture subtle nuances, such as whispers. It works in over 24 languages and can seamlessly switch between them.
