Gemini 2.5: Enhancing Our Most Advanced Models Further
New Gemini 2.5 Capabilities
Native Audio Output and Improvements to Live API
Today, the Live API is introducing a preview version of audio-visual input and native audio out dialogue, enabling you to build conversational experiences with a more natural and expressive Gemini.
This update allows users to adjust the model’s tone, accent, and speaking style. For instance, you can instruct the model to adopt a dramatic voice for storytelling. Additionally, it supports tool use to perform searches on your behalf.
You can experiment with a set of early features, including:
- Affective Dialogue: This feature enables the model to detect emotions in the user’s voice and respond appropriately.
- Proactive Audio: The model can ignore background conversations and recognize the right moments to respond.
- Thinking in the Live API: The model utilizes Gemini’s thinking capabilities to assist with more complex tasks.
We are also releasing new previews for text-to-speech in 2.5 Pro and 2.5 Flash. These features offer first-of-its-kind support for multiple speakers, allowing text-to-speech with two distinct voices via native audio out.
Similar to Native Audio dialogue, text-to-speech is expressive and capable of capturing subtle nuances, such as whispers. It works in over 24 languages and can seamlessly switch between them.
