Gemini Pioneers Innovation with Speedier Models, Extended Context, AI Agents, and Additional Features
1.5 Flash excels at summarization, chat applications, image and video captioning, data extraction from long documents and tables, and more. This is due to its training via a process called “distillation,” where the most essential knowledge and skills from a larger model are transferred to a smaller, more efficient model. For further details about 1.5 Flash, please refer to our updated Gemini 1.5 technical report and explore the Gemini technology page for information on its availability and pricing.
Significantly improving 1.5 Pro
Over the last few months, we’ve made substantial enhancements to 1.5 Pro, our top model for general performance across a wide range of tasks. In addition to extending its context window to 2 million tokens, we’ve improved its capabilities in code generation, logical reasoning and planning, multi-turn conversation, and audio and image understanding through data and algorithmic advancements. These enhancements have resulted in notable improvements on both public and internal benchmarks for each of these tasks.
1.5 Pro is now adept at following increasingly complex and nuanced instructions, including those that specify product-level behavior related to role, format, and style. We’ve enhanced control over the model’s responses for specific use cases, such as shaping the persona and response style of a chat agent or automating workflows through multiple function calls. Additionally, we’ve enabled users to guide model behavior by setting system instructions.
The Gemini API and Google AI Studio now include audio understanding, allowing 1.5 Pro to reason across image and audio for videos uploaded in Google AI Studio. We are also in the process of integrating 1.5 Pro into Google products, including Gemini Advanced and Workspace applications. For more information about 1.5 Pro, please consult our updated Gemini 1.5 technical report and visit the Gemini technology page.
Gemini Nano understands multimodal inputs
Gemini Nano is broadening its capabilities beyond text-only inputs to incorporate images as well. Starting with Pixel, applications using Gemini Nano with Multimodality will be able to comprehend the world in a manner similar to human perception — not only through text but also through sight, sound, and spoken language. For more insights into Gemini 1.0 Nano on Android, please refer to the relevant resources from Hotnchill.
