🤖 AI & Beyond

Revamped Gemini Models Now Production-Ready, Lowered Pricing for 1.5 Pro, Enhanced Rate Limits, and Additional Updates – Google Developers Blog

Sure! Here’s a revised version of the content with promotional calls-to-action removed or rephrased:


Today, we’re releasing two updated production-ready Gemini models: Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002, along with significant enhancements:

  • 50% reduced price on 1.5 Pro (for inputs and outputs of prompts <128K)

  • 2x higher rate limits on 1.5 Flash and ~3x higher on 1.5 Pro
  • 2x faster output and 3x lower latency
  • Updated default filter settings

These new models build on our latest experimental releases and include meaningful improvements to the Gemini 1.5 models introduced at Google I/O in May. Developers can access these latest models at no cost through Google AI Studio and the Gemini API. For larger organizations and Google Cloud customers, the models are also available on Vertex AI.

Overall quality has improved, particularly in areas such as math, long context, and vision. The Gemini 1.5 series is designed for robust performance across a variety of text, code, and multimodal tasks. For instance, Gemini models can synthesize information from lengthy PDFs, answer questions regarding large code repositories, and analyze hour-long videos to create useful summaries and content.

With the latest updates, 1.5 Pro and Flash are now better, faster, and more cost-effective for production use. We observe a ~7% increase in MMLU-Pro, a more challenging version of the popular MMLU benchmark. Both models show a significant ~20% improvement on math benchmarks like MATH and HiddenMath. They also excel (ranging from ~2-7%) in evaluations measuring visual understanding and Python code generation.

Additionally, we have enhanced the overall helpfulness of model responses while upholding stringent content safety standards. This results in fewer refusals and more useful answers across various topics. Both models now feature a more concise style, reflecting developer feedback aimed at simplifying usage and reducing costs. For applications such as summarization and question answering, the default output length of the updated models is approximately 5-20% shorter than prior versions. In chat-based products, where users might favor longer responses, additional guidance is available on how to encourage more verbose and conversational outputs.

For further information on how to transition to the latest versions of Gemini 1.5 Pro and 1.5 Flash, the Gemini API models page includes comprehensive details.

Gemini 1.5 Pro
We continue to be impressed by the creative applications of Gemini 1.5 Pro’s 2 million token long context window and multimodal capabilities. There are numerous new use cases yet to be explored, from video analysis to processing lengthy PDFs. As of October 1st, 2024, we are announcing a 64% reduction in input token prices, a 52% reduction in output tokens, and a 64% reduction in incremental cached tokens for the Gemini 1.5 Pro model on prompts with fewer than 128K tokens. Coupled with context caching, these changes are intended to lower costs associated with utilizing Gemini.

Increased rate limits
To enhance the building experience for developers, we are increasing the paid tier rate limits for 1.5 Flash to 2,000 RPM and for 1.5 Pro to 1,000 RPM, up from 1,000 and 360, respectively. In the coming weeks, we anticipate ongoing increases in Gemini API rate limits to enable more extensive development with Gemini.

2x faster output and 3x less latency
Alongside core improvements, we have significantly reduced latency with 1.5 Flash and increased the output tokens per second, facilitating new use cases with our most advanced models.

Updated filter settings
Since Gemini’s initial launch in December 2023, building a safe and reliable model has been a priority. The latest versions (the -002 models) feature enhancements in the model’s ability to adhere to user instructions while ensuring safety. Developers will continue to have access to a suite of safety filters that can be applied to our models. For the recently released models, filters will not be applied by default, allowing developers to configure settings to best suit their needs.

Gemini 1.5 Flash-8B Experimental updates
We are rolling out an improved version of the Gemini 1.5 model announced in August, called Gemini-1.5-Flash-8B-Exp-0924. This enhanced version includes considerable performance boosts across both text and multimodal use cases. It is now accessible through Google AI Studio and the Gemini API.

The positive feedback from developers regarding 1.5 Flash-8B has been remarkable, and we will continue to refine our experimental and production release pipeline based on this input.

We are enthusiastic about these updates and look forward to seeing the innovative solutions you will create with the new Gemini models! Additionally, Gemini Advanced users will soon have access to a chat-optimized version of Gemini 1.5 Pro-002.