Google DeepMind at the ICML 2024 Conference
Research
Published
19 July 2024
Exploring AGI, the challenges of scaling and the future of multimodal generative AI
Next week, the artificial intelligence (AI) community will convene for the 2024 International Conference on Machine Learning (ICML). Running from July 21-27 in Vienna, Austria, this conference serves as an international platform for showcasing the latest advancements, exchanging ideas, and shaping the future of AI research. This year, teams from across Google DeepMind will present more than 80 research papers. At our booth, we’ll showcase our multimodal on-device model, Gemini Nano, our new family of AI models for education called LearnLM, and demo TacticAI, an AI assistant designed to assist with football tactics. Here, we introduce some of our oral, spotlight, and poster presentations:
Defining the path to AGI
What is artificial general intelligence (AGI)? This term refers to an AI system that is at least as capable as a human at most tasks. As AI models continue to evolve, understanding what AGI may look like in practical terms will become increasingly vital. We’ll present a framework for classifying the capabilities and behaviors of AGI models. Depending on their performance, generality, and autonomy, our paper categorizes systems ranging from non-AI calculators to emerging AI models and other novel technologies. We’ll also demonstrate that open-endedness is essential for building generalized AI that exceeds human capabilities. While many recent AI advancements were propelled by existing Internet-scale data, open-ended systems can create new discoveries that expand human knowledge.
At ICML, we’ll demo Genie, a model that can generate a variety of playable environments based on text prompts, images, photos, or sketches.
Scaling AI systems efficiently and responsibly
Developing larger, more capable AI models necessitates more efficient training methods, closer alignment with human preferences, and enhanced privacy safeguards. We’ll showcase how employing classification instead of regression techniques simplifies the scaling of deep reinforcement learning systems and achieves state-of-the-art performance across various domains. Additionally, we propose a novel approach that predicts the consequences of a reinforcement learning agent’s actions, facilitating rapid evaluation of new scenarios. Our researchers will present an alignment-maintaining strategy that reduces the need for human oversight, alongside a new game theory-based method for fine-tuning large language models (LLMs), which better aligns LLM outputs with human preferences. We critique the training of models on public data combined only with “differentially private” fine-tuning, arguing that this method may not deliver the privacy or utility often claimed.
VideoPoet is a large language model designed for zero-shot video generation.
New approaches in generative AI and multimodality
Generative AI technologies and multimodal capabilities are broadening the creative possibilities within digital media. We’ll introduce VideoPoet, which utilizes an LLM to produce high-quality video and audio from multimodal inputs, including images, text, audio, and other video content. We’ll also present Genie (generative interactive environments), capable of generating various playable environments for training AI agents based on text prompts, images, photos, or sketches. Lastly, we introduce MagicLens, an innovative image retrieval system that employs text instructions to fetch images with more nuanced relationships beyond mere visual similarity.
Supporting the AI community
We’re proud to sponsor ICML and nurture a diverse AI and machine learning community by supporting initiatives related to Disability in AI, Queer in AI, LatinX in AI, and Women in Machine Learning. If you’re attending the conference, visit the Google DeepMind and Google Research booths to connect with our teams, watch live demos, and learn more about our research.
Learn more
