Google DeepMind Releases EmbeddingGemma 2, a New Multimodal Embedding Model
Google DeepMind has launched EmbeddingGemma 2, an open multimodal embedding model designed to process and unify text, images, video, and audio inputs into a single 768-dimensional vector space. This model, with 740 million parameters, combines a 270 million parameter text model with modular vision (170M) and audio (300M) encoders. It is built to run efficiently on consumer hardware like mobile devices and laptops, providing low-latency semantic representations for various on-device applications such as search, retrieval-augmented generation (RAG), classification, and clustering. EmbeddingGemma 2 supports over 100 languages and shows a 14% improvement in code tasks compared to its predecessor. It also features Matryoshka Representation Learning (MRL), allowing for truncated embeddings across 128d, 256d, 512d, and 768d, which can reduce vector storage costs by up to six times with minimal quality impact. The model has an 8K token context window, capable of processing minutes of audio or video, and uses light...