Google DeepMind has released EmbeddingGemma 2, an open-weight embedding model that puts text, code, images, video frames and audio into a single shared vector space small enough to run on a phone. The model shipped on October 6 under the commercially permissive Apache 2.0 license, with weights published to Hugging Face as google/embeddinggemma-2.
The full multimodal configuration contains 740 million parameters, which Google splits into a 270 million text component, a 170 million vision encoder and a 300 million audio encoder. The encoders are modular, so developers load only the modalities an application needs. A text-only workload does not have to carry the rest of the model. The system is built on the Gemma 4 architecture and shares its text tokenizer and audio encoder, which Google says lets it run in one pipeline with a generative Gemma 4 model at a lower combined memory footprint.
The practical target is on-device retrieval. Google reports about 191MB of active RAM for text-only weights and roughly 567MB for the full multimodal model on a Pixel 11 Pro. Those are Google measurements on its own reference hardware, so treat them as vendor figures rather than independent benchmarks, but they indicate the deployment class: phones, laptops and other edge devices running semantic search with no cloud round trip and no per-query API bill.
Context window grows to 8,192 tokens, four times the first EmbeddingGemma. Under Google's default settings that budget represents up to about 5.5 minutes of audio, 29 images or 58 video frames, or interleaved combinations. All modalities share the same context budget, so mixing them reduces what fits for each individual type.
The storage math is the quieter headline. Embeddings ship at 768 dimensions and, using Matryoshka Representation Learning, can be truncated to 512, 256 or 128 dimensions, cutting vector storage by up to six times. Google says quality holds close to lossless down to 256 dimensions and reserves 128 for text-only work. The model card warns that truncated vectors must be re-normalized before cosine similarity and that queries and documents need matching dimensions.
Code retrieval gets a substantial upgrade. Google reports a jump from 68.76 to 78.68 on MTEB Code compared with the previous generation, a gain of 9.92 points, which matters for coding agents that need to index a private repository without uploading it. The model card also lists support for more than 100 languages and describes a zero-shot intent-routing use case that matches inputs against classification labels in milliseconds with no fine-tuning.
Distribution is already moving. Google says EmbeddingGemma 2 will arrive as a service on Android through ML Kit in the coming weeks with NPU acceleration on supported devices, and it already ships inside the AI Edge Gallery's Instant Media Search and Video Moments Finder. The first EmbeddingGemma passed 20 million downloads, and this release is clearly aimed at the same audience: developers who want search and retrieval that never leaves the device.
Comments (0)
Log in to join the discussion
Log InNo comments yet