News Microsoft Puts a Petaflop in Your Lap: Surface Laptop Ultra Ships Oct 16 From $2,599 With 128GB Unified Memory and Local 120B-Parameter Models AI Agents Docker Is Now Shipping AI Agents: docker-agent Turns a YAML File and an OCI Registry Into an Agent Distribution Channel Gemini Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM Business Alphabet's Isomorphic Labs Is in Talks to Raise at a $40 Billion Floor, Five Months After a $2.1 Billion Series B News Boston Dynamics Names Ex-Alexa Chief Rohit Prasad as CEO as Hyundai Gears Up to Build 30,000 Humanoid Robots by 2028 AI Agents Manus Parent Butterfly Effect Raises More Than $500 Million From Boyu and IDG in Its First Outside Funding Since Returning to Independence Video Generators CDN Veteran Wangsu Puts 300 Million Yuan Into Video Model Startup Sand.ai - a $960 Million Implied Valuation for a Company That Lost $21 Million Last Half Security AI Agents Now Write One in Three Pull Requests - GitHub Is Rebuilding Secret Protection Around That Fact News Microsoft Puts a Petaflop in Your Lap: Surface Laptop Ultra Ships Oct 16 From $2,599 With 128GB Unified Memory and Local 120B-Parameter Models AI Agents Docker Is Now Shipping AI Agents: docker-agent Turns a YAML File and an OCI Registry Into an Agent Distribution Channel Gemini Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM Business Alphabet's Isomorphic Labs Is in Talks to Raise at a $40 Billion Floor, Five Months After a $2.1 Billion Series B News Boston Dynamics Names Ex-Alexa Chief Rohit Prasad as CEO as Hyundai Gears Up to Build 30,000 Humanoid Robots by 2028 AI Agents Manus Parent Butterfly Effect Raises More Than $500 Million From Boyu and IDG in Its First Outside Funding Since Returning to Independence Video Generators CDN Veteran Wangsu Puts 300 Million Yuan Into Video Model Startup Sand.ai - a $960 Million Implied Valuation for a Company That Lost $21 Million Last Half Security AI Agents Now Write One in Three Pull Requests - GitHub Is Rebuilding Secret Protection Around That Fact

Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM

Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM

Google DeepMind released EmbeddingGemma 2 on October 6, an open-weight Apache 2.0 embedding model with 740 million parameters that maps text, code, images, video frames and audio into a single 768-dimensional space. Google reports about 191MB of active RAM for a text-only setup on a Pixel 11 Pro, and vectors can be truncated to 128 dimensions for up to 6x storage savings.

Google DeepMind has released EmbeddingGemma 2, an open-weight embedding model that puts text, code, images, video frames and audio into a single shared vector space small enough to run on a phone. The model shipped on October 6 under the commercially permissive Apache 2.0 license, with weights published to Hugging Face as google/embeddinggemma-2.

The full multimodal configuration contains 740 million parameters, which Google splits into a 270 million text component, a 170 million vision encoder and a 300 million audio encoder. The encoders are modular, so developers load only the modalities an application needs. A text-only workload does not have to carry the rest of the model. The system is built on the Gemma 4 architecture and shares its text tokenizer and audio encoder, which Google says lets it run in one pipeline with a generative Gemma 4 model at a lower combined memory footprint.

The practical target is on-device retrieval. Google reports about 191MB of active RAM for text-only weights and roughly 567MB for the full multimodal model on a Pixel 11 Pro. Those are Google measurements on its own reference hardware, so treat them as vendor figures rather than independent benchmarks, but they indicate the deployment class: phones, laptops and other edge devices running semantic search with no cloud round trip and no per-query API bill.

Context window grows to 8,192 tokens, four times the first EmbeddingGemma. Under Google's default settings that budget represents up to about 5.5 minutes of audio, 29 images or 58 video frames, or interleaved combinations. All modalities share the same context budget, so mixing them reduces what fits for each individual type.

The storage math is the quieter headline. Embeddings ship at 768 dimensions and, using Matryoshka Representation Learning, can be truncated to 512, 256 or 128 dimensions, cutting vector storage by up to six times. Google says quality holds close to lossless down to 256 dimensions and reserves 128 for text-only work. The model card warns that truncated vectors must be re-normalized before cosine similarity and that queries and documents need matching dimensions.

Code retrieval gets a substantial upgrade. Google reports a jump from 68.76 to 78.68 on MTEB Code compared with the previous generation, a gain of 9.92 points, which matters for coding agents that need to index a private repository without uploading it. The model card also lists support for more than 100 languages and describes a zero-shot intent-routing use case that matches inputs against classification labels in milliseconds with no fine-tuning.

Distribution is already moving. Google says EmbeddingGemma 2 will arrive as a service on Android through ML Kit in the coming weeks with NPU acceleration on supported devices, and it already ships inside the AI Edge Gallery's Instant Media Search and Video Moments Finder. The first EmbeddingGemma passed 20 million downloads, and this release is clearly aimed at the same audience: developers who want search and retrieval that never leaves the device.

Comments (0)

Log in to join the discussion

Log In

No comments yet