AI & LLMs

Google's EmbeddingGemma 2 unifies multimodal search in a tiny on-device model

Google released EmbeddingGemma 2, an open 740M-parameter model that maps text, code, images, video and audio into one vector space for efficient local retrieval.

Smartphone screen showing a neural network connecting various media type icons.
Illustration created for this article

Google has released EmbeddingGemma 2, an open-source model designed to handle text, code, image, video, and audio search within a single lightweight framework. Published on October 6, 2026, this 740-million-parameter model allows developers to run complex retrieval tasks directly on mobile devices without relying on cloud infrastructure or extensive preprocessing steps like transcription.

What happened

The new model maps five distinct input types into a shared 768-dimensional vector space. This architecture eliminates the need to generate text captions for images or transcripts for audio files before they can be searched. In internal testing on a Pixel 11 Pro, the fully quantized model consumed approximately 567MB of active RAM. Google released the model weights under the Apache 2.0 license, making them freely available for commercial and private use. Deployment tools are already accessible via LiteRT and MediaPipe Tasks, with an Android ML Kit integration featuring NPU acceleration expected in the coming weeks.

A key architectural feature is its modularity. Developers do not need to load the entire 740-million-parameter model if their application only requires specific data types. The base encoder for text and code uses 270 million parameters and occupies about 191MB of RAM. Adding the vision encoder for images and video increases the count to 440 million parameters, while adding the audio encoder brings it to 570 million. Loading all encoders results in the full parameter count. Because every configuration projects into the same embedding space, teams can start with a text-only index and add image or audio capabilities later without re-embedding existing data.

Key details

  • The model supports a context window of 8,192 tokens, up from 2,048 in the previous version, allowing it to process up to 5.5 minutes of audio, 29 images, or 58 video frames in a single input.
  • Video processing defaults to sampling one frame per second, meaning the 58-frame limit covers just under a minute of footage.
  • Google trained the model using Matryoshka Representation Learning, which allows embeddings to be truncated to 512, 256, or 128 dimensions without retraining.
  • Truncating to 256 dimensions reduces a million-vector index from roughly 1.5GB to 500MB while retaining most quality for text and code, and about 95% quality for multimodal retrieval.
  • At 128 dimensions, retrieval quality drops to around 90% for text and code, and about 75% for image, video, and speech, requiring careful testing before deployment.
  • The model achieved an MTEB Code score of 78.68, a significant improvement over the original EmbeddingGemma’s score of 68.76.

Background

Retrieval-Augmented Generation (RAG) systems typically rely on embedding models to convert data into numerical vectors that represent semantic meaning. Traditionally, handling different media types required separate pipelines: optical character recognition for images, automatic speech recognition for audio, and standard tokenization for text. These intermediate steps add latency, cost, and potential points of failure. Embedding models map these inputs into a high-dimensional vector space where similar concepts are located close together, enabling search engines to find relevant information based on meaning rather than just keyword matches.

Matryoshka Representation Learning is a technique that addresses the storage burden of these vectors. By training the model to maintain useful information even when the vector dimensions are reduced, it allows developers to trade a small amount of retrieval accuracy for significant savings in storage and memory usage. This is particularly critical for edge devices like smartphones or local servers where resources are constrained compared to cloud data centers.

Why it matters

For teams running self-hosted software, this development reduces the complexity of building multimodal search features. Previously, creating a system that could search through meeting recordings, scanned documents, and code repositories required maintaining multiple specialized models and preprocessing services. EmbeddingGemma 2 consolidates these into a single dependency. The ability to load only the necessary encoders means that a documentation portal might only need the text encoder, keeping its memory footprint minimal, while a media asset management tool could activate the vision encoder without changing the underlying database schema.

The efficiency gains also impact how local agents operate. By using the model for classification tasks via MediaPipe Decision, applications can evaluate hundreds of options in milliseconds without invoking a large language model. This prevents unnecessary token consumption and reduces latency in decision-making loops. For IT managers concerned about data privacy, the ability to perform all indexing and searching on-device ensures that sensitive audio or visual data never leaves the local environment, aligning with strict compliance requirements.

What you can do

  • Evaluate your current RAG pipelines to identify if separate preprocessing steps for audio or images can be replaced by direct embedding with EmbeddingGemma 2.
  • Test the modular encoders to determine if your application can operate with only the text and code base, saving nearly 300MB of RAM compared to the full model.
  • Experiment with Matryoshka truncation at 256 dimensions to reduce index storage costs, verifying that the 5% drop in multimodal retrieval quality is acceptable for your use case.
  • Prepare for the upcoming Android ML Kit integration if you are developing mobile applications, ensuring your hardware supports NPU acceleration for optimal performance.
  • Consider using the model for classification tasks in agent workflows to reduce reliance on larger generative models for simple decision-making steps.
  • Review the Apache 2.0 license terms to confirm compliance with your organization’s open-source usage policies before integrating the weights into production systems.

More news

All news