Google launched EmbeddingGemma 2 on Monday, which is the new generation of its multimodal embedding model for device deployment. It can process text, images, audio, and video locally on hardware, with a focus on data privacy protection and offline inference capabilities, making it directly attractive to mobile AI application developers.
The new model has 740 million parameters and is built on the Gemma architecture. It is released under the Apache 2.0 commercial-friendly license as open-source software.
Google stated that this model improved by 9.92 points in the MTEB code evaluation compared to the previous generation, and achieved leading benchmark results in the sub-1B large-scale multimodal embedding models. In some indicators, it even surpassed specialized models that are more than twice its size.

EmbeddingGemma 2 will directly affect the application development landscape centered around scenarios that prioritize local search and privacy, such as the RAG (retrieval-enhanced generation) pipeline.
Since the release of the previous generation product EmbeddingGemma, it has accumulated over 20 million downloads. This upgrade extends its capabilities from pure text to multimodal formats, and is expected to further enhance Google's developer ecosystem advantages in the field of AI infrastructure on the client side.
Parameter simplification makes end-side inference costs controllable.
The core design logic of EmbeddingGemma 2 is to achieve high-quality multimodal reasoning with limited hardware resources. The model adopts a modular architecture; a pure text workload requires only about 270 million parameters. The visual encoder (170 million parameters) and the audio encoder (300 million parameters) can be loaded as needed, with the total number of parameters for the full-modal version being 740 million.
Google's published test data shows that after quantitative processing, when the model runs on Google Pixel 11 Pro, it requires only about 191MB of active memory for pure text and approximately 567MB of memory in full multimodal mode. For consumer-grade devices with limited memory, this metric is of practical significance for deployment.
In terms of storage efficiency, the model introduces Matryoshka to represent learning (MRL) technology, allowing developers to dynamically truncate the output vector from 768 dimensions to 512, 256, or 128 dimensions, achieving a maximum storage compression ratio of up to 6 times that of local vector databases.
The context window has been expanded by four times, enhancing the multimodal processing capabilities.
EmbeddingGemma Expanding the context window to 8K token is four times that of the previous generation, allowing for direct processing on local hardware of audio up to about 5.5 minutes long, 29 images, 58 frames of video, or a mixed input of these modalities.
This capability makes cross-modal semantic retrieval scenarios feasible on the client side, for example, locating specific video segments through voice memos, or retrieving multi-hour audio recordings based on text queries.
In terms of code retrieval capabilities, the model's score increased from 68.76 to 78.68 in the MTEB Code benchmark test, representing a rise of 9.92 points. This improvement is suitable for scenarios such as local code library indexing, semantic code search, and programming intelligence health checks.
Google stated that this model maintains comparable multi-language text embedding performance to its predecessor, while also seeing improvements in quality across dimensions such as images, videos, documents, and audio.
Cooperates with Gemma to support fully offline RAG pipelines.
EmbeddingGemma 2 and Google's generative model Gemma share a text tokenizer and audio encoder, which can be combined to run in the same inference pipeline, achieving lower overall memory usage.
This design provides developers with a path to build a complete RAG pipeline locally on the device, without the need to connect to a cloud server at any point; the data never leaves the terminal device.
Google points out that locally generated embeddings help to ensure data privacy, reduce pipeline latency, and enable cross-modal search and retrieval functions to operate in a completely offline state. This alignment coincides with the current focus of some companies and developers on data sovereignty and privacy compliance.
The model has been made open-source. Relevant evaluation metrics and model information can be found through EmbeddingGemma 2 model card.











