Google releases EmbeddingGemma 2: a 740-million-parameter multimodal embedding model, focusing on end-side privacy search
Wallstreetcn
34m ago
Ai Focus
Google launched EmbeddingGemma 2 on Monday, which is a multimodal embedding model designed for device deployment. It can process text, images, audio, and video locally on hardware, with an emphasis on privacy protection and offline inference. The model has 740 million parameters and is based on the Gemma architecture. It is released under the Apache 2.0 license.
Helpful
No.Help

Google launched EmbeddingGemma 2 on Monday, which is the new generation of its multimodal embedding model for device deployment. It can process text, images, audio, and video locally on hardware, with a focus on data privacy protection and offline inference capabilities, making it directly attractive to mobile AI application developers.

The new model has 740 million parameters and is built on the Gemma architecture. It is released under the Apache 2.0 commercial-friendly license as open-source software.

Google stated that this model improved by 9.92 points in the MTEB code evaluation compared to the previous generation, and achieved leading benchmark results in the sub-1B large-scale multimodal embedding models. In some indicators, it even surpassed specialized models that are more than twice its size.

EmbeddingGemma 2 will directly affect the application development landscape centered around scenarios that prioritize local search and privacy, such as the RAG (retrieval-enhanced generation) pipeline.

Since the release of the previous generation product EmbeddingGemma, it has accumulated over 20 million downloads. This upgrade extends its capabilities from pure text to multimodal formats, and is expected to further enhance Google's developer ecosystem advantages in the field of AI infrastructure on the client side.

Parameter simplification makes end-side inference costs controllable.

The core design logic of EmbeddingGemma 2 is to achieve high-quality multimodal reasoning with limited hardware resources. The model adopts a modular architecture; a pure text workload requires only about 270 million parameters. The visual encoder (170 million parameters) and the audio encoder (300 million parameters) can be loaded as needed, with the total number of parameters for the full-modal version being 740 million.

Google's published test data shows that after quantitative processing, when the model runs on Google Pixel 11 Pro, it requires only about 191MB of active memory for pure text and approximately 567MB of memory in full multimodal mode. For consumer-grade devices with limited memory, this metric is of practical significance for deployment.

In terms of storage efficiency, the model introduces Matryoshka to represent learning (MRL) technology, allowing developers to dynamically truncate the output vector from 768 dimensions to 512, 256, or 128 dimensions, achieving a maximum storage compression ratio of up to 6 times that of local vector databases.

The context window has been expanded by four times, enhancing the multimodal processing capabilities.

EmbeddingGemma Expanding the context window to 8K token is four times that of the previous generation, allowing for direct processing on local hardware of audio up to about 5.5 minutes long, 29 images, 58 frames of video, or a mixed input of these modalities.

This capability makes cross-modal semantic retrieval scenarios feasible on the client side, for example, locating specific video segments through voice memos, or retrieving multi-hour audio recordings based on text queries.

In terms of code retrieval capabilities, the model's score increased from 68.76 to 78.68 in the MTEB Code benchmark test, representing a rise of 9.92 points. This improvement is suitable for scenarios such as local code library indexing, semantic code search, and programming intelligence health checks.

Google stated that this model maintains comparable multi-language text embedding performance to its predecessor, while also seeing improvements in quality across dimensions such as images, videos, documents, and audio.

Cooperates with Gemma to support fully offline RAG pipelines.

EmbeddingGemma 2 and Google's generative model Gemma share a text tokenizer and audio encoder, which can be combined to run in the same inference pipeline, achieving lower overall memory usage.

This design provides developers with a path to build a complete RAG pipeline locally on the device, without the need to connect to a cloud server at any point; the data never leaves the terminal device.

Google points out that locally generated embeddings help to ensure data privacy, reduce pipeline latency, and enable cross-modal search and retrieval functions to operate in a completely offline state. This alignment coincides with the current focus of some companies and developers on data sovereignty and privacy compliance.

The model has been made open-source. Relevant evaluation metrics and model information can be found through EmbeddingGemma 2 model card.

Tip
$0
Like
0
Save
0
Views 19
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
The next infrastructure challenge for AI is electricity: DIGITIMES analyzes the transformation of 800VDC data centers
A new report from DIGITIMES Intelligence states that as the computing power of AI continues to expand, the challenges for data centers are shifting from processors to power supply. The report analyzes how the 800VDC architecture can reshape the power transmission in the next generation of AI data centers, and points out that GaN, SiC as well as the synergy with servers OEM and ODM will become key to competition.
PR Newswire
·2026-10-08 10:11:59
4
The LAPTOP token of Hunter Biden rose by 21% after the collapse forensic report.
The meme coin of Hunter Biden rose by 21.11% on the same day after the release of a forensic report regarding its historic 99% plunge. The report attributed the collapse mainly to the extremely low liquidity among market makers, stating that this made it "1,000 times easier to exit" the market.
Coinpedia
·2026-10-08 09:50:45
17
Intel network card driver exposed to be damaging connections; IPv6, OneNote, OneDrive, and Microsoft suffer for 365 days
According to a report by the German technology blog BornCity, a user discovered that the Intel network card driver installed through Windows Update contained a flaw Bug which silently damaged the network connection IPv6, resulting in the inability to use services such as Microsoft OneNote, Microsoft 365, OneDrive, Teams, and Azure portals. The issue was traced back to the traffic control driver in Intel Connectivity Performance Suite ( ICPS). After testing with Intel's official universal driver, the functionality returned to normal, but the version notes did not mention the fix for the defect related to IPv6.
The Block
·2026-10-08 09:50:43
14
Goldman Sachs warns: Rising interest rates will put pressure on the stock market and drag down consumer spending
Goldman Sachs stated that higher interest rates may weaken the "wealth effect" on the stock market, thereby suppressing U.S. consumer spending. The bank expects that if interest rates remain high, consumer growth, residential real estate investment, and capital expenditure could all be affected by 2027.
Businessinsider
·2026-10-08 09:22:43
18
View More