TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Google announced EmbeddingGemma 2 on Oct. 6, 2026, describing it as an open model that maps text, code, images, audio and video into a shared embedding space. The 740-million-parameter model is designed for on-device use and is released under the Apache 2.0 license; performance and memory figures cited so far come from Google.
Google announced EmbeddingGemma 2 on Oct. 6, presenting a 740-million-parameter multimodal embedding model that maps text, code, images, audio and video into one shared space. Released under the Apache 2.0 license and designed for on-device inference, it is aimed at developers building local search and retrieval tools without sending all input to a cloud service.
Google says the model is built on the Gemma 4 architecture and supports cross-modal searches, such as finding a video clip using a voice memo or searching audio recordings with a text query. The company describes it as natively multimodal, meaning the supported data types can be processed within a single model rather than requiring separate embedding systems for each type.
The model has an 8,192-token context window, which Google says is four times that of EmbeddingGemma. According to the company, it can process up to 5.5 minutes of audio, 29 images or 58 video frames within that window, including interleaved combinations. Google also says developers can use a text-only version requiring 270 million parameters, with optional vision and audio encoders listed at 170 million and 300 million parameters respectively.
Google reports a 9.92-point increase on MTEB Code compared with the original EmbeddingGemma, rising from 68.76 to 78.68. It also says vector outputs can be reduced from 768 dimensions to 512, 256 or 128 through Matryoshka Representation Learning, with up to sixfold storage reduction. On a Google Pixel 11 Pro, the company reports active RAM use of about 191 MB for quantized text-only weights and 567 MB for the full multimodal model. These are Google-provided figures; independent validation is not included in the announcement.
Local Search Across Media Types
The release targets developers who want apps to search across different kinds of personal or work data on a device. A shared embedding space could let an app connect a spoken note with a related image or video, or retrieve code and documents using natural-language queries. If the model performs as Google reports, that could simplify systems that currently need separate models and indexing steps for different media.
Running embeddings locally may also reduce the amount of user data sent to remote services and allow search to work when a device is offline. That does not by itself guarantee privacy: developers still determine what data an app stores, how it is protected and whether other parts of a workflow send information elsewhere. The model’s size and configurable vector dimensions are intended to make local deployment more practical, though actual resource use will vary by device and implementation.
multimodal embedding model for on-device AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Text Embeddings to Multimodal
Google introduced the first EmbeddingGemma as a lightweight text-embedding option. In its Oct. 6 announcement, the company said the earlier model had received more than 20 million downloads and had been used for on-device search and retrieval-augmented generation, or RAG. That download figure is Google’s report and does not establish how many active applications or users rely on the model.
EmbeddingGemma 2 broadens the product from text embeddings to combined text, code, image, video and audio representations. Google says it shares a text tokenizer and audio encoder with Gemma 4, which may help developers run the embedding model alongside a generative model in a local RAG pipeline. The company points developers to model weights on Hugging Face and Kaggle, as well as deployment options including MediaPipe, LiteRT, transformers.js and WebGPU. Availability through Gemini Enterprise Agent Platform Model Garden was described as coming soon.
“EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings.”
— Google, in its launch announcement
As an affiliate, we earn on qualifying purchases.
Independent Tests Still Needed
The announcement provides Google’s benchmark results and hardware memory estimates, but does not include independent evaluations or a full comparison methodology in the supplied material. The model card is identified as the source for full metrics, but the announcement’s broad claims about quality across image, video, document and audio tasks should be read as Google’s assessment until external testing is available.
It is also not yet clear how performance and memory use will vary across phones, computers and browser deployments, or what trade-offs appear when developers select smaller vector dimensions or quantize the model. Google says the model can process specified amounts of media within its context window; practical throughput and latency depend on the hardware and application. Model Garden availability was still pending at announcement time.
As an affiliate, we earn on qualifying purchases.
Model Access and Evaluations
Developers can access the model weights through Hugging Face and Kaggle, while Google says Model Garden availability is planned for a later date. The company also directs builders to guides for inference, deployment and fine-tuning, with support routes across Google AI Edge tools and third-party frameworks.
The next useful indicators will be independent benchmark results, hands-on measurements across consumer devices and details on Model Garden access. Those tests can clarify whether the reported quality, memory use and cross-modal search capabilities hold across real applications—not only Google’s stated evaluation conditions.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is EmbeddingGemma 2?
It is a 740-million-parameter embedding model from Google for representing text, code, images, audio and video in a shared space for search and retrieval.
Is EmbeddingGemma 2 open source?
Google released it under the Apache 2.0 license, a commercially permissive license. The model weights are listed on Hugging Face and Kaggle.
Can it run on a device without an internet connection?
Google designed the model for on-device inference and says local processing can support offline retrieval. Whether a specific app works offline depends on its implementation and other services it uses.
How much memory does it need?
Google reports that, with quantization on a Pixel 11 Pro, active RAM use can be about 191 MB for text-only weights and 567 MB for the full multimodal model. These are company-reported estimates, not independent measurements across devices.
When will it be available through Model Garden?
Google said Gemini Enterprise Agent Platform Model Garden availability was coming soon, but did not give a specific date in the announcement.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
