Google has launched EmbeddingGemma 2, a multimodal embedding model that connects text, images, audio, and video, in addition to code, within a unified representation space. The model targets developers who want to build semantic search, retrieval, and RAG applications locally on devices, without sending data to external servers.
The model has 740 million parameters and is released under the commercially permissive Apache 2.0 license. Google says it is built on the Gemma 4 architecture and can, for example, find a specific video clip using a voice memo or search through long audio recordings with a text query, using a single multimodal model.
What does the model offer in practice?
The most notable aspect of EmbeddingGemma 2 is its ability to bring different types of data together in a shared representation, enabling cross-modal search and retrieval operations. It can be used to search within a media library using text or an image, identify specific moments in a video through a text or audio query, or support local file retrieval before passing the results to a generative model such as Gemma 4 to provide context and reasoning.
- It has an 8,000-token context window, four times that of the previous EmbeddingGemma, with support for up to 5.5 minutes of audio, 29 images, or 58 video frames.
- Output vectors can be reduced from 768 dimensions to 512, 256, or 128 dimensions using Matryoshka Representation Learning, potentially reducing storage space and memory consumption in vector databases by up to six times.
- It provides a lighter configuration for text-only tasks, with optional vision and audio encoders to support multimodal use.
- After quantization, Google says active memory consumption is approximately 191 megabytes for the text-only weights and approximately 567 megabytes for the full multimodal model on a Google Pixel 11 Pro.
Improvements in code and local operation
According to Google, EmbeddingGemma 2 achieved a 9.92-point improvement on the MTEB Code benchmark, rising from 68.76 to 78.68, making it suitable for indexing codebases, performing semantic search within them, and supporting coding-agent retrieval. The company also says it delivers strong results relative to its size on text, vision, and audio tasks, outperforming larger specialized models by more than two times in some comparisons.
Local operation requires relatively limited resources and can help reduce latency, preserve data privacy, and enable offline operation. Sharing the text tokenizer and audio encoder with Gemma 4 may also reduce the total memory required when building a local pipeline that combines retrieval and reasoning.
Availability tools and limitations
Google has made the model weights available through Hugging Face and Kaggle, with subsequent availability expected through Model Garden in the Gemini Enterprise Agent Platform. It can be deployed using Google AI Edge MediaPipe or LiteRT, and supports environments and tools such as transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, in addition to transformers.js and WebGPU for the browser.
This means that EmbeddingGemma 2 is not limited to a ready-made search experience, but provides a component that can be integrated into various local applications. However, the performance and memory figures cited here were provided by Google, while the model’s actual suitability will remain dependent on the type of data, the required retrieval accuracy, hardware constraints, and the results of independent testing. The announcement also did not specify detailed conditions for every use case or a schedule for the model’s availability in Model Garden.