Google DeepMind has released EmbeddingGemma 2, a new open-weight artificial intelligence model designed to understand and connect text, code, images, video and audio within a single embedding space. The 740-million-parameter model is built on the Gemma 4 architecture and is designed specifically for efficient on-device inference rather than requiring every search or retrieval request to be sent to a cloud server.
Source check, 7 October 2026: Google DeepMind announced EmbeddingGemma 2 on 6 October and published a model card. The 740-million-parameter size, device-memory figures and benchmark gains below are Google-reported results. They are configuration-specific, not proof of equal speed or accuracy on every phone or language.
The launch expands Google’s EmbeddingGemma family from text-focused embeddings into native multimodal retrieval. Google says EmbeddingGemma 2 can power applications such as searching a media library with natural-language queries, finding a particular moment in a video using a voice recording, indexing codebases locally and building privacy-focused retrieval-augmented generation (RAG) systems. The model is available under the commercially permissive Apache 2.0 license.
Key takeaways
- Google DeepMind has launched EmbeddingGemma 2.
- The model has up to 740 million parameters.
- It supports text, code, images, video and audio.
- All supported modalities are mapped into a shared 768-dimensional vector space.
- The base text-and-code configuration uses 270 million parameters.
- Developers can selectively add vision and audio encoders depending on their application.
- It supports an 8,192-token context window.
- Matryoshka Representation Learning allows vectors to be reduced from 768 dimensions to 512, 256 or 128.
- Google says this can reduce vector-storage requirements by up to six times.
- The model is released under the Apache 2.0 license.
- Google says it can run locally on consumer hardware and is optimized for on-device applications.
What is EmbeddingGemma 2?
EmbeddingGemma 2 is not a conventional chatbot or text-generation model.
Instead, it is an embedding model. Embeddings convert information such as a sentence, image, audio recording or piece of code into a numerical representation called a vector.
Those vectors allow software to determine which pieces of information are semantically similar.
For example, a conventional keyword search might struggle to connect the phrase “dog playing in snow” with a photograph that contains a dog but has no text describing it.
An embedding model can represent the text and image in the same mathematical space, allowing software to retrieve the photograph based on the meaning of the query.
EmbeddingGemma 2 takes that concept further by placing text, code, images, video and audio into a unified 768-dimensional embedding space.
From text embeddings to multimodal search
The first EmbeddingGemma model was primarily designed for text embeddings.
Google says the original model had more than 20 million downloads and was used by developers for applications including on-device search and privacy-focused RAG systems. EmbeddingGemma 2 expands the same basic concept across additional forms of information.
The new model can process:
- Text
- Source code
- Images
- Video
- Audio
Instead of requiring separate embedding systems for each modality and another mechanism to align their outputs, EmbeddingGemma 2 projects them into a common vector space.
That creates the possibility of cross-modal search.
A user could type a description and retrieve an image. A voice recording could be used to locate a video clip. A developer could search a codebase using natural-language descriptions of what a function does.
These are examples of retrieval rather than generation, which is why embedding models are increasingly important in modern AI applications.
Why on-device AI is important
One of Google’s main goals with EmbeddingGemma 2 is to make sophisticated retrieval possible without continuously sending information to the cloud.
When embeddings are generated locally, sensitive documents, personal photographs, audio recordings and other information can remain on the user’s device.
That can improve privacy while also reducing network latency.
It can also make applications useful when an internet connection is weak or unavailable.
Google specifically highlights local media search, on-device RAG and other retrieval applications as use cases for the model. The company says EmbeddingGemma 2 can run on consumer hardware, including phones and other edge devices.
This matters because retrieval is often one of the first steps in an AI application.
An AI assistant may need to locate relevant information before a larger generative model can answer a question. If that retrieval step can happen locally, developers can potentially reduce cloud costs and keep private information on-device.
A modular model instead of one fixed configuration
Another important feature is EmbeddingGemma 2’s modular design.
The complete model has 740 million parameters, but developers do not necessarily need to load every component.
Google describes a 270-million-parameter configuration for text and code. Adding the vision encoder brings the configuration to about 440 million parameters, while adding audio produces a roughly 570-million-parameter configuration. Loading all supported modalities brings the model to 740 million parameters.
This allows developers to match the model to their hardware and application.
A company building a local code-search tool, for example, does not need to load image and audio components.
A media-search application can use the relevant visual and audio capabilities.
The approach can reduce memory requirements and improve deployment efficiency.
Google expands context to 8K tokens
EmbeddingGemma 2 also increases the context window substantially compared with its predecessor.
The new model supports up to 8,192 tokens for text and code, compared with a 2K-token context in the original EmbeddingGemma.
Google says the shared context capacity can also accommodate multimodal inputs such as images, video frames and audio.
The company’s launch material says EmbeddingGemma 2 can process up to roughly 5.5 minutes of audio, 29 images or 58 video frames under its specified input limits.
That makes the model more useful for applications where information cannot be represented by a short piece of text alone.
Storage efficiency could be a major advantage
Embedding models can become expensive to operate when applications need to index millions of documents, images or other pieces of information.
Every item stored in a vector database requires storage for its embedding.
EmbeddingGemma 2 addresses this through Matryoshka Representation Learning, or MRL.
The technique allows developers to truncate a 768-dimensional embedding to 512, 256 or 128 dimensions.
The smaller representations require less storage and can make similarity searches more efficient.
Google says reducing vectors to 128 dimensions can provide up to a sixfold reduction in vector-storage requirements. At 256 dimensions, Google says most of the original quality is retained for text and code, while image, video and speech retrieval retain about 95% of the full representation’s quality.
For developers building large local or cloud-based retrieval indexes, this can become a meaningful cost advantage.
Google claims major improvement in code retrieval
EmbeddingGemma 2 also focuses heavily on code.
Google says the model scores 78.68 on the MTEB Code benchmark, compared with 68.76 for EmbeddingGemma 1.
That represents a 9.92-point improvement, according to Google’s published evaluation. The developer guide describes this as a 14% improvement over the previous model on the benchmark.
Better code embeddings could be useful for developer tools that need to search large repositories.
For example, an AI coding agent could first use embeddings to locate potentially relevant functions, classes or documentation before a larger generative model performs deeper reasoning.
That can reduce the amount of information that needs to be processed by the more expensive model.
Multimodal RAG becomes easier
One of the most important applications could be multimodal retrieval-augmented generation.
RAG systems normally retrieve relevant information from an external knowledge base before giving that information to a generative AI model.
For text-heavy applications, that might mean searching documents and returning relevant paragraphs.
But real-world information is increasingly multimodal.
A company’s knowledge base could contain PDFs, photographs, diagrams, presentations, videos, audio recordings and source code.
EmbeddingGemma 2 allows those different formats to be represented in the same embedding space.
When paired with Gemma 4, Google says developers can build on-device RAG systems in which retrieval and generation operate together locally. The two model families also share components such as the text tokenizer and audio encoder, which Google says can reduce their combined memory footprint.
Practical examples for developers
Google has already demonstrated several applications around the model.
One is Instant Media Search, where users can search through local media using semantic queries.
Another is Video Moments Finder, which can locate specific moments in video using text or audio queries.
Google also highlights local file retrieval combined with Gemma 4 for contextual reasoning.
These examples illustrate an important difference between embedding AI and conventional chatbots.
The model does not necessarily need to generate a long answer itself.
Instead, it can make large collections of information searchable based on meaning.
That can become the underlying retrieval layer for many AI applications.
Open licensing could accelerate adoption
EmbeddingGemma 2 is released under the Apache 2.0 license.
That is significant for developers and businesses because the license permits commercial use, modification and deployment subject to the license terms.
Google has also made the model available through platforms including Hugging Face and Kaggle, while supporting deployment frameworks such as MediaPipe, LiteRT, Transformers, Sentence Transformers, vLLM, SGLang, MLX, Ollama and LM Studio.
This broad ecosystem support could make it easier for developers to experiment with the model without being locked into a single API.
It also fits Google’s wider strategy of making the Gemma family available as open models that developers can download and adapt.
EmbeddingGemma 2 versus cloud-based embeddings
EmbeddingGemma 2 does not necessarily replace cloud embedding APIs.
Cloud models can provide greater computing resources and may be easier for developers who do not want to manage local infrastructure.
The advantage of EmbeddingGemma 2 is different.
It gives developers an option for running retrieval locally, where privacy, latency, offline functionality and hardware efficiency are more important.
This could be particularly useful for mobile applications, enterprise devices, industrial systems and products that handle sensitive information.
The choice will ultimately depend on the workload.
For very large databases hosted in the cloud, a cloud embedding service may remain simpler. For private information stored on a phone or laptop, local embeddings could be considerably more attractive.
The bigger significance for Google’s AI strategy
EmbeddingGemma 2 shows that Google’s open-model strategy is moving beyond chatbots and generative models.
AI applications increasingly require multiple components: a model that retrieves information, another that reasons over it, tools that execute actions and infrastructure that stores data.
Embedding models are an important part of that stack.
By making a compact multimodal embedding model available under Apache 2.0, Google is effectively giving developers a local retrieval component that can sit underneath applications built with Gemma and other AI systems.
The fact that the model is designed to work across text, code, images, video and audio makes it particularly relevant as AI applications become more multimodal.
What this means for developers
For developers, the biggest opportunity is not simply that EmbeddingGemma 2 is another AI model.
Its importance comes from the combination of multimodality, small size, open licensing and local execution.
A developer can potentially build an application that understands a user’s documents, photos, recordings and videos without sending every piece of information to a remote server.
That could enable a new generation of privacy-focused AI assistants and search products.
It could also reduce the infrastructure required for smaller AI applications.
However, developers still need to test retrieval quality on their own data. Google’s benchmark results are useful indicators, but performance can vary considerably depending on the language, domain, database size and retrieval method.
The Bigger Picture
EmbeddingGemma 2 reflects a broader shift in AI development: intelligence is increasingly being distributed across the entire computing stack rather than concentrated only in enormous cloud models.
The model’s 740-million-parameter maximum configuration is tiny compared with frontier generative models, but that is precisely the point. Retrieval does not always require a massive model. A compact system that can understand different types of information and locate the right data can make larger AI systems more useful and less expensive.
Google’s release also strengthens the case for on-device AI. If models such as EmbeddingGemma 2 can provide accurate multimodal retrieval on phones and PCs, developers can build AI features that work offline, respond faster and keep sensitive data closer to the user.
Looking Ahead
The next stage will be adoption. The availability of EmbeddingGemma 2 through open-model platforms and popular developer frameworks gives developers a relatively low-friction way to experiment with multimodal retrieval, local RAG and AI-powered search.
The more important test will be whether developers use the model in real products and whether its quality-per-parameter advantage holds across practical workloads. If that happens, EmbeddingGemma 2 could become an important building block for the emerging generation of privacy-focused, multimodal and on-device AI applications.
Sources and what has been tested
Primary records: Google DeepMind’s 6 October launch, the model card and developer guide. Independent reporting: SiliconANGLE and The Next Web. Qdrant’s original tests offer a separate technical check of vector-storage trade-offs; its retention figures are Qdrant’s vendor-reported findings on five text datasets and do not independently verify Google’s multimodal benchmark claims.
For broader context, read our coverage of Indian-language open AI models, open robotics infrastructure and enterprise open-model development. For Indian builders, the practical question is whether local retrieval can keep private documents and media on the device while still meeting accuracy, latency and battery requirements across affordable phones; Google’s release alone does not establish that outcome.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



