Google DeepMind announced EmbeddingGemma 2, which Google says is its first natively multimodal open model for on-device embeddings. Google says the model goes beyond text by putting code, images, audio, and video into a shared embedding space.
Google's pitch is that developers can use that shared representation to build local search and retrieval tools across different kinds of files. In practical terms, the company says EmbeddingGemma 2 is optimized to run locally on mobile and desktop, including offline uses such as finding a specific clip in audio or video with a text query.
What Google says the model is for
Google says EmbeddingGemma 2 is built on the Gemma 4 architecture and released under an Apache 2.0 license. It also says the model can work alongside Gemma 4 in an on-device retrieval-augmented generation setup, where EmbeddingGemma 2 retrieves local files and Gemma 4 reasons over them.
That makes this more about indexing and retrieval than about a standalone chatbot. Embedding models turn inputs into numerical representations that software can compare and search, which is why Google is framing this around local files, privacy, and offline use.
The specs in Google's announcement
In the announcement thread, Google says EmbeddingGemma 2 has 740 million parameters and uses from about 191 MB to 567 MB of active RAM. Google also says the model has an 8K context window, which it describes as four times larger than the first generation.
For multimodal inputs, Google says one pass can handle up to 5.5 minutes of audio, 29 images, or 58 video frames.
Those details help explain the intended role. Google is positioning EmbeddingGemma 2 as a relatively small open model for organizing and searching local media across formats, rather than as a large general-purpose model.
The main limitation in the supplied evidence is that the substantive details about capabilities, memory use, and intended applications come from Google's own announcement posts. That supports reporting the launch and Google's stated specs, but not independently confirming how the model performs in broader real-world use.
Still, the release shows where Google wants the Gemma line to expand: from text-focused embeddings to a single local model that can help apps search and retrieve several kinds of data without sending them to the cloud.