Google has released EmbeddingGemma 2, which the company describes as its first natively multimodal open model engineered for on-device embeddings. Google says it is optimized to run locally on mobile and desktop for search and retrieval tasks, including offline use. Google on local use
Google says EmbeddingGemma 2 is built on the Gemma 4 architecture and released under an Apache 2.0 license. In the company’s description, it goes beyond text to unify images, video, audio and code in a single embedding space. Google announcement
That local-first pitch is the main selling point in Google’s launch posts. Google says the model is optimized to run on phones and desktops, and its examples describe finding a specific video clip from a voice memo or searching through hours of audio with a simple text query, all without an internet connection. Google also says that, when paired with Gemma 4, EmbeddingGemma 2 can support on-device retrieval-augmented generation by retrieving local files while Gemma 4 reasons over them for grounded answers. Google on local use Google on RAG pairing
A modular footprint for different setups
Google and related launch posts describe EmbeddingGemma 2 as modular rather than one fixed-size model. Those posts say configurations range from 270 million parameters for text-only use up to 740 million parameters for the full multimodal version. In an external launch summary of the release, Hugging Face’s Phil Schmid also listed 440 million parameters for text-and-vision and 570 million for text-and-audio. Osanseviero post Phil Schmid post
Google says active RAM use ranges from about 191MB to 567MB, depending on configuration. The company also says EmbeddingGemma 2 has an 8K context window, which it describes as four times larger than the first generation, and says that is enough to process up to 5.5 minutes of audio, 29 images or 58 video frames in one pass. Google specs post
Google’s benchmark claims
Google says the 740 million-parameter version is “best-in-class for its size” and that it outperforms some models more than twice its size. Sundar Pichai made a similar claim in his launch post, describing EmbeddingGemma 2 as a new open multimodal model focused on on-device efficiency. Google specs post Sundar Pichai post
Other launch-related posts point to claimed gains in code search. Phil Schmid’s post cites a 14% jump on MTEB Code, while Rohan Paul’s recap of the release gives a more specific claimed increase from 68.76 to 78.68. Phil Schmid post Rohan Paul summary
The release centers on a model that Google says can work across text, code, images, audio and video on local hardware under an Apache 2.0 license. For developers building private or offline search tools, the pitch is straightforward: keep the retrieval step on the device instead of sending personal files to a server.