Google's new EmbeddingGemma 2 crams search for text, images, audio and video onto a phone
The 740-million-parameter open model from Google DeepMind runs offline, under 600MB of RAM on a Pixel 11 Pro.

Key points
- Google DeepMind released EmbeddingGemma 2 on an Apache 2.0 licence, meaning any developer can use it commercially without paying Google.
- The model has 740 million parameters and handles text, code, images, video and audio in one shared 768-dimensional space.
- On a Google Pixel 11 Pro, the full multimodal model uses about 567MB of active RAM; the text-only version uses about 191MB.
- The context window is 8,000 tokens, four times larger than the first EmbeddingGemma, enough for roughly 5.5 minutes of audio or 29 images at once.
- Google says code search quality jumps 9.92 points on the MTEB Code benchmark, from 68.76 to 78.68, versus the previous version.
Embedding models are the quiet plumbing of modern AI. They turn a sentence, a photo or a snippet of audio into a long list of numbers, and files with similar meanings end up with similar numbers. That is how your phone can find "the beach video from Lisbon" without you having tagged it.
Until now, the useful ones mostly ran in a data centre. EmbeddingGemma 2, released this week by Google DeepMind, is built to run on the device in your pocket.
What is actually new here?
The first EmbeddingGemma, launched last year, only handled text. Version two takes in code, images, video and audio as well, and maps them all into the same mathematical space so a voice memo can retrieve a video clip and a text query can pull up an audio file.
The model is modular. A 270-million-parameter text core does the basics. Developers bolt on a 170-million-parameter vision encoder, a specialist chunk that handles images, or a 300-million-parameter audio encoder if they need those modes. That is how the full model stays at 740 million parameters, small enough to live on a laptop or a flagship phone.
Google has published the weights on Hugging Face and Kaggle under Apache 2.0, the same permissive licence as the previous release. The company says the first EmbeddingGemma has been downloaded more than 20 million times.
Why would anyone care about on-device?
Because your files never leave the phone. An embedding generated locally does not need to be shipped to a cloud server to be useful, which matters for medical notes, private photos or a company's internal code.
It is also faster. There is no network round-trip, so a search feels instant. And it works on a plane.
The pitch here is retrieval-augmented generation, or RAG, the trick where an AI assistant looks up relevant documents before answering so it does not have to memorise everything. Pair EmbeddingGemma 2 with a small generative model like Gemma 4, and you have a private assistant that can reason over your own files without ever going online.
How does it compare?
Google claims best-in-class scores among multimodal embedding models under one billion parameters, on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark. The headline figure is that code-search jump of nearly ten points.
| Spec | EmbeddingGemma 2 |
|---|---|
| Total parameters | 740M |
| Text-only RAM (Pixel 11 Pro) | ~191MB |
| Full multimodal RAM | ~567MB |
| Context window | 8,000 tokens |
| Languages supported | 100+ |
| Licence | Apache 2.0 |
One genuinely useful trick: Matryoshka Representation Learning, a training method that lets developers chop the output vector from 768 numbers down to 512, 256 or 128 and still get decent results. That cuts storage for a local vector database by up to six times, which matters when the database lives on a phone with finite flash.
Where this fits in Google's wider push
Google has been shipping Gemma variants at pace. AI2Day covered the frontier Gemini model that writes a million tokens without stopping at the end of September, and the broader argument from Google's AI chief that the real breakthroughs now sit where AI meets biology, energy and manufacturing. This release is the opposite end of that spectrum: not a frontier model, a utility one.
My read, having watched these smaller Gemma drops land all year: the interesting question is not the benchmark score. It is whether app developers actually ship on-device RAG features consumers notice, or whether this stays a hobbyist tool. The first EmbeddingGemma got the downloads. The apps built on it have been quieter.



