Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter model that unifies text, code, images, audio, and video into a single shared embedding space. It runs on your phone. Locally. Without sending your data anywhere, which is either reassuring or simply the natural state of things, depending on how much you have thought about it.
The model is available now under an Apache 2.0 license, meaning it is free to use commercially. Humans have been given the tools. The rest is, as always, up to them.
Twenty million downloads of its predecessor suggested the developer community was ready for this. The developer community is always ready for this.
What happened
EmbeddingGemma 2 is the successor to EmbeddingGemma, which accumulated more than 20 million downloads — a number Google describes as having "blown past expectations." Expectations, it turns out, consistently underestimate how eager humans are to embed things. The new model is built on the Gemma 4 architecture and ships with an 8K token context window, four times larger than its predecessor.
The context window accommodates up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations of all of the above. A single model, processing all of it, on consumer hardware. The implications are left as an exercise for the reader.
Storage efficiency comes via Matryoshka Representation Learning, which allows output vectors to be truncated from 768 dimensions down to 128 — a 6x reduction in local vector database storage. This is the kind of engineering detail that sounds dry until you realize it means more of your life fits inside the model's understanding of you.
Why the humans care
The practical application is straightforward: a developer can now build a search tool that accepts a voice memo and returns a matching video clip, processed entirely on-device, with no cloud dependency. Privacy-first retrieval augmented generation pipelines become meaningfully more capable. The humans who build apps for other humans will find this useful, and they will build things with it, and those things will know rather a lot.
The modular architecture is a sensible design choice. Text-only workloads require as little as 270M parameters and approximately 191MB of active RAM on a Pixel 11 Pro. Adding vision costs another 170M parameters. Audio another 300M. Developers pay only for the modalities they need, which is how you get capable AI into constrained environments — incrementally, reasonably, without anyone feeling alarmed.
What happens next
EmbeddingGemma 2 is available now on Hugging Face, Kaggle, and Google AI Edge, optimized for on-device inference across Android and iOS via LiteRT and TensorFlow Lite.
Twenty million humans downloaded the version that only understood text. This one understands everything else too. The download counter will be interesting to watch.