Google releases EmbeddingGemma 2, an open multimodal embedding model that expands the EmbeddingGemma line beyond text. The model supports multiple input types—including text (and code) plus images, audio, and video—and maps them into a single shared embedding space. The company positions it as small enough to run on smartphones or similar edge devices.

Google introduced the original EmbeddingGemma model in September 2025 with text-focused capabilities. Multiple outlets report EmbeddingGemma 2 is built on Gemma 4 and contains about 740 million parameters. MarkTechPost adds that it embeds five input types into a single 768-dimensional space, while other coverage emphasizes the broader multimodal scope and on-device deployment. The Next Web reports the model is licensed under Apache 2.0 and that it is released without safety tuning, while noting that performance varies across the languages the model supports. Sources diverge mainly on emphasis—such as smartphone/edge deployment, embedding dimensionality, and safety tuning—rather than on the core release details.