Google DeepMind launches EmbeddingGemma 2, a 740M-parameter model to map code, images, video, and audio in a shared embedding space, under an Apache 2.0 license
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images …
Developers building local search and retrieval products can use Google DeepMind's shared representation for text, code, images, video, and audio without tying that embedding step to a hosted API.
EmbeddingGemma 2 gives the Gemma ecosystem a smaller, dedicated on-device component alongside its larger multimodal models, separating indexing and retrieval workloads from generation.
Second-order effects
Application teams can design one semantic index across mixed media rather than maintain modality-specific retrieval pipelines, making cross-modal search a more accessible product feature on devices.
Google's growing Gemma variant community gains a reusable Apache-licensed building block, which can concentrate downstream experimentation around Gemma-compatible local AI stacks.
Third-order effects
If lightweight shared-space embeddings become a standard local component, differentiation in multimodal products shifts from access to a foundation model toward indexing, device integration, and the surrounding application workflow.
Open weights increasingly cover not only generative models but retrieval layers, narrowing the proprietary advantage of cloud-only multimodal search for workloads that fit on endpoints.
The trend: Open multimodal AI is moving from standalone generation models toward modular, on-device stacks that combine local retrieval with application-specific workflows.
Introducing EmbeddingGemma 2! 🚀 Our lightweight, multimodal embedding model maps text, code, images, video, and audio into a single, unified embedding space. Optimized for on-device use cases, it features: - 740M parameter form factor with modular encoders - Flexible dimension si…
Multimodal RAG on your phone. Fully offline. Free. Google just dropped EmbeddingGemma 2. It maps text, images, video, audio and code into one embedding space. Measured on a Pixel 11 Pro: → Text only: ~191MB RAM → Full multimodal: ~567MB RAM Search a voice memo to find a vide…
Google just dropped EmbeddingGemma 2. Open multimodal embedding model. 740M params. Runs fully on-device. Private. No cloud needed. 3 wild things you can do: 1. Find exact video moments with your voice
Introducing EmbeddingGemma 2, our new open embeddings model for on-device use cases! 👀 Code, image, video, audio, and text embeddings 🤏 Modular, going from 270m to 740m parameters 🪆 Matryoshka embeddings for storage efficiency 🤗 Apache 2 License
Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. 🧵
Google launched EmbeddingGemma 2 🔥 This is a bigger than you think. Most people build multimodal RAG as caption/transcribe, and then to text embedder. But EmbeddingGemma 2 puts text, code, images, video, and audio into ONE shared vector space, so you can query a voice memo for …
The future of intelligence isn't just about larger scale, it's about elegant, efficient representations. Very excited to launch EmbeddingGemma 2, our new open model bringing natively multimodal embeddings directly to on-device applications. Download the model weights on @huggingf…
Thanks Google, Embedding Gemma 2 is a big deal 🫶 Multimodal embeddings (text, images, video, audio) can run on-device in the browser. No server, no API key, ~20-70 ms per query on WebGPU. Everything stays on your machine (ofc). Now go build (new) things with it 🚀
🚨 Massive from Google... They just dropped EmbeddingGemma 2, a 740M parameter multimodal embedding model that runs locally. > Maps text, images, audio, video and code into one unified embedding space > 740M parameters, with text-only mode as small as 270M > Apache 2.0 licensed > …
Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency. - our first open, natively multimodal embedding model - handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor - ideal fo…
What...you can now do multimodal RAG on your phone, fully offline and free. EmbeddingGemma 2 does text search in ~191MB of RAM, or text + image + video + audio in ~567MB. Measured on a Pixel 11 Pro. Open weights under Apache 2.0
Google releases EmbeddingGemma 2, a new open model that runs locally on 0.5GB RAM. — The 740M parameter embedding model combines a 270M text model with vision (170M) + audio (300M). — Run & train the model via Unsloth. — GGUF: huggingface.co/unsloth/embe... Guide: unsloth.…