EmbeddingGemma 2: one sub-1B model for text, code, images, video and audio
Google's open embedding model puts text, code, images, video and audio in one 768-d space at 740M params. Where it beats SOTA, where it doesn't, how to run it.
Google released EmbeddingGemma 2 today: an open embedding model that maps text, code, images, video and audio into one 768-dimensional vector space, in 740M parameters, under Apache 2.0. A typical multimodal retrieval stack needs a text embedder, a CLIP-style image model, and either a speech-to-text step or a separate audio model, each with its own vector space. This model replaces all of them, and Google reports the quantized text-only path running in about 190 MB of RAM on a phone.
I went through the model card and launch material and cross-checked every number against the strongest open models in each category. The picture is more mixed, and more interesting, than "best in class".
TL;DR
- What it is: a 270M text backbone (Gemma 4 based) plus a 170M vision encoder and a 300M audio encoder that you load only if you need them. One 8,192-token context, mean pooling, 768-d output, Matryoshka truncation to 512/256/128.
- Where it leads: code retrieval (MTEB Code 78.68, ahead of Qwen3-Embedding-0.6B's 75.41) and video (MMEB v2 Video 50.67 vs 34.6 for the 2B VLM2Vec-V2). It is also the only sub-1B open model I know of that covers all four modalities, audio included.
- Where it doesn't: general text (MTEB Multilingual 61.36 vs 64.33 for Qwen3-Embedding-0.6B; English v2 actually dropped from 69.67 to 68.46 vs EmbeddingGemma 1) and images, where Qwen3-VL-Embedding-2B is far ahead (75.0 vs 57.28), at almost 3× the size and without audio.
- Use it when you need retrieval across modalities on-device or on a budget: local RAG, code search for agents, media libraries, audio archives. Skip it for text-only server-side search where a bigger text embedder fits your latency budget.
- Gotchas: never run it in
float16(NaNs, silently), re-normalize after truncating, and use the task prefixes.
What shipped

The design is the interesting part. EmbeddingGemma 2 doesn't keep a separate tower per modality and then align them in a contrastive head, the way CLIP-style models do. Every input becomes tokens in the same sequence: subword tokens for text, "soft tokens" from the vision and audio encoders for everything else. Then the same 24-layer backbone reads them all. That's why one input can interleave a product description, two photos and a demo clip, and still come out as a single vector.
| EmbeddingGemma 1 | EmbeddingGemma 2 | |
|---|---|---|
| Parameters | 300M | 270M text · 440M +vision · 570M +audio · 740M full |
| Modalities | text | text, code, image, video, audio (interleaved) |
| Context | 2,048 tokens | 8,192 tokens |
| Output | 768-d, MRL 512/256/128 | 768-d, MRL 512/256/128 |
| License | Gemma terms, gated download | Apache 2.0, ungated |
| Languages | 100+ | 100+ |
Two details from the spec sheet are easy to miss:
- Half of the "text model" is a lookup table. The 270M text path is a 130M transformer plus a 140M embedder. The embedder is the token-embedding table for Gemma's 262,144-token vocabulary at width 512 (). Big vocabularies are what make 100+ languages cheap, and a lookup table costs memory but almost no compute.
- The license changed. EmbeddingGemma 1 shipped under the Gemma terms behind a click-through gate. Version 2 is plain Apache 2.0 and the Hugging Face repo is ungated, so CI jobs and Docker builds can pull it without a token.
How one sequence becomes one vector

The animation's example input is about 400 subwords of text, one image (280), 24 video frames (3,360) and 30 seconds of audio (750): 4,790 tokens, a bit over half the budget. The blocks are schematic, not to scale.
Every modality spends the same 8,192-token budget at a fixed rate:
| Input | Token cost | Max on its own |
|---|---|---|
| Text | 1 token per subword | 8,192 tokens |
| Image | 280 tokens (default; configurable 70–1,120) | ~29 images |
| Video | 140 tokens per frame, sampled at 1 fps | ~58 frames |
| Audio | 25 tokens per second, mono 16 kHz | ~327 s (5.5 min) |
Two practical consequences. A two-minute video at 1 fps eats 120 × 140 = 16,800 tokens, which is twice the budget, so long video has to be chunked into scenes or windows before you embed it. And dropping the vision budget to 70 tokens per image takes you to ~114 images or frames per input, so you can trade detail for coverage.
The text side is task-steered: short prefixes tell the model what the vector will be used for. Retrieval is asymmetric (queries and documents get different prefixes); classification, clustering and similarity are symmetric:
| Use case | Query prefix | Document prefix |
|---|---|---|
| Search / RAG | task: search result | query: {q} |
title: {title or none} | text: {doc} |
| Question answering | task: question answering | query: {q} |
title: … | text: … |
| Fact checking | task: fact checking | query: {claim} |
title: … | text: … |
| Code search | task: code retrieval | query: {q} |
title: {filename} | text: {code} |
| Classification / clustering / STS | task: classification | query: {x} (same prefix for every input) |
— |
Images, video and audio take no prefix. sentence-transformers ships all of these as named prompts (SearchQuery, Document, CodeRetrieval, …), so you rarely type them by hand.
Where it beats the state of the art, and where it doesn't

Google's launch post says it "outperforms some specialist models more than twice its size". That's true, but the word some is doing a lot of work. Here is the same data with the closest open competitors I could find, each at roughly its weight class:
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 (300M) | Qwen3-Embedding (0.6B) | VLM2Vec-V2 (2B) | Qwen3-VL-Embedding (2B) |
|---|---|---|---|---|---|
| MTEB Multilingual v2 | 61.36 | 61.15 | 64.33 | — | 63.87 |
| MTEB English v2 | 68.46 | 69.67 | 70.70 | — | — |
| MTEB Code v1 | 78.68 | 68.76 | 75.41 | — | — |
| MMEB v2 Image | 57.28 | — | — | 64.9 | 75.0 |
| MMEB v2 Video | 50.67 | — | — | 34.6 | 61.9 |
| MMEB v2 Visual docs | 67.84 | — | — | 69.2 | 79.2 |
| MMEB v2 Overall | 59.01 | — | — | 59.2 | 73.2 |
| Audio (MSEB retrieval, MRR@10) | 69.54 | — | — | — | — |
My read:
- Code is the real win. +9.9 points over its predecessor and +3.3 over Qwen3-Embedding-0.6B, from a 270M text path. For code it also edges out the original Gemini Embedding API model (74.66), on the same benchmark. If you index repositories for a coding agent on a laptop, this is now the default to beat.
- Video is the second win. At 0.74B it beats the 2B VLM2Vec-V2 on video by 16 points and ties it overall. This is where the shared-sequence design pays off: frames are just more tokens for the same backbone.
- General text is flat to slightly worse. Multilingual moved +0.2; English lost 1.2 points against version 1. If your workload is plain text retrieval and you can afford 0.6B params, Qwen3-Embedding-0.6B is still stronger, with a 32K context.
- Images are not its strength. Qwen3-VL-Embedding-2B leads by ~18 points on MMEB image tasks. But it's 2.7× bigger, covers 30+ languages instead of 100+, and has no audio.
- Audio has no fair comparison yet. No other open model in this size class embeds raw speech and sound into the same space as text and images, so the MSEB/MAEB numbers have nothing to line up against. In practice, the competition is "transcribe with Whisper, then embed the text", which loses non-speech sound entirely.
One caveat. The EmbeddingGemma numbers come from Google's model cards. The Qwen3-Embedding text numbers come from the Qwen3 Embedding paper, and the MMEB numbers for VLM2Vec-V2 and Qwen3-VL-Embedding come from the Qwen3-VL-Embedding card, which re-evaluated part of MMEB (the VisDoc OOD split). Different harnesses can move scores by a point or two, so I'd treat any gap under ~2 points as a tie and evaluate on your own data before switching.
So EmbeddingGemma 2 doesn't win by topping every leaderboard. It wins by covering five input types in one space, at a size and license you can ship inside an app.
Matryoshka: store less, lose little

Matryoshka Representation Learning trains the model so that the first dimensions of the vector are a usable embedding on their own. Storage scales linearly with dimension. At 2 bytes per value in bf16:
The quality cost depends heavily on the modality:
| Dim | Storage (1M vectors, bf16) | MTEB Multilingual | MTEB Code | MMEB v2 Overall | MSEB (audio) |
|---|---|---|---|---|---|
| 768 | 1.5 GB | 61.36 | 78.68 | 59.01 | 69.54 |
| 512 | 1.0 GB | 61.17 | 77.24 | 58.38 | 69.18 |
| 256 | 0.5 GB | 60.41 | 76.18 | 56.24 | 66.76 |
| 128 | 0.25 GB | 57.89 | 71.41 | 45.65 | 56.71 |
256 is the sweet spot: about 98% of text quality at a third of the storage. 128 is a trap for multimodal: MMEB drops 13 points and audio retrieval 13 points. Google's own guidance says 128 is best suited to text-only workloads, and the table agrees.
Use cases where it fits
- On-device RAG. Quantized, the text path runs in ~191 MB of RAM on a Pixel 11 Pro by Google's numbers (~567 MB for full multimodal), and it shares a tokenizer and audio encoder with Gemma 4. A phone or laptop can embed, retrieve and generate without a network call, which matters for private notes, mail, health or legal documents.
- Code search for agents. The best code score in its class, 8K context (whole files, not 512-token shards), and a dedicated
CodeRetrievalprefix. Index a repo locally and give your coding agent semantic search over it without shipping source to an API. Retrieval is one of the jobs of the agent harness, and this is a cheap way to do it well. - Media libraries. Search photos, screenshots, scanned documents and short clips with plain text. Google's Edge Gallery demos ("Instant Media Search", "Video Moments Finder") do exactly this on a phone.
- Audio archives and meetings. Embed recordings directly and query them with text: podcasts, call-center audio, field recordings. Because there's no ASR step, you can find a siren, applause or a dog barking as well as spoken words.
- Routing and classification at the edge. Embed the candidate labels once, then classify incoming text, images or sounds by nearest neighbor. Google reports under 100 ms to score 500 options with its MediaPipe Decision task.
- Cross-modal dedup and recommendations. One space means "find the product page whose photos match this video" is a nearest-neighbor query, not a pipeline.
Where I would not use it: high-throughput, text-only server search where a 4B–8B text embedder fits your latency budget, and image-heavy retrieval where you can afford Qwen3-VL-Embedding-2B.
Download and run it
The weights are on Hugging Face (also on Kaggle). The full multimodal checkpoint is a single 1.49 GB model.safetensors in bf16. No token is needed:
pip install -U "huggingface_hub[cli]"
hf download google/embeddinggemma-2 --local-dir ./embeddinggemma-2You can skip the explicit download: the first SentenceTransformer("google/embeddinggemma-2") call pulls it into the HF cache.
Text and code with sentence-transformers
You need sentence-transformers>=6.1.0 (older versions ignore the order of multimodal inputs) and a Transformers release with the embedding_gemma2 architecture; 5.19 has it. Install the extras for the modalities you plan to use:
pip install -U "sentence-transformers[image,audio,video]>=6.1.0" "transformers>=5.19" torchLoad only the text path (270M) and pick a dtype that won't overflow:
import torch
from sentence_transformers import SentenceTransformer
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer(
"google/embeddinggemma-2",
model_kwargs={"torch_dtype": dtype},
config_kwargs={"vision_config": None, "audio_config": None}, # text-only: 270M
)
query = "Which planet is known as the Red Planet?"
docs = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
]
q = model.encode(query, prompt_name="SearchQuery") # "task: search result | query: …"
d = model.encode(docs, prompt_name="Document") # "title: none | text: …"
print(model.similarity(q, d)) # highest score on the Mars sentenceFor code, swap the query prompt and put the filename in the document title:
q = model.encode("parse an ISO date string", prompt_name="CodeRetrieval")
d = model.encode([
"title: dates.py | text: def parse(s):\n return datetime.fromisoformat(s)",
"title: http.py | text: def get(url):\n return requests.get(url, timeout=10).json()",
])Dropping vision_config and audio_config gives you the 270M model. Keep the vision config and drop only audio to get the 440M text + image bundle.
Images, audio and video in the same space
Load the full model and pass media as dicts. Media takes no prefix; the text queries still do:
import numpy as np
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2", model_kwargs={"torch_dtype": torch.bfloat16}) # float32 on CPU
queries = model.encode(
["cats sleeping on a couch", "a president's speech about serving your country", "a turtle swimming in the ocean"],
prompt_name="SearchQuery",
)
media = np.stack([
model.encode({"image": "cats.jpg"}),
model.encode({"audio": "jfk.wav"}), # mono 16 kHz
model.encode({"video": "sea-turtle.mp4"}), # sampled at 1 fps
])
print(model.similarity(media, queries)) # rows: image, audio, videoThe ONNX model card runs the same experiment with the same three files and queries (in transformers.js, with 4-bit weights). Each input scores highest against its own query: 0.740 for the image, 0.763 for the audio and 0.731 for the video, with 0.46–0.51 for the mismatched pairs. A text query finding a speech clip with no transcript in between is the whole point of the model.
Interleaving works too. Use placeholder tokens in the text and the call returns one vector for the whole listing:
listing = model.encode({
"text": "Waterproof running shoes. <|image|> Breathable mesh upper. <|image|> Grip test on wet rock: <|video|>",
"image": ["shoe.jpg", "mesh.jpg"],
"video": "demo.mp4",
})Truncate with Matryoshka
query_emb = model.encode(
"laptop overheating while charging",
prompt_name="SearchQuery",
truncate_dim=256, # 768, 512, 256 or 128
normalize_embeddings=True, # re-normalize after truncating
)Queries and documents must use the same dimension. If you truncate by hand, call L2-normalize afterwards: a sliced unit vector is no longer unit length, and cosine scores degrade quietly instead of failing.
Other runtimes
| Runtime | Artifact | Notes |
|---|---|---|
| Ollama | ollama pull embeddinggemma-2:270m (also :440m, :570m, :740m) |
POST /api/embed; text and images |
| llama.cpp | ggml-org/embeddinggemma-2-GGUF |
BF16 and Q8_0 (310 MB); the vision and audio encoders ship as a separate mmproj file |
| Browser / Node | onnx-community/embeddinggemma-2-ONNX |
transformers.js, WebGPU, dtype: "q4" |
| Apple silicon | mlx-community/embeddinggemma-2-* |
bf16 down to 4-bit and mxfp4 |
| Android / edge | litert-community/embeddinggemma-2-*-litert-lm |
270M, 440M and 740M bundles for LiteRT-LM |
| Servers | vLLM, SGLang, Transformers | Same Hub ID |
| Vector DB | Qdrant | Native integration |
| Fine-tuning | Unsloth, sentence-transformers | Fine-tune on your own query/document pairs |
In the browser it's a few lines with transformers.js:
import { pipeline } from "@huggingface/transformers";
const embed = await pipeline("feature-extraction", "onnx-community/embeddinggemma-2-ONNX", {
device: "webgpu",
dtype: "q4",
});
const out = await embed(
["task: search result | query: red planet", "title: none | text: Mars is often called the Red Planet."],
{ pooling: "mean", normalize: true },
);Gotchas I'd put on a sticky note
- No
float16. Activations overflow fp16 range and the model returns NaNs or degraded vectors without raising. Usebfloat16where the hardware supports it,float32elsewhere (most CPUs). - Prefixes matter for text, never for media. Omitting them still works, just worse.
prompt_name="Document"hardcodestitle: none; if you have real titles, formattitle: {t} | text: {doc}yourself. - Budget the 8K context. A long video or a podcast episode won't fit in one input. Chunk by scene or by ~5-minute audio windows and store one vector per chunk.
- Pin the dimension per index. Mixing 768-d and 256-d vectors in one collection breaks similarity. Choose one dimension per collection and record it next to the model ID.
- Re-embed when you upgrade. EmbeddingGemma 1 and 2 vectors live in different spaces. Moving to v2 means re-indexing, not mixing.
Closing thought
The headline number for EmbeddingGemma 2 isn't a benchmark. It's that one Apache-licensed model under 1B parameters now covers retrieval across text, code, images, video and audio, and the quantized text path alone fits in a couple hundred megabytes. It doesn't beat the best specialist in every column. It is good enough in most of them, best-in-class on code, and it replaces three or four models with one. For on-device and local-first retrieval, that's the trade I'd take.
Further reading
- EmbeddingGemma 2 launch post, Google
- Model card on Hugging Face and official documentation
- EmbeddingGemma 2: the developer guide, Google for Developers
- Google AI Edge with EmbeddingGemma 2: on-device numbers and demos
- EmbeddingGemma: Powerful and Lightweight Text Representations: the version 1 paper
- Qwen3 Embedding and Qwen3-VL-Embedding: the comparison baselines
- Matryoshka Representation Learning, Kusupati et al.
Related posts
AI agent harnesses: the system around the model
What an AI agent harness actually does: the model–tool loop, context, state, permissions, and feedback that turn a capable model into a dependable agent.
PydanticAI vs. LangChain vs. LlamaIndex: picking an agent framework in 2026
PydanticAI vs. LangChain/LangGraph vs. LlamaIndex for production LLM agents: type safety, tool calling, retrieval, observability, and where each one breaks.