A local EmbeddingGemma 2 deployment is not automatically a drop-in replacement for an existing embedding service. The runtime can change with little application work, but changing the representation contract usually means rebuilding vectors. The safest plan is to separate three decisions: how the model runs, whether existing vectors remain compatible, and whether mixed-modal retrieval is good enough for your corpus.
The migration decision in one page
Use EmbeddingGemma 2 when you need local text, code, image, video, or audio embeddings in one model family and can afford a controlled backfill. Do not switch a production query encoder first and “catch up” on documents later: an embedding model is part of the index schema, even when the vector dimension looks familiar.
| Decision | Practical answer |
|---|---|
| Local starting point | Sentence Transformers with the official checkpoint |
| Text-only footprint | 270M parameters when vision and audio are disabled |
| Full multimodal footprint | 740M parameters |
| Native output | 768 dimensions |
| Storage compromise | 256d is the first setting to test; 128d needs stronger validation for multimodal data |
| Existing vectors | Reuse only when the full representation contract is unchanged and compatibility is demonstrated |
| Production cutover | Build a second index or use versioned named vectors, then switch model and index together |
Google’s model card reports 61.36 on MTEB multilingual v2, 78.68 on MTEB code v1, 67.84 NDCG@5 on visual-document retrieval, 50.67 Hit@1 on video retrieval, and 69.54 MRR@10 on audio retrieval at 768 dimensions. Those are useful reference points, not a substitute for testing your own queries.
What changes—and what does not—when you move to EmbeddingGemma 2
EmbeddingGemma 2 maps text, code, images, video, and audio into a shared 768-dimensional space. The checkpoint is modular: the official developer guide describes a 270M text-only configuration, 440M text-plus-vision configuration, 570M text-plus-audio configuration, and 740M full configuration. Disabling an encoder reduces loaded weights and peak memory; it does not by itself create a new semantic space.
That distinction matters during migration. A text query produced with the 270M configuration can be compared with an EmbeddingGemma 2 document embedding produced with the full configuration because Google documents these configurations as sharing a compatible vector space. That does not mean an old EmbeddingGemma 1, Qwen, Nomic, or API-provider vector can be queried safely with EmbeddingGemma 2 merely because it also has 768 coordinates.
Task formatting is part of the contract too. For asymmetric retrieval, EmbeddingGemma 2 expects a search-query instruction such as task: search result | query: ... and a document form such as title: ... | text: .... Code retrieval has its own task instruction. If your old pipeline used different prefixes, chunking, normalization, or embedded fields, record those changes as a new representation version and validate them as a migration.
Runtime support: choose the narrowest local path
Start with Sentence Transformers for correctness
The official model card documents google/embeddinggemma-2 with Sentence Transformers and Transformers. Install the multimodal extras when you need media support:
pip install -U "sentence-transformers[image,audio,video]" transformers
This is the best reference path for a migration because prompt names, truncation, normalization, and multimodal input behavior follow the official examples. It is not necessarily the lowest-latency serving path, but it gives you a trustworthy baseline before optimizing.
A minimal text-only smoke test looks like this:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)
query = model.encode(
"embedding model migration",
prompt_name="SearchQuery",
truncate_dim=256,
normalize_embeddings=True,
)
document = model.encode(
"Rebuild vectors when the embedding representation changes.",
prompt_name="Document",
truncate_dim=256,
normalize_embeddings=True,
)
print(model.similarity(query, document).item())
Run this before introducing a server, quantization, or a vector database. It checks that the checkpoint, task prompts, dimension, and normalization path work together.
Use ecosystem runtimes only after checking feature parity
Google’s developer guide lists vLLM, Hugging Face Transformers, Sentence Transformers, SGLang, MLX, Ollama, LM Studio, and LiteRT as supported development or deployment tools. Treat that list as an availability signal, not proof that every runtime exposes the same combination of text, image, video, audio, interleaving, task prefixes, truncation, and batching.
For each candidate runtime, verify five items with a real request: the exact checkpoint revision, the modality inputs you use, the output dimensions, post-truncation normalization, and the query/document prefix behavior. A runtime that serves text quickly but ignores your visual-document path is not equivalent to the full model.
A compact native server is an optimization, not the migration plan
The public embeddinggemma.c repository provides a specialized C11/Metal-style server for EmbeddingGemma 300M, with CPU, Metal, CUDA, ROCm, and Intel XPU variants. Its README documents an OpenAI-compatible /v1/embeddings endpoint, 768/512/256/128 dimensions, and a 278 MB Q4_0 model download. The project reports a controlled Apple M5 Max comparison against llama.cpp build b8981 across 54 cells, with a 1.25× geometric-mean advantage; those are project-specific throughput results, not a quality comparison or proof of multimodal parity with the 740M checkpoint.
The useful migration property is API shape. If your application already speaks OpenAI-style embeddings, an endpoint-compatible local server can reduce adapter work. Still, keep the Sentence Transformers result as the correctness oracle until the server’s modality and prefix behavior match your production pipeline.
Index rebuild risk: dimension is only one migration axis
Re-embed when the source-to-vector function changes
Assume a full rebuild when you change the model family, model version, task prefix, normalization, chunking, truncation policy, embedded fields, or similarity semantics. Qdrant’s migration guide and the model-migration analysis from Nalar make the same operational point: document and query vectors must belong to the same representation version. Equal dimensions do not establish semantic compatibility.
Do not slice an old 768-dimensional vector and call it a 256-dimensional EmbeddingGemma 2 vector. EmbeddingGemma 2’s Matryoshka outputs are trained for supported truncation sizes and must be renormalized after truncation. The model card reports the following official reference scores:
| Dimension | Storage reduction | MTEB multilingual v2 | Code v1 | MIEB Lite | MMEB v2 overall |
|---|---|---|---|---|---|
| 768 | 1× | 61.36 | 78.68 | 64.64 | 59.01 |
| 512 | 1.5× | 61.17 | 77.24 | 64.32 | 58.38 |
| 256 | 3× | 60.41 | 76.18 | 63.13 | 56.24 |
| 128 | 6× | 57.89 | 71.41 | 59.06 | 45.65 |
The official model card also reports 67.84 NDCG@5 on visual-document retrieval and 50.67 Hit@1 on video retrieval at 768 dimensions; use those as full-width baselines rather than inventing reduced-dimension values. The reliable decision is directional: 256d is much closer to full quality than 128d, and multimodal scores fall more sharply at 128d. Use the official checkpoint to regenerate every vector at the selected dimension rather than truncating vectors from another model.
Shared EmbeddingGemma 2 space can reduce unnecessary work
There is one important exception. If your existing corpus was already embedded with EmbeddingGemma 2 and you are only loading a different subset of its encoders, Google’s developer guide says the configurations share one vector space. A text-only query can match a full-model document vector. In that case, you may not need to regenerate existing text vectors solely because the serving process now loads vision or audio support.
You still need new vectors for new media-bearing records. A text-only index cannot retrieve an image, video, or audio item that was never embedded. Adding multimodal retrieval is therefore an incremental corpus migration even when the model checkpoint is unchanged.
Use blue-green or named-vector cutover
For a live system, Qdrant’s migration pattern is the clearest template: create a new collection, dual-write new records, backfill from authoritative source data, compare Recall@10/MRR/nDCG@10, switch an alias, and keep the old collection for rollback. Qdrant’s guide uses version 1.19.0, a 512-dimensional example, and batches of 100 points; those values are examples, not requirements for EmbeddingGemma 2.
A named-vector design can keep old and new representations in one collection, but only if your vector database supports it and your update path writes both consistently. Weaviate’s vectorizer migration guide recommends collection aliases for production because the old collection can be retained for an immediate rollback and deleted after validation. Its alternative—adding a vector to the existing collection—can permanently increase storage and is better suited to comparison than a clean final state.
Mixed-modal quality: validate the slices that the model changes
EmbeddingGemma 2’s shared space is valuable only if the retrieval behavior matches your data. A text-only benchmark can confirm that the migration did not break text search while missing failures in PDF pages, charts, image captions, video frames, audio clips, or interleaved records.
Start with separate labeled slices:
- Text query → text chunk.
- Code query → code chunk.
- Text query → image or visual document.
- Text query → video frame or audio segment.
- Mixed text-plus-media query → mixed document.
- Cross-language query → document in the languages you serve.
Keep 768d or 512d as the first multimodal baseline. The official model card assigns 280 tokens per image, 140 per video frame, and 25 per second of audio within a shared 8,192-token context. Mixed inputs consume the same budget, so a record containing text, images, and video has less room for each component than a single-modality input.
The model card also reports that 128d causes a larger quality decline for multimodal tasks than for text-only tasks. That makes 128d a reasonable first-stage shortlist candidate for a large text index, not a default for a mixed-media archive. Test 256d against your actual visual-document and cross-modal queries before accepting its storage savings.
Compare a unified EmbeddingGemma 2 pipeline with your current separate text/image pipeline on identical queries; do not infer mixed-modal quality from the model’s shared-space architecture alone.
A staged deployment plan for an existing RAG system
- Inventory the current contract. Record model ID, checkpoint revision, prefixes, chunking, dimensions, metric, normalization, source fields, and every modality already indexed.
- Create a representative evaluation set. Include Recall@k, MRR, or nDCG targets plus separate text, code, visual-document, audio, video, language, and long-query slices.
- Establish the local baseline. Run the same corpus through Sentence Transformers first. Log embedding latency, search latency, memory, index size, failures, and score distributions.
- Build a versioned candidate index. Keep stable document IDs and authoritative source text/media outside the vector store so the backfill is reproducible.
- Reconcile writes during backfill. Use a source snapshot plus change replay, or dual-write new and updated records to both representation versions.
- Shadow production queries. Compare ranked results, empty-result rates, latency, and labeled relevance without changing user-visible answers.
- Cut over atomically. Bind the EmbeddingGemma 2 query encoder and its matching index under one version or alias. Never deploy the new query encoder against the old index as an intermediate state.
- Keep rollback available. Retain the old index and query path until representative traffic clears the acceptance thresholds; then stop dual writes and reclaim storage.
FAQ
Can EmbeddingGemma 2 run CPU-only?
Yes, with a CPU-capable runtime. The model card recommends float32 when bfloat16 is unavailable, and the text-only configuration is 270M parameters. CPU throughput depends on runtime, precision, batching, and hardware, so measure it on your corpus rather than borrowing GPU numbers.
The runtime choice is a trade-off between reference correctness, feature parity, serving efficiency, and the cost of validating a new retrieval version.