Searching for an EmbeddingGemma 2 API leads to an important distinction: Google’s hosted embedding API is Gemini Embedding 2, while EmbeddingGemma 2 is an open model designed mainly for local and edge inference. That makes it attractive for private multimodal search, but it also means you must choose and operate the serving layer yourself.
Is EmbeddingGemma 2 available as a Google API?
EmbeddingGemma 2 is officially released, but Google’s current managed Gemini API documentation names gemini-embedding-2, not embeddinggemma-2. The EmbeddingGemma 2 model card and Google’s developer guide describe a downloadable model used with local libraries such as Sentence Transformers.
| Need | Better fit | Access pattern |
|---|---|---|
| Managed Google endpoint | Gemini Embedding 2 | Google-hosted Gemini API |
| Private local inference | EmbeddingGemma 2 | Hugging Face/Sentence Transformers or another runtime |
| Local REST compatibility | EmbeddingGemma 2 | Ollama, LiteRT-LM, or a third-party server |
| Phone or edge retrieval | EmbeddingGemma 2 | Google AI Edge / device runtime |
A local /v1/embeddings endpoint is exposed by the runtime you deploy, not by Google Cloud. If you meant the managed service, the Gemini Embedding 2 documentation shows the cloud SDK and request formats; use gemini-embedding-2, not embeddinggemma-2.
from google import genai
client = genai.Client()
result = client.models.embed_content(
model="gemini-embedding-2",
contents="A private semantic search service",
)
print(result.embeddings)
The managed call uses Google’s hosted API, while the local model uses the credentials and limits imposed by the runtime you deploy.
What the local model actually contains
EmbeddingGemma 2 is a 740-million-parameter multimodal embedder. Its design separates a 270M text core from optional vision and audio encoders, allowing a deployment to load only the modalities it needs. Google and DeepMind position it for text, code, image, video, and audio retrieval rather than text generation.
| Specification | EmbeddingGemma 2 |
|---|---|
| Total parameters | 740M |
| Text core | 270M |
| Vision encoder | 170M |
| Audio encoder | 300M |
| Native vector size | 768 dimensions |
| Smaller MRL sizes | 512, 256, and 128 dimensions |
| Context window | 8,192 tokens |
| Modalities | Text, code, image, video, audio |
| License | Apache 2.0 |
The model card describes a shared vector space for cross-modal comparison. The 740M figure describes the full model; Google’s developer guide shows selective encoder use, so runtime memory and active computation can be lower for text-only paths that omit vision and audio.
Pick a serving path by deployment target
Sentence Transformers for a Python application
For a Python service, the official documented route is the google/embeddinggemma-2 checkpoint through Sentence Transformers. It gives you direct control over batching, device placement, prompts, normalization, and vector truncation.
A retrieval workflow should use separate query and document instructions. Google’s examples use a search-query prefix for the query and a document format such as title: none | text: .... The safest pattern is to use model.encode with the relevant prompt name rather than embedding both sides with a generic call.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
query_vector = model.encode(
"How do I rotate an API key?",
prompt_name="query",
normalize_embeddings=True,
)
document_vectors = model.encode(
[
"title: API keys | text: Rotate keys from the security settings page.",
"title: Billing | text: Download invoices from the billing page.",
],
prompt_name="document",
normalize_embeddings=True,
)
Choose this route for Python-level control; use an HTTP-serving runtime when multiple services need a stable contract.
Ollama for a quick local REST endpoint
Ollama’s EmbeddingGemma 2 page provides a simple local API at http://localhost:11434/api/embed:
ollama pull embeddinggemma-2
curl http://localhost:11434/api/embed \\
-d '{
"model": "embeddinggemma-2",
"input": "A private semantic search service"
}'
Ollama lists model tags such as 270m, 440m, 570m, and 740m; the visible package sizes range from roughly 378 MB to 1.3 GB. Treat these as separately packaged model variants, not interchangeable labels for the full 740M checkpoint. Verify the installed tag and supported input modalities before building a production contract: the family description is multimodal, but the visible variant listings do not document every modality equally clearly.
You still need consistent task prefixes, model and dimension parity, and an index rebuild when changing models.
Edge runtimes for device deployment
Google AI Edge documents EmbeddingGemma V2 in its Universal Embedder guide, while the LiteRT-LM embedding model documentation describes a local OpenAI-compatible /v1/embeddings serving pattern. Use this route when offline operation and device-side privacy outweigh the convenience of a conventional cloud deployment.
For a server on ordinary hardware, start with Sentence Transformers or Ollama. Move to an edge-specific runtime when offline operation, privacy, startup footprint, or device integration is a first-order requirement.
Multimodal use cases that justify the larger model
EmbeddingGemma 2 is most compelling when a project needs one retrieval space across media types.
| Use case | Why multimodal embeddings help |
|---|---|
| Cross-modal media search | Match natural-language queries with product photos, video clips, audio, and captions. |
| Visual-document retrieval | Combine OCR wording with page layout and embedded imagery when searching scans. |
| On-device intent routing | Route private text or media locally without sending raw inputs to a hosted service. |
Code search and developer retrieval
The published evaluation table reports a 78.68 MTEB Code score for EmbeddingGemma 2 versus 68.76 for EmbeddingGemma 1 on the cited code benchmark. That is a useful reason to test it for repository search, API documentation retrieval, and code-oriented RAG, but it is not a guarantee for your language mix or codebase.
When the larger model is not worth migrating to
For a pipeline that receives only ordinary OCR text, multimodal support may add complexity without adding retrieval quality. One Paperless-ngx user described the trade-off this way:
“I'm not sure embeddinggemma-2 is any better than regular embeddinggemma for the plain OCR that paperless-ngx sends to the model. seems like a lot more work for the same results.” — u/Great-Cow7256, Reddit
That is not a benchmark result, but it captures the right migration test: compare retrieval quality on your actual corpus before rebuilding a working text-only index.
Dimension choice: 768d, 512d, 256d, or 128d
Google’s embedding documentation documents Matryoshka-style truncation for EmbeddingGemma 2, so you can choose a smaller representation after encoding. Smaller vectors reduce index storage and transfer size, but quality falls at the most aggressive setting.
| Output | Compression ratio | MTEB multilingual v2 | MTEB code v1 | MSEB retrieval |
|---|---|---|---|---|
| 768d | 1× | 61.36 | 78.68 | 69.54 |
| 512d | 1.5× | 61.17 | 77.24 | 69.18 |
| 256d | 3× | 60.41 | 76.18 | 66.76 |
| 128d | 6× | 57.89 | 71.41 | 56.71 |
These figures are reproduced from the published evaluation table on Ollama’s model page. Start with 768d for a new multimodal index, 512d or 256d when storage matters, and 128d only after testing a text-heavy workload.
Do not mix dimensions inside one vector index. If an existing database stores 768-dimensional vectors, switching to 256d requires re-embedding the indexed documents and rebuilding the index. Query vectors must use the same model, prompts, normalization, and dimension as document vectors.
Make the decision by workflow, not by model size
Use Gemini Embedding 2 for a managed Google endpoint. Use Sentence Transformers for Python-level control, Ollama for a quick local HTTP service, and AI Edge/LiteRT-LM when offline device deployment matters.
Keep a smaller text-only model when the corpus is plain OCR and the current index meets its relevance target. EmbeddingGemma 2 can simplify a multimodal architecture, but it does not automatically improve a text-only one.
EmbeddingGemma 2 API FAQ
Is EmbeddingGemma 2 available through the Gemini API?
Google’s managed Gemini API documentation currently identifies gemini-embedding-2. EmbeddingGemma 2 is documented primarily as an open model for local inference, although local runtimes can expose API-compatible endpoints.
Can EmbeddingGemma 2 run on a CPU?
CPU inference can be used with local runtimes that explicitly provide a CPU backend; Google’s AI Edge embedding documentation is the relevant runtime reference. Performance still depends on hardware, quantization, batch size, and modality.
Do existing vectors need to be rebuilt?
Usually, yes, if you change the embedding model, task formatting, normalization policy, or vector dimension. Keep the model identifier, dimension, and preprocessing metadata with the index so a migration can be reproduced.