AIREITER

EmbeddingGemma 2 API: Local Deployment and Multimodal Use Cases

Last Updated: 2026-10-06 19:13:48

Searching for an EmbeddingGemma 2 API leads to an important distinction: Google’s hosted embedding API is Gemini Embedding 2, while EmbeddingGemma 2 is an open model designed mainly for local and edge inference. That makes it attractive for private multimodal search, but it also means you must choose and operate the serving layer yourself.

Is EmbeddingGemma 2 available as a Google API?

EmbeddingGemma 2 is officially released, but Google’s current managed Gemini API documentation names gemini-embedding-2, not embeddinggemma-2. The EmbeddingGemma 2 model card and Google’s developer guide describe a downloadable model used with local libraries such as Sentence Transformers.

NeedBetter fitAccess pattern
Managed Google endpointGemini Embedding 2Google-hosted Gemini API
Private local inferenceEmbeddingGemma 2Hugging Face/Sentence Transformers or another runtime
Local REST compatibilityEmbeddingGemma 2Ollama, LiteRT-LM, or a third-party server
Phone or edge retrievalEmbeddingGemma 2Google AI Edge / device runtime

A local /v1/embeddings endpoint is exposed by the runtime you deploy, not by Google Cloud. If you meant the managed service, the Gemini Embedding 2 documentation shows the cloud SDK and request formats; use gemini-embedding-2, not embeddinggemma-2.

from google import genai

client = genai.Client()
result = client.models.embed_content(
    model="gemini-embedding-2",
    contents="A private semantic search service",
)
print(result.embeddings)

The managed call uses Google’s hosted API, while the local model uses the credentials and limits imposed by the runtime you deploy.

What the local model actually contains

EmbeddingGemma 2 is a 740-million-parameter multimodal embedder. Its design separates a 270M text core from optional vision and audio encoders, allowing a deployment to load only the modalities it needs. Google and DeepMind position it for text, code, image, video, and audio retrieval rather than text generation.

SpecificationEmbeddingGemma 2
Total parameters740M
Text core270M
Vision encoder170M
Audio encoder300M
Native vector size768 dimensions
Smaller MRL sizes512, 256, and 128 dimensions
Context window8,192 tokens
ModalitiesText, code, image, video, audio
LicenseApache 2.0

The model card describes a shared vector space for cross-modal comparison. The 740M figure describes the full model; Google’s developer guide shows selective encoder use, so runtime memory and active computation can be lower for text-only paths that omit vision and audio.

Pick a serving path by deployment target

Sentence Transformers for a Python application

For a Python service, the official documented route is the google/embeddinggemma-2 checkpoint through Sentence Transformers. It gives you direct control over batching, device placement, prompts, normalization, and vector truncation.

A retrieval workflow should use separate query and document instructions. Google’s examples use a search-query prefix for the query and a document format such as title: none | text: .... The safest pattern is to use model.encode with the relevant prompt name rather than embedding both sides with a generic call.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query_vector = model.encode(
    "How do I rotate an API key?",
    prompt_name="query",
    normalize_embeddings=True,
)
document_vectors = model.encode(
    [
        "title: API keys | text: Rotate keys from the security settings page.",
        "title: Billing | text: Download invoices from the billing page.",
    ],
    prompt_name="document",
    normalize_embeddings=True,
)

Choose this route for Python-level control; use an HTTP-serving runtime when multiple services need a stable contract.

Ollama for a quick local REST endpoint

Ollama’s EmbeddingGemma 2 page provides a simple local API at http://localhost:11434/api/embed:

ollama pull embeddinggemma-2

curl http://localhost:11434/api/embed \\
  -d '{
    "model": "embeddinggemma-2",
    "input": "A private semantic search service"
  }'

Ollama lists model tags such as 270m, 440m, 570m, and 740m; the visible package sizes range from roughly 378 MB to 1.3 GB. Treat these as separately packaged model variants, not interchangeable labels for the full 740M checkpoint. Verify the installed tag and supported input modalities before building a production contract: the family description is multimodal, but the visible variant listings do not document every modality equally clearly.

You still need consistent task prefixes, model and dimension parity, and an index rebuild when changing models.

Edge runtimes for device deployment

Google AI Edge documents EmbeddingGemma V2 in its Universal Embedder guide, while the LiteRT-LM embedding model documentation describes a local OpenAI-compatible /v1/embeddings serving pattern. Use this route when offline operation and device-side privacy outweigh the convenience of a conventional cloud deployment.

For a server on ordinary hardware, start with Sentence Transformers or Ollama. Move to an edge-specific runtime when offline operation, privacy, startup footprint, or device integration is a first-order requirement.

Multimodal use cases that justify the larger model

EmbeddingGemma 2 is most compelling when a project needs one retrieval space across media types.

Use caseWhy multimodal embeddings help
Cross-modal media searchMatch natural-language queries with product photos, video clips, audio, and captions.
Visual-document retrievalCombine OCR wording with page layout and embedded imagery when searching scans.
On-device intent routingRoute private text or media locally without sending raw inputs to a hosted service.

Code search and developer retrieval

The published evaluation table reports a 78.68 MTEB Code score for EmbeddingGemma 2 versus 68.76 for EmbeddingGemma 1 on the cited code benchmark. That is a useful reason to test it for repository search, API documentation retrieval, and code-oriented RAG, but it is not a guarantee for your language mix or codebase.

When the larger model is not worth migrating to

For a pipeline that receives only ordinary OCR text, multimodal support may add complexity without adding retrieval quality. One Paperless-ngx user described the trade-off this way:

“I'm not sure embeddinggemma-2 is any better than regular embeddinggemma for the plain OCR that paperless-ngx sends to the model. seems like a lot more work for the same results.” — u/Great-Cow7256, Reddit

That is not a benchmark result, but it captures the right migration test: compare retrieval quality on your actual corpus before rebuilding a working text-only index.

Dimension choice: 768d, 512d, 256d, or 128d

Google’s embedding documentation documents Matryoshka-style truncation for EmbeddingGemma 2, so you can choose a smaller representation after encoding. Smaller vectors reduce index storage and transfer size, but quality falls at the most aggressive setting.

OutputCompression ratioMTEB multilingual v2MTEB code v1MSEB retrieval
768d1×61.3678.6869.54
512d1.5×61.1777.2469.18
256d3×60.4176.1866.76
128d6×57.8971.4156.71

These figures are reproduced from the published evaluation table on Ollama’s model page. Start with 768d for a new multimodal index, 512d or 256d when storage matters, and 128d only after testing a text-heavy workload.

Do not mix dimensions inside one vector index. If an existing database stores 768-dimensional vectors, switching to 256d requires re-embedding the indexed documents and rebuilding the index. Query vectors must use the same model, prompts, normalization, and dimension as document vectors.

Make the decision by workflow, not by model size

Use Gemini Embedding 2 for a managed Google endpoint. Use Sentence Transformers for Python-level control, Ollama for a quick local HTTP service, and AI Edge/LiteRT-LM when offline device deployment matters.

Keep a smaller text-only model when the corpus is plain OCR and the current index meets its relevance target. EmbeddingGemma 2 can simplify a multimodal architecture, but it does not automatically improve a text-only one.

EmbeddingGemma 2 API FAQ

Is EmbeddingGemma 2 available through the Gemini API?

Google’s managed Gemini API documentation currently identifies gemini-embedding-2. EmbeddingGemma 2 is documented primarily as an open model for local inference, although local runtimes can expose API-compatible endpoints.

Can EmbeddingGemma 2 run on a CPU?

CPU inference can be used with local runtimes that explicitly provide a CPU backend; Google’s AI Edge embedding documentation is the relevant runtime reference. Performance still depends on hardware, quantization, batch size, and modality.

Do existing vectors need to be rebuilt?

Usually, yes, if you change the embedding model, task formatting, normalization policy, or vector dimension. Keep the model identifier, dimension, and preprocessing metadata with the index so a migration can be reproduced.