The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →EmbeddingGemma 2 is Google DeepMind’s open multimodal embedding model for turning text, code, images, video and audio into vectors in a shared 768-dimensional space. That makes cross-media search possible—for example, using a text query to retrieve images or video, or matching an audio query to video content. It is a retrieval component, not a generative assistant, and the performance and device figures discussed here are Google-reported rather than independently tested.
What is EmbeddingGemma 2?
Google announced EmbeddingGemma 2 on October 6, 2026, describing it as a model built on the Gemma 4 architecture, released under the Apache 2.0 license and designed for local or edge inference. The launch announcement is by Google DeepMind Research Engineers Sahil Dua and Henrique Schechter Vera.
The “five modalities” framing counts text and code separately. The model handles text and code alongside images, video and audio, projecting their representations into one shared embedding space. An embedding is a numerical representation that can be compared with other embeddings, so an application can retrieve related items across media types. The model does not itself generate answers or provide a complete search product: developers still need an indexing and retrieval system, and often an interface for users.
Google calls it “the most capable model for on-device multimodal embeddings.” That is the company’s characterization, not an independent comparison against all competing models. Google also says the original EmbeddingGemma passed 20 million downloads; that figure refers to the first model, not downloads of EmbeddingGemma 2.
#1 Best Overall
How large is it, and what can you load?
The full checkpoint totals 740 million parameters. Its independent components let developers select a smaller configuration when they do not need every input type. Google’s model card and developer guide describe these parameter counts:
| Configuration | Components | Total parameters | Input coverage |
|---|---|---|---|
| Text-only | Text | 270 million | Text and code |
| Text and vision | Text plus vision | 440 million | Text, code and images; video uses visual frames |
| Text and audio | Text plus audio | 570 million | Text, code and audio |
| Full model | Text, vision and audio | 740 million | Text, code, images, video and audio |
Google lists the text component as 130 million parameters for the transformer backbone plus 140 million for the embedder. The vision component has 170 million parameters and the audio component 300 million. Selecting a partial configuration can reduce the model components an application loads, but the parameter count alone does not establish a device’s total memory requirement or runtime speed.
The model card also lists 24 layers, a vocabulary of 262,144 entries, mean pooling, a 512-to-768 projection layer, grouped-query/multi-query attention and 1,024-token sliding windows. The shared context length is 8,192 tokens.
How much image, video or audio can one input contain?
The 8,192-token context is shared across the input. At Google’s documented defaults, the maximums below assume an input made up of just one modality; text or another media type in the same request consumes part of that budget.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsRank #2
| Single-modality input | Google’s documented default maximum | Token allocation |
|---|---|---|
| Images | About 29 images | 280 tokens per image |
| Video | About 58 frames | 140 tokens per frame; default sampling is 1 frame per second |
| Audio | About 327 seconds, or roughly 5.5 minutes | 25 tokens per second |
These are approximate capacity figures from Google’s model card, not guarantees for every mixed input. Combining modalities reduces the available quantity for each. The card says a configurable lower vision-token budget can allow more images or frames, with a trade-off in visual detail or quality. Google’s documented audio input should be mono at 16 kHz.
What do Google’s benchmark results show?
Google’s model card reports the following scores for the full-precision checkpoint with native 768-dimensional outputs, except where a benchmark’s metric is specified. These are vendor-published results from 2026, not independent tests. Scores use different datasets and metrics, so they should not be ranked against one another unless the benchmark and metric are the same.
| Benchmark | Metric | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|---|
| MTEB multilingual v2 | Mean(Task) | 61.36 | 61.15 |
| MTEB Code v1 | Mean(Task), NDCG@10 | 78.68 | 68.76 |
Google also reports these EmbeddingGemma 2 results in its model card:
- MIEB lite: Mean(TaskType) 64.64.
- MMEB v2 image: Hit@1 57.28.
- MMEB v2 visual-document: NDCG@5 67.84.
- MMEB v2 video: Hit@1 50.67.
- MSEB retrieval: MRR@10 69.54.
- MAEB: Mean(Task) 49.39.
Google describes the model as leading among multimodal embedders under one billion parameters. That claim is the company’s assessment; the reviewed figures do not provide an independent, common-conditions head-to-head against named competitors.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →How should you choose an embedding size?
EmbeddingGemma 2 supports Matryoshka output dimensions of 768, 512, 256 or 128. Shorter vectors use less storage, but can reduce retrieval quality. Google’s developer guide gives the following approximate quality and storage comparisons; the storage example is for one million vectors stored in bfloat16.
| Dimensions | Google’s guide-reported quality guidance | Storage for one million bfloat16 vectors |
|---|---|---|
| 768 | Full-dimensional reference | About 1.5 GB |
| 512 | Not stated in the guide | Not stated in the guide |
| 256 | About 95% of full quality for image, video and speech retrieval | Not stated in the guide |
| 128 | About 90% for text and code, but about 75% for image, video and speech retrieval; Google describes it as best suited to text-only use | About 250 MB |
The percentages and storage quantities are Google’s approximate guide figures, not universal outcomes for every dataset or retrieval task. The guide does not state a quality estimate for 512 dimensions or a storage figure for 512 and 256 dimensions. Validate the chosen size against the application’s own queries and corpus, especially for multimodal use at 128 dimensions. After truncating a vector, L2-normalize it, and use the same dimensionality for query and corpus vectors.
How should developers prepare inputs and vectors?
Use task instructions for text
For text tasks, Google recommends adding task-specific instruction prefixes. In asymmetric retrieval—where a short query searches a collection of longer documents—use a query instruction for the query and document formatting for corpus items. For symmetric similarity or classification, apply the corresponding task instruction consistently to the items being compared. The model card gives examples for web or document search, question answering, fact-checking, code retrieval, classification, clustering and sentence similarity. Google says text inputs still work without a prefix, but precision is reduced.
These text prefixes are not used for media inputs. Keep the query and indexed content aligned to the task and output dimension; mixing incompatible formats or vector sizes undermines comparisons.
Recommended Free Tools
Rank #4
Choose a supported numerical precision
Google recommends bfloat16 where hardware supports it, or float32 where it does not, including on most CPUs. The model card warns against float16: its limited dynamic range can produce NaN values or silently degraded embeddings.
Plan for the full retrieval application
Embeddings are useful only as part of a workflow that encodes content, stores vectors and retrieves relevant matches. Google’s developer guide names Qdrant as a vector-storage option. The right database, indexing strategy and relevance checks depend on the application; the model’s shared space does not by itself guarantee good search results.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Can it run locally or in a browser?
Google reports that, with quantization on a Pixel 11 Pro, the text-only weights use about 191 MB of active RAM and the full multimodal model about 567 MB. These are Google’s figures for that specific phone and setup—not minimum requirements, general RAM estimates or guarantees for other devices.
Google’s October 6, 2026 launch says weights are available through Hugging Face and Kaggle, with on-device optimized versions via the LiteRT Community on Hugging Face. The launch described support in Gemini Enterprise Agent Platform Model Garden as “coming soon”; its availability may have changed since that announcement.
Best Value
Google lists MediaPipe and LiteRT for on-device deployment, transformers.js and WebGPU for browser use, and transformers, Sentence Transformers (version 6.1.0 or later in the developer guide), MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio among development or serving options. It also links to Unsloth fine-tuning guidance. These are integrations and resources named by Google, not a claim that every combination supports every modality or feature equally.
What are the model’s data and safety limitations?
Google’s model card says pretraining included web documents, code, images, video, audio and paired cross-modality examples, with a data cutoff of January 2025. The web-text portion included more than 140 languages. Google describes the model as supporting more than 100 languages, while warning that performance may be unequal across them.
The card says training-data filtering included multiple stages for child sexual abuse material and automated filtering for certain personal information and other sensitive data. It also says EmbeddingGemma 2 is a pretrained embedding model with no post-training alignment, safety tuning or output-level moderation. Developers are responsible for application-level protections such as retrieval filtering and fairness testing, and must follow Google’s Gemma Prohibited Use Policy.
How does EmbeddingGemma 2 compare with the first EmbeddingGemma?
The available direct comparison in Google’s model card is limited to the two MTEB results above: the multilingual v2 Mean(Task) score is 61.36 versus 61.15, and the Code v1 Mean(Task), NDCG@10 score is 78.68 versus 68.76. Those figures are useful within their named benchmarks, but they do not establish a universal ranking across tasks or products.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFor practical deployment, the more consequential choices may be which modalities to load, how much of the shared context each input uses, and whether reduced vector dimensions meet the application’s quality needs. The reviewed Google materials do not establish an independent head-to-head against competing embedding products under common conditions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




