Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
Mountain View desk6 min

Google DeepMind Launches EmbeddingGemma 2, Mapping Five Input Types Into One Vector Space

EmbeddingGemma 2 is Google’s open multimodal embedding model for cross-media retrieval. Here are its shared vector space, model configurations, input limits, reported results and practical trade-offs.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

EmbeddingGemma 2 is Google DeepMind’s open multimodal embedding model for turning text, code, images, video and audio into vectors in a shared 768-dimensional space. That makes cross-media search possible—for example, using a text query to retrieve images or video, or matching an audio query to video content. It is a retrieval component, not a generative assistant, and the performance and device figures discussed here are Google-reported rather than independently tested.

What is EmbeddingGemma 2?

Google announced EmbeddingGemma 2 on October 6, 2026, describing it as a model built on the Gemma 4 architecture, released under the Apache 2.0 license and designed for local or edge inference. The launch announcement is by Google DeepMind Research Engineers Sahil Dua and Henrique Schechter Vera.

The “five modalities” framing counts text and code separately. The model handles text and code alongside images, video and audio, projecting their representations into one shared embedding space. An embedding is a numerical representation that can be compared with other embeddings, so an application can retrieve related items across media types. The model does not itself generate answers or provide a complete search product: developers still need an indexing and retrieval system, and often an interface for users.

Google calls it “the most capable model for on-device multimodal embeddings.” That is the company’s characterization, not an independent comparison against all competing models. Google also says the original EmbeddingGemma passed 20 million downloads; that figure refers to the first model, not downloads of EmbeddingGemma 2.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How large is it, and what can you load?

The full checkpoint totals 740 million parameters. Its independent components let developers select a smaller configuration when they do not need every input type. Google’s model card and developer guide describe these parameter counts:

Configuration Components Total parameters Input coverage
Text-only Text 270 million Text and code
Text and vision Text plus vision 440 million Text, code and images; video uses visual frames
Text and audio Text plus audio 570 million Text, code and audio
Full model Text, vision and audio 740 million Text, code, images, video and audio

Google lists the text component as 130 million parameters for the transformer backbone plus 140 million for the embedder. The vision component has 170 million parameters and the audio component 300 million. Selecting a partial configuration can reduce the model components an application loads, but the parameter count alone does not establish a device’s total memory requirement or runtime speed.

The model card also lists 24 layers, a vocabulary of 262,144 entries, mean pooling, a 512-to-768 projection layer, grouped-query/multi-query attention and 1,024-token sliding windows. The shared context length is 8,192 tokens.

How much image, video or audio can one input contain?

The 8,192-token context is shared across the input. At Google’s documented defaults, the maximums below assume an input made up of just one modality; text or another media type in the same request consumes part of that budget.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Single-modality input Google’s documented default maximum Token allocation
Images About 29 images 280 tokens per image
Video About 58 frames 140 tokens per frame; default sampling is 1 frame per second
Audio About 327 seconds, or roughly 5.5 minutes 25 tokens per second

These are approximate capacity figures from Google’s model card, not guarantees for every mixed input. Combining modalities reduces the available quantity for each. The card says a configurable lower vision-token budget can allow more images or frames, with a trade-off in visual detail or quality. Google’s documented audio input should be mono at 16 kHz.

What do Google’s benchmark results show?

Google’s model card reports the following scores for the full-precision checkpoint with native 768-dimensional outputs, except where a benchmark’s metric is specified. These are vendor-published results from 2026, not independent tests. Scores use different datasets and metrics, so they should not be ranked against one another unless the benchmark and metric are the same.

Benchmark Metric EmbeddingGemma 2 EmbeddingGemma 1
MTEB multilingual v2 Mean(Task) 61.36 61.15
MTEB Code v1 Mean(Task), NDCG@10 78.68 68.76

Google also reports these EmbeddingGemma 2 results in its model card:

  • MIEB lite: Mean(TaskType) 64.64.
  • MMEB v2 image: Hit@1 57.28.
  • MMEB v2 visual-document: NDCG@5 67.84.
  • MMEB v2 video: Hit@1 50.67.
  • MSEB retrieval: MRR@10 69.54.
  • MAEB: Mean(Task) 49.39.

Google describes the model as leading among multimodal embedders under one billion parameters. That claim is the company’s assessment; the reviewed figures do not provide an independent, common-conditions head-to-head against named competitors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should you choose an embedding size?

EmbeddingGemma 2 supports Matryoshka output dimensions of 768, 512, 256 or 128. Shorter vectors use less storage, but can reduce retrieval quality. Google’s developer guide gives the following approximate quality and storage comparisons; the storage example is for one million vectors stored in bfloat16.

Dimensions Google’s guide-reported quality guidance Storage for one million bfloat16 vectors
768 Full-dimensional reference About 1.5 GB
512 Not stated in the guide Not stated in the guide
256 About 95% of full quality for image, video and speech retrieval Not stated in the guide
128 About 90% for text and code, but about 75% for image, video and speech retrieval; Google describes it as best suited to text-only use About 250 MB

The percentages and storage quantities are Google’s approximate guide figures, not universal outcomes for every dataset or retrieval task. The guide does not state a quality estimate for 512 dimensions or a storage figure for 512 and 256 dimensions. Validate the chosen size against the application’s own queries and corpus, especially for multimodal use at 128 dimensions. After truncating a vector, L2-normalize it, and use the same dimensionality for query and corpus vectors.

How should developers prepare inputs and vectors?

Use task instructions for text

For text tasks, Google recommends adding task-specific instruction prefixes. In asymmetric retrieval—where a short query searches a collection of longer documents—use a query instruction for the query and document formatting for corpus items. For symmetric similarity or classification, apply the corresponding task instruction consistently to the items being compared. The model card gives examples for web or document search, question answering, fact-checking, code retrieval, classification, clustering and sentence similarity. Google says text inputs still work without a prefix, but precision is reduced.

These text prefixes are not used for media inputs. Keep the query and indexed content aligned to the task and output dimension; mixing incompatible formats or vector sizes undermines comparisons.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a supported numerical precision

Google recommends bfloat16 where hardware supports it, or float32 where it does not, including on most CPUs. The model card warns against float16: its limited dynamic range can produce NaN values or silently degraded embeddings.

Plan for the full retrieval application

Embeddings are useful only as part of a workflow that encodes content, stores vectors and retrieves relevant matches. Google’s developer guide names Qdrant as a vector-storage option. The right database, indexing strategy and relevance checks depend on the application; the model’s shared space does not by itself guarantee good search results.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can it run locally or in a browser?

Google reports that, with quantization on a Pixel 11 Pro, the text-only weights use about 191 MB of active RAM and the full multimodal model about 567 MB. These are Google’s figures for that specific phone and setup—not minimum requirements, general RAM estimates or guarantees for other devices.

Google’s October 6, 2026 launch says weights are available through Hugging Face and Kaggle, with on-device optimized versions via the LiteRT Community on Hugging Face. The launch described support in Gemini Enterprise Agent Platform Model Garden as “coming soon”; its availability may have changed since that announcement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google lists MediaPipe and LiteRT for on-device deployment, transformers.js and WebGPU for browser use, and transformers, Sentence Transformers (version 6.1.0 or later in the developer guide), MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio among development or serving options. It also links to Unsloth fine-tuning guidance. These are integrations and resources named by Google, not a claim that every combination supports every modality or feature equally.

What are the model’s data and safety limitations?

Google’s model card says pretraining included web documents, code, images, video, audio and paired cross-modality examples, with a data cutoff of January 2025. The web-text portion included more than 140 languages. Google describes the model as supporting more than 100 languages, while warning that performance may be unequal across them.

The card says training-data filtering included multiple stages for child sexual abuse material and automated filtering for certain personal information and other sensitive data. It also says EmbeddingGemma 2 is a pretrained embedding model with no post-training alignment, safety tuning or output-level moderation. Developers are responsible for application-level protections such as retrieval filtering and fairness testing, and must follow Google’s Gemma Prohibited Use Policy.

How does EmbeddingGemma 2 compare with the first EmbeddingGemma?

The available direct comparison in Google’s model card is limited to the two MTEB results above: the multilingual v2 Mean(Task) score is 61.36 versus 61.15, and the Code v1 Mean(Task), NDCG@10 score is 78.68 versus 68.76. Those figures are useful within their named benchmarks, but they do not establish a universal ranking across tasks or products.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For practical deployment, the more consequential choices may be which modalities to load, how much of the shared context each input uses, and whether reduced vector dimensions meet the application’s quality needs. The reviewed Google materials do not establish an independent head-to-head against competing embedding products under common conditions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.