Start with an official Gemma 4 QAT checkpoint if Google provides one for your model size and deployment runtime, and reducing memory is your main goal. Google reports that its QAT models deliver higher overall quality than its standard PTQ baselines while using less memory. That is a vendor-reported overall result, not evidence that QAT beats every PTQ method on every task or device. Choose PTQ when its format or runtime better fits your deployment, then compare both approaches on your own workload.
What QAT and PTQ mean for Gemma 4
Post-training quantization (PTQ) compresses a model after it has been trained. Quantization-aware training (QAT) incorporates simulated quantization into training, giving the model an opportunity to adapt to the precision limits. Google describes this distinction in its Gemma 4 model overview.
As an Amazon Associate I earn from qualifying purchases.
Google’s launch article says, “While PTQ is already effective at preserving quality, our QAT results yield even higher overall quality compared to standard PTQ baselines.” Treat that as Google’s overall finding, not a guarantee for every quantizer, bit width, task, model variant, or runtime. The reviewed official material does not provide a numerical Gemma 4 QAT-versus-PTQ quality advantage across controlled, matched task and hardware tests. There is no supported universal percentage to cite.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteChoose by runtime and checkpoint availability
For Gemma 4, the practical choice often starts with the deployment tool: the official QAT artifacts are tied to specific formats and workflows. Google’s model overview documents these directions:
#1 Best Overall
| Deployment target | Documented QAT direction | Qualification |
|---|---|---|
| Local inference with llama.cpp or LM Studio | Q4_0 GGUF checkpoints | Google lists E2B, E4B, 12B, 26B-A4B, and 31B variants. |
| vLLM or SGLang serving | W4A16 compressed-tensors checkpoints | Google’s overview lists E2B, E4B, 12B, and 31B. The vLLM recipe excludes 26B-A4B from 4-bit W4A16 because of excessive quality loss in that recipe. |
| Mobile or edge deployment | Mobile-optimized QAT checkpoints | Google’s overview lists E2B and E4B; the format uses specialized low-bit components and optimized KV caches. |
| Conversion to another format | Unquantized QAT checkpoint | Intended for downstream conversion or compilation; whether it works depends on the destination toolchain. |
| Speculative decoding | QAT target with a matching QAT assistant | The official model card says to use assistant and target checkpoints at the same precision. |
For the 26B-A4B model in the documented vLLM recipe, the project suggests int8 per-channel weight-only quantization instead of 4-bit W4A16. This is recipe-specific guidance, so check current runtime support and evaluate the option for your setup. See the vLLM Gemma 4 recipe and Google’s format overview.
Understand what the memory figures include
Quantized weights are only part of runtime memory. Google’s overview says its base-weight estimates exclude software overhead and KV-cache memory. KV-cache use grows with prompt and generated tokens, so longer contexts and outputs require more memory; serving concurrency and runtime overhead also affect the total.
Rank #2
The vLLM recipe reports these estimated W4A16 memory changes:
Recommended Free Tools
| Gemma 4 model | Recipe estimate before | Recipe estimate with W4A16 |
|---|---|---|
| E2B | 9.8 GB | 7.3 GB |
| E4B | 15.2 GB | 9.8 GB |
| 12B | 22.8 GB | 8.3 GB |
| 31B | 59.0 GB | 19.8 GB |
These are estimates from the vLLM recipe, not universal device requirements. They do not establish total memory for every context length, serving load, or runtime configuration.
Rank #3
Mobile figures describe specific configurations
Google’s June 5, 2026 article says its mobile-specialized format reduces Gemma 4 E2B’s memory footprint to 1 GB. In a separate configuration, the text-only E2B model without Per-Layer Embeddings requires less than 1 GB. These figures refer to the stated configurations, not a guarantee of total memory for every runtime, context, or workload. Google’s article describes static activations, channel-wise quantization, targeted 2-bit layers, and embedding and KV-cache optimization as parts of the mobile approach. See Google’s Gemma 4 QAT announcement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Compare quality and speed on your workload
Neither the overall Google quality claim nor a smaller weight estimate tells you which checkpoint will work best for your particular prompts and hardware. Compare candidates using the same base model, task set, context length, runtime version, and device. Include the capabilities you actually need—such as coding, reasoning, factual answers, or multimodal input—and measure both quality and latency or throughput.
For serving, include total memory and the concurrency you expect, rather than evaluating weights alone. The vLLM recipe’s throughput and speculative-decoding guidance is tied to documented runtime and hardware scenarios; it says its speculative-decoding settings were benchmarked on NVIDIA A100 and H100 and that optimal settings may vary. Do not assume those settings or results transfer unchanged to different hardware.
Quick Recap
Best Value
A practical decision path
- Identify the model and runtime you need. Find whether Google documents a matching QAT checkpoint for your Gemma 4 variant and deployment tool in the Gemma 4 overview.
- Try the matching QAT artifact first when memory is the constraint. Google says its QAT checkpoints preserve quality similar to bfloat16 and outperform standard PTQ baselines overall; treat both as Google’s reported findings, not universal guarantees.
- Use PTQ if it fits your deployment better. It may be the practical option when the needed format or runtime is not served by an official QAT artifact, or when your own evaluation favors it for memory, quality, or speed.
- Measure the whole workload. Test representative prompts and context lengths, account for generated tokens and KV cache, and track the quality, latency or throughput, and memory that matter to your use case.
- Check variant-specific compatibility. Do not assume every Gemma 4 model has identical 4-bit support. For speculative decoding, pair assistant and target checkpoints at matching precision.
Sources for Gemma 4 formats and compatibility
- Google’s June 5, 2026 QAT announcement describes the mobile-specific approach and Google’s overall quality comparison.
- Google AI for Developers’ Gemma 4 model overview documents QAT formats, memory-estimate qualifications, and deployment routes.
- Google’s Gemma 4 E2B QAT Q4_0 GGUF model card gives checkpoint-specific information, including the same-precision guidance for speculative decoding.
- The vLLM Project’s Gemma 4 recipe documents its W4A16 estimates and model-specific serving guidance.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




