Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Deliver E2B in 2.86 GiB at 2.30× bf16
embeddings gemma google
| Source: Dev.to | Original article
Google's Gemma 4 model, run on a Tesla T4, reduces its embedding size from 6.33 GiB to 2.86 GiB using int4 quantization, while maintaining identical output and boosting vLLM throughput by up to 37%.
Google has demonstrated that its quantization‑aware‑trained (QAT) Gemma 4 E2B model can be compressed to 4‑bit integer (int4) embeddings without losing output fidelity, while slashing memory use and boosting inference speed on a single NVIDIA Tesla T4.
The experiment, detailed in a step‑by‑step guide, repacked the model’s bf16 embedding tables into int4 using the grid QAT pipeline. The resulting checkpoint occupies 2.86 GiB, down from the original 6.33 GiB, yet every greedy‑decoded token matches the bf16 reference exactly. When served with vLLM on a Compute Engine VM, the int4‑embedding build delivers 11‑37 % higher token throughput than Google’s own W4A16 export, and the overall model loading time is cut by more than half.
Why it matters: embedding tables typically dominate the memory footprint of large language models, especially on modest GPUs such as the T4. By compressing them to int4, developers can run Gemma 4 E2B on cheaper, lower‑power hardware while retaining the same quality of output. The throughput gains also translate into lower latency and operating costs for cloud‑based inference services. This follows earlier findings that QAT weights decode 1.79 × faster than bf16 on the same hardware, underscoring the practical benefits of end‑to‑end 4‑bit quantisation.
What to watch next: the community will likely test the int4‑embedding approach on larger Gemma 4 variants and on alternative accelerators, while Google may extend the technique to multimodal inputs and the upcoming 128 K context window. Observers should also keep an eye on how these compression gains influence pricing and accessibility of AI services across the Nordic cloud market.
Sources
Back to AIPULSEN