Repacked QAT Gemma 4 runs on One TPU v5e, delivering 12B at 675 tokens per second
chips gemma google tpu
| Source: Dev.to | Original article
Google has unveiled a new quantization‑aware‑trained (QAT) version of its Gemma 4 family, repacked into int4 and int8 formats for vLLM inference on a single TPU v5e chip. The repack delivers a full suite of model sizes—from the 2 billion‑parameter “E2B” up to 26 billion—while the 12‑billion‑parameter variant runs at 11.31 GiB of memory and produces 675 output tokens per second. In head‑to‑head tests the int4/int8 Gemma 4 scores up to 2.4 points higher than Google’s own 4‑bit exports at the same speed, confirming that the QAT checkpoints improve both efficiency and quality.
The breakthrough matters because it pushes the limits of what a single accelerator can handle. By halving the precision of weights without sacrificing accuracy, the model fits comfortably on a v5e core, cutting hardware costs and energy consumption for high‑throughput workloads such as conversational agents, coding assistants, and multimodal reasoning. The result aligns with the broader industry push for token‑efficient inference, a theme we explored earlier this week in our coverage of GPT‑6 Astra robot agents that achieved higher success rates while using 65 % fewer tokens.
Looking ahead, developers will be watching how quickly the repacked Gemma 4 is integrated into cloud‑based inference services and whether similar QAT pipelines can be applied to larger models. Further performance gains may emerge from upcoming TPU generations or from refinements to the int4/int8 embedding strategy. The next milestone will be real‑world deployment at scale, especially in Nordic enterprises that rely on low‑latency, cost‑effective AI for agentic workflows.
Sources
Back to AIPULSEN