Qwen 3.8 Flash Next (125B) Achieves 100 T/s on Consumer‑Grade RTX 4090
nvidia open-source qwen
| Source: HN | Original article
A 125‑billion‑parameter Qwen 3.8 Flash Next model can run on a consumer RTX 4090 GPU, delivering roughly 100 trillion operations per second.
A new open‑source inference engine called **Strata** has demonstrated that the 125‑billion‑parameter Qwen 3.8 Flash Next model can run on a single consumer‑grade NVIDIA RTX 4090 GPU, reaching roughly 100 tera‑operations per second. The benchmark, posted on GitHub and highlighted on Hacker News, shows the model delivering 21 tokens per second during decoding and 364 tokens per second for pre‑fill, all while using a 250 k‑token context window on the card’s 24 GB of VRAM.
The achievement matters because it shatters the long‑standing “VRAM barrier” that has kept MoE (Mixture‑of‑Experts) models of this size confined to data‑center hardware. By activating only about 6 billion parameters per token, Qwen 3.8 Flash Next can fit within the memory limits of a high‑end gaming GPU, opening the door to truly local, high‑capacity AI for developers, researchers and power users. The result is a potential shift away from cloud‑only inference, with implications for privacy, latency and cost.
The breakthrough follows a wave of hardware‑focused AI news, including OpenAI’s recent Jalapeño inference chip co‑designed with Broadcom, underscoring a broader trend toward democratizing large‑scale models.
What to watch next: the Strata project will likely expand support to AMD GPUs and explore further speed‑ups through quantisation or kernel optimisations. Industry observers will monitor whether other 100‑billion‑plus models can be similarly ported, and how cloud providers respond to a growing ecosystem of “datacenter‑grade” AI running on consumer hardware. The community‑licensed Qwen 3.8 weights, released under the Qwen Community License 1.0, make it easy for anyone to experiment, suggesting rapid iteration and broader adoption in the months ahead.
Sources
Back to AIPULSEN