Laya's 421M-Parameter Decision Model Speed Benchmarked Across NVIDIA GPUs
benchmarks nvidia
| Source: Mastodon | Original article
A benchmark of the 421 M‑parameter Laya decision model on an NVIDIA H100 NVL shows it can process over 15 million decisions per day.
A new benchmark shows that Laya, an open‑source 421‑million‑parameter decision model, can sustain massive throughput on a single NVIDIA H100 NVL GPU. In a recent test the model delivered 15.1 million decisions per day while keeping the 99th‑percentile latency under 130 ms. The result highlights how a non‑autoregressive “System 1” engine can combine modest memory requirements—about 1 GB—with sub‑30 ms per‑decision speeds, rivaling cloud‑based inference services.
Laya, released under an Apache 2.0 licence, is positioned as a free alternative to commercial offerings such as TypeSafe’s Jev. It supports more than 100 languages and charges nothing per token, making it attractive for developers seeking locally hosted, cost‑effective decision‑making. Earlier community tests reported average decision times of roughly 33 ms and a best‑case 32.8 ms on CPU‑based setups, underscoring the model’s efficiency across hardware tiers.
The benchmark matters because it demonstrates that high‑throughput, low‑latency inference is no longer exclusive to large cloud providers. Enterprises in the Nordics, where data‑sovereignty and energy costs are key concerns, can now contemplate deploying sophisticated decision models on‑premise without sacrificing performance. The ability to process over 15 million decisions daily on a single H100 also suggests that scaling to multi‑GPU clusters could push throughput into the hundreds of millions, potentially reshaping cost structures for real‑time AI applications.
Going forward, observers will watch for broader performance data on other NVIDIA GPUs and alternative accelerators, as well as any software optimisations that further shrink latency. Adoption metrics—especially in sectors like finance, logistics and public services—will indicate whether Laya can translate its technical promise into real‑world impact.
Sources
Back to AIPULSEN