Mercury 2.5 LLM achieves 770 tokens per second
reasoning
| Source: HN | Original article
Mercury 2.5 LLM reaches 770 tokens per second, delivering output roughly seven times faster than comparable models.
Inception’s newest diffusion‑based large language model, Mercury 2.5, has demonstrated a striking speed advantage in early tests. Through the company’s public API the model generated 770.4 tokens per second, a rate roughly seven times faster than the 108.6‑token‑per‑second median recorded for comparable reasoning models. In‑house deployments have pushed the figure even higher, with Inception reporting live‑run speeds north of 1,100 tokens per second.
The boost is more than a headline number. Faster token throughput translates into lower latency for interactive applications, tighter integration loops for developers, and the ability to handle larger workloads without scaling compute resources proportionally. Inception also claims a 40 % intelligence uplift over its predecessor, Mercury 2, positioning the model as the most capable diffusion LLM currently available and on par with cost‑optimised frontier offerings.
Speed has become a decisive factor as enterprises move from experimental AI pilots to production‑grade services. A model that can reason at this pace could reshape real‑time use cases such as conversational assistants, code‑generation tools, and on‑the‑fly data analysis, where every millisecond counts.
The next steps will reveal whether Mercury 2.5’s performance holds up across diverse workloads and pricing tiers. Industry observers will be watching benchmark releases, developer adoption rates, and any pricing disclosures from Inception. If the model lives up to its claims, it could pressure rival providers to accelerate their own inference optimisations, tightening the race for the fastest, most cost‑effective reasoning engines.
Sources
Back to AIPULSEN