Qwen 3.8 27B achieves 1,500 tokens per second on Cerebras
qwen
| Source: HN | Original article
Cerebras now offers the Qwen 3.8 27B model, delivering roughly 1,500 tokens per second.
Cerebras has added Alibaba’s Qwen 3.8 27B to its public inference catalog, promising generation speeds of roughly 1 500 tokens per second and a context window of 128 k tokens. The dense, multimodal model—capable of processing both text and images—targets “agentic coding, tool use, research and long‑running workflows,” according to the provider’s API description.
The announcement marks the first time the 27‑billion‑parameter Qwen 3.8 model is offered as a managed service, removing the need for users to maintain their own hardware. By delivering high‑throughput inference on a platform built for large‑scale AI, Cerebras positions the model as a practical option for developers building complex pipelines that require sustained reasoning over very long inputs, such as autonomous‑driving assistants or multi‑step research agents.
The move matters because it expands the competitive landscape beyond the recently launched Gemini 3.8 Flash series, which has been highlighted for its speed and cost profile. While Gemini Flash focuses on rapid, short‑context generation, Qwen 3.8’s 128 k token window and multimodal capabilities cater to use cases where breadth of context and visual understanding are essential. For Nordic enterprises exploring AI‑driven automation, the combination of Cerebras’ hardware efficiency and Qwen’s extended context could lower barriers to deploying sophisticated agents at scale.
What to watch next includes pricing details for the Cerebras offering, real‑world benchmark results against peers such as Gemini Flash, and any extensions of the context length beyond 128 k tokens. Observers will also be keen to see how quickly developers adopt the model for tool‑augmented workflows and whether additional multimodal features are rolled out through the Qwen Cloud API.
Sources
Back to AIPULSEN