OpenAI and Cerebras Confirm 750 MW AI Inference Deployment Through 2028
inference openai
| Source: Mastodon | Original article
OpenAI and Cerebras have confirmed a multi‑year partnership to deploy 750 MW of wafer‑scale AI inference capacity through 2028.
OpenAI and Cerebras have formalised a multi‑year partnership to add 750 megawatts of wafer‑scale AI compute to OpenAI’s inference platform. The deal, announced on June 23, is valued at more than $20 billion and will be rolled out in phases beginning in 2026, with the final tranche expected by the end of 2028. The first 100 MW of capacity is already operating at Cerebras’s Memphis facility, marking the start of a staged build‑out that promises ultra‑low‑latency inference for OpenAI’s services.
The deployment is significant because it represents the largest high‑speed inference infrastructure ever announced. By leveraging Cerebras’s wafer‑scale chips, OpenAI can deliver faster response times for applications that demand real‑time processing, from conversational agents to code‑generation tools such as the recently reported Dot agent. The scale of the investment also underscores a strategic shift toward dedicated, on‑premise inference hardware, reducing reliance on generic cloud GPUs and potentially lowering operational costs for high‑throughput workloads.
Industry observers will watch how the phased rollout impacts OpenAI’s product performance and pricing, especially as the company expands its enterprise offerings. Key questions include the actual latency gains achieved, the energy efficiency of the wafer‑scale systems, and whether the partnership will spur similar deals between other AI developers and specialised hardware vendors. Further announcements are expected as each tranche comes online, offering concrete metrics that could reshape expectations for AI inference speed and scalability across the sector.
Sources
Back to AIPULSEN