Cerebras launches the “world's fastest” AI inference service with “GPU-impossible performance”, with costs starting at $0.10 per million tokens, to rival Nvidia
Cerebras had previously focused its wafer-scale hardware claims on running and training very large language models, including a single-device NLP model record. This launch turns that hardware positioning into a metered inference service, putting its value proposition in terms developers and AI product teams can compare: speed and token cost.
The move foreshadowed a broader inference challenge to Nvidia from Cerebras, Groq and larger platforms targeting inference workloads. Later plans for AWS to deploy Cerebras hardware for inference would extend that route to market, while preserving a lower-cost alternative for some workloads.
First-order effects
Cerebras gains a commercial inference offering with published entry pricing, giving customers a direct way to evaluate its wafer-scale system against GPU-backed services.
Nvidia faces a more explicit application-serving challenge: Cerebras is marketing performance that it says conventional GPUs cannot deliver, rather than selling hardware only as a training alternative.
Second-order effects
AI teams with latency-sensitive or high-volume workloads can benchmark specialized inference against GPU capacity on both throughput and token economics, increasing pressure on providers to differentiate by workload rather than accelerator brand alone.
The service model makes Cerebras's hardware claims subject to operational comparison—availability, model support and delivered cost—not just chip-level benchmarks.
Third-order effects
If specialized systems repeatedly deliver lower latency or cost for serving models, inference can fragment into workload-specific infrastructure instead of remaining centered on general-purpose GPUs.
The durable competitive layer shifts toward integrated hardware, serving software and distribution: a faster accelerator matters commercially only when customers can consume it as a reliable service.
The trend: AI inference is becoming a separate infrastructure battleground, where specialized accelerators compete with GPUs on delivered tokens, latency and operating cost.
Cerebras has set a new record for AI inference speed, serving Llama 3.1 8B at 1,850 output tokens/s and 70B at 446 output tokens/s. @CerebrasSystems has just launched their API inference offering, powered by their custom wafer-scale AI accelerator chips. Cerebras Inference is [im…
Verified by @ArtificialAnlys, Cerebras Inference achieves 1,850 tokens/sec on Llama 3.1 8B and 450 tokens/sec on Llama 3.1 70B! By dramatically reducing processing time, we're enabling more complex AI workflows and enhancing real-time LLM intelligence. This includes a new class […
“ https://deeplearning.ai/ has multiple agentic workflows that require prompting an LLM repeatedly to get a result. Cerebras has built an impressively fast inference capability which will be very helpful to such workloads,” said Andrew Ng about Cerebras Inference. Learn more
We used Cerebras to train @llm360 models and throughout the process thinking how it would be a perfect fit for inference. Nice to see it come to fruition, and congratulations to the @CerebrasSystems team!
Cerebras is 10x faster than GPUs 10x = 3 Moore's Law doublings = 10 years It's basically like getting a GPU from 2034. When has this ever happened in tech? [image]
Cerebras says their new API is the fastest at LLM inference. We compared five different Llama 70B providers and we found the same results. You can explore it here: https://wandb.me/... Or, read the full comparison here: https://wandb.ai/... For more details, I interviewed [image]
Huge launch from @CerebrasSystems. We've known for decades how important speed is for end user experiences. Now you can get 20x faster inference speed than GPUs with @CerebrasSystems new launch. Game on.
Nice to see more inference competition from @CerebrasSystems I thought the “dial-up to broadband” analogy was compelling - tokens per second makes a massive difference to the user experience. And the price is right at $.10 per million API available right now for developers.
I have been playing with the @CerebrasSystems inference API for the last few days. Of course, it is really fast; in my testing, I am getting around 400 tok/s on LLama 70b. The endpoint is openai compatible, which makes it handy when testing different providers, as shown in [image…
Congrats to Cerebras on the impressive results! How SRAM-only ASICs like it stack up against GPUs? Spoiler: GPUs still rock for throughput, custom models, large models and prompts (common “prod” things). SRAM ASICs shine for pure generation. Long 🧵 https://x.com/...
Cerebras Inference is the fastest Llama3.1 inference API by far: 1,800 tokens/s for 8B and 450tokens/s for 70B. We are ~20x faster than NVIDA GPUs and ~2x faster than Groq. [image]
Lastly, we didn't just build a fast demo - we have capacity to serve hundreds of billions of tokens per day to developers and enterprises. We will be adding new models (eg. Llama3.1-405B) and ramp even greater capacity in the coming months. [image]
Cerebras Inference is just 10c per million tokens for 8B and 60c per million tokens for 70B. Our price-performance is so strong, we practically broke the chart on Artificial Analysis. [image]
Introducing Cerebras Inference ‣ Llama3.1-70B at 450 tokens/s - 20x faster than GPUs ‣ 60c per M tokens - a fifth the price of hyperscalers ‣ Full 16-bit precision for full model accuracy ‣ Generous rate limits for devs Try now: https://inference.cerebras.ai/ [video]