/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Cerebras launches the “world's fastest” AI inference service with “GPU-impossible performance”, with costs starting at $0.10 per million tokens, to rival Nvidia

Mike Wheatley / SiliconANGLE :

SiliconANGLE Mike Wheatley

Context & Ripple Effects

Cerebras had previously focused its wafer-scale hardware claims on running and training very large language models, including a single-device NLP model record. This launch turns that hardware positioning into a metered inference service, putting its value proposition in terms developers and AI product teams can compare: speed and token cost.

The move foreshadowed a broader inference challenge to Nvidia from Cerebras, Groq and larger platforms targeting inference workloads. Later plans for AWS to deploy Cerebras hardware for inference would extend that route to market, while preserving a lower-cost alternative for some workloads.

First-order effects

  • Cerebras gains a commercial inference offering with published entry pricing, giving customers a direct way to evaluate its wafer-scale system against GPU-backed services.
  • Nvidia faces a more explicit application-serving challenge: Cerebras is marketing performance that it says conventional GPUs cannot deliver, rather than selling hardware only as a training alternative.

Second-order effects

  • AI teams with latency-sensitive or high-volume workloads can benchmark specialized inference against GPU capacity on both throughput and token economics, increasing pressure on providers to differentiate by workload rather than accelerator brand alone.
  • The service model makes Cerebras's hardware claims subject to operational comparison—availability, model support and delivered cost—not just chip-level benchmarks.

Third-order effects

  • If specialized systems repeatedly deliver lower latency or cost for serving models, inference can fragment into workload-specific infrastructure instead of remaining centered on general-purpose GPUs.
  • The durable competitive layer shifts toward integrated hardware, serving software and distribution: a faster accelerator matters commercially only when customers can consume it as a reliable service.

The trend: AI inference is becoming a separate infrastructure battleground, where specialized accelerators compete with GPUs on delivered tokens, latency and operating cost.

Discussion

  • @artificialanlys @artificialanlys on x
    Cerebras has set a new record for AI inference speed, serving Llama 3.1 8B at 1,850 output tokens/s and 70B at 446 output tokens/s. @CerebrasSystems has just launched their API inference offering, powered by their custom wafer-scale AI accelerator chips. Cerebras Inference is [im…
  • @iancutress @iancutress on x
    Day 2 at @hotchipsorg, Sean Lie from @CerebrasSystems to the stage talking about inference at the wafer scale, a🧵 [image]
  • @aiatmeta @aiatmeta on x
    Verified by @ArtificialAnlys, @CerebrasSystems Inference is capable of serving Llama 3.1 70B at 450 tokens/sec and Llama 3.1 8B at 1,850 tokens/sec!
  • @cerebrassystems @cerebrassystems on x
    Verified by @ArtificialAnlys, Cerebras Inference achieves 1,850 tokens/sec on Llama 3.1 8B and 450 tokens/sec on Llama 3.1 70B! By dramatically reducing processing time, we're enabling more complex AI workflows and enhancing real-time LLM intelligence. This includes a new class […
  • @deeplearningai @deeplearningai on x
    “ https://deeplearning.ai/ has multiple agentic workflows that require prompting an LLM repeatedly to get a result. Cerebras has built an impressively fast inference capability which will be very helpful to such workloads,” said Andrew Ng about Cerebras Inference. Learn more
  • @aurickq Aurick Qiao on x
    We used Cerebras to train @llm360 models and throughout the process thinking how it would be a perfect fit for inference. Nice to see it come to fruition, and congratulations to the @CerebrasSystems team!
  • @soumithchintala Soumith Chintala on x
    very cool stuff coming out of Cerebras!
  • @draecomino James Wang on x
    Cerebras is 10x faster than GPUs 10x = 3 Moore's Law doublings = 10 years It's basically like getting a GPU from 2034. When has this ever happened in tech? [image]
  • @l2k Lukas Biewald on x
    Cerebras says their new API is the fastest at LLM inference. We compared five different Llama 70B providers and we found the same results. You can explore it here: https://wandb.me/... Or, read the full comparison here: https://wandb.ai/... For more details, I interviewed [image]
  • @sarahtavel Sarah Tavel on x
    Huge launch from @CerebrasSystems. We've known for decades how important speed is for end user experiences. Now you can get 20x faster inference speed than GPUs with @CerebrasSystems new launch. Game on.
  • @gavinsbaker Gavin Baker on x
    Nice to see more inference competition from @CerebrasSystems I thought the “dial-up to broadband” analogy was compelling - tokens per second makes a massive difference to the user experience. And the price is right at $.10 per million API available right now for developers.
  • @capetorch Thomas Capelle on x
    I have been playing with the @CerebrasSystems inference API for the last few days. Of course, it is really fast; in my testing, I am getting around 400 tok/s on LLama 70b. The endpoint is openai compatible, which makes it handy when testing different providers, as shown in [image…
  • @dzhulgakov Dmytro Dzhulgakov on x
    Congrats to Cerebras on the impressive results! How SRAM-only ASICs like it stack up against GPUs? Spoiler: GPUs still rock for throughput, custom models, large models and prompts (common “prod” things). SRAM ASICs shine for pure generation. Long 🧵 https://x.com/...
  • @ylecun Yann LeCun on x
    Inference on Cerebras: Llama 3.1 70B at 450 tokens/s, Llama 3.1 8B at 1,700 tokens/s.
  • @cerebrassystems @cerebrassystems on x
    Cerebras Inference is the fastest Llama3.1 inference API by far: 1,800 tokens/s for 8B and 450tokens/s for 70B. We are ~20x faster than NVIDA GPUs and ~2x faster than Groq. [image]
  • @cerebrassystems @cerebrassystems on x
    Lastly, we didn't just build a fast demo - we have capacity to serve hundreds of billions of tokens per day to developers and enterprises. We will be adding new models (eg. Llama3.1-405B) and ramp even greater capacity in the coming months. [image]
  • @cerebrassystems @cerebrassystems on x
    Cerebras Inference is just 10c per million tokens for 8B and 60c per million tokens for 70B. Our price-performance is so strong, we practically broke the chart on Artificial Analysis. [image]
  • @cerebrassystems @cerebrassystems on x
    Introducing Cerebras Inference ‣ Llama3.1-70B at 450 tokens/s - 20x faster than GPUs ‣ 60c per M tokens - a fifth the price of hyperscalers ‣ Full 16-bit precision for full model accuracy ‣ Generous rate limits for devs Try now: https://inference.cerebras.ai/ [video]