This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?
I think the frontier providers keep the size of their models carefully hidden.
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
Not necessarily, there could be diminishing returns on mere parameters count .
There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*.
if you have the money as an "individual user" to purchase one of their racks... save your money and retire.
* Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.
Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.
Ah but if you have the kind of money where this is a reasonable retirement hobby purchase, you aren't bothered by representing yourself as a "enterprise" :P
And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.
Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C…
More realistically, you need much more cooling water.
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
Without having any inside information, one possible theory:
All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.
or
The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.
imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
And, the software side isn't finished being optimized, either. We've seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve.
It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.
Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.
There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.
As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.
TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
Yeah, but there's certainly a part of the curve where price drops by X OOMs and demand increases by much more than X OOMs. (Presumably some of that is substitution and some of that is new use cases.)
I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?
I'd think the infrastructure would tend towards factories, smelters, and so on. Industrial things that have reasonably high power demands, can use the building, and don't care about the lack of windows.
They're typically not built where you want housing, and the buildings are distinctly the wrong shape.
If you can't use the power infrastructure profitably my next thought would be warehousing.
But also... we've seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.
Huh, why I'm not surprised that HN is full of opinions confidently stated without any numbers or resources to back up?
> built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.
Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.
GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.
Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.
If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.
But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity.
If they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.
Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).
I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training.
I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
Is it just me or is it bizarre that they're advertising old open-weight models.
GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7.
Kimi K2.7 (April) not K2.7-code (June) or K3 (July).
Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models).
Meanwhile the closed source GPT 5.6 sol is up to date (June)...
Should potential purchasers take away from this that they're not going to be able to run recent models unless they front the cost of developing software or something?
I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product.
But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".
I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
(Where did you see that?)
This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?
I think the frontier providers keep the size of their models carefully hidden.
It's rumored fable is around that 10T number
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
Not necessarily, there could be diminishing returns on mere parameters count .
There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Oops did they just out GPT-5.6 sol’s parameter count?
Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
It would be even better if a version available to individual users were released soon.
They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*.
if you have the money as an "individual user" to purchase one of their racks... save your money and retire.
* Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.
> save your money and retire.
Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.
Ah but if you have the kind of money where this is a reasonable retirement hobby purchase, you aren't bothered by representing yourself as a "enterprise" :P
I’ll get that 250kW home power service dropped in next week!
Conspicuously missing: power consumption figures
162 kW
I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.
And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.
God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.
I presume per rack?
Can you imagine something radiating that much energy into a space in your home?
It's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building.
This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.
Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C…
More realistically, you need much more cooling water.
I guess because I have actually set foot in a data center I don't imagine literally every product in my home.
"10x more throughput per watt than CS-3"
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
Without having any inside information, one possible theory:
All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.
or
The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.
If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.
it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.
imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
> enabling massive clusters and models with more than 50 trillion parameters
Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Wow!
And, the software side isn't finished being optimized, either. We've seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve.
It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.
Hence why taalas was one of the best strategic acquisitions of the year.
I'm honestly baffled they were not acquired by somebody else (sorry AMD).
Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.
There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.
As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.
TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
Yeah, but there's certainly a part of the curve where price drops by X OOMs and demand increases by much more than X OOMs. (Presumably some of that is substitution and some of that is new use cases.)
I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?
I'd think the infrastructure would tend towards factories, smelters, and so on. Industrial things that have reasonably high power demands, can use the building, and don't care about the lack of windows.
They're typically not built where you want housing, and the buildings are distinctly the wrong shape.
If you can't use the power infrastructure profitably my next thought would be warehousing.
But also... we've seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.
Huh, why I'm not surprised that HN is full of opinions confidently stated without any numbers or resources to back up?
> built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.
Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.
GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.
I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?
In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?
And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.
So unless there are some other bottlenecks, we'll use a lot of simulation servers.
Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.
If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.
But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
On the plus side, lots of cheap servers to swoop up :)
But power hungry.
In that 5+ year timeline, the compute per watt could change by three orders of magnitude.
GPUs are to LLMs what CPUs are to gaming — not a good fit.
> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity.
Did nobody proofread this?
If they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.
That's unusually honest and the sharpest thing in this thread.
You're underselling it, and here's why.
Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.
There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection.
It’s still incredibly obvious.
Maybe it is just part of their "compact design".
An error no frontier LLM would make, eh
KV caching status?
What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4
44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).
I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training.
I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
Cerebras is only claiming ~2x the performance of Groqvidia which usually isn't enough for people to switch.
I wonder what are the benchmarks of hashcat on different hashes.
Is it just me or is it bizarre that they're advertising old open-weight models.
GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7.
Kimi K2.7 (April) not K2.7-code (June) or K3 (July).
Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models).
Meanwhile the closed source GPT 5.6 sol is up to date (June)...
Should potential purchasers take away from this that they're not going to be able to run recent models unless they front the cost of developing software or something?
I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.
I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product.
But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".
I feel like 6-figures would be the clearance price on it...
6 figures is a single mid range Xeon or Epyc server these days.