At the expected price point it would be nice if it came with a free pied-à-terre in Kensington (https://en.wikipedia.org/wiki/Kensington). Just a small studio apartment, nothing special.
I'm just barely starting to wrap my head around mapping model sizes and quants to hardware components and constraints.
I don't get what this device is for.
I a million percent understand wanting 252GB-VRAM, that would get me a 280-320B model like glm-5.3 or deepseek-v4-flash, which would be a massive improvement over my gpt-oss:20b 16GB toy. I would gladly pay a grand for this, I would never pay ten grand for this, and it seems to be priced around a hundred grand.
So obviously the customer is commercial not consumer.
Can anyone planning a project around one of these at work share what their workload is shaped like and how they're modeling price/performance?
For instance, I don't get the 512GB of system memory, I'd gladly drop that to 128 to save money. Am I missing something about commercial workloads? Is a 1T paramater model at 20 tok/s more important to your workload than a 300B one at 60? Is it simply a co-dependency of not the model but the other software you're running on the machine thats using/interacting-with/being-driven-by the model?
Whats your napkin math to justify 100k? Actually thats not even really the question, its more like - whats your napkin math to determine between the "dual linked GB10" use case vs this product's use case vs an 8U supermicro with 4 cards use case.
Quite possibly, depending on grid-scale renewable deployment. I also am about to install a set of solar panels, so a base level of power will cost me only the depreciation of the PV hardware.
We could condition datacenter installs to providing power to the grid - you want to build a datacenter, you also need to build a wind or solar farm that can fully power its peak load.
There's a lot more going on here because the host machine has a boatload (496GB) of expensive LPDDR5x (which is stupidly expensive) that can also be used as unified (but slower) memory to the GPU, 72 ARM64 cores, stupid fast NVlink/QSFP networking, the PSU to support all that etc. etc.
Basically it's the same as a tray in a GB300 NVL72, but in workstation/desktop form. Niche would be AI researchers.
They will... not sell a lot of these. But what a beast.
I am too lazy to price out 496GB of DDR5 but, um, mostly because it is terrifying to see what today's prices look like.
You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.
> You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.
That's kind of the nature of capitalism and price elasticity - the manufacture price only limited the minimum sale price, and sale price usually reflects how much is the market willing to pay for the good.
Where you might save a lot of money is on building something that's very targeted to your needs that matches them better than a GB300 workstation, for a lower price.
HP’s spec sheet fills in details the platform announcements skipped. The CPU memory is four 128GB SOCAMM modules delivering 396GB/s, and the Grace CPU is soldered to the host processor module rather than socketed. The two pools add up to the 748GB coherent space that lets the GPU address CPU memory directly, which is what makes trillion-parameter inference and fine-tuning of models in the 100 billion parameter class possible on a single box. HP’s footnote on those model sizes is that the harness quantizes at FP4.
I find it just misleading by advertising 700GB RAM as AI headline. I could plug a 32GB GPU to my 1.5TB ram server and call it “AI station with 1.5TB+ RAM” just so that you find it actually useless compared to the headline
I don't know about "useless" (it seems quite useful to me) but I do feel mislead. It's unified memory in the same way that my current dGPU has unified memory. I guess nvlink-c2c probably (?) doesn't introduce a bottleneck but it's still two distinct arenas with very different performance characteristics.
7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models.
It's a workstation, not a rack. It's for AI researchers. I'd love to have one on my desk.
Someone should edit the title to make it clear it only has 252GB of “AI” memory
Looks like multiple vendors are seeling these:
https://www.gigabyte.com/in/Enterprise/Tower-Server/W775-V10...
https://www.msi.com/Landing/NVIDIA-DGX-STATION
> a Kensington slot
I think you might need a bit more than that at this price point ...
At the expected price point it would be nice if it came with a free pied-à-terre in Kensington (https://en.wikipedia.org/wiki/Kensington). Just a small studio apartment, nothing special.
You'd need to spent ~£500K+ to get a studio in Kensington these days that isn't a shoebox.
One listing is 410K for 280 sq/ft coming out at £1464 per square foot, almost exactly 10x the price per sq/ft we paid for our house a few years ago.
So $100K would be £74K which would get you ~50 sq/ft.
A kensington lock in this context is like a sign that says "do not move", which is still kinda useful
I'm just barely starting to wrap my head around mapping model sizes and quants to hardware components and constraints.
I don't get what this device is for.
I a million percent understand wanting 252GB-VRAM, that would get me a 280-320B model like glm-5.3 or deepseek-v4-flash, which would be a massive improvement over my gpt-oss:20b 16GB toy. I would gladly pay a grand for this, I would never pay ten grand for this, and it seems to be priced around a hundred grand.
So obviously the customer is commercial not consumer.
Can anyone planning a project around one of these at work share what their workload is shaped like and how they're modeling price/performance?
For instance, I don't get the 512GB of system memory, I'd gladly drop that to 128 to save money. Am I missing something about commercial workloads? Is a 1T paramater model at 20 tok/s more important to your workload than a 300B one at 60? Is it simply a co-dependency of not the model but the other software you're running on the machine thats using/interacting-with/being-driven-by the model?
Whats your napkin math to justify 100k? Actually thats not even really the question, its more like - whats your napkin math to determine between the "dual linked GB10" use case vs this product's use case vs an 8U supermicro with 4 cards use case.
252GB of HBM3e at 100k vs. a multi A6000 96GB setup. GB300 seem expensive in comparison at ~$100k. Am I missing something?
Without looking it up and doing the math, I bet GB300 has higher Tflops and memory bandwidth, especially when used with e.g. NVFP4.
Plus lower peak power draw
How long until the higher power draw nullifies the higher acquisition price of the GB300 machine?
Are you in a world where energy prices are goimg down?
Quite possibly, depending on grid-scale renewable deployment. I also am about to install a set of solar panels, so a base level of power will cost me only the depreciation of the PV hardware.
We could condition datacenter installs to providing power to the grid - you want to build a datacenter, you also need to build a wind or solar farm that can fully power its peak load.
There's a lot more going on here because the host machine has a boatload (496GB) of expensive LPDDR5x (which is stupidly expensive) that can also be used as unified (but slower) memory to the GPU, 72 ARM64 cores, stupid fast NVlink/QSFP networking, the PSU to support all that etc. etc.
Basically it's the same as a tray in a GB300 NVL72, but in workstation/desktop form. Niche would be AI researchers.
They will... not sell a lot of these. But what a beast.
I am too lazy to price out 496GB of DDR5 but, um, mostly because it is terrifying to see what today's prices look like.
You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.
> You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.
That's kind of the nature of capitalism and price elasticity - the manufacture price only limited the minimum sale price, and sale price usually reflects how much is the market willing to pay for the good.
Where you might save a lot of money is on building something that's very targeted to your needs that matches them better than a GB300 workstation, for a lower price.
HP’s spec sheet fills in details the platform announcements skipped. The CPU memory is four 128GB SOCAMM modules delivering 396GB/s, and the Grace CPU is soldered to the host processor module rather than socketed. The two pools add up to the 748GB coherent space that lets the GPU address CPU memory directly, which is what makes trillion-parameter inference and fine-tuning of models in the 100 billion parameter class possible on a single box. HP’s footnote on those model sizes is that the harness quantizes at FP4.
How many arms and legs does one of these cost?
3.14 kidneys.
But can it run crysis?
if not yet then it may be achievable:
https://www.guru3d.com/story/nvidia-dgx-spark-achieves-175-f...
Can I get 32 GB of RAM for a sane price instead?
I guess I'm too poor to even know the price
MSI’s DGX Station (what this is) was listed for $99K
Dell's equivalent starts at £180k.
Furious
What is it with those stupid names?
No price, so of course this is not for the smelly working class.
Imagine a beowolf cluster of these!
[dead]
[dead]
>> 252GB of HBM3e at 7.1TB/s,
So it has only 252GB of actual ”AI” memory making it “useless”/toy for actual real world AI workloads(I.e it can’t replace something like opus 5)
I find it just misleading by advertising 700GB RAM as AI headline. I could plug a 32GB GPU to my 1.5TB ram server and call it “AI station with 1.5TB+ RAM” just so that you find it actually useless compared to the headline
Should work just fine for MoE models where active set fits into 252GB.
Can’t you do that already more or less with a Mac Studio with 256 or better 512fb of ram?
Where are you getting 7 TB/s of memory bandwidth?
I don't know about "useless" (it seems quite useful to me) but I do feel mislead. It's unified memory in the same way that my current dGPU has unified memory. I guess nvlink-c2c probably (?) doesn't introduce a bottleneck but it's still two distinct arenas with very different performance characteristics.
7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models.
It's a workstation, not a rack. It's for AI researchers. I'd love to have one on my desk.
What even is this comment?
[dead]