18 points | by el_duderino 3 hours ago ago
17 comments
Discussed yesterday https://news.ycombinator.com/item?id=49920932
I got a 64gb 6000mt/s kit for 250 recently. Not a humble brag.
How?
More like they're praying it does while waiting for more options to vest
yeah, at some point memory footprint becomes an architectural constraint again, not just a number you throw more hardware at.
This is driven by AI and at a certain point you just need a minimum amount of memory to load all the floating points representing Neural Net weights.
But still it is about GiB instead of KiB.
It's about 1TiB of HBM going into a GPU server eating about 3TiB of DDR that could have gone to consumer electronics.
The memory capacity wall has been coming at us at high speed in plain sight for decades. We've been ignoring it.
Doesn’t seem like it was coming at a high speed if it took decades to manifest.
Can someone please explain to me how the recent llm cache breakthrough doesn't alleviate this memory shortage issue? https://intl.cloud.baidu.com/en/article/8937874
Why use less memory when you can just have more LLM?
See Jevons paradox. https://en.wikipedia.org/wiki/Jevons_paradox
Demand outstrips supply. Efficient software is great, but it doesn’t directly resolve insufficient supply of hardware.
I have this wild theory that this has not just todo with AI but with weapons production overall.
Not everyone is using baidu?
If training and inference is hardware constrained, and you can train and server better and bigger models on the same hardware with memory optimizations, that's exactly what I would expect companies to do.
Discussed yesterday https://news.ycombinator.com/item?id=49920932
I got a 64gb 6000mt/s kit for 250 recently. Not a humble brag.
How?
More like they're praying it does while waiting for more options to vest
yeah, at some point memory footprint becomes an architectural constraint again, not just a number you throw more hardware at.
This is driven by AI and at a certain point you just need a minimum amount of memory to load all the floating points representing Neural Net weights.
But still it is about GiB instead of KiB.
It's about 1TiB of HBM going into a GPU server eating about 3TiB of DDR that could have gone to consumer electronics.
The memory capacity wall has been coming at us at high speed in plain sight for decades. We've been ignoring it.
Doesn’t seem like it was coming at a high speed if it took decades to manifest.
Can someone please explain to me how the recent llm cache breakthrough doesn't alleviate this memory shortage issue? https://intl.cloud.baidu.com/en/article/8937874
Why use less memory when you can just have more LLM?
See Jevons paradox. https://en.wikipedia.org/wiki/Jevons_paradox
Demand outstrips supply. Efficient software is great, but it doesn’t directly resolve insufficient supply of hardware.
I have this wild theory that this has not just todo with AI but with weapons production overall.
Not everyone is using baidu?
If training and inference is hardware constrained, and you can train and server better and bigger models on the same hardware with memory optimizations, that's exactly what I would expect companies to do.