The memory trade doubts
By Markos
First, let’s be clear. The memory trade will be a bumpy ride. But for my thesis nothing has changed. It is all about perception and questioning.
I have moved the Rubin Ultra de-spec to the top of my memory notes, and not because it is bearish. It is a important read on the latent HBM demand thesis. So I want to walk through what it actually tells us, and then I want to walk through every reason it might be wrong, because I have spent the last weeks stress testing this against my own thesis rather than looking for reasons to feel good about it.
Start with what is not in dispute as Jukan pointed out correctly. Nvidia’s incentive is to ship as many Rubin racks as possible, and lowering HBM content in favour of SOCAMM does exactly that, because of the wafer usage.
Every stack of HBM consumes wafers that could have gone to the module memory sitting on the other side of the rack, and a rack that is short of module memory does not ship at all. From the supply side, the memory players sell all their bits anyway, just spread across more units, and they pick up some incremental module-memory revenue on top. There is no margin conflict here. Both sides win. Nvidia ships more racks, and the memory makers sell everything they can make.
But winning is not the same as getting what you wanted. Nvidia is optimising under a constraint it did not choose. It is making the best of the situation. And that distinction is the whole article.
What the buyer told us without meaning to
Because in this case Nvidia has effectively handed us both numbers.
What it originally wanted is on the record: the announced 1 TB Rubin Ultra configuration. What it is reportedly settling for is around 192 GB, still in testing and not finalised. That is roughly a fivefold reduction, and it is the difference between design intent and delivered reality, published by the buyer itself in two separate acts.
That is revealed behaviour, not a survey, and it is the basis for how I estimate wish-list demand against the small deficit the sell side currently models. Most latent demand is inferred. Original intent was announced and then reduced.
In June, Jensen visited SK hynix at Computex, signed their first HBM4E wafer sample, and left a message on it asking them to make more. clear signal from Jensen.
Now let me be precise about the mechanism, because the usual description of this is wrong and the difference matters.
A part of the people say HBM simply got too expensive for the current architecture. In a shortage, price is not the thing that binds. Availability is. Nvidia is not spec. down because the bill got large. It is specking down because the cost that moved is not the price of HBM. It is the opportunity cost of the wafers standing behind it. And an opportunity cost does not disappear if prices come down. Simple wafer math.
Why the want does not decay in my eyes.
That underlying want does not decay while it is being suppressed, and it does not decay because of NPO either. The workload driver keeps rising. Context windows have grown by orders of magnitude and are still growing. Whatever wanted 1 TB in 2027 will want more than that later. So even if supply reached parity, a part of unserved demand accumulates through the whole de-spec period and releases at the next platform. a potential memory increase as a drop-in when it fits the existing mechanical, thermal and power budget, and forces a platform refresh is also unlikely. And doesnt impact the thesis. Going from 192GB to a 1TB-class configuration with 16-high stacks at relaxed package height is very unlikely to fit.
So what happens at the next architecture? Here I think most people have the mechanism backwards.
The usual framing is that the compute upgrade pays for the extra memory. I do not read it that way. Token generation is bandwidth bound, not compute bound. Add compute without the memory to feed it and you have bought an accelerator that waits. There is a published measurement of exactly this: push the working set one tier down the hierarchy and nearly all the latency ends up sitting in transfers, while the accelerator runs at a small fraction of its power envelope. That is a machine paying for compute it cannot use.
So the next generation does not carry more HBM because it can suddenly afford to. It carries more because the compute it is being sold on is not fully usable without it. Memory is not competing with compute for the budget. It is the thing that unlocks the budget, and of course you pay for the stage that sets your output.
The generational data says the same thing from the other side. Compute per accelerator has been growing two to three times per generation, while bandwidth per memory stack grows maybe half again. That gap is the memory wall, and it widens every cycle. Just holding the bandwidth-to-compute ratio constant requires more stacks per accelerator each generation, not fewer. Which means the real ceiling on HBM content is not demand at all. It is geometry, how many stacks physically fit around a package. And a binding geometry cap is an argument for more accelerators, not for less memory.
Can the architectures just walk away from HBM?
This is the question I spent the most time on, because it is the one that would actually end the thesis. I stress tested it against every alternative on the table, and the physics keeps answering the same way.
Start with SRAM, because it is the cleanest case. The dream solution has always been to put the memory on the logic die itself. The problem is that SRAM has effectively stopped scaling. Logic density still improves meaningfully every node. SRAM bit cells have barely moved across the last two nodes, a few percent against roughly forty for logic. Which means every new node makes on-die memory relatively more expensive, not less. The escape route that would most cleanly remove external HBM looks like it is closing. and it is closing for physics reasons.
Then pooling and tiering, and here is the point almost everyone misses. HBM is bought in stacks, and every stack delivers bandwidth and capacity together, in fixed proportion. Nearly every substitution story you will read is a capacity story: pool the cold weights, tier the cache, offload to flash. But if bandwidth is what binds, and for token generation it is, then removing capacity requirement removes zero stacks. The buyer still needs the stacks for their bandwidth and simply ends up with spare capacity. And this is important i will come back to that later. When you run the interconnect numbers. Pooled memory over the current standard moves on the order of a hundred gigabytes per second per link. One HBM stack moves on the order of two terabytes per second. That is more than an order of magnitude, and it is a property of the interface, not a maturity problem that another software release fixes. Pooling is real, and I watch its deployment dates. What it serves is the capacity-bound edge of the market. It does not reach the bandwidth-bound core.
The LPDDR designs carrying no HBM at all are real too. Shipping, with named customers. I model them as a deduction rather than pretending they do not exist. But look at what they serve: capacity-bound inference at lower throughput. They grow the accelerator market at its edge. They do not displace the centre.
And I want to be honest about the shape of what I just gave you. That is not four arguments. It is one argument with four supports: everything cheaper than HBM is an order of magnitude short on bandwidth, and the gap is physical.
But what if
There is a version of this thesis that says every buyer wants more HBM than it can get, and that version does not survive contact with the record.
At least one major custom-silicon programme has specced its current generation on the previous HBM generation across both of its configurations, and the read there is a deliberate yield and cost choice rather than a supply constraint. That buyer did not want more capable memory. It wanted cheaper memory. In this case Google TPU.
But look at why, because the reason is structural rather than a difference in taste. An internal programme is optimising a task-specific workload. It knows exactly what it will run, it controls both the model and the silicon, and it can spec precisely to that. So tailor made to their own workflow.
Nvidia is selling into a workload it has never seen. Frontier training, frontier inference, and sovereign build-outs like the Japan programme, where the buyer itself does not yet know what it will be running in three years. You cannot spec down for a workload you have not seen. You want your best product there.
What would change my mind… and then Jevon kicks in
I would rather write this down now than be asked for it later.
Everything I said about physics defends against hardware substitutes. It defends against nothing on the software side. The fastest-growing piece of memory demand is the KV cache, and its size is a direct function of the attention mechanism the models use. Today’s frontier models use full attention, which makes that cache grow with context length, and that growth is a big part of my demand case. If the frontier moves to sub-quadratic attention and holds quality on long-range work, the cache stops growing the way I assume. No new silicon required, and one model generation to propagate across the installed base. It has not happened. The hybrids that exist still show quality gaps at frontier scale, and history says a cheaper cache gets spent immediately on longer context. This is Jevons paradox. And all major names strengthen this direction in commentary. We will have much bigger models in the future.
The second is the design question. If the next architecture also lands below its own design intent, under looser supply, then the reduction was a preference and not a supply response, and my reading of this whole episode is wrong.
The third is efficiency at the fleet level. If the next generation makes each accelerator productive enough that buyers need fewer of them for the same tokens, then content per unit rises while unit counts fall and total bits go nowhere. Same answer as above. Every efficiency gain in this field so far has been reinvested in bigger models and longer context. And will be for the foreseable future. Which equals more HBM demand
None of those “tripwires” is shown yet. Until one does, what the buyers wanted is a higher HBM spec than what they were able to get.
And that is what makes the case for the memory suppliers, on the shortage, the pricing power and the earnings durability, a much longer-term one than the next 4 quarters potential earnings visibility. Their wishlist has not changed.
Final conclusion
Almost everyone modelling this memory cycle does the same thing. Supply on one side, order demand on the other, subtract, publish the deficit. Which of course gives you a pretty small shortage.
But order demand in a supply-constrained market is a censored number. It is simply what buyers ask for knowing what they can get. So it tells you a lot about how the shortage is being administered, and nothing about what happens when supply meets the wish list.
The wish list is the only thing that speaks to the years nobody can really see. And building it means you have to model and understand why the memory is needed at all. That is what I explained above. The bandwidth per token. The cache that keeps growing with context. And with that, the wall that widens every generation. Not what was allocated, but what would be bought.
And if you look at the wish list, your question for seeing beyond 2028 should be simple. Either the new capacity meets a satisfied market, or an unsatisfied one.
With all the points I have presented, and the behaviour of Nvidia, the choice to optimise for bandwidth so you can get the maximum out of your newer generation, which you then have to drop because you need more SOCAMM, for me that signs unsatisfied.
Thank you all for reading.
Markos


