19 comments

[ 0.19 ms ] story [ 50.4 ms ] thread
At least one graph showing global total datacenter compute over time and projected to come online would have been helpful.
[dead]
If instead of us training the models to be human utility maximizers, the models had figured out how to train us to maximize their own utility, could we tell the difference?
What on earth is that H100 price trend graph. The equally spaced x-axis points are 2x6 months, 6x3 months, 9x1 month. The whole visualisation of the trend is ruined on the back of that.

There are liars, damned liars, and people who play silly buggers with scales.

It will be fascinating to watch this play out. Particularly if AI-compute satellite constellations become a reality - We're trending towards a matrioshka brain and I'll be happy if I live to see the beginnings of that future
GPUs aren’t scarce. There is just a lot of demand, so prices have gone up.

Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.

Yes but how much of that compute shortage is from demand that is subsidized? We’ve seen companies like Uber drastically cut how much they are willing to spend on AI because they are paying actual usage costs, while at the same time OpenAI and Anthropic increase the limits on their fixed cost plans for individuals meaning people not paying usage costs are using it more and more… doesn’t this show that the compute shortage is because OpenAI and Anthropic are paying for it, not their customers? And the moment OpenAI and Anthropic stop paying for it, demand will collapse.
Meh. That’s like saying we have a parking shortage in cities. No we have more car journies than we need.

We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park)

Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam

It’s all choices

We just are making bad ones mostly

The choices now are would you like ads or more ads. Shall I waste compute seeing if the web page you clicked on can be summarised ?

The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out.

It might take the US stock market with it …

Isn’t the compute shortage temporary?

I don’t believe in 5-10yrs we’ll be in a shortage anymore

We're in such a bull market, even shortages grow.
When gas prices went up in the 70s because of fuel shortage, smaller (more fuel efficient) cars became more popular to use less fuel to do the same thing. I wonder if the same will happen with compute, by making software more efficient, and do the same thing with less compute.
I feel this post is blind to many of the secondary side effects of this "shortage". The rapid increase in prices and delivery times is having deleterious effects on all things tech - everything from phones to smart appliances and all sorts of gadgets has moved into unreachable price levels.

Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range.

So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.

Four things are currently correct:

1. There is a huge demand for compute, specifically GPU compute

2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills

3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying.

4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others).

All while advances in open weight models are making it appear that the major labs truly have no model moat.

Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.

> Compute Is Becoming a Competitive Moat

No, it is not a moat.

The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over.

> At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027.

MS bought more GPU's than they had rack space for: https://www.datacenterdynamics.com/en/news/microsoft-has-ai-...

Open AI bought out the memory: https://x.com/kwharrison13/status/2029248559388746168 but they have no means to consume anywhere close to their order.

Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them.

Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows).

At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.

To preface, this article presents valid points that I agree with. I would have also liked to know their findings on computational memory (CXL) and not just HBM or DDR. Also, a theoretical background on the Von Neumann and Harvard architectures would be helpful.

Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.

This article is about silicon. But the other shortage that matters is meat-compute.

AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.

But.. "good design taste" has no compiler.

Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.

In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.

While building Fabs takes time you can only imagine that the amount of capital flooding in means there will be an explosion in infrastructure supply over the next decade and eventually prices will crater