This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
Existing CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
We already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.
Reducing everything to national aggregates provides no insight into the strong negative externalities, imposed on the immediate surrounding communities, of unregulated industrial facilities. That's literally the reason we have zoning laws in the first place,
There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)
this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
I cannot fathom tbe mindspace that leads to this being anythning but a simple observation. It might even make it all the way to obscure trivia or interesting observation, but paradox? Certainly not.
If you make the thing more accessible, more people are going to use it. If it consumes a resource, the use of that resource will increase in relation to the increased adoption.
Hydrogen engines use hydrogen. Making hydrogen engines cheaper will increase adoption. Increased adoption will increase consumption of hydrogen.
Like, who'd have ever thought "oh wow, we've gotten to the point that people can have a computer in their own home, surely electricity use will plummet." or "oh wow, more than 50% of the population can now feasibly purchase an internal combustion engine, surely fuel demand will plummet."
In the original context they'd decreased the cost and complexity of steam engines. Anyone who'd seen the amount of money people were making with the old steam engines would be clearly incentivized now that they have the same economic opportunity available for less capital up front. Therefore more steam engines, therefore more fuel demand. Who in their right mind would really be surprised that resource consumption went up when people could and did build more machines?
There is something counter-intuitive about the idea that making an engine that accomplishes the same amount of work with half the fuel will result in MORE fuel usage overall. You might expect it to be the same, or decline slightly, but the paradoxical element is that overall consumption goes up.
And you can say of course, it's so obvious, how could a dumdum not see that! But then there are lots of examples of things where increased efficiency results in less usage overall, because demand is inelastic, etc. Jevon's paradox doesn't apply to everything.
I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
> I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
Adding onto it, I feel as if this relates to some points regarding predictions of future in general. It is easier for us to look from the future to the past and think that it must be very obvious (as you also mention) but its also very counter-intuitive at the same time and there are just so so much nuance about basically any situation within it that its hard to really capture it all, and even then, be prepared for surprises and counter-intuitiveness.
I really like the Peter Drucker quote about it.
“The only thing we know about the future is that it will surprise us.” — Peter Drucker
and, “The future is fundamentally different from the past.”
— Frank Knight, Risk, Uncertainty and Profit (1921)
If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.
All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).
Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.
I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.
With the corollary that old hardware valuations will plummet with them.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens).
It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
It isn't about bunk takes or not. It is about the motivation behind doing something.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
> "I'm not so worried about the motivation behind it besides how it biases their takes."
Troi oi. Read what you wrote again.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.
The guy that was part of FTX, fired from openai for alleged theft, got billions in a hedge fund somehow then lost billions. Why are all these people so scummy? It is like voting Trump three times in a row.
They're very intelligent people who do very clever things at a young age, which draws the attention of very rich people who can exploit them to get richer, and no one tells the young person they're being exploited. They're heaped with praise and 'wealth' (millions, but crumbs compared to what they're making for other people), and told they're geniuses who can do no wrong, mostly by the media that happens to be owned by the rich.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
> I love how now you have to consider the possible s*** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods
I mean, previously you could have said something much the same except substitute "frat boys".
The semianalysis people have scripts which incorrectly count their numerators and Denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.
for example, their people think that GB300s are twice as fast as B300s, when really their benchmarks just incorrectly divide GB300 instances on azure by 8 instead of 4, since they don't read or verify any of the code that executes their benchmarks.
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
The "industry news and research" part of the AI industry feels very... suspect to me. My intuition is telling me that it's a bunch of people with influencer-y type social media skills and no actual credentials just grifting because there's so much money floating around.
When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
I don't agree. At the moment companies like NVIDIA take several times what it costs to make a chip. I think the fair split for the technology contribution is more like 50-50, maybe even 30-70 in favour of the manufacturer.
With competition we will actually have the fair split, whatever that is, and thus much lower prices.
At the moment, to have a big AI firm you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.
Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.
One could hope so, or it would be best for me if it were. I don't think I can count on that though. I don't think I can hope for anything better than 30-70 in my lifetime, and I think 10-90 almost requires you to buy the design firm.
I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.
And the brain is literally only producing electrochemical signals.
I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
> I don’t see how tokens can’t produce speech or track metabolic needs.
It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).
I am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.
I couldn't source the parameters from the screenshot or the nearby graphs, but from the nearby graphs you can see that at concurrency C=1, tokens/Joule (vertical axis) has totally plummeted, and obviously concurrent inference is much more efficient by batching. Divide the memory by the bandwidth and thats how long it takes to dump the full RAM contents through the chip. Do you want to do this once per token for a single conversation, or do you want to progress multiple conversations if you're going through all the weights anyway? The peak in the graphs is easily 22x more efficient than the low bottom right part on the graphs. So in batched mode its already more efficient than human speech.
I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.
Just checked wikipedia page... they have like 100K neurons only.. wtf!!! How can nature cramp all senses, including spatial, motion, life maintenance and general thinking into just 100K neurons?
A lot of the low level stuff is outsourced to biochemistry: the physical properties of proteins, and the various self-regulating biochemical systems of an animal, can "encode" a lot of intelligence, easing up on the computational demands of the brain proper.
Our brain is more like a MoE model activating only a few neurons for specific activities making it more efficient unlike a dense model activating all the params.
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
22 times more efficient is not like 22 times more powerful. It’s extremely harder to close the gap in power efficiency than in raw power. Simply because the power efficiency we see is the result of billion years evolution optimization.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
Well Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras?
I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq
Cerebras is targeting a distinctly different point on the cost/latency curve. They are betting that there will be some high value applications where latency and not just throughput is super important.
It is being used as part of a combined system. For example AWS is pushing for Trainium + WSE 3. The WSE 3 does the decode and the Trainium does the prefill.
Even in nvidia land rubin + LPU does a similar thing.
It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
> Well Sam Altman finally has built a moat against Chinese open weight AI
Hes got a press release.
The issue is, baking something to silicon requires discipline and about 2 years.
This isn't something you can just change your mind on halfway through. Trust me, I know. You need a clear vision of what you want to support, why and what bits of a chip you need to achieve that.
> and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies
That sounds quite like...nonsense?
Chip companies work on years-long cycles. They know today what are they launching 4-5 years from now.
I think the Chinese are going to be building their own chips aided with AI. DeepSeek, z.AI, MiniMax, Moonshot, etc, it's a race. The take off has really started.
How can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?
NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.
One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.
"How can OpenAI produce a LLM at scale more economically than Google, Amazon, and Microsoft which have experience in planet scale software and scale efficiencies unlike them".
One answer is they're quite good at poaching talent.
For datacenters specifically I've never understood what specifically consumes the water. Arent the water-cooling loops closed, so the water just cycles around and around and around?
The system which runs coolant over the chips can be closed but the part which uses an evaporative system to cool that is still open loop and vents water into the air, no?
Evaporated water is condensed, and in the process transfers its heat into another place that removes it. Another simple example is a pot of boiling water with a lid on it.
That's the problem, removing heat at a sufficient rate. Of course it can be done, but the most efficient way (in terms of cost) is just open loop evaporation.
I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be specific) needs to be colder to make snow.
Anyways, most air compression stations use water to cool the air, and then evaporative coolers to cool the water. The water is reused, but a ton (not sure of the percentage) is lost into the air. It's more or less a tower with a big fan on top, and water percolates down from the top, being cooled by the air as it goes. The water is then collected and pumped through the system again (but of course has to be always topped up to counteract what was lost to evaporation).
Anyways, long story short is it's most cost effective to just spray water into the air to cool water, as long as water is free/cheap.
Oh like one of those scenic cone towers like on a nuclear power plant?
Iiuc youre saying: its more cost-efficient to waste water using evaporative cooling so thats what we'll get, not that a closed loop with a heat exchanger is technically infeasible?
Closed loop is of course feasible, and closed loop with an evaporation step is also feasible, but open loop evaporative is the most efficient in terms of total energy usage (cost) and most cost effective in general, so that's what we mostly get right now.
Exactly. A car is a closed loop system. Coolant (which is water with some chemicals) cools the engine, then the hot coolant flows through a radiator which transfers that heat to the air, and then it just keeps cycling through. If there's no leaks then no coolant is lost.
The above method could be scaled up to data center levels, but is less efficient in terms of energy usage and cost. Cheaper/easier to just spray water into the air and get "free" cooling that way, as long as you have a source of free/cheap water to replenish what is lost to evaporation.
I'm not a nuclear engineer either, but I assume this is what is happening inside of the big iconic nuclear stacks. That's just cooling water being evaporated off to keep it cool.
At the datacenter side, it depends on the method of cooling. You can chill the air or the chips (or both), doesn't matter, you still need to cool, and that still needs water. The question is, where is the water used?
- If they use either evaporative cooling or a liquid-cooled heat exchanger, that uses tons of water consistently. This requires less energy (it's mostly passive) so you use more water.
- If they use closed-loop water cooling and/or heat pumps/electric chillers, that uses much less water - at the DC. But it does require more energy to circulate the water, run fans, etc. If you are using more energy, where is the energy coming from? It's coming from power plants, which require... you guessed it... more water (e.g. thermoelectric, hydroelectric, geothermal, concentrated solar). They need water in order to generate the power, and lots of it. Coal, natural gas, and nuclear, all use steam to generate energy. Nuclear also uses water to cool the reactor. And water is used extensively to extract coal, oil, and natural gas. Geothermal uses water in the ground. Concentrated solar uses solar to heat water to make steam.
You can't not use a ton of water in one fashion or another. It just depends what method, and on what end the water is used. And the crazy thing is, most new datacenters are being built in places with extremely little water. They're doing it Guess how that's gonna work out as the planet gets hotter?
I don't know why I got downvoted to hell for stating facts every datacenter architect knows. HN be HN'in.
If the populist campaign is to Make Affordable DRAM Again, then it's not a terrible solution.
The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?
If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.
I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts.
Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower.
It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
ill give you that the way we talk about this ppl seem to think wed do this tomorrow, but in the 70s a kb of ram took an entire rack and tons of power also. Its seems equally plausible that we could go into a cycle of iterative refinement of baked model hardware that would end up in "personal ai" just like we got to personal computing.
Baked model hardware is not the next step in the chain here, in the next 2-3 years we might hopefully see some HBF (high bandwidth flash) hardware to try and get at same memory bandwidth today at a somewhat lower price point and much lower power dissipation.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.
> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
Taalas needed a giant chip (6nm) for an 8B model.
At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
Nope. But we are hitting some pretty impressive levels with 128B models.
The other thing is, a lot of the time, model performance is improved with more 'thinking' time.
The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol?
Yeah, but then there's the size of KV cache needing to be read through that HBM interface for each token, putting a hard limit on the tok/s based on the memory bandwidth.
On some models a large context can be a notable proportion of the size of the weights themselves.
For example, qwen 3.8 27b uses ~64kb/token for the kv cache - so for a 256k token context that's ~16gb of the kv cache for a ~54gb model (assuming 2 bytes-per-param/f16 for both).
So if the current non-baked-in chip is already memory bandwidth bound, as is often the case for current hardware and models, and the "only KV cache in HBM" chip has the same total memory bandwidth, it can only ever be (54/16)=~3.4x faster for the baked in-silicon model.
> Taalas needed a giant chip (6nm) for an 8B model.
You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon (nope, done all the time) or with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.
"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal
That's only half the problem. OpenAI is contractually obligated, if you will, to believe that models will continue improving at an impressive rate for the foreseeable future (otherwise their valuation makes no sense).
If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.)
And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.
Assume a $800 billion valuation. $100 billion ad network. $30 billion op income. 26x price to op income ratio. It's right there for them to grab, or someone else to grab.
Their valuation does make sense if you believe: 1) they can retain a massive user base and 2) a massive user base can be monetized. Future value is almost always pulled forward these days for high growth tech companies.
An LLM the size of Google search in users is even more valuable than Google search. The ad market for LLMs will be even larger than search was (no matter what HN prefers).
The monetization part is the easier part. Silicon Valley understands extraordinarily well how to build ad networks. If OpenAI maintain their gigantic user base, a $100 billion ad network is a given bolt-on. They'd have to screw that up in an epic way to not get there.
Facebook - Insta - WhatsApp is an absolute dogshit tandem with a gigantic user base. $228 billion in ad sales and still expanding 10% per year.
Google knows this is what's happening, that's why they don't care about chasing Anthropic very much. They're busy completely remaking how their core search business works.
If a strange and quirky architecture can increase their margins, they could start becoming so wildly profitable that they don't care about their investor driven valuation anymore.
The whole point is that it's supposed to be more efficient. But models are also still getting absurdly more efficient every year. 18 months is a long time right now (and 18 is only time to tape out, not operational in data centers).
Even if they stopped get more efficient right now, you would also not be able to train them against new tools/harnesses or knowledge.
The good enough level isn't ever arriving. We're in the first or second inning for LLMs. They will rapidly subdivide in complexity, they will not stagnate in the next decade.
Beyond the model, when would you freeze processor performance, such that it was good enough? Because that's exactly what freezing on Talaas is premised around.
The semiconductor technology will also continue to improve. You lose twice. Talaas is one of the dumbest ideas I've seen in semiconductors in decades.
Good enough will hit when the tech stops advancing quickly. You could have a "good enough" model but in 2 years if the general purpose chip can run it just as fast, there is no point having the single purpose one.
There are other options. I worked for a startup called NVXL and we were programming DNNs into FPGAs using OpenCL, on custom boards we built to plug into NVME. It worked great, but it wasn't fast enough at the time to compete with Nvidia, or even Intel AVX512. Ultimately the company failed, but maybe some hybrid like that could work for LLMs? I haven't been up to date on how DNNs and LLMs look like under the hood these days, but there's got to be someone doing something similar.
Closer to 2 weeks, as long as the new model fits: your model bits are entirely in mask rom, so you can change them with a metal only ECO that only touches two metal layers. You probably don't even need to run any timing analysis, etc. since all of the changes are going to be isolated in a gigantic square that's isolated from everything else.
I think it would make sense to upload the weights like a firmware into the chip rather than baking the weights itself directly onto the chip. If there is an option of easily updating the firmware from time to time it would work well.
> For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
but you trade updatability, which I don't think is worth it yet.
Maybe! (1) Would SOL level intelligence be useful 3 years from now? 5 years? (2) would dedicated chips be the most affordable way to run this model in 3-5 years?
I suspect the answer to both of these questions is yes right now, but I agree it’s borderline.
LoRAs are a parameter efficient way to update models. The fine tuning stopgap only has to work long enough to extend the lifespan of a model until the next chip comes out.
Lol i'm surprised this hasn't happened already. Minecraft is a great platform for scripting as long as you think every single other platform is too fast and convenient.
I’m not sure that’s viable because of the time it takes to spin a bunch of chips and then get them installed in datacenters. Hard to imagine that would take less than half a year at the most optimistic. That’s a lot of added latency for a business moving this fast.
Yeah this feels like Altman not knowing engineering well enough to realize where the focus should be.
Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement.
My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.
AI is bigger than just LLMs. The real place model-hardwired chips will find value is in robotics. There you need very local, low latency inferencing with relatively stable models to handle motor control and navigation tasks. Higher level reasoning can be delegated to LLMs and run asynchronously.
I don't think any underestimates how powerful it would be. It's the social aspect that is difficult. I wouldn't want to be sat in a pub with you and your always-on notetaker.
> Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
This often happened, though not always.
An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs.
Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued.
It didn't happen for the original fixed function graphics though, CPUs kept eating their lunch every year or so with fun rendering techniques. We still use some of these techniques for 2d graphics.
Only when they became more general with shaders, and then added support for GPGPU, did it truly take off.
I think this generality is the lesson here, not the fact that GPUs are not CPUs.
These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be
Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though.
I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb.
Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.
EAX was very powerful in its heyday, but it has died because of a thousand cuts.
First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features.
Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity.
Then Microsoft changed the Windows driver model, cutting the driver's direct access to the card. All of the timing sensitive effects were gone in an instant. I remember installing the new drivers and getting literally nothing. Sound Blaster was the only card with an hardware mixer, and Microsoft didn't feel like enabling them. Mixing at the DirectX layer killed the cards.
Soundblaster's very closed stance didn't help them either. None of the cards after Audigy2 worked with Linux when I had my desktop system.
After my Audigy2ZS, I moved to Asus Xonar D2X. Its positional audio capabilities were nice, but I mostly bought it for its Linux support and sound quality, and that was top notch in that regards.
Then sound cards became commodity. Everybody stopped making good cards. Musicians moved to audio interfaces, audiophiles moved to DACs.
Just looked to the SoundBlaster website. Internal cards are very limited. One DAC, one DTS enabled 7.1 sound card for PC cinema systems, three game oriented lower end cards, nothing else.
As you mention the 90's you're probably referring to their embrace, extend, extinguish strategy. This situation with the sound card hardware mixing was quite different and a lot later.
Windows Vista moved away from kernel mode drivers where possible to help address the issue of BSODs which were not uncommon on Windows XP. It turns out that Microsoft was incorrectly getting the blame for BSODs when it was actually the fault of buggy video and sound card drivers.
While Vista was regarded as a "bad" Windows, it laid most of the foundations which allowed Windows 7 to be regarded as "really good" (by Windows standards).
From a hardware perspective, by the time Windows 7 arrived pretty much all drivers had been updated for Vista and had their kinks worked out (mostly, I would still have my NVidia drivers crash on occasion, but they wouldn't cause a BSOD anymore to being forced to run in user mode). Windows 7 also did optimizations to make things less resource intensive and more RAM in PCs had become the norm which helped too.
On the UAC front Windows 7 was also much better, they calibrated the UAC prompts to come up less often, but what also happened is a lot of the 3rd party software which was needlessly requiring admin rights for no good reason (except that it was badly written), had finally been fixed by the time Windows 7 was released.
Creative Labs also turned itself into a brand I actively avoid for various reasons.
In early 2000s for example, friend with something like a Sound Blaster Live, but couldn't use it anymore as they lost the drivers, their website only had downloads for driver updates, requiring you to still have a driver CD, so no more CD meant no way to get the drivers.
They had not done very much meaningful stuff since the release of EAX and used patents to prevent anyone else from competing with them. Onboard sound cards were by and large indistinguishable from a quality perspective as Creative Labs ones, but cheaper which Creative Labs combated mostly with lawyers as opposed to upping their game.
Some guy after being frustrated with a long-standing bug in drivers for their Creative Labs sound card, dug into the binaries and made a fix for it. During which they also discovered that you could simply flip a switch in the driver to unlock features only meant to be available on more expensive hardware. Creative Labs of course went straight to lawyers to shut them down.
My brother bought the Creative Labs WoW headset which would have its mic get progressively softer until he would leave and re-join the call/voice chat room. They never released a driver update to fix this.
By the time that Microsoft announced no more "hardware acceleration" for sound cards, I had zero sympathy for Creative Labs, I was already convinced that they made pretty shoddy hardware/software and were mostly riding on their reputation from the 90s and some patents they managed to get.
Eeeeeh idk about the Sound Blaster comparison. Creative earned their place in the early-mid 90s solely because they were the one company making a sound card with drivers that actually worked properly.
It wasn't really the cool reverb effects or wave tables, though those were a nice bonus. It was just "I can tell my computer to make sound and it actually makes sound without days of troubleshooting."
Granted, similar things could be said about 3dfx. It's was a 3D card with drivers that actually worked.
I recently heard, for the first time, what Space Quest sounded like with a Roland board attached. It was mindblowingly amazing. And all most anyone ever experienced was bleep, bloop.
This couldn't have been easy. The team at OpenAI has worked a miracle.
For example, Meta and Microsoft’s AI ASIC programs not getting off the ground despite being at it for much longer shows that cost is only one part of the equation.
I bet cost is of no issue with the capx where it is at. It is almost certainly organizational. Meta throws money at every problem and it never seems to workout for them.
There is drastically more power and profit in the software ultimately.
Apple is in the software first, the hardware second. Everyone at Apple has been trained to understand this for decades, and Jobs pointed it out endlessly. Apple's real moat is software (services, iOS, experience, MacOS).
Windows, Office, Azure, et al. Microsoft accumulated approximately one zillion dollars in profit on the back of software. It's a vastly superior business to anything hardware has traditionally seen. Nvidia is the first true juggernaut hardware profit machine, and the AI boom in extended hardware (RAM, storage) will prove temporary (even if there is a feast during that time). Microsoft's advantage and moat was Windows-Office for decades. It was a far better business than Intel's chip biz.
Google is a software company first. Every aspect of what made them and maintains them is software first, hardware second. They're a $400 billion software company. Their ad machine is software. Search is software.
Facebook is software. Instagram is software. WhatsApp is software. A $200 billion software company. They're not selling hardware, they're selling ads via software, they're monetizing users that use their software.
AWS is at least half software as an entity in terms of complexity, competitive advantage, et al. That's a two trillion dollar business.
LLMs can run successfully with various hardware approaches. The software is the value at the end of this, regardless of the hardware under it. The sole exception so far that may be sustainable is Nvidia, and we'll see if the bottom falls out from under that margin monster (China, specialized AI chips, whatever it happens to be that cuts under them massively).
Hardware always gets its margin squeezed eventually because it's a manufactured good (with inventory, fabs, etc). Software is hyper margin by default, you have to layer a lot of garbage on top of it to kill the margin. Nvidia is 33 years old, they have had a rich business for three years, that's it.
The AI boom is the sole reason anything in hardware has looked great in the past 20 years. Check the margins & op income for the top 20 hardware companies, from TI to AMD to Intel to Nvidia to Micron to Sandisk to Samsung to TSMC to ASML, prior to the AI boom of the past couple years. It won't last indefinitely. And after the return to a more normal environment happens, the hyper margins in software will persist.
> Will be interesting to see if inference chips are here to stay
To me, the efficiency gains of inference chips are so significant that they are certainly here to stay — barring a revolution of sorts that leads to a world devoid of AI as we know it.
I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
My point is, that if companies develop ASICs for LLM work, then GPUs will stop being the go-to computation device for these workloads, and that will mean, hopefully, that their architectures will stop being warped so as to cater to LLM work.
I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.
So, I started digging while waiting for various day-job inference calls to return, ha.
Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]
In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]
Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.
The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]
And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]
I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.
The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.
Not publicly acknowledging how misallocation of capital may be happening today shows he is corrupt. He’s not that dumb to not know it’s a major risk to the whole story, and is certainly financially incentivised to write as he does.
If you research the origins of Dwarkesh, even more conspiracy level points emerge. I do not think even in the handful videos post Leopold fund collapse he has addressed it in any ways. That is the point of so called observers, they pretend to be impartial but everyone is a hustler in some way.
everyone's silicon beats everyone else's benchmarks until it has to run someone's actual production workload. the real test is six months of your own inference traffic, not a vendor's chart.
The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.
As opposed to closed-source models? Benchmarks for GPT Sol wouldn’t be particularly meaningful, as no one else can run the benchmark, and we don’t know what the exact model specs are.
Picking the best open source models is really the best they can do.
> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.
How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?
In the slides on twitter you can see Jalapeno CAN do speculative decoding. In fact they explicitly mention how compute is disaggregated 3 ways now: prefill, predict, decode, and how a huge Jalapeno advantage is that it uses dark sillicon to switch between these without having to move the KV cache which remains local.
Then the article contradicts the slides because it states that OpenAI choose not to disaggregate prefill and decode. Idk you men with "predict"---conventional LLM serving comprises only two phases.
276 comments
[ 0.26 ms ] story [ 14.2 ms ] threadToken prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
If you make the thing more accessible, more people are going to use it. If it consumes a resource, the use of that resource will increase in relation to the increased adoption.
Hydrogen engines use hydrogen. Making hydrogen engines cheaper will increase adoption. Increased adoption will increase consumption of hydrogen.
Like, who'd have ever thought "oh wow, we've gotten to the point that people can have a computer in their own home, surely electricity use will plummet." or "oh wow, more than 50% of the population can now feasibly purchase an internal combustion engine, surely fuel demand will plummet."
In the original context they'd decreased the cost and complexity of steam engines. Anyone who'd seen the amount of money people were making with the old steam engines would be clearly incentivized now that they have the same economic opportunity available for less capital up front. Therefore more steam engines, therefore more fuel demand. Who in their right mind would really be surprised that resource consumption went up when people could and did build more machines?
And you can say of course, it's so obvious, how could a dumdum not see that! But then there are lots of examples of things where increased efficiency results in less usage overall, because demand is inelastic, etc. Jevon's paradox doesn't apply to everything.
I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
Adding onto it, I feel as if this relates to some points regarding predictions of future in general. It is easier for us to look from the future to the past and think that it must be very obvious (as you also mention) but its also very counter-intuitive at the same time and there are just so so much nuance about basically any situation within it that its hard to really capture it all, and even then, be prepared for surprises and counter-intuitiveness.
I really like the Peter Drucker quote about it.
“The only thing we know about the future is that it will surprise us.” — Peter Drucker
and, “The future is fundamentally different from the past.” — Frank Knight, Risk, Uncertainty and Profit (1921)
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens).
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
And the cycle continues.
https://fortune.com/2026/03/16/peter-thiel-giving-pledge-bil...
I mean, previously you could have said something much the same except substitute "frat boys".
Yeah, those guys aren't biased at all.
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
lol. lmao even.
Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
With competition we will actually have the fair split, whatever that is, and thus much lower prices.
At the moment, to have a big AI firm you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.
Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.
I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
Take that, Jalapeno!
Productivity is not the only reason to let these meatbags burn energy.
Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Patience pays.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq
Even in nvidia land rubin + LPU does a similar thing.
It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
Hes got a press release.
The issue is, baking something to silicon requires discipline and about 2 years.
This isn't something you can just change your mind on halfway through. Trust me, I know. You need a clear vision of what you want to support, why and what bits of a chip you need to achieve that.
Man, if only someone made like, chips that could lots of different calculations all at the same time!
That sounds quite like...nonsense?
Chip companies work on years-long cycles. They know today what are they launching 4-5 years from now.
One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.
One answer is they're quite good at poaching talent.
Example of such a system being used specifically for datacenters: https://blog.vantage-dc.com/2026/04/22/cooling-without-the-d...
Evaporated water is condensed, and in the process transfers its heat into another place that removes it. Another simple example is a pot of boiling water with a lid on it.
I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be specific) needs to be colder to make snow.
Anyways, most air compression stations use water to cool the air, and then evaporative coolers to cool the water. The water is reused, but a ton (not sure of the percentage) is lost into the air. It's more or less a tower with a big fan on top, and water percolates down from the top, being cooled by the air as it goes. The water is then collected and pumped through the system again (but of course has to be always topped up to counteract what was lost to evaporation).
Anyways, long story short is it's most cost effective to just spray water into the air to cool water, as long as water is free/cheap.
Iiuc youre saying: its more cost-efficient to waste water using evaporative cooling so thats what we'll get, not that a closed loop with a heat exchanger is technically infeasible?
The above method could be scaled up to data center levels, but is less efficient in terms of energy usage and cost. Cheaper/easier to just spray water into the air and get "free" cooling that way, as long as you have a source of free/cheap water to replenish what is lost to evaporation.
I'm not a nuclear engineer either, but I assume this is what is happening inside of the big iconic nuclear stacks. That's just cooling water being evaporated off to keep it cool.
- If they use either evaporative cooling or a liquid-cooled heat exchanger, that uses tons of water consistently. This requires less energy (it's mostly passive) so you use more water.
- If they use closed-loop water cooling and/or heat pumps/electric chillers, that uses much less water - at the DC. But it does require more energy to circulate the water, run fans, etc. If you are using more energy, where is the energy coming from? It's coming from power plants, which require... you guessed it... more water (e.g. thermoelectric, hydroelectric, geothermal, concentrated solar). They need water in order to generate the power, and lots of it. Coal, natural gas, and nuclear, all use steam to generate energy. Nuclear also uses water to cool the reactor. And water is used extensively to extract coal, oil, and natural gas. Geothermal uses water in the ground. Concentrated solar uses solar to heat water to make steam.
You can't not use a ton of water in one fashion or another. It just depends what method, and on what end the water is used. And the crazy thing is, most new datacenters are being built in places with extremely little water. They're doing it Guess how that's gonna work out as the planet gets hotter?
I don't know why I got downvoted to hell for stating facts every datacenter architect knows. HN be HN'in.
Or at least Nvidia GPUs will become slightly cheaper for regular consumers again
There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.
The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?
If the chips weren't this compelling they would have something different to announce.
These are paperclip maximizers who just happen to wear human skin - there is no underlying premise nor ideological goal.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower.
It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.
Taalas needed a giant chip (6nm) for an 8B model.
At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
The other thing is, a lot of the time, model performance is improved with more 'thinking' time.
The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol?
The more problem like these they solve the more they will look like GPU.
On some models a large context can be a notable proportion of the size of the weights themselves.
For example, qwen 3.8 27b uses ~64kb/token for the kv cache - so for a 256k token context that's ~16gb of the kv cache for a ~54gb model (assuming 2 bytes-per-param/f16 for both).
So if the current non-baked-in chip is already memory bandwidth bound, as is often the case for current hardware and models, and the "only KV cache in HBM" chip has the same total memory bandwidth, it can only ever be (54/16)=~3.4x faster for the baked in-silicon model.
You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon (nope, done all the time) or with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.
It is.
As I said, they could have shrunk it with a smaller process node, but that's not at all close to what would be required for a GPT Sol size model.
If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.)
And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.
Their valuation does make sense if you believe: 1) they can retain a massive user base and 2) a massive user base can be monetized. Future value is almost always pulled forward these days for high growth tech companies.
An LLM the size of Google search in users is even more valuable than Google search. The ad market for LLMs will be even larger than search was (no matter what HN prefers).
The monetization part is the easier part. Silicon Valley understands extraordinarily well how to build ad networks. If OpenAI maintain their gigantic user base, a $100 billion ad network is a given bolt-on. They'd have to screw that up in an epic way to not get there.
Facebook - Insta - WhatsApp is an absolute dogshit tandem with a gigantic user base. $228 billion in ad sales and still expanding 10% per year.
Google knows this is what's happening, that's why they don't care about chasing Anthropic very much. They're busy completely remaking how their core search business works.
Even if they stopped get more efficient right now, you would also not be able to train them against new tools/harnesses or knowledge.
Beyond the model, when would you freeze processor performance, such that it was good enough? Because that's exactly what freezing on Talaas is premised around.
The semiconductor technology will also continue to improve. You lose twice. Talaas is one of the dumbest ideas I've seen in semiconductors in decades.
but you trade updatability, which I don't think is worth it yet.
I suspect the answer to both of these questions is yes right now, but I agree it’s borderline.
1. https://matx.com/
2. https://www.d-matrix.ai/
3. https://www.etched.com/
4. https://www.positron.ai/
5. https://hyperaccel.ai/
6. https://axelera.ai/
7. https://www.enchargeai.com/
8. https://furiosa.ai/
Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement.
My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.
I think people underestimate how much of a revolution having an always-on, privacy-preserving personal notetaker / secretary would be.
Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
This often happened, though not always.
An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs.
Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued.
---
[1] https://en.wikipedia.org/w/index.php?title=Larrabee_(microar...
Only when they became more general with shaders, and then added support for GPGPU, did it truly take off.
I think this generality is the lesson here, not the fact that GPUs are not CPUs.
Maybe the money will still flow into this industry after all
I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb.
Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.
In fairness on board (depending on the board but on the whole) is pretty good.
First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features.
Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity.
Then Microsoft changed the Windows driver model, cutting the driver's direct access to the card. All of the timing sensitive effects were gone in an instant. I remember installing the new drivers and getting literally nothing. Sound Blaster was the only card with an hardware mixer, and Microsoft didn't feel like enabling them. Mixing at the DirectX layer killed the cards.
Soundblaster's very closed stance didn't help them either. None of the cards after Audigy2 worked with Linux when I had my desktop system.
After my Audigy2ZS, I moved to Asus Xonar D2X. Its positional audio capabilities were nice, but I mostly bought it for its Linux support and sound quality, and that was top notch in that regards.
Then sound cards became commodity. Everybody stopped making good cards. Musicians moved to audio interfaces, audiophiles moved to DACs.
Just looked to the SoundBlaster website. Internal cards are very limited. One DAC, one DTS enabled 7.1 sound card for PC cinema systems, three game oriented lower end cards, nothing else.
Windows Vista moved away from kernel mode drivers where possible to help address the issue of BSODs which were not uncommon on Windows XP. It turns out that Microsoft was incorrectly getting the blame for BSODs when it was actually the fault of buggy video and sound card drivers.
While Vista was regarded as a "bad" Windows, it laid most of the foundations which allowed Windows 7 to be regarded as "really good" (by Windows standards).
From a hardware perspective, by the time Windows 7 arrived pretty much all drivers had been updated for Vista and had their kinks worked out (mostly, I would still have my NVidia drivers crash on occasion, but they wouldn't cause a BSOD anymore to being forced to run in user mode). Windows 7 also did optimizations to make things less resource intensive and more RAM in PCs had become the norm which helped too.
On the UAC front Windows 7 was also much better, they calibrated the UAC prompts to come up less often, but what also happened is a lot of the 3rd party software which was needlessly requiring admin rights for no good reason (except that it was badly written), had finally been fixed by the time Windows 7 was released.
In early 2000s for example, friend with something like a Sound Blaster Live, but couldn't use it anymore as they lost the drivers, their website only had downloads for driver updates, requiring you to still have a driver CD, so no more CD meant no way to get the drivers.
They had not done very much meaningful stuff since the release of EAX and used patents to prevent anyone else from competing with them. Onboard sound cards were by and large indistinguishable from a quality perspective as Creative Labs ones, but cheaper which Creative Labs combated mostly with lawyers as opposed to upping their game.
Some guy after being frustrated with a long-standing bug in drivers for their Creative Labs sound card, dug into the binaries and made a fix for it. During which they also discovered that you could simply flip a switch in the driver to unlock features only meant to be available on more expensive hardware. Creative Labs of course went straight to lawyers to shut them down.
My brother bought the Creative Labs WoW headset which would have its mic get progressively softer until he would leave and re-join the call/voice chat room. They never released a driver update to fix this.
By the time that Microsoft announced no more "hardware acceleration" for sound cards, I had zero sympathy for Creative Labs, I was already convinced that they made pretty shoddy hardware/software and were mostly riding on their reputation from the 90s and some patents they managed to get.
It wasn't really the cool reverb effects or wave tables, though those were a nice bonus. It was just "I can tell my computer to make sound and it actually makes sound without days of troubleshooting."
Granted, similar things could be said about 3dfx. It's was a 3D card with drivers that actually worked.
This couldn't have been easy. The team at OpenAI has worked a miracle.
Apple is in the software first, the hardware second. Everyone at Apple has been trained to understand this for decades, and Jobs pointed it out endlessly. Apple's real moat is software (services, iOS, experience, MacOS).
Windows, Office, Azure, et al. Microsoft accumulated approximately one zillion dollars in profit on the back of software. It's a vastly superior business to anything hardware has traditionally seen. Nvidia is the first true juggernaut hardware profit machine, and the AI boom in extended hardware (RAM, storage) will prove temporary (even if there is a feast during that time). Microsoft's advantage and moat was Windows-Office for decades. It was a far better business than Intel's chip biz.
Google is a software company first. Every aspect of what made them and maintains them is software first, hardware second. They're a $400 billion software company. Their ad machine is software. Search is software.
Facebook is software. Instagram is software. WhatsApp is software. A $200 billion software company. They're not selling hardware, they're selling ads via software, they're monetizing users that use their software.
AWS is at least half software as an entity in terms of complexity, competitive advantage, et al. That's a two trillion dollar business.
LLMs can run successfully with various hardware approaches. The software is the value at the end of this, regardless of the hardware under it. The sole exception so far that may be sustainable is Nvidia, and we'll see if the bottom falls out from under that margin monster (China, specialized AI chips, whatever it happens to be that cuts under them massively).
Hardware always gets its margin squeezed eventually because it's a manufactured good (with inventory, fabs, etc). Software is hyper margin by default, you have to layer a lot of garbage on top of it to kill the margin. Nvidia is 33 years old, they have had a rich business for three years, that's it.
The AI boom is the sole reason anything in hardware has looked great in the past 20 years. Check the margins & op income for the top 20 hardware companies, from TI to AMD to Intel to Nvidia to Micron to Sandisk to Samsung to TSMC to ASML, prior to the AI boom of the past couple years. It won't last indefinitely. And after the return to a more normal environment happens, the hyper margins in software will persist.
To me, the efficiency gains of inference chips are so significant that they are certainly here to stay — barring a revolution of sorts that leads to a world devoid of AI as we know it.
I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw
I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.
So, I started digging while waiting for various day-job inference calls to return, ha.
Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]
In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]
Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.
The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]
And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]
I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.
The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.
[1] https://www.dwarkesh.com/p/dylan-jon
[2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w...
[3] https://www.theinformation.com/articles/dylan-patel-semianal...
[4] https://www.latent.space/p/dylanpatel-cooking
[5] https://www.dwarkesh.com/p/dylan-patel-3
[6] https://www.theinformation.com/briefings/exclusive-semianaly...
[7] https://news.ycombinator.com/item?id=31065646
Picking the best open source models is really the best they can do.
> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.
How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?
> it states that OpenAI choose not to disaggregate prefill and decode
They disaggregate INSIDE the chip, not by having separate machines for the 3 phases. the slides:
https://x.com/beffjezos/status/2092416851737518190