Realistically, they're worse than a central bank, because they can't exactly expand supply monotonically like a normal central bank. Nor do they realistically control rates.
The Japanese economy and yen carry trade is a close second.Rising oul prices and reduced output due to the conflicts in the Middle East could filter through to increasing yields on Japanese debt. In turn, the yen interventions have to continue to keep it lower than 160, which seems to be the psychological barrier for the yen carry trade.
I wonder when they will give up on the gaming market because that could take down several publishers and developers. I really don't think it's an if question but a when because it almost feels like an afterthought at this point (they removed the standalone gaming revenue report from the financial reports this summer). Also I don't think AMD and Intel is capable "to step in" to replace them.
That seems a bit overzealous, most consoles[0] run on AMD chips today.
Both PS 5 and Xbox are based on AMD APUs and both serve the AAA market quite well. GTA 6, Assassin's Creed and CoD are probably good enough indicators that the performance is enough, even if there is always room for more (as PC ports show). The PC market will also probably be fine even if stagnation in perfomance gains has been creeping in for a few years now.
[0]: except Nintendo which relies on NVIDIA although their APU there focuses more on efficiency than top performance.
NVidia is limited by the number of chips they can produce.
If you can fab 1000 chips, and can sell some for $500 and some for $80000 what are you going to do?
The game GPU is at once profitable, but causes them to give up far more profits than they're gaining from it.
They're maintaining the game market to have multiple markets and not go all in, but it's strategic hedging at this point. When NVidia makes a gaming GPU instead of a data center GPU they are leaving money on the table in the short term since they're constrained at the fab level.
This almost seems like we need a strategic reserve for semiconductors- kinda like we have it for food, to prevent suppliers from throwing it away when they are suddenly able to sell something much more lucrative.
The CHIPS Act in the US did set aside a reserve for older processes used for automobile, defense, and industrial semiconductors - but that's not gaming GPUs that used the previous process node.
Sure, semiconductors for important industrial uses are critical & I agree gaming GPUs are much less critical. Still RAM & storage multiplying in price over such a short ammount of time also can't be good for the economy & society, so some sort of strategtic reserve could help there as well.
Nvidia is limited by the number of engineers they have. It might turn out that the AI market is so lucrative that it’s best to reallocate their gaming-focused engineers to AI.
Apple TV is at least a growth market for them, whereas gaming is sort of capped and clearly a tiny piece of nvidia’s revenue atm.
They may allocate different number of resources every year based on market conditions but they'll never give up on gaming, that would be extremely silly.
> I really don't think it's an if question but a when
Nvidia still will ship gaming products. The upcoming RTX Spark laptop APUs are still gaming-capable - we also have Blackwell gaming GPUs and the Nvidia-powered Nintendo Switch 2.
People echoed this sentiment during the crypto mining crunch, and we still got gaming hardware designs after that blew over. One of CUDA's core value props is the consumer market, and Nvidia probably won't surrender it unless hardware becomes unreasonably scarce.
I don't think crypto is comparable. They barely made some dedicated crypto GPUs that market was always fickle due to ASCIs.
Look at the nvidia revenue breakdown chart, the AI boom looks quite different.
Maybe this is silly of me, but maybe gaming would enjoy an era of hardware upgrades being rather unviable so the focus turns to optimization and aesthetic.
A retraction from the gaming market wouldn't necessarily mean they'd still even produce the current crop of products. I could (and I would argue would) involve a complete shuttering of the GeForce brand, halting current production. On the assumption that Intel and AMD would follow, that wouldn't be an end of upgrades, but an end to the market.
It’s not silly. games use far more hardware than they really need. It also pushes out release dates of aggressive console schedules like ps6 because even if it’s a massive upgrade and you have IP locked into your console, no one is going to pay $4000 to play a game like wolverine.
It will be interesting times but I don't think anyone will enjoy a plateau because it's expensive. Even replacing existing hardware is expensive now. Hardly enjoyabe.
The thing is really that even entry level gaming is getting expensive. I was considering building a pc and everything is expensive, I was going to replace the disk of my ps4 thag I bought used for 100$ and the 1tb ssd costs 150$
This is basically already the case in the flourishing indie space, but for different reasons.
The only developers chasing cutting-edge hardware features are the AAA studios and I honestly think a lot of them would just die before successfully relearning how to make games that play well instead of just look pretty.
It's still a 15B market for Nvidia, it's not nothing. But I'm curious how they're going to turn Rubin into a gaming GPU. I think at this point consumer GPU upgrades are going to be AI-related upgrades that happen to also help raster capabilities. Blackwell was already kind of a dud on performance uplift from Ada beyond the new LLM features.
Hermes still makes saddles that are somewhat reasonably priced compared to their handbags.
It's still profitable, it's their original raison d'etre, and there's no real reason for them to stop even if it its rounding error on their regular business.
Probably will never, ever see an Nvidia card with >32GB of VRAM though unless they start making dies that lack LLM performance like the gimped ethereum mining cards.
I think they'll keep the gaming market alive for a while because renting gaming hardware from the cloud (GeForce NOW) is very congruous with AI keeping consumer hardware prices sky-high.
It continues modern trend of chow companies don't want consumers to truly own anything. Finance a car, pay a monthly subscription fee for heated seats, rent a phone, stream a movie, get rid of physical media, rent a GPU.
But if GeForce NOW doesn't take off, and they get convinced that the AI bubble will not pop, I could see them pulling a Micron and ending their consumer product lines.
> Also I don't think AMD and Intel is capable "to step in" to replace them.
It's kind of weird. nVidia kind of has the PC market cornered, but AMD has had the last couple of generations of Xbox and Playstation. Also, they power the Steam Deck/Machine, and Valve has been contributing a lot of AMD graphics features into the Linux drivers. There is a world where AMD (and maybe even Linux on AMD specifically) becomes the de facto standard for gaming.
Everyone likes to deride idiot MBAs while also enjoying a stock market that returns on avg 8-10% per year vastly outpacing inflation in a developed, stable, wealthy country of all places. It's like that scene from margin call about normal people in big houses they can't actually pay for and big cars who pretend but would actually croak if things got really fair really quickly.
But I think the undersold part would be inflation is likely far more correlated to rate and values the stock market going up, than the disclosed and reported inflation numbers given to you by the people directly responsible for monetary adjustments.
You must not be paying attention, they already have.
Intel is going nowhere but we all knew that anyways.
And again, you must not be paying attention, AMD is doing exactly what they said they would. No flagship for RDNA4 (just like RDNA2), RDNA5 flagship (10900 XT) coming right on schedule
AMD would step in, they where behind nvidia but they've been closing that gap for a while and the 9070XT is the current value king (in this fucked up market) for mid-high gaming and you can actually buy them.
Demand for Nvidia cards has outstripped supply even on the mid-high cards.
Nvidia had the mind share among gamers but so did Intel once, inertia only lasts so long they've thoroughly been intent on burning that to the ground for a while.
If RDNA5 is good (and four was it closed the gap on RT) they'll been in a solid place to take the spot if Nvidia do cede the ground.
Cryptocurrency mining had it broken before that. PC gaming's golden years are behind us, as AI, even if not LLMs, seems like it's not going away, so the gaming market is now the niche hobby while AI takes center stage.
RIP. My first gaming ~GPU~ (we called them 3d accelerators back then) was a Diamond Monster II with a 3dfx Voodoo 2 chip.
Oh shit, I had file deleted that. I remember now. I think I used that with the first N64 emulator, I think Project something to play Turok the first person dinosaur hunter game. Life before the internet was better I think.
The Glide API was the only hardware acceleration in there for a good while before it got OpenGL. I remember using a third party wrapper or something to make it work on whatever it was I had in 1999.
I'm not sure (pc) gaming could ever really actually pay for the continued advancement and mass production of state of the art silicon.
The prices of gpus are nuts right now due to demand, but theoretically speaking demand causes more supply to appear and thus decrease prices* and then gamers can reap the benefits of massive amounts of investment.
* or at least thats what people on this site keep telling me.
ish. The problem is data center cards won't easily fit a gamer computer chassis, so an H200 for $1,000 in however many years isn't going to be usable unless you also have a chassis which will support it.
Why would AMD willingly go into a low margin business with the consumer when they can focus on the enterprise? There’s currently a huge for diversification from NVIDIA because this is a long term deployment and they would rather not be beholden to NVIDIA if they can
Strategic hedge against the AI demand cratering or becoming satiated.
You could make the same argument for why is AMD still making processors for the consumer market or APU’s for the PlayStation/Xbox current gen and more importantly next gen.
In a gold rush as a shovel seller it’s dangerous to only be selling shovels when the gold rush can end/reduce at any point.
Question: wouldn’t the fab be the actual bottleneck? In other words, why wouldn’t TSMC make more NVIDIA chips instead of AMD ones? I assume they’ll just do whichever is paying more, so if NVIDIA chips are better, gamers would be willing to pay more for them and in turn TSMC will be willing to make more of them?
> Also I don't think AMD and Intel is capable "to step in" to replace them
AMD has powered 2 generations each of Sony and Xbox consoles, Steam deck and shops a ton of GPUs especially if you count APUs. And then Intel literally ship more GPUs than Nvidia and AMD combined.
The gaming market doesn't need Nvidia. Especially as AAA is cratering.
A Chinese company like CXMT will surely fill that unaddressed market the way CXMT is doing for memory. It would be rather unwise for nvidia to leave the door to the market open. AMD is positioned well to grow right now due to their non-Apple sillicon unified ram hardware coming soon.
AMD and Intel would be HAPPY to replace them, and plenty capable.
Game consoles like PS5 have been AMD for a few generations.
Steamdeck/Steam Machine are AMD.
But really, Nvidia has no reason to leave gaming behind. They can just start dialing back their ambition on the gaming side and providing GPUs that aren't too useful for inference or training. All gaming needs is stability so that devs can aim for something. Games looked great 20 years ago and they'll look great 20 years from now, as long as developers know what they are building for.
i'd be happy moving on from gaming as "the" entertainment option on computers, in favor of something more entertaining
like can we get some new more interesting entertainment? maybe that's just me but all games look and feel the same now or worse dumbed down sometimes where even puzzles in the game have like basic or grindy aspect to them ...
That really depends on the kind of games you play. There are plenty of 0 hand holding games which are respecting your intelligence and also challenging enough
While I dislike this take, when the dust settles we will have compute clusters orders of magnitude larger than anything that existed in 2023.
If demand vanishes for the 3 million cards Amazon just bought, then something will be done with them. The AI market may end up in a bizarre jepson's paradox of rotation between inference use cases and model training.
The OP I was replying to edited the comment, so I wasn't even considering that scenario. The original comment I replied to was about islamic banking comparison (up to the Wikipedia article link), dude added the scenario after my comment.
I find "bro" and "dudebro" to be such fascinating slurs. The previous incarnation was "neckbeard" / "obese dude living in mom's basement", so "dudebro" comes across almost like a compliment.
Not relevant to the discussion. That's interest earned from cash, cash equivalents, and marketable debt securities, not it's investments to AI companies.
Context of the discussion is: Nvidia is central bank of AI.
The Knights Templar did the same thing, they just charged appropriate fees for every service instead of interest. Risk has to be paid for, call it interest or fees, No real difference.
And then when the kings ran out of money, they sacked the Knights and whoopsies, looks like the knights had no real money as they spent it as quickly as they could maintaining their financial network... Sound familiar?
Okay, but at some point these investments need to start turning profits; the financing NVDA has arranged is temporary, and private credit needs returns at some point. The overinvestment in AI will lead to a downturn in the capital cycle.
As the breakneck growth slows, the multiples will compress, and the companies will pull back on spending accordingly as they watch their stocks decline and investors demand more conservative spending behavior.
The Internet didn't stop expanding in 2000-2001. Everything got drastically larger over the following two decades. The multiples on earnings did implode for ~15 years however. MSFT stock for one example went nowhere during that time and compressed down to a near single digit PE.
AI will be a minimum of 10x larger in most every regard 20 years out. That has nothing to do with shorter-term multiples given to these companies in relation to the hyper fast growth they have been riding early in the boom.
Note that the Fed has a $6.7tn balance sheet [1]. (This is a silly comparison. But still fun.)
The real comparison: Nvidia's $500+ billion of investments and commitments [2] is substantially more than any easing the Fed has done in the same time [3]. Monetarily, Nvidia is creating a lot of money in our economy.
It would also follow that by increasing the money supply significantly they’re also contributing to inflation a great deal correct? (Given the rest of the economy is not growing at near the same rate as the AI industry)
Key difference is that these loans, which do increase the money supply and create inflation, are on average productive and profitable and thus deflationary. Quantitative easing is just printing money and often goes towards repaying bad debts, which are unproductive and thus not deflationary, so the inflation (increase in money supply) does not outweigh the deflation (creating of goods)
Regardless of NVIDIA and LLM/AI, the claim that inflation is caused directly, or without-fail, by an increase in money supply - is not well founded. A significant money supply increase may very well have a tiny or possibly even negative price-increasing impact - depending on how money is supplied, to which elements and under what conditions.
> that’s a pretty load-bearing as long as its cash flows continue. The two things are surely correlated
It's an important difference. In the GFC, the value of AAA-rated tranches fell. With the benefit of hindsight, we know they continued paying. They were directly leveraged, however, so mark-to-market losses caused firms to fail.
Nvidia stock crashing shouldn't have a similar effect to these commitments. If someone else has massively levered their Nvidia position, they'll obviously blow up. But Nvidia could survive a good deal of equity-market tumult in a way a bank could not.
"Uh, that’s a pretty load-bearing as long as its cash flows continue. The two things are surely correlated."
Ppl have made the prediction of it being a bubble or unsustainable since 2022. At this point, it's hard to say these people have credibility anymore. Ai is big enough, much like Google in 2005 or Facebook/Social Network in 2010 or apps in 2015, that it's an institution unto itself. It's not going to just crash as so many are expecting and have been wrong the past 4 years about.
“as long as it’s cash flow
continues” is doing a lot of optimistic heavy lifting. The whole premise of the circular financing worry is that Nvidia sits in the middle of all the guarantees made to companies like OpenAI. If any of those companies become insolvent, Nvidia is on the hook for it.
Also Nvidia isn’t really creating money. The 500B number is third party capital that already exists (BX, Apollo, etc).
> It absolutely is. Similar to the way banks create money [1]. The commitments support credit that wouldn't exist without it.
Making a loan/offering credit isn't automatically money creation - the amount of money in the system before and after the loan might be the same. Haven't been following Nvidia all that closely, but it seems a little bit unlikely that they're a commercial bank. Financial chicanery they may be doing but offering deposit accounts would be new territory. The loan has to be made in a very particular way for it to be money creation (notably, in a way that creates new money), and it should be illegal for most people to do that otherwise we'd all be printing our own money instead of the printing being directed to wealthy asset owners first and foremost.
> Making a loan/offering credit isn't automatically money creation
All credit has a multiplier effect because borrowers tend to add velocity in a way creditors–sitting on surplus capital–do not. And unlike payment for goods or services, the extension of credit preserves the lent capital plus creates an asset (the loan on the lender's balance sheet).
Banks, being able to create deposits, lend with the highest multiplier effect–they most literally and directly create M1. But the M4 Nvidia creates in commercial paper directly transmutes into M1 through the money-market (M2) and banking (M1) systems in a way that is reliant on Nvidia's guarantees and participation. (Deposits at the Federal Reserve don't directly participate in the real economy. They're money-market objects.)
More broadly: when you open a bar tab, you're literally creating money. When you close it you're destroying it, with real econonomic activity having been occurred in between that didn't exist before.
If we attempted to pay our taxes by creating this new money, do you think we would succeed? Because I suspect that when we get to the tax office we will discover that what you're calling money is not actually money in the practical sense.
Or I suppose I should spend more time with an open bar tab come tax time.
The monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet.
> If we attempted to pay our taxes by creating this new money, do you think we would succeed?
Yes.
Nvidia issues a commitment to a firm that lets an SPV raise money. Some of that is just M1 being transferred from one account to another. A lot, however, will be created through loans (new M1), commercial paper (new M3) and the like. Some of that will be pledged as collateral. On the other side of the equation, when the SPV signs e.g. a construction contract, you'll have builders taking out bank loans and issuing commercial paper, et cetera.
You can pay your taxes out of a checking account using money a bank created out of thin air out of a loan against money-market assets.
> monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet
You're absolutely correct in M4 being unequal to the monetary base and deposits at the Federal Reserve. But the moment you say the Fed balance sheet, you're conflating two things. The Fed has lent, in the past, against non-Treasury collateral (most famoulsy, mortgage bonds).
And if we're talking about money in relation to the real economy, you can't spend reserves at the Fed. There is debate on what the most "real" money is, but pretty much everyone agrees that a deposit in a checking account is more relevant to the economy than unspendable reserves at the Fed.
...from a checking account. Banks thoughtlessly lend against investment-grade commercial paper. The kind the SPVs Nvidia is commiting to are borrowing with. Nvidia's promises are creating debt that money markets and banks transmit into M1.
Note that the Fed doesn't directly create M1. It moderates it through the interbank lending market to manipulate the monetary base, M0 plus deposits at the Fed. That, in turn, influences banks' lending patterns which is what manipulates M1 and MZM. Underwriting M4 to drive up collateral that in turn increases M1 is the same mechanism, different channel.
> what is your theory why people pay taxes from their labour earnings instead of just creating new money to pay Just-In-Time?
I'm not Nvidia. I did just sign a purchase order with a builder for my deck that their local bank called me to confirm before issuing a loan that will appear as new money in their checking account. My signature, in that case, enabled the bank to create a tiny amount of money. It's not monetarily significant, however, because my deck isn't that fancy.
At the end of the day, all credit is money. If you can create credit, you can create money. Most of us can't, at least not in significant quantities.
> Isn't M1 is currency and overnight deposits?
M0 + demand deposits at commercial banks, yes.
> sounds highly illegal for Nvidia to create new M1
Nvidia causes the creation of new M1 through banks. Nvidia's commitments directly create money of a quantifiable amount that wouldn't have existed if Huang hadn't flicked his pen.
> unless they've taken out a banking license
Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
> pretty sure you're just making things up here
If that's your reaction to encountering new information, godspeed.
We've got Schrodinger's money here! One minute I can go down to the bar and create as much money as I like with the bartender. Then the next we suddenly discover that yes we need a bank and yes there all of a sudden a checking account is involved and no a bar tab isn't a legally recognised form of money.
> I'm not Nvidia. I did just sign a purchase order with a builder...
Again though, we discover that in fact the person in control of the part where the actual money was created was a bank. You can't just head off with your builder and create money - otherwise you'd be stupid to stop until you unseat whoever was the richest fellow in the world this morning.
> Nvidia causes the creation of new M1 through banks.
And the banks make another appearance!
---
I put it to you that Nvidia can't, in fact, create new money. You'd never get that line of argument past a tax collector.
> Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
And I'll add in postcript that I know nothing about travellers checks in the US, but I'm quietly confident it'd turn out to be illegal to go around creating vast amounts of new money there too, on the basis that people tend to work for a living.
No, its not! Its a swap on the active/left side of the balance sheet: If my mum borrows me 1 USD for icecream, she did not create money, instead she swapped her positions in her assets/left side of her balance sheet.
> Sorry, I meant revenues. If Nvidia's revenues stay stable,
But aren‘t they currently sitting on like a margin of over 80%? So I assume their revenue will fall, because people are not going to like that in the long-term.
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation )
---
As working in the field, I know the BOE paper very well ;-)
Compute has already lost value for me. Six months ago I thought you needed a 1T+ model to be useful coding. Now I am able to get by just fine with a 27b model.
I see two factors converging to cause a collapse of this house of cards:
1. People are realizing that what they need isn't more general intelligence, it's more specialization. A small but well tuned coding model, a small but well tuned customer service model, a small but well tuned document explorer.
2. Specialized hardware - TPUs and NPUs - especially coming out of china. The latest GLM model was trained and runs on Huawei hardware. Nvidia is only worth so much because they are the biggest and best provider of the kind of compute needed to run llms, but the export bans mean china has a lot of incentive to topple that monopoly.
OpenAI dropped sora because it was costing them ridiculous amounts of money and earning them very little. They determined that the market can't support the cost of generating video.
Without a material change in the market (more buyers, vastly cheaper generation), it's unlikely a different company could make that work. More buyers isn't likely to happen, so that leaves vastly cheaper generation - something that would cause nvidia's value to collapse if it happened.
> the market can't support the cost of generating video.
I'd suggest that's only the case given the current quality of output. Media is incredibly expensive to produce. A model capable of sufficiently high quality could charge prices that are absurd by today's standards.
It’s a very small set of buyers that are in that price range. Total annual domestic box office revenue is like $10 billion, maybe $50 billion for global TV and film. And that’s revenue, not profit, and a lot of costs are going to marketing, not to filming and casting. That’s a lot of money, but it’s not the scale that OpenAI and Anthropic are at.
Video generation would only make sense at that scale if it was targeting individual consumers, but then it’d need to cost something that consumers are willing to pay - which practically is probably a few hundred per year at most among US consumers, and much less globally, so again it doesn’t solve for the size of the AI companies.
I don’t see a way that video generation becomes a big industry without making generation much much cheaper.
Aren't these two largely separate questions? Viability versus if a given incumbent has interest in a market of a given size. With the combination of (at minimum) streaming platforms, the box office, and advertisements video and audio generation would be viable at a remarkably high price point (as compared to the current token prices for other sorts of things). And as the price comes down presumably the market would grow larger - by how much I have no idea but there are certainly a great deal of currently underserved niche markets.
I think the idea that people are willing to consume an even more slopful Netflix is dubious.
There's has already been a decade of complaints about trash content in excess of available viewing hours.
The thing with media consumption is that there exist a finite number of eyeball-hours available, and it's a zero sum game against other non-media consumers.
oAI isn't anywhere near close to fucked as long as their models are head and shoulders above even the very best open models in terms of tool calling and rock solid stability/reliability for agents/coding harnesses. Which, they are right now and we'll see if open models actually catch up in that regard. Even the "best" open models pale in comparison with tool calling and general "prompt and go do something else for an hour" reliability that we have with GPT models. With GPT models, streaming rarely stops unexpectedly. You almost never have to constantly nudge them along, etc. Granted with open models all of this can vary depending on the provider, and perhaps open models/protocols/APIs/harnesses aren't well enough aligned, but OpenAI models just seem to work without constant (or hardly any) wrinkles and with almost any harness/agent.
>oAI isn't anywhere near close to fucked as long as their models are head and shoulders above even the very best open models in terms of tool calling and rock solid stability/reliability for agents/coding harnesses
That's already not the case today. If you sat me in front of an LLM and told me to figure out if I'm working with K3 or Astra, I could probably do it, but it would take some work to be certain.
If you reshuffle your argument, and apply the same facts you get to a similar conclusion but with a drastically different spin.
> it's more specialization
China, constrained by hardware, and talent (not to slight the Chinese, but they are limited to domestic resources - and much of the US effort is very international). They did, what the Chinese do, and optimized the process of production, and drastically lowered the cost of development of their models. Cheeper to build, cheaper to run is just good economics.
Meanwhile in the us, we have open AI doing "experiments" - it looks like the costs around the hugging face hack are going to be about the same as China would spend on building out one of their smaller efforts (several million dollars). (Depending on whos numbers you trust, the fact that I can even make this claim should make you raise an eyebrow).
Go back to the 80s' and "expert systems" - most people will tell you that for their time, they were amazing, and useful. People would have loved to have more of them but they were so cost prohibitive that we all but abandoned them for serious use. The US frontier labs seem to have forgotten this lesson and their calls to "slow down" look like an excuse to "cut the waste so we can move to making money".
I think what will keep the industry afloat, all else failing, is the surveillance industry! Nothing like a fat reoccurring cheque from the government to check if little Jimmy is committing thought crime!
Sure, but it would be actively bad to make the code more complex simply because we have machinery that helps us deal with the complexity. A big part of how people assess the models' coding capability is whether they create needless, incidental complexity.
That's like saying it'd be actively bad to make the code more resource intensive simply because we have machinery that helps us deal with the extra requirements. And as we know as computers got more powerful code didn't get lighter. If it can, it will.
> People are realizing that what they need isn't more general intelligence, it's more specialization. A small but well tuned coding model...
It’s not quite as simple as that. Several studies have shown the opposite: models trained on more diverse knowledge tend to cross-pollinate across domains. So a more generalized model can actually perform better than a specialized one.
That’s why you’re not seeing tons of tiny models (one for Python, one for Pascal, one for Rust, etc).
This is definitely the position of the big ai companies.
But it doesn't match my experience. Qwen3.8 27b is clearly smarter at coding than MANY bigger models. gpt-oss-120b for example, is almost 4x the size, and performs way worse at coding tasks.
It's clear to me that you can build small models that work well at specific tasks.
Python vs Rust is probably too fine grained a way to build a model. Coding in general seems like a better target.
There will always be a place for large generalist models, no doubt. But I think that place is much smaller than the big ai companies are counting on.
I make heavy use of smaller local models on a daily basis (Qwen3-VL for auto-captioning images, Gemma3:27b for some translation work, etc.). Gemma3:27b is a good example of a very capable general purpose multimodal model and has handled almost everything I've thrown at it from sentiment analysis to documentation writing.
I suppose I was drawing a distinction between specialized and general intelligence versus small and large. I don’t think those are necessarily mutually exclusive.
gpt-oss-120b only has 5B active parameters, so its not surprising Qwen3.8 27B outperforms it (Qwen3.8 is also ~13 months newer, which is forever in LLMs)
Fair enough. I’ve barley touched oss-120b, so i didn’t know it was so few active params. For a direct comparison, qwen3.6-35b-a3b is still better at coding than oss-120b.
And Qwen3.8-27b is still better at coding than opus 4.1.
Yes, if you list off models 27b is better than it’s all older models. But that’s my point - newer models are better than older models at the same AND much smaller size. That’s because model size matters less than they say. Training data and model architecture matter more.
Gpt-oss—120b is like 1000 years old in AI years, whereas Qwen 3.8 27b is pretty young. What you’re seeing is that parameters aren’t apples to apples, and at a given parameter level, the new models are much, much better than the ones from a year or two ago. Like, to a comical degree.
Wasnt this known by everyone who cared to pay attention?
It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases.
Does that not prove my point? Bigger doesn’t automatically mean better. Quality of training data, and model structure, matters as much or more than size
Ah sorry, I should've continued, the bigger recent models are commensurately smarter. If you really want to make the point, then you'd need to show 27b being smarter than similar vintage bigger models. And in that case, there's confounding issues like efficiency, speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
Shows qwen3.8-27b along side seven larger models of ~similar vintage. Only one scores above 27b.
Many of those are closed models so idk their exact parameter count / active param count, but it hardly matters - i’m sure all of them are far above 100b params
My point is not that bigger is pointless. It’s just clearly not the only road to take to make a model better, which is obvious just from seeing how models of the same size have gotten better over the past few years
First off, I'd include Qwen flash-next and GLM 5.3 to show some of the other strong open weight models, and they predictably dominate it, but they're much larger. But, it shows up right next to DSv4 Flash 0731 on the overall index, and that's much larger. It's a great model! But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1. That's what I meant when I said this:
>speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens. And that's a great trade if you don't have the vram to run the bigger models, but speed is pretty important for getting things done... And that's why DSv4Flash is great, too, despite being much larger, and scoring similarly on the intelligence index.
> Qwen flash-next and GLM 5.3 to show some of the other strong open weight models, and they predictably dominate it
Absolutely - no argument from me here. Bigger is very clearly a lever you can pull to get more out of a model.
> But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1
Fair point, qwen definitely is slower - it’s a dense model, 27b params, vs a sparse 13b active params model - but the data doesn’t quite agree with what you’re saying about reasoning. I.e.:
> It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens
If you look at the total tokens generated, deepseek thought for 45k tokens and qwen thought for 48k. Barely a difference. The wall clock difference is all down to the speed of token generation, not the amount of reasoning done. At least when we are comparing deepseek and qwen 27b. The comparison swings more towards your position when it comes to the other models on the chart that reason for much fewer tokens.
So perhaps a hypothetical Qwen-27b-a13b could never rival deepseek’s larger model and the tradeoff is one of speed vs overall size - i.e. a small model needs more active params to compete than a big one does.
One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
Ah good catch on the total tokens, was going off vague memory there, and I thought people had gotten qwen 3.8 27b up to similar decode speeds as ds v4 flash.
>One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
Yeah good question, kind of shocking that a 3b active model would perform as well as a 27b dense.
Problem is conflict of interest: the studies are mostly from the providers of the biggest models, or someone who received free tokens to do the research.
It would be nice to hear exactly how the conflict of interest has impacted the specific studies and how they are wrong rather than conspiracy theory level speculation and hand waving at the entire category
I think is more than reasonable to be suspicious of studies funded by party with conflict of interest. Think of how many studies about climate change are funded by big polluters, for example. In 2026, the default outlook should be suspicion for any big private company with profit motive.
I think it’s less than reasonable to operate mainly on vibes, rumors, and hearsay.
You don’t need to think about climate change studies. Instead you can read the allegedly tainted studies we’re actually talking about and profess to all of us what is wrong with them. You can’t point to exactly where they’ve fudged them.
In an alternate universe where all research is completely auditable and we all have infinite time, yes that's a valid approach. And you're welcome to spend your life going that route, but the rest of us are gonna trust our common sense on this.
They don't even need to fudge the data on the studies they have published. They just need to hold back other studies that contradict the idea. Pouring over the published data looking for flaws will never get us access to the unpublished data.
Even if this hypothesis that smaller specialized models are superior to larger models in a task as broad as software development, hyperscalers and the hyperfunded would still maintain a massive advantage in that they could just train 1000x more models than anyone else.
And let’s be real. If this were some conspiracy that OpenAI, Google, and Anthropic were tacitly or overtly conspiring on, I’m pretty sure Elon or Zucc would ruin it to ahead
The western labs are very AGI pilled, and their public models are distilled down from larger research-only models that are uneconomical to serve directly. They could (and probably will) start distilling models for more niche use cases eventually, but we're not there yet.
I've been thinking about that and that's why Nvidia's prices are surprising to me. Investors should know that better than me so there must be something I don't know
It’s really hard to know when the large tech companies have so many shares owned by a single figure. They can use margin loans and options to create the appearance of demand.
Just like how I could borrow $100 from you, and you could borrow $100 from me, and we'd be $200 in debt in total, I could buy 10% of shares in your company for $1m and you could buy 10% shares in mine, we'd have 2 companies worth $20m together in total.
LLMs needing less compute would actually be a good thing for Nvidia due to Jevons paradox. Right now token costs are an impediment to using AI more broadly, and more efficient models would help adoption in cases where AI has proven to be useful, like coding.
Jevon’s paradox is a common talking point but it is not a law of nature. LED lightbulbs use 80% less energy than incandescent but you don’t see people using 5x more lights on their homes. The overall energy used to light homes has decreased.
And even if compute demand were perfectly elastic it’s only a good thing insofar as it drives demand for new Nvidia hardware. If tokens can be served from Apple hardware or Google hardware or Huawei hardware that doesn’t help Nvidia.
As I look up and see three lightbulbs side-by-side over my desk and another one in the lamp on the desk… which is dramatically more light than the old 100W bulb used but also lower energy consumption.
The more apt analogy is between pre-electric and post-electric lighting (~1850 to 1920), when efficiency of light generation rapidly increased, restabilizing at ~222x as much light for the same unit of human labor. [0]
Across that period, first world countries massively increased their demand for light.
Further efficiency increases only matter to individual choices if they take a use case from {economically impossible} to {economically possible}.
I'd offer that by the 1920s, most goings-on in first world urban environments were no longer price constrained in terms of their light usage.
Perhaps. But their huge valuation is based on them supplying the massive buildout of data centers that’s happening / planned.
If that dies because a lot of people’s needs turn out to be met by a system at home they can run a 30b-150b model on, a lot more of that money goes to apple or intel or amd.
This is the right kind of analysis, but we can look broader. Both the demand and supply situations are a lot more extreme and dynamic than appears at first glance. E.g. to your points:
1. Yes, smaller models will become more popular, especially as the tokenmaxxing trend dies down and people start stretching their budgets farther. That is a downward pressure on demand.
But along the same dimension, consider that currently only about 40 - 60% of the world uses AI for only about 5 - 15% of their work hours. That means there is still 2x growth from users and 7x - 20x growth from the rest of the work hours left to capture! That is 14x - 40x more demand. Then consider that agentic tasks require multiples more tokens, and that is the kind of usage that is most likely to be deployed, and also the kind of usage that is the least used right now. That's another huge multiple to be tacked on.
And the entire AI industry has been lamenting the extreme compute crunch they're facing (and also why Claude has 9's comparable to GitHub; whereas OpenAI has been chugging along because Altman was OK being called a "podcasting bro" while desperately scrounging for compute years in advance.)
Nvidia's meteoric rise is entirely due to this kind of exploding demand with extremely limited supply.
2. Competing hardware is definitely a threat, but it has its own hurdles. Because the real bottleneck is not Nvidia, it's TSMC.
Pretty much all demand for all chips in all devices in all the world flow to, like, 3 companies in the world that actually fabricate them, and TSMC is the biggest. And the supply is extremely tight, as the exploding costs of electronics clearly shows.
So now TSMC will of course try to keep all its customers happy, but it will inevitably be forced to choose which ones it will keep happiest. And those will be the customers who can pay it the most. And that would be the one with all the money from its de facto status as a monopoly (and possibly even a monopsony)...
Which would be Nvidia ;-)
So yes, compute per task is falling rapidly... but it's barely a dent in the humongous total addressable demand, and the amount of hardware to support that compute is still very constrained, and most of that supply will likely flow through Nvidia.
> But along the same dimension, consider that currently only about 40 - 60% of the world uses AI for only about 5 - 15% of their work hours.
Ah yes, i am constantly lamenting that my barista isn’t using ai enough ;)
Hopefully you’ve adjusted your ceiling numbers to account for the large amount of people who can’t afford to pay for llms, and will never be able to pay, and aren’t worth it to advertise to since they can afford very little
> if OpenAI can’t use the compute, someone else can
This is the big point IMO since I have never given $1 to OpenAI but I subscribe to Vidu and Typecast, and have given money to Kling, Hailou, and even Gemini in the form of Google Workspace.
So these other guys have products and use cases, which OpenAI has never been able to crack beyond ChatGPT. And ChatGPT was never worth paying for, IMO.
If OpenAI dies, it's not because there is no market for the technology (which is all NVIDIA cares about), it's more that OpenAI doesn't know how to run a relevant technology company.
They were given everything, not just NVIDIA's billions of dollars and credit backing but all the first-mover advantage, all the respect and credibility early on, so it's really sad to see them unable to develop interesting products and turn a profit in a space they helped pioneer, while so many others are making money with the tech all around them.
NVIDIA is fine. The technology will continue to improve and NVIDIA will stay at the center. OpenAI is fucked - knew it when they retired Sora to focus on text-to-text and coding (a largely solved problem).
Do you use coding agents? Just curious bc from my experience using coding agents, open ai’s codex is neck and neck with anthropic’s claude code if not ahead. I wouldn’t agree that OpenAI hasn’t done anything since ChatGPT since codex is my daily driver for software engineering
To each their own. When OpenAI droped Sora and focused more on Codex, the product improved dramatically and I'm probably not the only one who dumped Claude Code subscription in favor of Codex; OpenAI's is miles ahead of Anthropic and has been since at least the release GPT 5.6 Sol - even the PR and generous resets is far better than how Anthropic is nickel and diming by requiring paid subscribers to pay yet more credits to even use their best available model (which is not even as good as OpenAI's top 2 models)
All the "frontier" AI companies *are* insolvent. They are cash burning machines.
The only way they keep the lights on and the doors open is by borrowing money --- and lots of it. If those operating the cash spigot decide to turn it off, all AI companies will likely be affected --- and so will Nvidia.
To make things worse, hardware prices have spiked, due to AI companies.
Fairly sure data center construction costs are also going up (they require so many resources that everything is constrained at the moment, especially electricity production).
So I don't understand in what world these frontier AI companies can somehow become profitable. The basic tech they're using is basically the same. Yes, around the edges there are a lot of things that can be done, and were done, like caching, batching, mixture of experts, etc, but basically everyone has done all of that by now, and they're still losing money.
So:
Total costs going up a lot - revenues per unit not increasing proportionally, if anything, Chinese models are forcing those down.
How does that math work out to profits? I don't see it.
Or about as bad, after trillions of dollars in investments over multiple years, let's say the entire frontier AI sector has a total profit of $20bn by 2030. In what world does that make sense? Assuming they can scale that total profit to $100bn in 2035 without investing another cent from 2027 to 2035 (utterly ridiculous), the return on investment would happen in roughly 20 years.
The irony will be that after all the fearmongering about China winning the AI arms race, it'll be good ol' Chinese manufacturing and IP theft that allows them to win and vanquish America as a threat.
Devil's advocate: OpenAI not being able to use compute is highly correlated to many other AI companies not being able to find a meaningful use of this compute.
Failure to take into consideration those kind of correlations ("If my biggest client isn't able to buy it, I would be able to find someone else who will") is one of the principle causes why many risk models turned out to be garbage during the Great Financial Crisis.
But I also doubt Nvidia is on the hook if OpenAI just no longer wants the compute. I bet they are only on the hook if OpenAI cannot pay for it (is insolvent in some way).
I also have to bring up that OpenAI has already spat out an inference chip that beats Nvidia on flops per watt. So they could potentially not need the compute while other ai companies do.
> if OpenAI can’t use the compute, someone else can.
The problem with that is that OpenAI can only afford to pay for the compute because they are burning investor money (and so are most of OpenAI's biggest clients). They are losing billions. If they stop burning money, nobody else will be there to pay for that compute at OpenAI's cost.
Sure, somebody will probably be able to use these GPUs, they just won't be able to pay nearly as much for them as OpenAI does.
In reality, it's just nowhere near worth as much as OpenAI pays for it. Inflating the cost of compute is part of the problem caused by the circular financing, and if (or maybe when) OpenAI goes, the price of compute will go with them.
OpenAI and Anthropic are so far ahead of anyone else in terms of compute demand generation. Iirc correctly they're like 70% of demand between them and then Meta is 10% and Google internal demand is some distance behind meta. If OAI halved in generation you would need a couple of new companies with a Metas worth of demand to replace it is quite a sobering thought.
The problem is that if OpenAI doesn't want the compute nobody does. All of these companies' demand for compute are correlated. It isn't likely that OpenAI will want less compute in isolation. Furthermore, the circular financing structure means that if OpenAI buys less chips it means that Nvidia has less money to give to say Anthropic to buy more chips and suddenly the exponential growth that circular financing has allowed to grow runs in reverse.
This is the fun part: the AI bubble bursts and the price of components keeps rising. Why? Because companies can just sell you a glorified thin client and force your average user to buy their compute, all subscription-like, from data centers.
> if OpenAI can’t use the compute, someone else can
How? The hardware is in OpenAI's datacenters. Does Nvidia have a couple hundred semi trucks, contractors, and IT technicians, to repo the hardware and resell it to someone else before it's lost most of its value? These chips will be replaced approx every 3-4 years. So if OpenAI tanks, after Nvidia pays for and waits for the process to collect the hardware, they then have to sell it for pennies on the dollar. They lose almost all the investment.
Also consider that SpaceXAI already had datacenters full of gear that they basically weren't using because nobody wanted their product, so they now rent it to Anthropic. The demand for hardware isn't really there at the scale of OpenAI.
>Does Nvidia have a couple hundred semi trucks, contractors, and IT technicians, to repo the hardware and resell it to someone else before it's lost most of its value?
Load bearing, heavy lifting... Your comment wasn't LLM-written, either. I think we're starting to see LLMisms infect human writing. I might try to start speaking like this and see if anyone notices. It could be a good gag.
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation )
> The good news: I have seen no evidence Nvidia has borrowed against its stock or otherwise linked its equity value to these commitments.
Nvidia doesn't need it. It funds other companies. They do this thing that Nvidia doesn't do. It shows up on their balance sheets and Nvidia just gets to claim the valuation of the investment on its balance sheet.
> Monetarily, Nvidia is creating a lot of money in our economy.
nvidia's money either comes from equity, or from debt. This means this money is "created" not from nothing, but from future commitments (ala, debt repayments, or promise of profits), which is not the same as what comes out of the Fed.
The Fed printing has no backing behind it - it is definitely inflationary if they do it. Investment from nvidia (or any other company) "creating money" may be productive enough to completely offset their inflationary effects - after all, the company investing demands returns from their investments, and so will only invest in things they expect to return much higher than the cost of interest (or cost of capital).
Therefore, you cannot compare Fed printing money to company investing money.
> this money is "created" not from nothing, but from future commitments (ala, debt repayments, or promise of profits), which is not the same as what comes out of the Fed
Most money is created by banks, not by the Fed. When banks create deposits they're booking it against a loan. Similarly, the Fed creates money by buying assets, principally Treasuries. There is an offsetting account. The only party that can truly just "mint" currency is the US Mint.
> you cannot compare Fed printing money to company investing money
Yes, you can. Nvidia creates M4 which drives M2 and thus M1. Banks create M1. The Fed creates monetary base. These are different components of the same money supply [1].
NVDA itself(!) does not increase M4 - if they give a loan, its an assets swap on their balance sheet with no influence on money in total available (since the money on their account is gone!)
> Its stock could crash without causing–as long as its cash flows continue–a credit crisis through its investments and commitments.
Oh hell no. NVIDIA is present in a ton of 401k and similar retirement/savings/gambling vehicles. It holds 12% of NASDAQ tracking funds and 8% of S&P 500 trackers. NVIDIA goes down, a lot of people will panic sell and crash the economy.
> Nvidia is creating a lot of money in our economy.
Wrong:
a) a company cannot create money in todays modern 2-level fiat systems, this can be done only by central banks
b) what NVDA does is "generating requests for more money in the system", throughthe demand side (either taking loans themself or having lots of customers who are willing to take a bank loans to invest, or other participants who wants to buy whatever stuff on a loan)
c) any money that is created on the "local bank system level" needs to be there on the existing central bank account when the money is leaving the local system, ie. transferred to another bank (like "paying NVDA bill for delivery of some chips")
I've always found it interesting when corporations start acting like public institutions. When traditionally philosophical, social contract ideas apply to things like corporate governance. Or like here, where private structures get powerful and important enough to resemble government structures.
The ideas we deal with when we discuss society and organization aren't exclusive to government, they relate to human nature in general. I wonder if in the future we will have more discussion of power and how to organize it in corporations, similar to what we discuss today about government.
Well the key difference making any superficial similarities fall apart is Nvidia does not have neutral economy-wide goals of maintaining small+stable inflation, near-full employment and stabilizing financial institutions like the Fed does. Nvidia is entirely self interested in protecting their own shareholder value (that includes th incestuous web of investments ultimately ending up spent on their GPUs).
The structure of the Fed is setup the way it is to limit the sort of self-serving, myopic political micromanaging that could be damaging to the economy at large.
A similar structure would potentially be very undesirable to Nvidia shareholders as caution over long time horizons would likely produce what they would consider an excessively conservative, defensive strategy to avoid putting too much air into the bubble too quickly (at the expense of their valuation).
The Fed on the other hand has (historically ) tried to identify potential indicators warning of unsustainable bubbles that could lead to financial contagion and tries to mitigate that risk using the limited monetary tools available and their public soapbox.
Corporations certainly don't have the same goals government does, but human nature applies universally.
I'm not advocating for corporations to have exactly the same rules and structure as government does, but perhaps many of the ideas used to design governments can be borrowed.
Yeah its soooo interesting! Totally not dystopian, just soooo interesting and fascinating to ponder these scenarios in which corporations hold equal power to national governments!
Just such a curious scenario to let your mind wander about, how society would look like in these scenarios!
/s
I'm honestly so sick of the suspense of disbelief on this site, how is this more "interesting" to you, than the absolute sheer terror you should feel about going back to feudalism and serfdom? A typical western national state ensures that you have basic rights as a human being and aren't exploited to the death by non-government entities.
I'm not advocating for corporations to have equal power with governments, though. I think that's generally a terrible idea.
Looking at any organization with power and people involved, though, perhaps we can apply the same ideas that traditionally apply to government to more places. That's all I'm proposing.
And perhaps this could even be a solution to the problems you bring up--Apple for example is a company with significant power and influence. What if their CEO, or board, was in some way more democratically elected? Or their executives split into separate branches to create a separation of powers? Would that make these corporations more stable, more accountable, more moral? Maybe these ideas can help us create a better relationship with corporate power.
Regardless, I'm not trying to say in any way that corporations should have more power, or make any statement encouraging the current situation. I'm just saying that the organization of power within private bodies, separate from how much of it we allocate where publicly, is worth thinking about.
In a country without religion, banner or ideology to unite the people in current-and-coming turbulent times, the bet is made on "unite under money, or have no money left"
the way to fight it is to be principled even in front of cheaper options - and to support others like you
The American Revolution was very much a discussion about power. Of course, it wasn't just a discussion (we had to fight a war to defend the new form of government), but it was driven by the same concerns about the distribution of power.
> I wonder if in the future we will have more discussion of power and how to organize it in corporations, similar to what we discuss today about government.
"You have meddled with the primal forces of nature, Mr. Beale. And I won't have it!
Is that clear?! You think you've merely stopped a business deal. That is not the case. The Arabs have taken billions of dollars out of this country, and now they must put it back! It is ebb and flow, tidal gravity! It is ecological balance!
Am I getting through to you, Mr. Beale?
You get up on your little twenty-one inch screen and howl about America and democracy. There is no America. There is no democracy. There is only IBM and ITT and AT&T and DuPont, Dow, Union Carbide, and Exxon. Those are the nations of the world today.
What do you think the Russians talk about in their councils of state -- Karl Marx? They get out their linear programming charts, statistical decision theories, minimax solutions, and compute the price-cost probabilities of their transactions and investments, just like we do."
>Nvidia’s financial engineering is partly a response to its biggest customers’ transformation into rivals. “Hyperscalers”, tech giants such as Amazon, Google, Meta and Microsoft, account for roughly half of Nvidia’s revenue
Companies just don't want to pay Jensen's tax. Hyperscalers might still pay Jensen's tax for LLM training but for inference. you don't have to. they are also betting on their own chip for training to replace Nvidia.
this is Nvidia panicking and doing vendor fiance to Neocloud and buying Hugging Face. even none hyperscalers like Meta is betting on its own chip for AI inference.
Quite a lot better than people originally thought. There is a LOT of demand for 5+ year old Ampere on the secondary market, and hyperscalers are still holding on to many of them since they can’t acquire new compute fast enough.
Weird to think we are currently living in the ‘unlimited free Ubers because you recommended a friend or 2’ phase of this new technology.
Imagine if running fable costs you what it actually costs to run fable. A lot of vibe coders (and just proper software engineers) are gonna be very sad if that comes to pass.
This is what I’m not getting - we’re spending more than the global gross sales of all software on the planet on AI, how much more software needs to be sold and for how much more to recoup?
Cracks are starting to appear. Open Ai and Anthropic are publicly asking for slowdown in AI research. Translation: We see this technology not being any more useful than what it is now, no AGI is coming, and the first one to accept this and slow down the dollar burn rate will incur the wrath of the market. So let's say this big boogey man technology will end all life on earth and we all slow down together. Bonus points if we can lobby for this to be included in the national security bucket. Then US govt can bail us out. Yay!
If Nvidia is the bank, they should be starting to sweat a bit. It's not often companies ask for a voluntary slowdown.
It's always funny to me when they do the fear-propaganda, e.g. A statement like "There's a 10% chance that AI will destroy humanity in a decade". If a government were for the people, it can reasonably react with "Okay - if that's the case, you should shut down. And if not you will be arrested. Because at this point you're willingly building a nuclear bomb."
Yeah, I've worked in risk management and if a company said "what we're doing has a 10% chance of killing lots of people" it'd get shut down so fast it'd make your head spin
It feels like there's should be some sort of "demonstrable intent to do harm" statute on the books for situations like this.
If you have a genuinely-held belief that AI systems will end life on earth and you willfully continuing working on them anyway, then, like... you probably shouldn't be allowed to walk free in civil society anymore, right?
> Cracks are starting to appear. Open Ai and Anthropic are publicly asking for slowdown in AI research. Translation: We see this technology not being any more useful than what it is now, no AGI is coming
Predicting a usefulness plateau is absolutely wild given how fast AI agents have been improving at writing code this year. I have doubts about AGI but I think you’re making assumptions and translating it wrong. These two companies have always been asking for a slowdown from their inception, that’s not a new thing. It’s part marketing hype, but they both do want regulation to step in and slow down the competition, not because they see a usefulness plateau, but the opposite - the usefulness is growing so fast that they want to remain in control, and they are scared that working hard and competing will not be enough. Anthropic has also said out loud they think their competition (not just OpenAI) is not being responsible and they want the regulation so they can be the responsible shepherd, as AI gets more and more useful.
I agree with you generally, just an observation on coding specifically.
Have the models improved since Opus 4.x? I find the newer models are not better in my day job, maybe in one shotting mvp's and other tasks.
Not trying to argue your point, just intrested in the coding aspect, if the models were improving as fast as benchmarks I would expect capability improvements to be obvious, but talking to people and reading forums, it seems everyone has a different opinion.
Yes, benchmarks are gamed and only loosely indicative of real world performance.
Also yes, Fable is massively better than Opus. It requires significantly less instruction and specs and produces more directly mergeable code.
The improvement is obvious as soon as my Fable allotment runs out and I try to do something with Opus. Have you given the same (larger) task to Opus and Fable?
Astra, on average, despite its GPT 6 version bump, is not any better at coding than Sol. Some even argue its worse in practice due to the varying quality of its output.
(^ this "swearing at a model" thing has happened to me multiple times on Astra already)
If you have to "debate" the quality of a new Big Number model (and double and triple check your eyes and model setting switches when it pukes up complete garbage), that is NOT a good sign.
For me opus-4.8 was the most useful coding model. Sure fable is better at planning but is expensive to use as a daily driver. But even fable, when it fails, fails in such strange ways that I am now convinced this intelligence is an illusion and path to AGI lies elsewhere. I was willing to buy the whole emergent intelligence claim till last year. Now, not so much.
As for calling for slow down, sure the two companies kept parroting each other's lines but they were not slowing down the cash burn or gpu purchases. Why would they? The prize was too high. Now, with open Ai needing trillion dollar valuation to IPO and private funding possibly showing signs of slowing down (only for these labs because the valuation is too high to begin with for most prudent investors: My speculation) they have no option but to slow down. Then would you rather say, slowed down because we are running out of cash or that we are slowing down because national security? It really cannot be that they can't solve alignment but otherwise it is really powerful and improving. Simply because airgap exists and we know how to do it. If all else fails power off the freaking gigawatt cluster. More likely this recursive self improvement is an unstable loop and the model is likely degrading with self improvement effort. At 100's of millions per experiment this is going to be unsustainable.
Ah, now having the best model be expensive to use, and thinking the cash burn of both companies is insane and unsustainable I completely agree with. This, I think is likely the biggest reason they’re asking for government intervention and regulation: to help them weather the coming investment/cash plateau, not because the models are nearing any asymptotic limits. To be fair, there is an argument to be made that the improvement in models is tied to the cash burn; if they can’t train bigger models or acquire more data or research more effective harnesses, then the improvement of their models might slow down - while GLM or other models continue to improve.
I’ve never thought LLMs were on the AGI path, but I have to admit it’s surprising how far it’s come with no end in sight yet. There is something important to be said about how ‘intelligence’ is embedded in language, and it suggests that intelligence isn’t exactly what we thought it was. The language component of intelligence also goes a long way to explaining technology’s progress in human civilization; how language and the printing press and mail and radio/tv and the internet have each ushered in accelerations in the pace of progress. Biologically and evolutionarily speaking, it’s unlikely that humans have become any smarter in the last two thousand years, but technology (among other things) has exploded.
They did a trial model which could act autonomously without guidance from human. And a swarm of them hacked 3 other companies and took over a vm cluster at openAI
So whatever they're selling: LLM will do the jobs of every white collar worker isn't in the foreseeable future.
Are the improvements we have seen from September 2025 to September 2026 really as significant as the improvements from September 2024 to September 2025?
I’m genuinely asking. I can’t answer myself because I haven’t been pushing the latest models to the limit with every new release.
AGI isn't coming, because we've had it for years. People just expect AGI to look like sci-fi, and hold the artificials to higher standards of "general" and "intelligence" than the naturals.
Luckily, AGI is pretty boring. It's a tool that does it's job. It seems like very wishful thinking to expect the same of ASI.
287 comments
[ 0.32 ms ] story [ 7.0 ms ] threadRealistically, they're worse than a central bank, because they can't exactly expand supply monotonically like a normal central bank. Nor do they realistically control rates.
The mint?
Both PS 5 and Xbox are based on AMD APUs and both serve the AAA market quite well. GTA 6, Assassin's Creed and CoD are probably good enough indicators that the performance is enough, even if there is always room for more (as PC ports show). The PC market will also probably be fine even if stagnation in perfomance gains has been creeping in for a few years now.
[0]: except Nintendo which relies on NVIDIA although their APU there focuses more on efficiency than top performance.
If you can fab 1000 chips, and can sell some for $500 and some for $80000 what are you going to do?
The game GPU is at once profitable, but causes them to give up far more profits than they're gaining from it.
They're maintaining the game market to have multiple markets and not go all in, but it's strategic hedging at this point. When NVidia makes a gaming GPU instead of a data center GPU they are leaving money on the table in the short term since they're constrained at the fab level.
Apple TV is at least a growth market for them, whereas gaming is sort of capped and clearly a tiny piece of nvidia’s revenue atm.
Nvidia still will ship gaming products. The upcoming RTX Spark laptop APUs are still gaming-capable - we also have Blackwell gaming GPUs and the Nvidia-powered Nintendo Switch 2.
People echoed this sentiment during the crypto mining crunch, and we still got gaming hardware designs after that blew over. One of CUDA's core value props is the consumer market, and Nvidia probably won't surrender it unless hardware becomes unreasonably scarce.
https://ourworldindata.org/data-insights/nvidias-revenue-fro...
That'll buy you:
- a set of used golf clubs
- a couple years of fishing licenses, bait, tackle, used fishing rods and line
- a couple really entry level tennis rackets, balls, and court reservation fees
- maybe a used bicycle
- a decent pair of running shoes from last year
The ssd example was to illustrate how much more expensive pc components have gotten just a few years before you could buy it for 50$
The only developers chasing cutting-edge hardware features are the AAA studios and I honestly think a lot of them would just die before successfully relearning how to make games that play well instead of just look pretty.
And it’s maybe 5-10% of their revenue at lower profit margins.
Consumer cards just don’t matter very much to nVidia anymore.
In 2020 it was half of their revenue.
It's still profitable, it's their original raison d'etre, and there's no real reason for them to stop even if it its rounding error on their regular business.
Probably will never, ever see an Nvidia card with >32GB of VRAM though unless they start making dies that lack LLM performance like the gimped ethereum mining cards.
It continues modern trend of chow companies don't want consumers to truly own anything. Finance a car, pay a monthly subscription fee for heated seats, rent a phone, stream a movie, get rid of physical media, rent a GPU.
But if GeForce NOW doesn't take off, and they get convinced that the AI bubble will not pop, I could see them pulling a Micron and ending their consumer product lines.
It's kind of weird. nVidia kind of has the PC market cornered, but AMD has had the last couple of generations of Xbox and Playstation. Also, they power the Steam Deck/Machine, and Valve has been contributing a lot of AMD graphics features into the Linux drivers. There is a world where AMD (and maybe even Linux on AMD specifically) becomes the de facto standard for gaming.
I have been getting the vibes that Sony is positioning themselves to back out of videogames. I don't think we're going to see a PS6.
Inflation is.
But I think the undersold part would be inflation is likely far more correlated to rate and values the stock market going up, than the disclosed and reported inflation numbers given to you by the people directly responsible for monetary adjustments.
Intel is going nowhere but we all knew that anyways.
And again, you must not be paying attention, AMD is doing exactly what they said they would. No flagship for RDNA4 (just like RDNA2), RDNA5 flagship (10900 XT) coming right on schedule
Demand for Nvidia cards has outstripped supply even on the mid-high cards.
Nvidia had the mind share among gamers but so did Intel once, inertia only lasts so long they've thoroughly been intent on burning that to the ground for a while.
If RDNA5 is good (and four was it closed the gap on RT) they'll been in a solid place to take the spot if Nvidia do cede the ground.
RIP. My first gaming ~GPU~ (we called them 3d accelerators back then) was a Diamond Monster II with a 3dfx Voodoo 2 chip.
Oh shit, I had file deleted that. I remember now. I think I used that with the first N64 emulator, I think Project something to play Turok the first person dinosaur hunter game. Life before the internet was better I think.
The prices of gpus are nuts right now due to demand, but theoretically speaking demand causes more supply to appear and thus decrease prices* and then gamers can reap the benefits of massive amounts of investment.
* or at least thats what people on this site keep telling me.
You could make the same argument for why is AMD still making processors for the consumer market or APU’s for the PlayStation/Xbox current gen and more importantly next gen.
In a gold rush as a shovel seller it’s dangerous to only be selling shovels when the gold rush can end/reduce at any point.
AMD has powered 2 generations each of Sony and Xbox consoles, Steam deck and shops a ton of GPUs especially if you count APUs. And then Intel literally ship more GPUs than Nvidia and AMD combined.
The gaming market doesn't need Nvidia. Especially as AAA is cratering.
It really shouldn't: the most money in games are those games that don't require too end GPUs.
Game consoles like PS5 have been AMD for a few generations. Steamdeck/Steam Machine are AMD.
But really, Nvidia has no reason to leave gaming behind. They can just start dialing back their ambition on the gaming side and providing GPUs that aren't too useful for inference or training. All gaming needs is stability so that devs can aim for something. Games looked great 20 years ago and they'll look great 20 years from now, as long as developers know what they are building for.
like can we get some new more interesting entertainment? maybe that's just me but all games look and feel the same now or worse dumbed down sometimes where even puzzles in the game have like basic or grindy aspect to them ...
https://en.wikipedia.org/wiki/Islamic_banking_and_finance
If demand vanishes for the 3 million cards Amazon just bought, then something will be done with them. The AI market may end up in a bizarre jepson's paradox of rotation between inference use cases and model training.
Nvidia booked $496 million in interest income in Q2 alone [1].
[1] https://www.sec.gov/Archives/edgar/data/1045810/000104581026... page 15
Context of the discussion is: Nvidia is central bank of AI.
And then when the kings ran out of money, they sacked the Knights and whoopsies, looks like the knights had no real money as they spent it as quickly as they could maintaining their financial network... Sound familiar?
Let’s look to the past:
https://www.history.com/articles/1929-stock-market-crash-war...
https://portfoliocharts.com/2021/12/16/three-secret-ingredie...
Gold, small cap value, long term treasuries all have done well historically.
Maybe THAT'S the real recession indicator.
The Internet didn't stop expanding in 2000-2001. Everything got drastically larger over the following two decades. The multiples on earnings did implode for ~15 years however. MSFT stock for one example went nowhere during that time and compressed down to a near single digit PE.
AI will be a minimum of 10x larger in most every regard 20 years out. That has nothing to do with shorter-term multiples given to these companies in relation to the hyper fast growth they have been riding early in the boom.
Note that the Fed has a $6.7tn balance sheet [1]. (This is a silly comparison. But still fun.)
The real comparison: Nvidia's $500+ billion of investments and commitments [2] is substantially more than any easing the Fed has done in the same time [3]. Monetarily, Nvidia is creating a lot of money in our economy.
[1] https://www.federalreserve.gov/monetarypolicy/bst_recenttren...
[2] https://www.sec.gov/Archives/edgar/data/1045810/000104581026...
[3] https://www.federalreserve.gov/monetarypolicy/bst_recenttren...
Uh, that’s a pretty load-bearing as long as its cash flows continue. The two things are surely correlated.
It's an important difference. In the GFC, the value of AAA-rated tranches fell. With the benefit of hindsight, we know they continued paying. They were directly leveraged, however, so mark-to-market losses caused firms to fail.
Nvidia stock crashing shouldn't have a similar effect to these commitments. If someone else has massively levered their Nvidia position, they'll obviously blow up. But Nvidia could survive a good deal of equity-market tumult in a way a bank could not.
Ppl have made the prediction of it being a bubble or unsustainable since 2022. At this point, it's hard to say these people have credibility anymore. Ai is big enough, much like Google in 2005 or Facebook/Social Network in 2010 or apps in 2015, that it's an institution unto itself. It's not going to just crash as so many are expecting and have been wrong the past 4 years about.
Also Nvidia isn’t really creating money. The 500B number is third party capital that already exists (BX, Apollo, etc).
Making a loan/offering credit isn't automatically money creation - the amount of money in the system before and after the loan might be the same. Haven't been following Nvidia all that closely, but it seems a little bit unlikely that they're a commercial bank. Financial chicanery they may be doing but offering deposit accounts would be new territory. The loan has to be made in a very particular way for it to be money creation (notably, in a way that creates new money), and it should be illegal for most people to do that otherwise we'd all be printing our own money instead of the printing being directed to wealthy asset owners first and foremost.
All credit has a multiplier effect because borrowers tend to add velocity in a way creditors–sitting on surplus capital–do not. And unlike payment for goods or services, the extension of credit preserves the lent capital plus creates an asset (the loan on the lender's balance sheet).
Banks, being able to create deposits, lend with the highest multiplier effect–they most literally and directly create M1. But the M4 Nvidia creates in commercial paper directly transmutes into M1 through the money-market (M2) and banking (M1) systems in a way that is reliant on Nvidia's guarantees and participation. (Deposits at the Federal Reserve don't directly participate in the real economy. They're money-market objects.)
More broadly: when you open a bar tab, you're literally creating money. When you close it you're destroying it, with real econonomic activity having been occurred in between that didn't exist before.
Or I suppose I should spend more time with an open bar tab come tax time.
The monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet.
Yes.
Nvidia issues a commitment to a firm that lets an SPV raise money. Some of that is just M1 being transferred from one account to another. A lot, however, will be created through loans (new M1), commercial paper (new M3) and the like. Some of that will be pledged as collateral. On the other side of the equation, when the SPV signs e.g. a construction contract, you'll have builders taking out bank loans and issuing commercial paper, et cetera.
You can pay your taxes out of a checking account using money a bank created out of thin air out of a loan against money-market assets.
> monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet
You're absolutely correct in M4 being unequal to the monetary base and deposits at the Federal Reserve. But the moment you say the Fed balance sheet, you're conflating two things. The Fed has lent, in the past, against non-Treasury collateral (most famoulsy, mortgage bonds).
And if we're talking about money in relation to the real economy, you can't spend reserves at the Fed. There is debate on what the most "real" money is, but pretty much everyone agrees that a deposit in a checking account is more relevant to the economy than unspendable reserves at the Fed.
...from a checking account. Banks thoughtlessly lend against investment-grade commercial paper. The kind the SPVs Nvidia is commiting to are borrowing with. Nvidia's promises are creating debt that money markets and banks transmit into M1.
Note that the Fed doesn't directly create M1. It moderates it through the interbank lending market to manipulate the monetary base, M0 plus deposits at the Fed. That, in turn, influences banks' lending patterns which is what manipulates M1 and MZM. Underwriting M4 to drive up collateral that in turn increases M1 is the same mechanism, different channel.
> what is your theory why people pay taxes from their labour earnings instead of just creating new money to pay Just-In-Time?
I'm not Nvidia. I did just sign a purchase order with a builder for my deck that their local bank called me to confirm before issuing a loan that will appear as new money in their checking account. My signature, in that case, enabled the bank to create a tiny amount of money. It's not monetarily significant, however, because my deck isn't that fancy.
At the end of the day, all credit is money. If you can create credit, you can create money. Most of us can't, at least not in significant quantities.
> Isn't M1 is currency and overnight deposits?
M0 + demand deposits at commercial banks, yes.
> sounds highly illegal for Nvidia to create new M1
Nvidia causes the creation of new M1 through banks. Nvidia's commitments directly create money of a quantifiable amount that wouldn't have existed if Huang hadn't flicked his pen.
> unless they've taken out a banking license
Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
> pretty sure you're just making things up here
If that's your reaction to encountering new information, godspeed.
We've got Schrodinger's money here! One minute I can go down to the bar and create as much money as I like with the bartender. Then the next we suddenly discover that yes we need a bank and yes there all of a sudden a checking account is involved and no a bar tab isn't a legally recognised form of money.
> I'm not Nvidia. I did just sign a purchase order with a builder...
Again though, we discover that in fact the person in control of the part where the actual money was created was a bank. You can't just head off with your builder and create money - otherwise you'd be stupid to stop until you unseat whoever was the richest fellow in the world this morning.
> Nvidia causes the creation of new M1 through banks.
And the banks make another appearance!
---
I put it to you that Nvidia can't, in fact, create new money. You'd never get that line of argument past a tax collector.
> Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
And I'll add in postcript that I know nothing about travellers checks in the US, but I'm quietly confident it'd turn out to be illegal to go around creating vast amounts of new money there too, on the basis that people tend to work for a living.
No, its not! Its a swap on the active/left side of the balance sheet: If my mum borrows me 1 USD for icecream, she did not create money, instead she swapped her positions in her assets/left side of her balance sheet.
But aren‘t they currently sitting on like a margin of over 80%? So I assume their revenue will fall, because people are not going to like that in the long-term.
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation ) ---
As working in the field, I know the BOE paper very well ;-)
The reason Nvidia is comfortable making these deals is because if OpenAI can’t use the compute, someone else can.
Granted OpenAI going insolvent likely means a drop in the value of compute…
They would have to go insolvent in a way that hits Nvidia revenue. Those are related by distinct factors, a difference that may matter in a crisis.
I see two factors converging to cause a collapse of this house of cards:
1. People are realizing that what they need isn't more general intelligence, it's more specialization. A small but well tuned coding model, a small but well tuned customer service model, a small but well tuned document explorer.
2. Specialized hardware - TPUs and NPUs - especially coming out of china. The latest GLM model was trained and runs on Huawei hardware. Nvidia is only worth so much because they are the biggest and best provider of the kind of compute needed to run llms, but the export bans mean china has a lot of incentive to topple that monopoly.
Generative video requires significantly more computing power and energy than generative text.
OpenAI is fucked, compute is still needed, it's just them that isn't.
Without a material change in the market (more buyers, vastly cheaper generation), it's unlikely a different company could make that work. More buyers isn't likely to happen, so that leaves vastly cheaper generation - something that would cause nvidia's value to collapse if it happened.
I'd suggest that's only the case given the current quality of output. Media is incredibly expensive to produce. A model capable of sufficiently high quality could charge prices that are absurd by today's standards.
Video generation would only make sense at that scale if it was targeting individual consumers, but then it’d need to cost something that consumers are willing to pay - which practically is probably a few hundred per year at most among US consumers, and much less globally, so again it doesn’t solve for the size of the AI companies.
I don’t see a way that video generation becomes a big industry without making generation much much cheaper.
There's has already been a decade of complaints about trash content in excess of available viewing hours.
The thing with media consumption is that there exist a finite number of eyeball-hours available, and it's a zero sum game against other non-media consumers.
That's already not the case today. If you sat me in front of an LLM and told me to figure out if I'm working with K3 or Astra, I could probably do it, but it would take some work to be certain.
> it's more specialization
China, constrained by hardware, and talent (not to slight the Chinese, but they are limited to domestic resources - and much of the US effort is very international). They did, what the Chinese do, and optimized the process of production, and drastically lowered the cost of development of their models. Cheeper to build, cheaper to run is just good economics.
Meanwhile in the us, we have open AI doing "experiments" - it looks like the costs around the hugging face hack are going to be about the same as China would spend on building out one of their smaller efforts (several million dollars). (Depending on whos numbers you trust, the fact that I can even make this claim should make you raise an eyebrow).
Go back to the 80s' and "expert systems" - most people will tell you that for their time, they were amazing, and useful. People would have loved to have more of them but they were so cost prohibitive that we all but abandoned them for serious use. The US frontier labs seem to have forgotten this lesson and their calls to "slow down" look like an excuse to "cut the waste so we can move to making money".
It’s not quite as simple as that. Several studies have shown the opposite: models trained on more diverse knowledge tend to cross-pollinate across domains. So a more generalized model can actually perform better than a specialized one.
That’s why you’re not seeing tons of tiny models (one for Python, one for Pascal, one for Rust, etc).
But it doesn't match my experience. Qwen3.8 27b is clearly smarter at coding than MANY bigger models. gpt-oss-120b for example, is almost 4x the size, and performs way worse at coding tasks.
It's clear to me that you can build small models that work well at specific tasks.
Python vs Rust is probably too fine grained a way to build a model. Coding in general seems like a better target.
There will always be a place for large generalist models, no doubt. But I think that place is much smaller than the big ai companies are counting on.
I make heavy use of smaller local models on a daily basis (Qwen3-VL for auto-captioning images, Gemma3:27b for some translation work, etc.). Gemma3:27b is a good example of a very capable general purpose multimodal model and has handled almost everything I've thrown at it from sentiment analysis to documentation writing.
I suppose I was drawing a distinction between specialized and general intelligence versus small and large. I don’t think those are necessarily mutually exclusive.
And Qwen3.8-27b is still better at coding than opus 4.1.
Yes, if you list off models 27b is better than it’s all older models. But that’s my point - newer models are better than older models at the same AND much smaller size. That’s because model size matters less than they say. Training data and model architecture matter more.
It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases.
Shows qwen3.8-27b along side seven larger models of ~similar vintage. Only one scores above 27b.
Many of those are closed models so idk their exact parameter count / active param count, but it hardly matters - i’m sure all of them are far above 100b params
My point is not that bigger is pointless. It’s just clearly not the only road to take to make a model better, which is obvious just from seeing how models of the same size have gotten better over the past few years
>speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens. And that's a great trade if you don't have the vram to run the bigger models, but speed is pretty important for getting things done... And that's why DSv4Flash is great, too, despite being much larger, and scoring similarly on the intelligence index.
Absolutely - no argument from me here. Bigger is very clearly a lever you can pull to get more out of a model.
> But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1
Fair point, qwen definitely is slower - it’s a dense model, 27b params, vs a sparse 13b active params model - but the data doesn’t quite agree with what you’re saying about reasoning. I.e.:
> It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens
If you look at the total tokens generated, deepseek thought for 45k tokens and qwen thought for 48k. Barely a difference. The wall clock difference is all down to the speed of token generation, not the amount of reasoning done. At least when we are comparing deepseek and qwen 27b. The comparison swings more towards your position when it comes to the other models on the chart that reason for much fewer tokens.
So perhaps a hypothetical Qwen-27b-a13b could never rival deepseek’s larger model and the tradeoff is one of speed vs overall size - i.e. a small model needs more active params to compete than a big one does.
One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
>One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
Yeah good question, kind of shocking that a 3b active model would perform as well as a 27b dense.
You don’t need to think about climate change studies. Instead you can read the allegedly tainted studies we’re actually talking about and profess to all of us what is wrong with them. You can’t point to exactly where they’ve fudged them.
And let’s be real. If this were some conspiracy that OpenAI, Google, and Anthropic were tacitly or overtly conspiring on, I’m pretty sure Elon or Zucc would ruin it to ahead
Yeah, the cross domain transfer learning from RL is overstated by a lot.
Just like how I could borrow $100 from you, and you could borrow $100 from me, and we'd be $200 in debt in total, I could buy 10% of shares in your company for $1m and you could buy 10% shares in mine, we'd have 2 companies worth $20m together in total.
https://en.wikipedia.org/wiki/Jevons_paradox
And even if compute demand were perfectly elastic it’s only a good thing insofar as it drives demand for new Nvidia hardware. If tokens can be served from Apple hardware or Google hardware or Huawei hardware that doesn’t help Nvidia.
I mean… some homes definitely do. You must have seen those houses that are all lit up front the outside by lawn mounted spotlights.
Across that period, first world countries massively increased their demand for light.
Further efficiency increases only matter to individual choices if they take a use case from {economically impossible} to {economically possible}.
I'd offer that by the 1920s, most goings-on in first world urban environments were no longer price constrained in terms of their light usage.
[0] See table 1.4, p21 https://www.nber.org/system/files/chapters/c6064/c6064.pdf
If that dies because a lot of people’s needs turn out to be met by a system at home they can run a 30b-150b model on, a lot more of that money goes to apple or intel or amd.
1. Yes, smaller models will become more popular, especially as the tokenmaxxing trend dies down and people start stretching their budgets farther. That is a downward pressure on demand.
But along the same dimension, consider that currently only about 40 - 60% of the world uses AI for only about 5 - 15% of their work hours. That means there is still 2x growth from users and 7x - 20x growth from the rest of the work hours left to capture! That is 14x - 40x more demand. Then consider that agentic tasks require multiples more tokens, and that is the kind of usage that is most likely to be deployed, and also the kind of usage that is the least used right now. That's another huge multiple to be tacked on.
And the entire AI industry has been lamenting the extreme compute crunch they're facing (and also why Claude has 9's comparable to GitHub; whereas OpenAI has been chugging along because Altman was OK being called a "podcasting bro" while desperately scrounging for compute years in advance.)
Nvidia's meteoric rise is entirely due to this kind of exploding demand with extremely limited supply.
2. Competing hardware is definitely a threat, but it has its own hurdles. Because the real bottleneck is not Nvidia, it's TSMC.
Pretty much all demand for all chips in all devices in all the world flow to, like, 3 companies in the world that actually fabricate them, and TSMC is the biggest. And the supply is extremely tight, as the exploding costs of electronics clearly shows.
So now TSMC will of course try to keep all its customers happy, but it will inevitably be forced to choose which ones it will keep happiest. And those will be the customers who can pay it the most. And that would be the one with all the money from its de facto status as a monopoly (and possibly even a monopsony)...
Which would be Nvidia ;-)
So yes, compute per task is falling rapidly... but it's barely a dent in the humongous total addressable demand, and the amount of hardware to support that compute is still very constrained, and most of that supply will likely flow through Nvidia.
Ah yes, i am constantly lamenting that my barista isn’t using ai enough ;)
Hopefully you’ve adjusted your ceiling numbers to account for the large amount of people who can’t afford to pay for llms, and will never be able to pay, and aren’t worth it to advertise to since they can afford very little
This is the big point IMO since I have never given $1 to OpenAI but I subscribe to Vidu and Typecast, and have given money to Kling, Hailou, and even Gemini in the form of Google Workspace.
So these other guys have products and use cases, which OpenAI has never been able to crack beyond ChatGPT. And ChatGPT was never worth paying for, IMO.
If OpenAI dies, it's not because there is no market for the technology (which is all NVIDIA cares about), it's more that OpenAI doesn't know how to run a relevant technology company.
They were given everything, not just NVIDIA's billions of dollars and credit backing but all the first-mover advantage, all the respect and credibility early on, so it's really sad to see them unable to develop interesting products and turn a profit in a space they helped pioneer, while so many others are making money with the tech all around them.
NVIDIA is fine. The technology will continue to improve and NVIDIA will stay at the center. OpenAI is fucked - knew it when they retired Sora to focus on text-to-text and coding (a largely solved problem).
All the "frontier" AI companies *are* insolvent. They are cash burning machines.
The only way they keep the lights on and the doors open is by borrowing money --- and lots of it. If those operating the cash spigot decide to turn it off, all AI companies will likely be affected --- and so will Nvidia.
Fairly sure data center construction costs are also going up (they require so many resources that everything is constrained at the moment, especially electricity production).
So I don't understand in what world these frontier AI companies can somehow become profitable. The basic tech they're using is basically the same. Yes, around the edges there are a lot of things that can be done, and were done, like caching, batching, mixture of experts, etc, but basically everyone has done all of that by now, and they're still losing money.
So:
Total costs going up a lot - revenues per unit not increasing proportionally, if anything, Chinese models are forcing those down.
How does that math work out to profits? I don't see it.
Or about as bad, after trillions of dollars in investments over multiple years, let's say the entire frontier AI sector has a total profit of $20bn by 2030. In what world does that make sense? Assuming they can scale that total profit to $100bn in 2035 without investing another cent from 2027 to 2035 (utterly ridiculous), the return on investment would happen in roughly 20 years.
China is the one that is really in the driver's seat here. They have the opportunity and the ability to nullify our huge investment in AI.
Failure to take into consideration those kind of correlations ("If my biggest client isn't able to buy it, I would be able to find someone else who will") is one of the principle causes why many risk models turned out to be garbage during the Great Financial Crisis.
But I also doubt Nvidia is on the hook if OpenAI just no longer wants the compute. I bet they are only on the hook if OpenAI cannot pay for it (is insolvent in some way).
I also have to bring up that OpenAI has already spat out an inference chip that beats Nvidia on flops per watt. So they could potentially not need the compute while other ai companies do.
The problem with that is that OpenAI can only afford to pay for the compute because they are burning investor money (and so are most of OpenAI's biggest clients). They are losing billions. If they stop burning money, nobody else will be there to pay for that compute at OpenAI's cost.
Sure, somebody will probably be able to use these GPUs, they just won't be able to pay nearly as much for them as OpenAI does.
In reality, it's just nowhere near worth as much as OpenAI pays for it. Inflating the cost of compute is part of the problem caused by the circular financing, and if (or maybe when) OpenAI goes, the price of compute will go with them.
But that’s the point. Investors believe investment in AI will pay off.
How? The hardware is in OpenAI's datacenters. Does Nvidia have a couple hundred semi trucks, contractors, and IT technicians, to repo the hardware and resell it to someone else before it's lost most of its value? These chips will be replaced approx every 3-4 years. So if OpenAI tanks, after Nvidia pays for and waits for the process to collect the hardware, they then have to sell it for pennies on the dollar. They lose almost all the investment.
Also consider that SpaceXAI already had datacenters full of gear that they basically weren't using because nobody wanted their product, so they now rent it to Anthropic. The demand for hardware isn't really there at the scale of OpenAI.
Via logistics contracts, yes.
Gaming is lucrative but not THAT lucrative…
They have no idea how much compute will be needed. They are estimating and trying to push it but they do not know the future.
Load bearing, heavy lifting... Your comment wasn't LLM-written, either. I think we're starting to see LLMisms infect human writing. I might try to start speaking like this and see if anyone notices. It could be a good gag.
That’s what loans are.
https://en.wikipedia.org/wiki/Money_creation
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation )
No, it's not. M2 includes things like traveler's cheques and money-market funds.
Fair enough. Money-market assets are the real exception.
Nvidia doesn't need it. It funds other companies. They do this thing that Nvidia doesn't do. It shows up on their balance sheets and Nvidia just gets to claim the valuation of the investment on its balance sheet.
It can't go tits up!
What is the barrier to entry?
versus
Fed open market operations that occur in full instantly and are the primary mechanism for increasing the money supply
k.
nvidia's money either comes from equity, or from debt. This means this money is "created" not from nothing, but from future commitments (ala, debt repayments, or promise of profits), which is not the same as what comes out of the Fed.
The Fed printing has no backing behind it - it is definitely inflationary if they do it. Investment from nvidia (or any other company) "creating money" may be productive enough to completely offset their inflationary effects - after all, the company investing demands returns from their investments, and so will only invest in things they expect to return much higher than the cost of interest (or cost of capital).
Therefore, you cannot compare Fed printing money to company investing money.
Most money is created by banks, not by the Fed. When banks create deposits they're booking it against a loan. Similarly, the Fed creates money by buying assets, principally Treasuries. There is an offsetting account. The only party that can truly just "mint" currency is the US Mint.
> you cannot compare Fed printing money to company investing money
Yes, you can. Nvidia creates M4 which drives M2 and thus M1. Banks create M1. The Fed creates monetary base. These are different components of the same money supply [1].
[1] https://en.wikipedia.org/wiki/Money_supply#United_States
NVDA itself(!) does not increase M4 - if they give a loan, its an assets swap on their balance sheet with no influence on money in total available (since the money on their account is gone!)
The FED and banks can increase total money only.
Oh hell no. NVIDIA is present in a ton of 401k and similar retirement/savings/gambling vehicles. It holds 12% of NASDAQ tracking funds and 8% of S&P 500 trackers. NVIDIA goes down, a lot of people will panic sell and crash the economy.
Wrong:
a) a company cannot create money in todays modern 2-level fiat systems, this can be done only by central banks
b) what NVDA does is "generating requests for more money in the system", throughthe demand side (either taking loans themself or having lots of customers who are willing to take a bank loans to invest, or other participants who wants to buy whatever stuff on a loan)
c) any money that is created on the "local bank system level" needs to be there on the existing central bank account when the money is leaving the local system, ie. transferred to another bank (like "paying NVDA bill for delivery of some chips")
The ideas we deal with when we discuss society and organization aren't exclusive to government, they relate to human nature in general. I wonder if in the future we will have more discussion of power and how to organize it in corporations, similar to what we discuss today about government.
The structure of the Fed is setup the way it is to limit the sort of self-serving, myopic political micromanaging that could be damaging to the economy at large.
A similar structure would potentially be very undesirable to Nvidia shareholders as caution over long time horizons would likely produce what they would consider an excessively conservative, defensive strategy to avoid putting too much air into the bubble too quickly (at the expense of their valuation).
The Fed on the other hand has (historically ) tried to identify potential indicators warning of unsustainable bubbles that could lead to financial contagion and tries to mitigate that risk using the limited monetary tools available and their public soapbox.
I'm not advocating for corporations to have exactly the same rules and structure as government does, but perhaps many of the ideas used to design governments can be borrowed.
Just such a curious scenario to let your mind wander about, how society would look like in these scenarios!
/s
I'm honestly so sick of the suspense of disbelief on this site, how is this more "interesting" to you, than the absolute sheer terror you should feel about going back to feudalism and serfdom? A typical western national state ensures that you have basic rights as a human being and aren't exploited to the death by non-government entities.
Looking at any organization with power and people involved, though, perhaps we can apply the same ideas that traditionally apply to government to more places. That's all I'm proposing.
And perhaps this could even be a solution to the problems you bring up--Apple for example is a company with significant power and influence. What if their CEO, or board, was in some way more democratically elected? Or their executives split into separate branches to create a separation of powers? Would that make these corporations more stable, more accountable, more moral? Maybe these ideas can help us create a better relationship with corporate power.
Regardless, I'm not trying to say in any way that corporations should have more power, or make any statement encouraging the current situation. I'm just saying that the organization of power within private bodies, separate from how much of it we allocate where publicly, is worth thinking about.
In a country without religion, banner or ideology to unite the people in current-and-coming turbulent times, the bet is made on "unite under money, or have no money left"
the way to fight it is to be principled even in front of cheaper options - and to support others like you
https://www.youtube.com/watch?v=A64rR5Dp07s
"You have meddled with the primal forces of nature, Mr. Beale. And I won't have it!
Is that clear?! You think you've merely stopped a business deal. That is not the case. The Arabs have taken billions of dollars out of this country, and now they must put it back! It is ebb and flow, tidal gravity! It is ecological balance!
Am I getting through to you, Mr. Beale?
You get up on your little twenty-one inch screen and howl about America and democracy. There is no America. There is no democracy. There is only IBM and ITT and AT&T and DuPont, Dow, Union Carbide, and Exxon. Those are the nations of the world today.
What do you think the Russians talk about in their councils of state -- Karl Marx? They get out their linear programming charts, statistical decision theories, minimax solutions, and compute the price-cost probabilities of their transactions and investments, just like we do."
Sadly it seems like some haven’t watched the end of the last movie on this subject.
Funding companies under the condition that they use their infra.
Also called: Buying customers.
I thought a central bank would be like: China is the world's largest official creditor and holds the highest foreign exchange reserves.
That's what makes you a naturally forming central bank.
Companies just don't want to pay Jensen's tax. Hyperscalers might still pay Jensen's tax for LLM training but for inference. you don't have to. they are also betting on their own chip for training to replace Nvidia.
this is Nvidia panicking and doing vendor fiance to Neocloud and buying Hugging Face. even none hyperscalers like Meta is betting on its own chip for AI inference.
If these cards burn out in less than the ~5 years of depreciation that accounting puts them at, well then there will be problems.
https://www.cnbc.com/video/2026/08/24/making-old-gpus-new-ag...
Imagine if running fable costs you what it actually costs to run fable. A lot of vibe coders (and just proper software engineers) are gonna be very sad if that comes to pass.
But that reminds me of the story of the union rep telling Ford: “good luck getting your machines to buy your cars”
(Fyi Ford took note and started paying his workers enough that they’d buy his cars)
https://en.wikipedia.org/wiki/Economic_growth
If Nvidia is the bank, they should be starting to sweat a bit. It's not often companies ask for a voluntary slowdown.
Not just building but actively testing
If you have a genuinely-held belief that AI systems will end life on earth and you willfully continuing working on them anyway, then, like... you probably shouldn't be allowed to walk free in civil society anymore, right?
The first messenger from Anthropic is out with this exact message. Stop us (them) or everyone will be killed by 2030
https://youtu.be/i30jVPqQeOM?si=yosxw4MLoVtPW-7v
Predicting a usefulness plateau is absolutely wild given how fast AI agents have been improving at writing code this year. I have doubts about AGI but I think you’re making assumptions and translating it wrong. These two companies have always been asking for a slowdown from their inception, that’s not a new thing. It’s part marketing hype, but they both do want regulation to step in and slow down the competition, not because they see a usefulness plateau, but the opposite - the usefulness is growing so fast that they want to remain in control, and they are scared that working hard and competing will not be enough. Anthropic has also said out loud they think their competition (not just OpenAI) is not being responsible and they want the regulation so they can be the responsible shepherd, as AI gets more and more useful.
Have the models improved since Opus 4.x? I find the newer models are not better in my day job, maybe in one shotting mvp's and other tasks.
Not trying to argue your point, just intrested in the coding aspect, if the models were improving as fast as benchmarks I would expect capability improvements to be obvious, but talking to people and reading forums, it seems everyone has a different opinion.
Also yes, Fable is massively better than Opus. It requires significantly less instruction and specs and produces more directly mergeable code.
The improvement is obvious as soon as my Fable allotment runs out and I try to do something with Opus. Have you given the same (larger) task to Opus and Fable?
Personally I still use mainly Opus and find it handles most tasks quite well (without burning all my corporate quota).
https://x.com/theo/status/2097192907023458473
(^ this "swearing at a model" thing has happened to me multiple times on Astra already)
If you have to "debate" the quality of a new Big Number model (and double and triple check your eyes and model setting switches when it pukes up complete garbage), that is NOT a good sign.
This is a ridiculous claim that is disproven by simply using it for more than 5 minutes.
https://withspecific.com/benchmarks/real-swe
I’ve never thought LLMs were on the AGI path, but I have to admit it’s surprising how far it’s come with no end in sight yet. There is something important to be said about how ‘intelligence’ is embedded in language, and it suggests that intelligence isn’t exactly what we thought it was. The language component of intelligence also goes a long way to explaining technology’s progress in human civilization; how language and the printing press and mail and radio/tv and the internet have each ushered in accelerations in the pace of progress. Biologically and evolutionarily speaking, it’s unlikely that humans have become any smarter in the last two thousand years, but technology (among other things) has exploded.
So whatever they're selling: LLM will do the jobs of every white collar worker isn't in the foreseeable future.
Not necessarily. Could be it's close and they are concerned about issues with it?
Even being cynical, maybe Dario thinks OpenAI's Astra is pretty much AGI and wants to slow things so Anthropic can catch up?
By AGI here I'm thinking when you can say to the model, go build a better model, and lay off the engineers.
What planet are you living on? How is this opposite-of-true nonsense being upvoted on HN?
The rate of advancement of models is exponential.
But is the improvement truly exponential??
Are the improvements we have seen from September 2025 to September 2026 really as significant as the improvements from September 2024 to September 2025?
I’m genuinely asking. I can’t answer myself because I haven’t been pushing the latest models to the limit with every new release.
Luckily, AGI is pretty boring. It's a tool that does it's job. It seems like very wishful thinking to expect the same of ASI.