Wow did you encapsulate millenia of management-labor disputes by saying don't worry be happy?
Let's play the same game with totalitarianism!
It's the fear they are watching everything
It's the fear nobody is watching at all
Oh wow, I totally understand the threat of totalitarianism from that.
And I bring up totalitarianism quite in particular, because aside from vastly empowering the elites in the war against labor, AI vastly empowers the elites for totalitarian monitoring and control.
> You didn’t need a bar chart to recognize that GPT-4 had leaped ahead of anything that had come before.
You did though. I remember when GPT-4 was announced, OpenAI downplayed it and Altman said the difference was subtle and wouldn't be immediately apparent. For a lot of the stuff ChatGPT was being used for the gap between 3 and 4 wasn't going to really leap out at you.
In the lead up to the announcement, Altman has set the bar low by suggesting people will be disappointed and telling his Twitter followers that “we really appreciate feedback on its shortcomings.”
OpenAI described the distinction between GPT-3.5—the previous version of the technology—and GPT 4, as subtle in situations when users are having a “casual conversation” with the technology. “The difference comes out when the complexity of the task reaches a sufficient threshold—GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5,” a research blog post read.
In the years since we got a lot more demanding of our models. Back then people were happy if they got models to write a small simple function and it worked. Now they expect models to manipulate large production codebases and get it right first time. So, the difference between GPT-3 and GPT-4 would be more apparent. But at the time, the reaction was somewhat muted.
The title is irritating, conflating AI with LLMs. LLMs are a subset of AI. I expect future systems will be mobs of expert AI agents rather than relying on LLMs to do everything. An LLM will likely be in the mix for at least the natural language processing but I wouldn't bet the farm on them alone.
The title annoys me more because if doesn't mention anything about time. AI will almost certainly get a good bit better eventually. The questions will it in the next couple of years or will we have to wait for some breakthrough.
I'm amused they seem to refer to Marcus and Zitron as "these moderate views of A.I". They are both pretty much professional skeptics who seem to fill their days writing AI is rubbish articles.
AI is so new and so powerful, that we don't really know how to use it yet. The next step is orchestration. LLMs are already powerful but they need to be scaled horizontally. "One shotting" something with a single call to an LLM should never be expected to work. That's not how the human brain works. We iterate, we collaborate with others, we reflect... We've already unlocked the hard and "mysterious" part, now we just need time to orchestrate and network it.
Powerful but we don't know how to use it?
If it is as powerful as all you true believers spout the usefulness would be self evident and that would be the display of its power.
But apparently it is powerful just because you say so, and then something, something ... business model ...
OpenAi has 700+ million users. Sam recently said only 7% of Plus users were using thinking (o3)!!! That means 93% of their users were using nothing but 4o!
Clearly the OpenAi leadership saw these stats and understood the main initial goal of GPT5 is to introduce this auto-router, and not go all in on intelligence for the 3-7% who care to use it.
This is a genius move IMO, and will get tons of users to flood to ChatGPT over competitors. Grok, Gemini, etc are now fighting over scraps of the top 1% while OpenAi is going after the blue ocean of users.
What happens is they go out of business: "these firms spent five hundred and sixty billion dollars on A.I.-related capital expenditures in the past eighteen months, while their A.I. revenues were only about thirty-five billion."
DeepSeek (and the like) will prevent the kind of price increases necessary for them to pay back hundreds of billions of dollars already spent, much less pay for more. If they don't find a way to make LLMs do significantly more than they do thus far, and a market willing to pay hundreds of billions of dollars for them to do it, and some kind of "moat" to prevent DeepSeek and the like from undercutting them, they will collapse under the weight of their own expenses.
My current intuition on this topic is that they are right about scaling but they are training on the wrong data.
LLMs were not intended to be the core foundation of artificial intelligence but an experiment around deep learning and language. Its success was an almost accidental byproduct of the availability of large amount of structured data to train from and the natural human bias to be tricked by language (Eliza effect).
But human language itself is quite weak from a cognitive perspective and we end up with an extremely broad but shallow and brittle model. The recent and extremely costly attempts to build reasoning around don't seem much more promising than using a lot of hardcoded heuristics, basically ignoring the bitter lesson.
I've seen many argue that a real human level AI should be trained from real-world experience, I am not sure this is true, but training should likely start from lower-level data than language, still using tokens and huge scale, and probably deeper networks.
They didn't answer much of the "What if," though... Am just imagining the massive financial losses taken by so many, and if a bailout becomes necessary, because too-big-to-fail now means Microsoft, Google, Facebook et al since we transferred so much of financial engineering economics onto them since '08.
AI doesn't need to get better than this. It is already saving millions of hours of previously wasted human productivity. The biggest threat to these companies if their products do not improve is the local running of LLMS. That would finally justify consumers buying more memory and processor speed.
It appears that Cal Newport has decided to be the one to most publicly initiate the inevitable Trough Of Disillusionment stage of the Hype Cycle. I'm not sure it'll last very long, though, considering (for starters) Google DeepMind's gold medal at the recent International Math Olympiad. Also, while he criticizes the cost-cutting measure which is ChatGPT 5, he doesn't even mention ChatGPT 5 Pro, which is performing excellently.
AI getting better is like maybe 50% or less of the equation. The other part is the infrastructure supporting AI applications. The infrastructure and interfaces that need to be built to fully take advantage of whats already here already has a long way to catch up.
I always expect things like this to eventually deliver about 90% of what they promise, which turns out to be 100% for some niche uses, and for the rest it just gets abandoned because 90% isn't good enough. Like when voice recognition became super hyped in the late 90's, it was going to change how the world interacts with machines, and eventually it turned into "Hey Siri"
We did a test of GPT5 yesterday. We asked it to generate a synopsis of a scientific topic and cite sources. We then checked those sources. GPT5 still hallucinated 65% of the citations.
It did things like:
Make up the paper title
Make up the authors for a real paper title
Mix a real title and a real journal
If it can't even reference real papers it certainly can't be trusted to match up claims of fact with real sources.
Current AI tools generate citations that LOOK real but ARE fake. This might not be solvable inside the LLM. If anyone could do it, it'd be OpenAI. (OK maybe I'm giving them too much credit, but they have a crap-ton of money and seem to show a real interest in making their AI better)
If it can't be done in the LLM we can't trust LLMs basically ever.
I suppose there's a pretty big loophole here. Doing it outside the LLM but INSIDE the LLM product would be good enough.
The first AI tool to incorporate that (internal citation and claim checking) will win because if the AI can check itself and prevent hallucinated garbage from ever reaching the user we can start to trust them and then they can do everything we've been promised.
Until that day comes we can't trust them for anything.
Google already did this, give free gemini deepresearch a spin. It's not perfect, but I have a feeling you'll be surprised if this is your honest impression.
28 comments
[ 2.2 ms ] story [ 58.0 ms ] threadWhat if it does?
There's a certain type of fear . . .
-- David FahlSame fear, different day.
Let's play the same game with totalitarianism!
It's the fear they are watching everything
It's the fear nobody is watching at all
Oh wow, I totally understand the threat of totalitarianism from that.
And I bring up totalitarianism quite in particular, because aside from vastly empowering the elites in the war against labor, AI vastly empowers the elites for totalitarian monitoring and control.
You did though. I remember when GPT-4 was announced, OpenAI downplayed it and Altman said the difference was subtle and wouldn't be immediately apparent. For a lot of the stuff ChatGPT was being used for the gap between 3 and 4 wasn't going to really leap out at you.
https://fortune.com/2023/03/14/openai-releases-gpt-4-improve...
In the lead up to the announcement, Altman has set the bar low by suggesting people will be disappointed and telling his Twitter followers that “we really appreciate feedback on its shortcomings.”
OpenAI described the distinction between GPT-3.5—the previous version of the technology—and GPT 4, as subtle in situations when users are having a “casual conversation” with the technology. “The difference comes out when the complexity of the task reaches a sufficient threshold—GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5,” a research blog post read.
In the years since we got a lot more demanding of our models. Back then people were happy if they got models to write a small simple function and it worked. Now they expect models to manipulate large production codebases and get it right first time. So, the difference between GPT-3 and GPT-4 would be more apparent. But at the time, the reaction was somewhat muted.
I'm amused they seem to refer to Marcus and Zitron as "these moderate views of A.I". They are both pretty much professional skeptics who seem to fill their days writing AI is rubbish articles.
But apparently it is powerful just because you say so, and then something, something ... business model ...
Clearly the OpenAi leadership saw these stats and understood the main initial goal of GPT5 is to introduce this auto-router, and not go all in on intelligence for the 3-7% who care to use it.
This is a genius move IMO, and will get tons of users to flood to ChatGPT over competitors. Grok, Gemini, etc are now fighting over scraps of the top 1% while OpenAi is going after the blue ocean of users.
DeepSeek (and the like) will prevent the kind of price increases necessary for them to pay back hundreds of billions of dollars already spent, much less pay for more. If they don't find a way to make LLMs do significantly more than they do thus far, and a market willing to pay hundreds of billions of dollars for them to do it, and some kind of "moat" to prevent DeepSeek and the like from undercutting them, they will collapse under the weight of their own expenses.
LLMs were not intended to be the core foundation of artificial intelligence but an experiment around deep learning and language. Its success was an almost accidental byproduct of the availability of large amount of structured data to train from and the natural human bias to be tricked by language (Eliza effect).
But human language itself is quite weak from a cognitive perspective and we end up with an extremely broad but shallow and brittle model. The recent and extremely costly attempts to build reasoning around don't seem much more promising than using a lot of hardcoded heuristics, basically ignoring the bitter lesson.
I've seen many argue that a real human level AI should be trained from real-world experience, I am not sure this is true, but training should likely start from lower-level data than language, still using tokens and huge scale, and probably deeper networks.
https://www.currentmarketvaluation.com/models/s&p500-mean-re...
https://www.cell.com/fulltext/S0092-8674(00)80089-6
How can you say progress has stalled without having visibility on the compute costs of gpt-5 relative to o3?
How can you say progress has stalled by referring to changes in benchmarks at the frontier over just 3.5 months?
In the one side I read stuff about exponential gains with every new model. On the other side, the coding improvements look logarithmic to me.
Current AI tools generate citations that LOOK real but ARE fake. This might not be solvable inside the LLM. If anyone could do it, it'd be OpenAI. (OK maybe I'm giving them too much credit, but they have a crap-ton of money and seem to show a real interest in making their AI better)
If it can't be done in the LLM we can't trust LLMs basically ever. I suppose there's a pretty big loophole here. Doing it outside the LLM but INSIDE the LLM product would be good enough.
The first AI tool to incorporate that (internal citation and claim checking) will win because if the AI can check itself and prevent hallucinated garbage from ever reaching the user we can start to trust them and then they can do everything we've been promised. Until that day comes we can't trust them for anything.