Buckmaster (the mathematician) and Alpöge used and credit AI substantially for their proof. Even if OpenAI did copy their ideas, it still wouldn't show that this didn't come from AI improving. OpenAI's proof is…
You're conflating two very different meanings of "the speed of light". Confusingly, "speed of light" can refer to the universal constant, c, which does not change in glass, or to the speed that light travels in a…
Why? 1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used…
There are several hundred thousand mathematicians producing hundreds of thousands of new results in math each year. Why can't most people name any human contributions to mathematics from the past decade? What is the…
The Kevin Buzzard post linked at the top says they budgeted £1M over 5 years for a smaller proof.
[dead]
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's…
I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.
The last version to fail on those questions was GPT 4.5. Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the…
It also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home) The test wasn't made to accurately measure IQs that high.
The paper also fails to show that their central example, Einstein, relied on sensory experience for his intuition leaps rather than general reasoning. They just kind of claim that thought experiments require sensory…
Only a tiny, tiny fraction of the parameters are encoding information that's specific to a particular programming language. Even if you could remove those without degrading performance, it would have a negligible effect…
The problem is Claude Fable is now better than most programmers I know at software architecture and performance optimization as well.
That's a good reason to use a dishwasher, and you should keep doing it. But the overall waste is small and people aren't going to care. I'm not saying people who already have a dishwasher will throw it away. But a robot…
Sure, but then there's no such thing as a network that isn't a classifier. Every physically computable function that terminates in finite time will map an input to a fixed set of outputs. And it goes against the common…
The LLM has processed two data modalities derived from the apple (text and vision). Your brain processed a third (taste). But it is still just a data stream, sensing compounds and chemical properties of the apple and…
LLMs are not classifiers. A classifier is an algorithm or neural net that assigns a label from a fixed set of labels to an input. You can broaden the definition of classifier to anything that internally divides its…
I have a dishwasher and I still usually just hand wash. It takes about 10 seconds to wash a dish. The side benefit is all your dishes are always available. With the dishwasher, up to one full dishwasher load are dirty…
Unitree's R1 humanoid robot is only about $6000, and it's still a nascent, smallish scale technology. They will come down. If the future home robots are any good, it saves you from buying a dishwasher and robot vacuum.…
Yes, much like that. If he hadn't destroyed evidence, he could have argued it was malicious prosecution.
What you're describing is malicious prosecution or abuse of process. It's illegal and it would destroy the prosecution's case. Not only that, but the victim could sue for damages.
The irony is in this case the in-context and classifier "guardrails" would have almost certainly stopped the attack while their attempts at your definition of guardrails (the sandboxing) failed. In general, people keep…
It's a strange experiment. Claude and GPT aren't generating the video. They're directing and editing it, and they request video from a generative video model using mainly text-to-video. Neither Claude nor GPT can…
The UNESCO/World Bank literacy rate is basically defined how you thought. But high income countries don't usually report this because literacy by this measure is nearly universal. So they often report at higher…
Questions like that cost a tiny fraction of a cent. "What's the capital of Sri Lanka?" cost a fifth of a cent at GPT 5.5 API price, and would cost a fraction of that if the question were routed to a more suitable,…
Buckmaster (the mathematician) and Alpöge used and credit AI substantially for their proof. Even if OpenAI did copy their ideas, it still wouldn't show that this didn't come from AI improving. OpenAI's proof is…
You're conflating two very different meanings of "the speed of light". Confusingly, "speed of light" can refer to the universal constant, c, which does not change in glass, or to the speed that light travels in a…
Why? 1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used…
There are several hundred thousand mathematicians producing hundreds of thousands of new results in math each year. Why can't most people name any human contributions to mathematics from the past decade? What is the…
The Kevin Buzzard post linked at the top says they budgeted £1M over 5 years for a smaller proof.
[dead]
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's…
I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.
The last version to fail on those questions was GPT 4.5. Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the…
It also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home) The test wasn't made to accurately measure IQs that high.
The paper also fails to show that their central example, Einstein, relied on sensory experience for his intuition leaps rather than general reasoning. They just kind of claim that thought experiments require sensory…
Only a tiny, tiny fraction of the parameters are encoding information that's specific to a particular programming language. Even if you could remove those without degrading performance, it would have a negligible effect…
The problem is Claude Fable is now better than most programmers I know at software architecture and performance optimization as well.
That's a good reason to use a dishwasher, and you should keep doing it. But the overall waste is small and people aren't going to care. I'm not saying people who already have a dishwasher will throw it away. But a robot…
Sure, but then there's no such thing as a network that isn't a classifier. Every physically computable function that terminates in finite time will map an input to a fixed set of outputs. And it goes against the common…
The LLM has processed two data modalities derived from the apple (text and vision). Your brain processed a third (taste). But it is still just a data stream, sensing compounds and chemical properties of the apple and…
LLMs are not classifiers. A classifier is an algorithm or neural net that assigns a label from a fixed set of labels to an input. You can broaden the definition of classifier to anything that internally divides its…
I have a dishwasher and I still usually just hand wash. It takes about 10 seconds to wash a dish. The side benefit is all your dishes are always available. With the dishwasher, up to one full dishwasher load are dirty…
Unitree's R1 humanoid robot is only about $6000, and it's still a nascent, smallish scale technology. They will come down. If the future home robots are any good, it saves you from buying a dishwasher and robot vacuum.…
Yes, much like that. If he hadn't destroyed evidence, he could have argued it was malicious prosecution.
What you're describing is malicious prosecution or abuse of process. It's illegal and it would destroy the prosecution's case. Not only that, but the victim could sue for damages.
The irony is in this case the in-context and classifier "guardrails" would have almost certainly stopped the attack while their attempts at your definition of guardrails (the sandboxing) failed. In general, people keep…
It's a strange experiment. Claude and GPT aren't generating the video. They're directing and editing it, and they request video from a generative video model using mainly text-to-video. Neither Claude nor GPT can…
The UNESCO/World Bank literacy rate is basically defined how you thought. But high income countries don't usually report this because literacy by this measure is nearly universal. So they often report at higher…
Questions like that cost a tiny fraction of a cent. "What's the capital of Sri Lanka?" cost a fifth of a cent at GPT 5.5 API price, and would cost a fraction of that if the question were routed to a more suitable,…