Employees at the companies that do this stuff know exactly what these regulations are for and are violating them quite deliberately. They don't actually think it's "government bureaucracy" whatever they tell the public.
The article claims that "for the first time, string theory is testable," when in reality: * the tests here concern particles that aren't actually known to exist yet * lots of variants of string theory can be invalidated…
Yeah people don't seem to get that the whole point of having a tightly checked kernel is so you don't have to care so much about the rest of it. Tactic heavy proofs have been "slop" long before LLMs got involved, and…
Yes, clearly it's completely absurd to think that shredding rare books to more cheaply train LLMs is anything but a moral good, which is why there is this entire thread is full of people jumping through hoops to explain…
I think maybe a better way of explaining it would be that an uninformative proof by definition needs to be based on proving that the set under consideration must be inhabited without ever defining an object in that set.…
Ah, I didn't realize this was a generational thing. I am definitely a "new" intuitionist, so that probably greatly influences my perspective. I suppose that before results like this, the setoid model, etc. were known…
A direct counterexample is more "informative" in a very literal sense (its truth value doesn't collapse). But the extra proof relevant content we can use here is not that large -- all it means in this case is that we…
I think the constructive position is basically that people's entire issue with lack of excluded middle being absent is just that people like being able to say "P" instead of "~~P" because it sounds better, considering…
As a constructivist: we don't disagree :) We just distinguish between "don't disagree" and "agree." Constructive mathematics says it's fine if you want to claim that there's not no counterexample -- you just can't use…
You can go through my commenter history and know I'm no fan of LLMs. I don't overstate LLM capabilities and am highly skeptical of them in general. 5.6 Pro is genuinely pretty good at certain kinds of math problems that…
> Why is it only GPT doing this, why not Claude? Because Claude can't do it. Anyone who tells you that Fable is better than GPT 5.6 at pure math is lying to you.
If you really believe this, try to use GPT 5.6 to prove an open problem you know nothing about. You might get lucky, but if you don't, you will soon discover that 5.6 can make "progress" towards a theorem without…
Oh boy this is a cool blog post. Encourage everyone to read it.
I'm honestly not familiar enough with how well-developed graph theory is in Lean to be able to say. The paper is mostly using pretty old results, so it's mostly a matter of whether that stuff has already been formalized…
"But LLMs are prone to hallucinations which can really impact a string of interdependent logic like a proof. So I’m assuming it would respond with something that’s not complete nonsense to this proof most of the time."…
I absolutely think that with the rise of LLM generated theorems we need mechanization more than ever, yeah. But I felt that was already pretty important for human proofs, too, and people are just more amenable to the…
Frontier labs have had multiple major announcements in the past about supposedly novel LLM generated theorems that turned out to be vastly overstating what actually happened. That's part of why they were so…
As someone who's used proof checkers a fair amount, if you don't have some high level idea about the proof, it's an open problem, and the hard part isn't some extremely tedious finite case analysis, it's extremely…
Yeah it's a very very short proof that uses no mathematics developed within the last 30 years. Which doesn't necessarily make it wrong, but in the absence of mechanization in Lean or proper peer review I think this it…
Mostly true on Earth, but not on other planets with lower gravity, and AFAIK it depends on the rock type. Hence why you have Olympus Mons on Mars (or insanely tall ice mountains on Pluto, when that material couldn't…
"Less healthy" is distinct from "shortens lifespan" which was pretty much the entire point I was making. I understand that the idea that there are well-controlled massive studies with enough power to detect differences…
Yeah confirmed with more usage today. It seems ~ on par with (if not slightly worse than) 5.5 on math-oriented stuff.
Sounds like absolute BS to me. Even in very large scale studies specifically designed for studying mortality, only morbid obesity has been negatively correlated with lifespan. There is even some evidence that being a…
Not super impressed, but I doubt my requests are getting routed to Opus -- it just doesn't seem to be as good at mathematics as it is at code (I found this to be the case last time it was released as well).
Given their "shape stability" design, not necessarily. The three ways that multithreaded access can cause UB are: * changing the type of the underlying memory (e.g. because it's part of an enum variant and you changed…
Employees at the companies that do this stuff know exactly what these regulations are for and are violating them quite deliberately. They don't actually think it's "government bureaucracy" whatever they tell the public.
The article claims that "for the first time, string theory is testable," when in reality: * the tests here concern particles that aren't actually known to exist yet * lots of variants of string theory can be invalidated…
Yeah people don't seem to get that the whole point of having a tightly checked kernel is so you don't have to care so much about the rest of it. Tactic heavy proofs have been "slop" long before LLMs got involved, and…
Yes, clearly it's completely absurd to think that shredding rare books to more cheaply train LLMs is anything but a moral good, which is why there is this entire thread is full of people jumping through hoops to explain…
I think maybe a better way of explaining it would be that an uninformative proof by definition needs to be based on proving that the set under consideration must be inhabited without ever defining an object in that set.…
Ah, I didn't realize this was a generational thing. I am definitely a "new" intuitionist, so that probably greatly influences my perspective. I suppose that before results like this, the setoid model, etc. were known…
A direct counterexample is more "informative" in a very literal sense (its truth value doesn't collapse). But the extra proof relevant content we can use here is not that large -- all it means in this case is that we…
I think the constructive position is basically that people's entire issue with lack of excluded middle being absent is just that people like being able to say "P" instead of "~~P" because it sounds better, considering…
As a constructivist: we don't disagree :) We just distinguish between "don't disagree" and "agree." Constructive mathematics says it's fine if you want to claim that there's not no counterexample -- you just can't use…
You can go through my commenter history and know I'm no fan of LLMs. I don't overstate LLM capabilities and am highly skeptical of them in general. 5.6 Pro is genuinely pretty good at certain kinds of math problems that…
> Why is it only GPT doing this, why not Claude? Because Claude can't do it. Anyone who tells you that Fable is better than GPT 5.6 at pure math is lying to you.
If you really believe this, try to use GPT 5.6 to prove an open problem you know nothing about. You might get lucky, but if you don't, you will soon discover that 5.6 can make "progress" towards a theorem without…
Oh boy this is a cool blog post. Encourage everyone to read it.
I'm honestly not familiar enough with how well-developed graph theory is in Lean to be able to say. The paper is mostly using pretty old results, so it's mostly a matter of whether that stuff has already been formalized…
"But LLMs are prone to hallucinations which can really impact a string of interdependent logic like a proof. So I’m assuming it would respond with something that’s not complete nonsense to this proof most of the time."…
I absolutely think that with the rise of LLM generated theorems we need mechanization more than ever, yeah. But I felt that was already pretty important for human proofs, too, and people are just more amenable to the…
Frontier labs have had multiple major announcements in the past about supposedly novel LLM generated theorems that turned out to be vastly overstating what actually happened. That's part of why they were so…
As someone who's used proof checkers a fair amount, if you don't have some high level idea about the proof, it's an open problem, and the hard part isn't some extremely tedious finite case analysis, it's extremely…
Yeah it's a very very short proof that uses no mathematics developed within the last 30 years. Which doesn't necessarily make it wrong, but in the absence of mechanization in Lean or proper peer review I think this it…
Mostly true on Earth, but not on other planets with lower gravity, and AFAIK it depends on the rock type. Hence why you have Olympus Mons on Mars (or insanely tall ice mountains on Pluto, when that material couldn't…
"Less healthy" is distinct from "shortens lifespan" which was pretty much the entire point I was making. I understand that the idea that there are well-controlled massive studies with enough power to detect differences…
Yeah confirmed with more usage today. It seems ~ on par with (if not slightly worse than) 5.5 on math-oriented stuff.
Sounds like absolute BS to me. Even in very large scale studies specifically designed for studying mortality, only morbid obesity has been negatively correlated with lifespan. There is even some evidence that being a…
Not super impressed, but I doubt my requests are getting routed to Opus -- it just doesn't seem to be as good at mathematics as it is at code (I found this to be the case last time it was released as well).
Given their "shape stability" design, not necessarily. The three ways that multithreaded access can cause UB are: * changing the type of the underlying memory (e.g. because it's part of an enum variant and you changed…