what are you even yapping about.
Adding the position vector is basic sure, but it's naive to think the model doesn't develop its own positional system bootstrapping on top of the barebones one.
fmri's are correlational nonsense (see Brainwashed, for example) and so are any "model introspection" tools.
So you think that this blog post would make it into any of the mainstream conferences? I doubt it.
peer review would encourage less hand wavy language and more precise claims. They would penalize the authors for bringing up bizarre analogies to physics concepts for seemingly no reason. They would criticize the fact…
didn't hackers used to be for piracy?
This suggests people should pre-register benchmarks. Because currently it feels like there is little incentive to publish benchmarks that models saturate.
I completely agree this shit is so depressing. When I saw the AlphaProof paper I basically spent 3 days in mourning basically, because their approach was so simple.
I think the whole paper is a satire lol.
Does it really? If you want an LLM to edit code you need to feed it every single line of code in a prompt. Is it really that surprising that having just learnt it has been timed out, and then seeing code that has an…
Hmm well the reason a pre-trained transformer is a fancy sentence completion engine is because that is what it is trained on, cross entropy loss on next token prediction. As I say, if you train an LLM to do math proofs,…
Isn't RL the algorithm we want basically?
How about you want to solve sudoku say.And you simply specify that you want the output to have unique numbers in each row, unique numbers in each column, and no unique number in any 3x3 grid. I feel like this is a very…
"AlphaProof is a system that trains itself to prove mathematical statements in the formal language Lean. It couples a pre-trained language model with the AlphaZero reinforcement learning algorithm, which previously…
If you're going to suggest something you think an LLM can't do I think at the very least as a show of good faith you should try it out. I've lost count of the number of times people have told me LLMs can't do shit that…
Yes but you are not taking an uncountable union. You are taking a finite union.
The only source I can find for this estimate is from a year ago. I feel like efficiency has gone up by a lot since then
anabolic steroids will kill you idk why you'd want to mess with them.
And yet it doesn't rule out that it can't. See new york times lawsuit
I don't see how the point about the typical human is relevant. Either you can reason or you can't, the ARC test is supposed to be an objective way to measure this. Clearly a vanilla LLM currently cannot do this, and…
The curious thing is if you can ever hit a snowball type point, because by proving theorems you are generating training data, and maybe you can get to a point where you effectively have limitless data.
Doubtful. Apple can get away with tiny iterations on smartphones because they have the brand and they know people will always buy their latest product. LLMs aren't physical products so there is no cost to switching…
Agree on hoping for an intelligence update, but I think it was clear from teasers that this was not gonna be GPT-5. I'm not sure how fair it is to classify the new multimodal capabilities as just a gimmick though. I…
Lol there's a massive difference between a framework that generates javascript, a language which has existed for 30 years at this point, and a magic LLM that no one on earth understands the internals of.
Given a fixed level of precision for input and output, and a function on this discrete space, we can construct a continuous extension of this function to the reals. Now using the paper we know that there is a neural…
what are you even yapping about.
Adding the position vector is basic sure, but it's naive to think the model doesn't develop its own positional system bootstrapping on top of the barebones one.
fmri's are correlational nonsense (see Brainwashed, for example) and so are any "model introspection" tools.
So you think that this blog post would make it into any of the mainstream conferences? I doubt it.
peer review would encourage less hand wavy language and more precise claims. They would penalize the authors for bringing up bizarre analogies to physics concepts for seemingly no reason. They would criticize the fact…
didn't hackers used to be for piracy?
This suggests people should pre-register benchmarks. Because currently it feels like there is little incentive to publish benchmarks that models saturate.
I completely agree this shit is so depressing. When I saw the AlphaProof paper I basically spent 3 days in mourning basically, because their approach was so simple.
I think the whole paper is a satire lol.
Does it really? If you want an LLM to edit code you need to feed it every single line of code in a prompt. Is it really that surprising that having just learnt it has been timed out, and then seeing code that has an…
Hmm well the reason a pre-trained transformer is a fancy sentence completion engine is because that is what it is trained on, cross entropy loss on next token prediction. As I say, if you train an LLM to do math proofs,…
Isn't RL the algorithm we want basically?
How about you want to solve sudoku say.And you simply specify that you want the output to have unique numbers in each row, unique numbers in each column, and no unique number in any 3x3 grid. I feel like this is a very…
"AlphaProof is a system that trains itself to prove mathematical statements in the formal language Lean. It couples a pre-trained language model with the AlphaZero reinforcement learning algorithm, which previously…
If you're going to suggest something you think an LLM can't do I think at the very least as a show of good faith you should try it out. I've lost count of the number of times people have told me LLMs can't do shit that…
Yes but you are not taking an uncountable union. You are taking a finite union.
The only source I can find for this estimate is from a year ago. I feel like efficiency has gone up by a lot since then
anabolic steroids will kill you idk why you'd want to mess with them.
And yet it doesn't rule out that it can't. See new york times lawsuit
I don't see how the point about the typical human is relevant. Either you can reason or you can't, the ARC test is supposed to be an objective way to measure this. Clearly a vanilla LLM currently cannot do this, and…
The curious thing is if you can ever hit a snowball type point, because by proving theorems you are generating training data, and maybe you can get to a point where you effectively have limitless data.
Doubtful. Apple can get away with tiny iterations on smartphones because they have the brand and they know people will always buy their latest product. LLMs aren't physical products so there is no cost to switching…
Agree on hoping for an intelligence update, but I think it was clear from teasers that this was not gonna be GPT-5. I'm not sure how fair it is to classify the new multimodal capabilities as just a gimmick though. I…
Lol there's a massive difference between a framework that generates javascript, a language which has existed for 30 years at this point, and a magic LLM that no one on earth understands the internals of.
Given a fixed level of precision for input and output, and a function on this discrete space, we can construct a continuous extension of this function to the reals. Now using the paper we know that there is a neural…