I don't understand what you mean. Test error is literally out of sample error. Marginal likelihood is designed to estimate out of sample error. The whole discussion is about out of sample; nothing has been about…
Ah, that is done in the usual Gaussian process context (not in my first example), and yes, it's a dirty idea, but you can justify it using differential privacy arguments (basically you are optimizing few parameters and…
Just to add on top of the quality reference provided by srean, I like to first drill in Bayesian principles and then use this article to derive PAC-Bayes from that: https://arxiv.org/abs/1605.08636 Regular PAC falls out…
No, I am talking about out of sample error and estimates thereof. It is "overfitting" to data, but it also has lower out of sample error than the case where you do not "overfit". This is why the notion of overfitting is…
I agree that this is a good nuanced take. However, I find that students who have learned PAC (which usually takes quite some time) often have to unlearn certain principles to do PAC-Bayes, so my comments come from a…
You can choose the prior according to any selection rule that does not see the data (actually, you can do more, but justifying this is the realm of empirical Bayes and requires some more precise arguments). In this…
Apologies, I'm skipping details, because that's how I speak with my colleagues, but I realize this is an external environment without context. No references since this is folklore (you can look at Hastie et al's…
Absolutely not. This link is a reference on PAC learning, which is thoroughly misleading in the land of deep learning and inevitably leads to vacuous bounds. This is common knowledge in deep learning. I would not…
This is provably not true, and you can use the marginal likelihood / PAC-Bayes to prove it (or any other framework for measuring model quality). Increase the number of parameters in a linear model way beyond the point…
This is an insane thing to read. Bubeck had a reputation even before he started with OpenAI. Of course it was him that was involved in this drama. This is such a sad mess, and it really didn't have to be this way.
It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it. So to say that it is unlikely is extremely suspicious. No, they did not…
Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.
Touche, I used to work in pure probability where that ridiculous Hardy-Littlewood rule used to cause all sorts of problems, but now work in statistics, where it is no longer an issue. To be clear, the colleagues I am…
Uh, care to explain? I have several colleagues that stopped submitting to journals once they reached full professor. They only submit papers from their students for the benefit of their careers. First-author papers, not…
Exactly, and the advantage is that checking that the problem is "formalized" here is essentially isolated to verifying that the final theorem statement matches the claim. If there are no 'sorry's and the program…
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly worked well. The hope was that if RLVR was used quite a bit,…
Why would they send it out for "expert review"? Every time, they have just made the AI generate a Lean proof. In fact, it seems like the most plausible direction to NS is computationally assisted detection of a blowup…
It basically is a formality at this level. Many top math researchers now hardly even submit to journals at all and just put up a preprint. At this scale, peer review happens by the audience. They don't need a journal to…
I also believe this. Post-training LLMs with vague metrics can only be achieved with RLHF, which is not impossible, but extremely costly and difficult. Instead, companies will opt for RLVR, focusing on math and…
I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench. Imagine if CRISPR…
It's good to see validated numerical proofs seeing a resurgence now that they are substantially easier to achieve. Others might be able to chime in, but my experience is that AI is effectively taking proofs that were…
[dead]
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex…
I hope you understand the context in which that was said. The point of that statement is that the only way to rigorously verify correctness of a program is by using formal methods. Those are often too difficult to…
No it really is about the test suite, and provably so. As another poster pointed out, speed is a superoptimization problem and the test suite provides the constraints. If the constraints are appropriately set, even a…
I don't understand what you mean. Test error is literally out of sample error. Marginal likelihood is designed to estimate out of sample error. The whole discussion is about out of sample; nothing has been about…
Ah, that is done in the usual Gaussian process context (not in my first example), and yes, it's a dirty idea, but you can justify it using differential privacy arguments (basically you are optimizing few parameters and…
Just to add on top of the quality reference provided by srean, I like to first drill in Bayesian principles and then use this article to derive PAC-Bayes from that: https://arxiv.org/abs/1605.08636 Regular PAC falls out…
No, I am talking about out of sample error and estimates thereof. It is "overfitting" to data, but it also has lower out of sample error than the case where you do not "overfit". This is why the notion of overfitting is…
I agree that this is a good nuanced take. However, I find that students who have learned PAC (which usually takes quite some time) often have to unlearn certain principles to do PAC-Bayes, so my comments come from a…
You can choose the prior according to any selection rule that does not see the data (actually, you can do more, but justifying this is the realm of empirical Bayes and requires some more precise arguments). In this…
Apologies, I'm skipping details, because that's how I speak with my colleagues, but I realize this is an external environment without context. No references since this is folklore (you can look at Hastie et al's…
Absolutely not. This link is a reference on PAC learning, which is thoroughly misleading in the land of deep learning and inevitably leads to vacuous bounds. This is common knowledge in deep learning. I would not…
This is provably not true, and you can use the marginal likelihood / PAC-Bayes to prove it (or any other framework for measuring model quality). Increase the number of parameters in a linear model way beyond the point…
This is an insane thing to read. Bubeck had a reputation even before he started with OpenAI. Of course it was him that was involved in this drama. This is such a sad mess, and it really didn't have to be this way.
It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it. So to say that it is unlikely is extremely suspicious. No, they did not…
Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.
Touche, I used to work in pure probability where that ridiculous Hardy-Littlewood rule used to cause all sorts of problems, but now work in statistics, where it is no longer an issue. To be clear, the colleagues I am…
Uh, care to explain? I have several colleagues that stopped submitting to journals once they reached full professor. They only submit papers from their students for the benefit of their careers. First-author papers, not…
Exactly, and the advantage is that checking that the problem is "formalized" here is essentially isolated to verifying that the final theorem statement matches the claim. If there are no 'sorry's and the program…
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly worked well. The hope was that if RLVR was used quite a bit,…
Why would they send it out for "expert review"? Every time, they have just made the AI generate a Lean proof. In fact, it seems like the most plausible direction to NS is computationally assisted detection of a blowup…
It basically is a formality at this level. Many top math researchers now hardly even submit to journals at all and just put up a preprint. At this scale, peer review happens by the audience. They don't need a journal to…
I also believe this. Post-training LLMs with vague metrics can only be achieved with RLHF, which is not impossible, but extremely costly and difficult. Instead, companies will opt for RLVR, focusing on math and…
I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench. Imagine if CRISPR…
It's good to see validated numerical proofs seeing a resurgence now that they are substantially easier to achieve. Others might be able to chime in, but my experience is that AI is effectively taking proofs that were…
[dead]
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex…
I hope you understand the context in which that was said. The point of that statement is that the only way to rigorously verify correctness of a program is by using formal methods. Those are often too difficult to…
No it really is about the test suite, and provably so. As another poster pointed out, speed is a superoptimization problem and the test suite provides the constraints. If the constraints are appropriately set, even a…