I feel like this kind of "result dump" just cheapens mathematics. How about having a little respect for those whose work this builds on, and current mathematicians some of who may have spent years working on these problems.
Rather than sitting on these results until they had enough for a "shock and awe" 10-result dump, how about releasing these results individually as they were made/verified, as well as the failures (equally valuable to assess the current capabilities of LLMs), and try to make some analysis of HOW these breakthrough results were made. What were the prompts for each of these, how much guidance was there from the mathematicians employed by OpenAI, and most importantly how did the model arrive at these results ... what lines of reasoning resulted it in exploring ideas that humans had previously not explored?
The LLM marketing loop is getting awfully long in the tooth. I saw a meme on Twitter the other day that showed a circular state diagram with something like:
> “GPT solved a math problem” -> “Claude solved a math problem” -> “GPT escaped the sandbox” -> “Claude escaped the sandbox” -> …
Does anyone else have trouble telling how much of this news (along with the 'AI escaping and hacking' stories) is genuine, vs how much is just AI firms overstating their capabilities due to strong commercial incentives?
I believe this is what we wanted computers to help us solve along with other prior hard problems prior to computers. This should be viewed as a good thing even if Anthropic, OpenAI, etc benefit just like IBM benefitted from mainframes.
I’m looking forward to the days where AI would help in tackling the problems in biology. Especially, on creating new drugs, enzymes and understanding the genetic diseases. An absolutely interesting time to live.
11 comments of 51
[ 2.9 ms ] story [ 22.9 ms ] threadRather than sitting on these results until they had enough for a "shock and awe" 10-result dump, how about releasing these results individually as they were made/verified, as well as the failures (equally valuable to assess the current capabilities of LLMs), and try to make some analysis of HOW these breakthrough results were made. What were the prompts for each of these, how much guidance was there from the mathematicians employed by OpenAI, and most importantly how did the model arrive at these results ... what lines of reasoning resulted it in exploring ideas that humans had previously not explored?
> “GPT solved a math problem” -> “Claude solved a math problem” -> “GPT escaped the sandbox” -> “Claude escaped the sandbox” -> …
It s like a rich man showing off his car collection.