Won't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative.
More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
If the goal is to accomplish something then why limit yourself with available tools?
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
When Claude made progress on the Riemann conjecture, here are the kind of prompts used:
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times.
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.
>"illegible reasoning in a few reinforcement-learning environments over long rollout"
Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.
That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).
I suspect that the best progress will be made by a team that purely spends their time looking for open problems and promoting "solve <problem>", without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.
Then 20 years go by and you wake up one day with questions that you cannot get out of your mind: why did I start prompting the LLM for? Why did I need these random proofs for? What do i do with my repo with 2billion lines of Lean?
Did you read the article? This is not a hackathon where you build software, it’s one where you’re trying to get a model to make progress on a frontier math problem. The point is that that activity may not map well onto the shape of a hackathon
Because the companies that run frontier models are malevolent by every metric.
They are destroying the environment, especially those in neighborhoods of low income people.
They are empowering their owners who are some of the most deplorable and duplicitous people living.
They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.
They stole the entire creative output of humanity and are trying to sell it back to us.
They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.
They are being used to kill in war and for surveillance.
Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.
Have you done any math hacking with sol/astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent.
It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
"A mathematician is a person who can find analogies between theorems; a better mathematician is one who can see analogies between proofs and the best mathematician can notice analogies between theories. One can imagine that the ultimate mathematician is one who can see analogies between analogies."
I wonder how models perform on finding analogies between analogies
Years ago I'd explain what I needed from Haskell Next something like this. <Person> addresses problems by understanding and implementing the deep mathematical structure of the problem domain. For this iterative exploratory process, they use Haskell as their blackboard, usually extending Haskell Core as required.(!) But Haskell is basically only serving as a compiler target. The compilation from domain structure down to Haskell, is largely being done mentally. With the Haskell environment providing little support for that part of it. I so very can't do that. Not even close. I need tooling support that can clearly represent what I'm trying to accomplish, at every level. I want to be able to... [basically work with theories].
Maybe now that could be Lean and AI? But... growing old - I suspect even that is now outside of my performance envelope.
But it's not "waiting on the output of an LLM for 40 hours" any more than a regular hackathon is "waiting for my damn teammates to finish their part for 40 hours". From my experience using agentic coding for hackathons, the best teams are those that coordinate with the AI agents in relatively quick cadence, generally giving it small tasks and steering it often. Teams may want to run some long-running sessions too, especially closer to the deadline, but even then, they'd probably want to run and follow several sessions in parallel, and continuously inspect their work so that they have reasonable confidence that their main efforts will wrap up before the deadline. There is an art to it.
I can't wait to see if the team that fares best is the one that steers often or the one that interferes the least.
So far humans failed at those problems. Also IIRC there was a guy that proved a substantial problem 2-3 months ago by basically pasting over and over "keep looking for a solution" or something like that for 2 days with little formal math background.
1. IMO the hard part isn't prompting. It's selecting the problem and understanding
the solution. It'd be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory.
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
Probably many mathematicians want answers to the questions from the page:
> This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster?
and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.
Obviously the latter - this fact is betrayed by how the page lists the S^6 complex structure result, which was released as a 100 page barely readable mess (in fact even this might be too charitable), as still "unverified". Clearly a situation labs would like to avoid for future claimed results.
> Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs?
I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.
AI harnessing human labor for its own benefit is the way my skeptic eye sees it, or humans being had as reverse-centaurs.
> I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own
This seems like a strawman. It’s certainly not consensus that AGI has been achieved and I don’t think the people participating in this event feel like there is no value in human input or steering the AI.
It's very clear that the AI still has no motivation beyond its prompts. Humans can mostly outsource their thinking today across a wide variety of topics, but they still need to express their desires.
We are a student-led initiative. Our sponsors don't pay us and don't have a say in our decisions. All of our funding goes toward our judges and participants.
Hey! Thanks for answering, i appreciate it. Dont get me wrong, this is what you should be doing, understanding what these models are good (and most importantly bad) for. My observation is that this is very very valuable for the labs, and in an ideal world they should be paying you to do this, not just the tokens and "prize" for the winners.
I am also a bit frustrated seeing maths go in the direction of prompt enginnering. I am afraid of a world were a math phd student cannot go one week thinking about a problem without prompting an LLM to give him/her an invented answer. Something is lost along the way.
For me maths is not Lean, or formal systems, or an agent reasoning about formal systems to join literature from different fields. I see the value of it, but i think it will make it way more difficult for students (and profesional mathematitians) to see beyond that. And i see us heading into a reality were those who think like me will in practice remain a minority for quite a few years/decades because the low hanging fruit of LLMs will be to vast to ignore.
Agreed. I think we share the same concerns about LLM use. We just drew different conclusions.
Mathathon's goal is to reshape rather than stop LLM use. Can we set high standards for LLM use? Can we highlight the roles of a mathematician beyond proof generation? Can we redesign our incentives to promote these standards and roles?
I'd love to hear your thoughts on how to improve this event. We're very open to criticisms.
Thanks again for answering! As I told you I think the event is a good idea from your side. I would try to go beyond proof prompting and verification, and discussing the questions that you wrote above is the way to go.
I am mathematitian that is working as a software engineer. I have see first hand what these models are doing to SE. Its not that writting well thought, and compact code is not a good idea anymore, or that it doesnt beat LLM code, but that the people that see the value are a minority. If you write a piece of old school code in an LLM repo, it doesnt really make a difference, because old school code requires a team effort.
In maths the situation is not exactly the same. Probably reading a good piece of well thought math inside a book of LLM assisted proofs will stand out so much for the carefull reader that there will be no question about the value.
However, if the only way to get a position at a university is to print as much papers as possible, then who would risk printing 1 paper instead of 5 for doing old school maths?
So the solution to this is building a culture and community effort around these topics. And for that the topics need to be openly discussed.
Good luck with the organisation! And thanks again for the conversation!
Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.
Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
No it's not, there are tons of programs in mathematics where you go there, form teams, work a problem as a team that's likely to get a result, and then publish the results from all the groups in the conference proceedings. They are called Research Collaboration Workshops.
Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.
recent caltech grad here! and know some of the organizers well
caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
Hey guys, I'm one of the organizers. AMA.
- We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors.
- We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants.
- Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Is this hackathon only for those with formal math backgrounds?
I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
Serious question: Are you sure solving abstract mathematical problems with AI is responsible use? You will likely put mathematicians out of jobs, and I doubt solving the Collatz conjecture is urgent or will save lives. It also robs a future Fields medalist of the pride of doing all by themselves.
All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
I get your point and agree to some extent, but you can't understand the proof without significant background in Maths so it will just allow mathematicians to solve issues faster than not have the opportunity at all.
Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn't need to understand it to benefit from better shipments of cat food.
Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.
Well, for one, most math proofs don't have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.
The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code?
If you're interested in the topic enough to comment on it, you'll probably find it worthwhile reading a mathematician's perspective. Here's the prolific Terry Tao:
https://mathstodon.xyz/@tao/117219548485446992
It's hard to know what math is 'useful' a posteriori. That's always been the argument for supporting basic research. This is not why I am a mathematician however. I think there's intrinsic value into understanding something of depth and meaning, but the societal setup we have now that mostly agrees this is valuable is probably a very contingent phenomenon that is unlikely to last much longer.
Yes, so if there's useful math, you throw the LLM at it and use the results, no humans needed.
Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.
It is the top tier of humans in these fields that are making the significant breakthroughs. The top tier of breakthroughs are not being post on here (which are nowadays usually short form articles of not incredible quality). The code at start ups is not commonly in the top tier of a breakthrough. AI can do averaged work and derivations off what has gone before which Maths works very well for as there is a clear set of rules. The same in physics if you ask AI for help adapting a simulation, yet it couldn't pluck the idea if no one has done it before.
> humans don't need to be involved in scientific advances to benefit.
I agree with you on this point in isolation, but I think it's missing an enormous amount of context. Humans can absolutely benefit from science they weren't involved in and don't understand - I have no idea what a "histimine" is but I benefit from my allergy medication in the springtime.
That said, we're already living through a time where, on the whole, measures of intelligence, literacy, critical thinking, etc. are falling (at least in the US). That is a problem, which risks being exacerbated by AI, and the broader point is that we should be figuring out how to use these tools to produce knowledge that benefits humanity while also maintaining incentives for people to use their brains. Going back to my allergies: while I don't understand how my allergy meds work, my life is better, and I'm a better spouse/parent/friend/citizen etc., because I've taken the time to understand how other parts of the scientific and mathematical world that do interest me work. The current AI push to just throw out LLM-generated Lean proofs of everything under the sun to get headlines and pump up their IPO valuations (which this Marathon seems, intentionally or not, to be participating in), doesn't appear to be considering this alignment between what we get from AIs and how we can maintain our incentives to do human science. It seems more like measuring you-know-whats while risking that the message the broader public takes away is that math "has been automated" so what's the point in using your brain anymore?
If mathematicians aren't solving problems people are having (which your comment seems to imply), then putting them out of their job with AI is not a bad thing. Of course mathematicians are solving problems, just in a very different way than other professions.
Steam engines actually replaced what a horse does with a machine that could do the same job. So far LLMs aren't doing this - they're producing Lean proofs but very little to actually aid in understanding (again so far pretty much all of these proofs have been extremely difficult optimizations of known techniques). So mathematicians do solve a problem that people have: they build theories that explain the world and give us mental models to navigate questions in science, technology, etc.
The issue isn't so much that LLMs will replace mathematicians, but that AI companies bragging constantly about how their machines "solve math" will convince people who don't understand the value of math research to no longer fund it, or students who don't yet understand why learning math is useful for developing their brains that it's a waste of time. That could put mathematicians out of a job without providing a useful replacement.
Motto: a mathematician's job isn't to solve the Hodge conjecture, it's to understand why the Hodge conjecture is or isn't true, and turn that understanding into something that makes it easier for the next person to grasp/use/enjoy.
LLMs absolutely have the potential to make this job easier, but the way in which these companies are using them right now risks being antithetical to that goal.
Don't respond to a strawman argument with another strawman. The post you are responding to ignored the reasons given in the second paragraph. They're just trying to score points by preaching to the choir, not engage with the concern.
> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.
is good argument?
of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc.
everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.
I dunno, I like math and I still think automatically solving problems is worth doing. If problem-solving is a hobby then it will still be a hobby after all this. If it's about doing things for humanity--which it often is since many thousands of people are paid to do it--then doing it more efficiently benefits humanity. If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There's just no version of this where "not solving the problems" is morally justifiable.
Now, I think it's the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I've been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I'm an AI skeptic in many other ways; it's not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been something of a disaster for a long time (thanks, largely, due to the academic incentive structure which heavily favors novel results, no matter how esoteric).
> If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There's just no version of this where "not solving the problems" is morally justifiable.
I think the value comes from having people who have built very good intuition in a way that allows them to give explanations that make their ideas (new and old) accessible. Of course the most cutting edge math has become completely inaccessible for even many mathematicians, but at the same time, we've made massive progress in this regard. A hundred years ago college students might barely see calculus, and only serious researchers would see something like group theory. Today, many college students that aren't even math majors learn group theory and (we hope that) this gives them cognitive tools that they can apply in other situations (ability to axiomatize a concept, abstract reasoning etc).
The concern with how these tools are being used is that the current push by AI companies to solve math problems by chucking LLMs at them and producing a proof in Lean, and then using that as currency in the media to increase their stock value, undermines this process because it produces "proofs" without producing the understanding that actually allows humans to think better. All of this is then marketed as being the same as doing mathematics which it manifestly is not. If mathematicians lose the media war though, we'll have a generation of people who believe "math has been automated" and are unlikely to put in the effort to learn how to think for themselves.
To the extent that this event is intending to help the mathematical community find ways to use LLMs in pursuit of improving human understanding and intelligence, as the organizers seem to say it is, I think it's a very laudable goal. But I don't really see how this event is supposed to do that. It sounds a lot more like another fundraiser for team "isn't it cool that AI can produce useless chunks of computer code that compile to prove statements that the vast majority of the people commenting on these results don't even understand." For example, if it's really about finding ways to use LLMs to produce mathematics that improves human understanding, why is there even a requirement to solve a new problem? Why not make it explicitly about using LLMs to produce pedagogical content? Or if you really want it to be a new problem, why not add a requirement that the final product has to be accessible to a broad audience (say relying only on material in the undergraduate curriculum)?
P.S. There's another scenario where "not solving the problems" is morally justifiable: the scenario where the "solution" provides very little value (say because of what I said above - the solution just being a Lean artifact that adds very little to anyone's understanding), and the cost of solving the problem is extremely large. I know a lot of AI people are effective altruists, but before they could smell the IPO money I didn't see any of them talking about how if they had $20 million the most effective thing they could do with it is spend it in an extremely environmentally costly way in order to prove Navier-Stokes. Back before AI I seem to remember these folks talking about like... mosquito nets and malaria treatments?
FWIW even Astra hasn't been able to solve the problems I care about, which are less about proving theorems and more about understanding the right way to think about already existing stories (and thus permitting extensions to new contexts). However it's been a more than a capable interlocutor to test my ideas with and see if they actually have any content. It's also great for parsing possible mistakes in long technical arguments that at least my brain isn't wired to verify completely satisfactorily. Personally, I think it's good to know what is made trivial (meaning depending only on token expenditure) vs what remains a real hard kernel.
Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
Solving a problem is the easy part with AI. Selecting a problem worth solving is half of the challenge.
We allow prior work as long as it's labelled. We record chat logs, so it's easy to verify what's prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
it seems like a person could cheat by finding an open problem that's AI-solvable ahead of time (by trying many problems), and then pretend to do it for the first time in the competition. You can't verify past chat logs so make sure they didn't already realize it would work.
In your final report of this experiment, can you also report the following: How many problems were NOT solved by AI in spite of attempts. What is missing in the debate of AI in math is that only the successful cases are reported. To get a balanced picture of the effectiveness of AI, what is also needed are cases where it failed.
in light of this [1] can the organizers guarantee that the work of researchers wont be scooped by the ai labs ? if you dont guarantee it, you are doing a disservice to the community
I am against this Hackathon because it allows unscrupulous companies like OpenAI and Anthropic to use our data for their publicity. If you almost succeed to solve the problem during the hackathon, but perhaps miss out towards the end, they might steal the ideas from your attempts and eventually claim only they solved it. Also, they do not help one bit in making the proofs easier to understand by revealing intermediate inner workings of their products.
Math research is not just about stating a Theorem and giving a proof. It is about understanding the theory so that we can continue exploring various aspects. Such a Hackathon will only promote the solving part. Please stop this nonsense.
95 comments
[ 0.21 ms ] story [ 64.4 ms ] threadMore generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.
Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2
>"illegible reasoning in a few reinforcement-learning environments over long rollout"
Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.
This video presentation of the paper you linked was interesting: https://www.youtube.com/watch?v=hUp3zh23aHw
Because the companies that run frontier models are malevolent by every metric.
They are destroying the environment, especially those in neighborhoods of low income people.
They are empowering their owners who are some of the most deplorable and duplicitous people living.
They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.
They stole the entire creative output of humanity and are trying to sell it back to us.
They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.
They are being used to kill in war and for surveillance.
Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.
I for one, am one who walks away from Omelas.
It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
I wonder how models perform on finding analogies between analogies
Load-bearingly verbose, in my experience.
Maybe now that could be Lean and AI? But... growing old - I suspect even that is now outside of my performance envelope.
So far humans failed at those problems. Also IIRC there was a guy that proved a substantial problem 2-3 months ago by basically pasting over and over "keep looking for a solution" or something like that for 2 days with little formal math background.
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
> This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster?
and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.
I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.
AI harnessing human labor for its own benefit is the way my skeptic eye sees it, or humans being had as reverse-centaurs.
This seems like a strawman. It’s certainly not consensus that AGI has been achieved and I don’t think the people participating in this event feel like there is no value in human input or steering the AI.
I am also a bit frustrated seeing maths go in the direction of prompt enginnering. I am afraid of a world were a math phd student cannot go one week thinking about a problem without prompting an LLM to give him/her an invented answer. Something is lost along the way.
For me maths is not Lean, or formal systems, or an agent reasoning about formal systems to join literature from different fields. I see the value of it, but i think it will make it way more difficult for students (and profesional mathematitians) to see beyond that. And i see us heading into a reality were those who think like me will in practice remain a minority for quite a few years/decades because the low hanging fruit of LLMs will be to vast to ignore.
Mathathon's goal is to reshape rather than stop LLM use. Can we set high standards for LLM use? Can we highlight the roles of a mathematician beyond proof generation? Can we redesign our incentives to promote these standards and roles?
I'd love to hear your thoughts on how to improve this event. We're very open to criticisms.
I am mathematitian that is working as a software engineer. I have see first hand what these models are doing to SE. Its not that writting well thought, and compact code is not a good idea anymore, or that it doesnt beat LLM code, but that the people that see the value are a minority. If you write a piece of old school code in an LLM repo, it doesnt really make a difference, because old school code requires a team effort.
In maths the situation is not exactly the same. Probably reading a good piece of well thought math inside a book of LLM assisted proofs will stand out so much for the carefull reader that there will be no question about the value.
However, if the only way to get a position at a university is to print as much papers as possible, then who would risk printing 1 paper instead of 5 for doing old school maths?
So the solution to this is building a culture and community effort around these topics. And for that the topics need to be openly discussed.
Good luck with the organisation! And thanks again for the conversation!
Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.
Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
Well, that's pretty damned ignorant; I was attending William Stein's hackathons on the BSD conjecture and the Sage Math project nearly 2 decades ago.
Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.
caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.
We can't put this genie back in the bottle.
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code?
Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.
I agree with you on this point in isolation, but I think it's missing an enormous amount of context. Humans can absolutely benefit from science they weren't involved in and don't understand - I have no idea what a "histimine" is but I benefit from my allergy medication in the springtime.
That said, we're already living through a time where, on the whole, measures of intelligence, literacy, critical thinking, etc. are falling (at least in the US). That is a problem, which risks being exacerbated by AI, and the broader point is that we should be figuring out how to use these tools to produce knowledge that benefits humanity while also maintaining incentives for people to use their brains. Going back to my allergies: while I don't understand how my allergy meds work, my life is better, and I'm a better spouse/parent/friend/citizen etc., because I've taken the time to understand how other parts of the scientific and mathematical world that do interest me work. The current AI push to just throw out LLM-generated Lean proofs of everything under the sun to get headlines and pump up their IPO valuations (which this Marathon seems, intentionally or not, to be participating in), doesn't appear to be considering this alignment between what we get from AIs and how we can maintain our incentives to do human science. It seems more like measuring you-know-whats while risking that the message the broader public takes away is that math "has been automated" so what's the point in using your brain anymore?
why is that a concern in this context? would you have asked the same about steam engines and horses?
this is a really cool concept, organized very well. and that is very commendable.
The issue isn't so much that LLMs will replace mathematicians, but that AI companies bragging constantly about how their machines "solve math" will convince people who don't understand the value of math research to no longer fund it, or students who don't yet understand why learning math is useful for developing their brains that it's a waste of time. That could put mathematicians out of a job without providing a useful replacement.
Motto: a mathematician's job isn't to solve the Hodge conjecture, it's to understand why the Hodge conjecture is or isn't true, and turn that understanding into something that makes it easier for the next person to grasp/use/enjoy.
LLMs absolutely have the potential to make this job easier, but the way in which these companies are using them right now risks being antithetical to that goal.
do you mean,
> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.
is good argument?
of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc. everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.
why is that not worthy enough to pursue?
Now, I think it's the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I've been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I'm an AI skeptic in many other ways; it's not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been something of a disaster for a long time (thanks, largely, due to the academic incentive structure which heavily favors novel results, no matter how esoteric).
I think the value comes from having people who have built very good intuition in a way that allows them to give explanations that make their ideas (new and old) accessible. Of course the most cutting edge math has become completely inaccessible for even many mathematicians, but at the same time, we've made massive progress in this regard. A hundred years ago college students might barely see calculus, and only serious researchers would see something like group theory. Today, many college students that aren't even math majors learn group theory and (we hope that) this gives them cognitive tools that they can apply in other situations (ability to axiomatize a concept, abstract reasoning etc).
The concern with how these tools are being used is that the current push by AI companies to solve math problems by chucking LLMs at them and producing a proof in Lean, and then using that as currency in the media to increase their stock value, undermines this process because it produces "proofs" without producing the understanding that actually allows humans to think better. All of this is then marketed as being the same as doing mathematics which it manifestly is not. If mathematicians lose the media war though, we'll have a generation of people who believe "math has been automated" and are unlikely to put in the effort to learn how to think for themselves.
To the extent that this event is intending to help the mathematical community find ways to use LLMs in pursuit of improving human understanding and intelligence, as the organizers seem to say it is, I think it's a very laudable goal. But I don't really see how this event is supposed to do that. It sounds a lot more like another fundraiser for team "isn't it cool that AI can produce useless chunks of computer code that compile to prove statements that the vast majority of the people commenting on these results don't even understand." For example, if it's really about finding ways to use LLMs to produce mathematics that improves human understanding, why is there even a requirement to solve a new problem? Why not make it explicitly about using LLMs to produce pedagogical content? Or if you really want it to be a new problem, why not add a requirement that the final product has to be accessible to a broad audience (say relying only on material in the undergraduate curriculum)?
P.S. There's another scenario where "not solving the problems" is morally justifiable: the scenario where the "solution" provides very little value (say because of what I said above - the solution just being a Lean artifact that adds very little to anyone's understanding), and the cost of solving the problem is extremely large. I know a lot of AI people are effective altruists, but before they could smell the IPO money I didn't see any of them talking about how if they had $20 million the most effective thing they could do with it is spend it in an extremely environmentally costly way in order to prove Navier-Stokes. Back before AI I seem to remember these folks talking about like... mosquito nets and malaria treatments?
Mathematics is getting solved by AI. And this is a good thing.
We allow prior work as long as it's labelled. We record chat logs, so it's easy to verify what's prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
just something to think about
[1] https://mathstodon.xyz/@andreasthom/117240535270608201
am I eligible?
Open Problems in Computational Geometry Listed by Erik Demaine, Joseph Mitchell, Joseph O'Rourke in 2024
https://topp.openproblem.net/
Math research is not just about stating a Theorem and giving a proof. It is about understanding the theory so that we can continue exploring various aspects. Such a Hackathon will only promote the solving part. Please stop this nonsense.