I (author of the paper) can't speak for others, but I have no affiliation whatsoever with OpenAI and have not received anything from them (I even pay for my subscription lol). I've thrown this problem at each model with…
clear specifications for what counts as completing the tasks, and an explicit list of what does not count as completing the task, and clearly stating to not return until the task has been completed. I've had agents run…
I (author of the original post and paper) can add a few things here: 1. My previous approaches with GPT 5.5 were really not very sophisticated in terms of my input. I threw the problem at it, and just kept encouraging…
I would agree with your take. I (author of the post & paper) learned a ton from working on small parts of problems my PhD advisor was doing a lot of the heavy lifting on, and later also from getting some results that…
Yes, order d is the minimal number of evaluations of gradients needed for the same problem! That has actually been known since 1979 (Nemirovsky and Yudin showed that), and there are methods with the same complexity so…
I (author of the paper) can't speak for others, but I have no affiliation whatsoever with OpenAI and have not received anything from them (I even pay for my subscription lol). I've thrown this problem at each model with…
clear specifications for what counts as completing the tasks, and an explicit list of what does not count as completing the task, and clearly stating to not return until the task has been completed. I've had agents run…
I (author of the original post and paper) can add a few things here: 1. My previous approaches with GPT 5.5 were really not very sophisticated in terms of my input. I threw the problem at it, and just kept encouraging…
I would agree with your take. I (author of the post & paper) learned a ton from working on small parts of problems my PhD advisor was doing a lot of the heavy lifting on, and later also from getting some results that…
Yes, order d is the minimal number of evaluations of gradients needed for the same problem! That has actually been known since 1979 (Nemirovsky and Yudin showed that), and there are methods with the same complexity so…