> It's better then incompetents code, worse then a motivated average dev... That's well put. > But good enough hence the real question is value aka time& money invested/quality. There's time invested SO FAR and time…
> I think "absolute crap but still useful to me" is a pretty high value and not worth neglecting If I extrapolate this example to my professional life, this code now manages millions of dollars, a single mistake can…
AI is great at reducing the search space and using human-like reasoning (in a brute-force way) to carry out the brute-force search. I'm not surprised by this result. This is exactly what AI should excel at, with human…
There's a big difference between one shotting a counterexample using AI and using AI to find a counter-example by brute-force. Both are impressive, of course, but they're hardly comparable.
> has fewer bugs who claimed that? Are you suggesting that the rewrite did not introduce any new bugs? The correct answer is, by the way, that no one knows since it's millions of lines of code no one has properly read.…
It's not even about requirements. It's about responsibility. Someone has to take responsibility for the code and the product. Someone has to hotfix a bug that's costing millions of dollars an hour and someone has to be…
Ever heard of the infinite monkey theorem? This is basically what LLMs do on really hard tasks. Prompt it a million times on a really hard problem and it might output the correct answer once.
People will downvote you because this comment is "not appropriate" for HN, but there were countless conversations on HN about how important these benchmarks are. I am literally LOLing at HN right now
> AI is also scarily good at writing tests :-) I hope you read those tests before claiming it's "scary good"
> This is a forum filled with experts Half of HN commentators probably work on basic CRUD. Armchair experts, maybe.
This is such a bad take. I'm impressed how often this gets parroted online. Next time, please check how many Poles left Poland for western EU since they joined.
yeah. In the future when? 2, 3, 5 years from now on? Do you think current LLMs don't get confused by shitty code? That code they're writing now will need fixing tomorrow, not 5 years later.
> I'll make more progress than mentally wearing myself out reading a bunch of LLM generated code trying to figure out how to solve the problem manually. Most engineers realize that there's currently more tech debt being…
> I'll make more progress than mentally wearing myself out reading a bunch of LLM generated code trying to figure out how to solve the problem manually. I feel sorry for whoever has to work on that codebase. This is the…
Good luck arguing with SWE benchmark purists
> what's needed is knowledge of solution techniques That's definitely in the training data
Ah right! Reminds me of AGI by 2025 :D
> The idea that code is something sacred and only devs can somehow do it is dying, and I personally love it, as I am watching it enable so many of my friends and family who have no idea how to code. People on HN are…
Yeah, good luck trusting the output!
Yeah, that's indeed a hot take. I am curious what kind of code you write for a living to have an opinion like this.
> well it actually implemented a normalization pipeline and a tax computing engine which then did the taxes, but close enough You can't seriously believe laymen will try to implement their own tax calculators.
Are you one of those naive people that still take these coding benchmarks seriously?
If you randomly sample letters from the alphabet and those letters make up actual words, then actual sentences. Did you think about it? Probably not
You should read the article you posted before you write a comment. Hint: check P_F=0 in tables 2, 3 and 4. "Factored" is doing a lot of lifting here and is borderline deceptive. Plenty of researchers have long ago…
To me, Github has always seemed well positioned to be a one-stop solution for software development: code, CI/CD, documentation, ticket tracking, project management etc. Could anyone explain where they failed? I keep…
> It's better then incompetents code, worse then a motivated average dev... That's well put. > But good enough hence the real question is value aka time& money invested/quality. There's time invested SO FAR and time…
> I think "absolute crap but still useful to me" is a pretty high value and not worth neglecting If I extrapolate this example to my professional life, this code now manages millions of dollars, a single mistake can…
AI is great at reducing the search space and using human-like reasoning (in a brute-force way) to carry out the brute-force search. I'm not surprised by this result. This is exactly what AI should excel at, with human…
There's a big difference between one shotting a counterexample using AI and using AI to find a counter-example by brute-force. Both are impressive, of course, but they're hardly comparable.
> has fewer bugs who claimed that? Are you suggesting that the rewrite did not introduce any new bugs? The correct answer is, by the way, that no one knows since it's millions of lines of code no one has properly read.…
It's not even about requirements. It's about responsibility. Someone has to take responsibility for the code and the product. Someone has to hotfix a bug that's costing millions of dollars an hour and someone has to be…
Ever heard of the infinite monkey theorem? This is basically what LLMs do on really hard tasks. Prompt it a million times on a really hard problem and it might output the correct answer once.
People will downvote you because this comment is "not appropriate" for HN, but there were countless conversations on HN about how important these benchmarks are. I am literally LOLing at HN right now
> AI is also scarily good at writing tests :-) I hope you read those tests before claiming it's "scary good"
> This is a forum filled with experts Half of HN commentators probably work on basic CRUD. Armchair experts, maybe.
This is such a bad take. I'm impressed how often this gets parroted online. Next time, please check how many Poles left Poland for western EU since they joined.
yeah. In the future when? 2, 3, 5 years from now on? Do you think current LLMs don't get confused by shitty code? That code they're writing now will need fixing tomorrow, not 5 years later.
> I'll make more progress than mentally wearing myself out reading a bunch of LLM generated code trying to figure out how to solve the problem manually. Most engineers realize that there's currently more tech debt being…
> I'll make more progress than mentally wearing myself out reading a bunch of LLM generated code trying to figure out how to solve the problem manually. I feel sorry for whoever has to work on that codebase. This is the…
Good luck arguing with SWE benchmark purists
> what's needed is knowledge of solution techniques That's definitely in the training data
Ah right! Reminds me of AGI by 2025 :D
> The idea that code is something sacred and only devs can somehow do it is dying, and I personally love it, as I am watching it enable so many of my friends and family who have no idea how to code. People on HN are…
Yeah, good luck trusting the output!
Yeah, that's indeed a hot take. I am curious what kind of code you write for a living to have an opinion like this.
> well it actually implemented a normalization pipeline and a tax computing engine which then did the taxes, but close enough You can't seriously believe laymen will try to implement their own tax calculators.
Are you one of those naive people that still take these coding benchmarks seriously?
If you randomly sample letters from the alphabet and those letters make up actual words, then actual sentences. Did you think about it? Probably not
You should read the article you posted before you write a comment. Hint: check P_F=0 in tables 2, 3 and 4. "Factored" is doing a lot of lifting here and is borderline deceptive. Plenty of researchers have long ago…
To me, Github has always seemed well positioned to be a one-stop solution for software development: code, CI/CD, documentation, ticket tracking, project management etc. Could anyone explain where they failed? I keep…