The biggest jump was Opus 4.6. Since then they have gradually gotten better at finding issues in your reasoning, not hallucinating, and being rigorous with the code, but much much worse at explaining things and…
Yep. Simple answer is they want to IPO in the fall, and a new Haiku does literally nothing for them
That would actually be a reasonable thing to do with a model that is good at writing though. I'm sure if you used that method on Opus 4.6 you'd get a pretty decent result, because Claude models used to be quite good at…
I believe AI technical prose peaked with Opus 4.6. I still use it (mostly for that purpose) and I think it's legitimately great. I'm hoping they are able to reverse the trends that benchmaxxing and RLAIF have wrought on…
I assure you it won't be as bad as this (answer to FAQ: "Will OpenLogi support Logitech Flow?") > It's on the roadmap, at the far end: a cross-computer pointer and clipboard bridge is a very large feature. The half that…
The entire ecosystem of CC is designed to facilitate burning tokens. You have to ask the LLM to write a script for the app to tell you which folder you're working in and which branch you're on. There are commands that…
What output style have you found to actually fix Opus 5's grating prose then? I find it leans hard into its preferred grammatical structures and rote sayings no matter what I include in the output style.
> because it's really hard to learn anything by skimming code that's being pooped out by Claude. I think this is pretty dubious. It's true that you have to do more proactive learning with an agentic workflow, but it…
You're right—and it's worth calling out explicitly, because that completely changes the approach here.
Here is one data point for cost: https://artificialanalysis.ai/models?cost=intelligence-vs-co... Here is another data point for output token efficiency: https://artificialanalysis.ai/models?cost=intelligence-vs-co...
I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better…
> So I would rather share a match with the occasional cheater than run un-auditable ring-0 software on the same machine I use for anything private. The article makes an argument that anti-cheat is not worth the…
I'm basically restating what you said, but it's amazing to me that the vast majority of people you will meet, even educated people, are casual dualists and free-will libertarians. If they happen to acknowledge…
Because, as the OP said, working hours are powered by norms. There are salaried positions and companies and teams that certainly will make you work 6 days a week, or make you feel like a bad worker if you don't do…
[dead]
The real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?
But there is a way for even an aligned federal government to fight back against the slide into authoritarianism, even with an authoritarian president expanding the powers of the executive, and that is for the other…
> “In my trips to Wall Street,” Dyer told the panel, “one of my analyst friends took me to lunch one day and said, ‘Joe, you have to get iRobot out of the defense business. It’s killing your stock price.’ And I…
Sure, but the author is arguing that the outcome you're describing is tightly coupled to the perverse incentives that he describes in the article. Investors pushed the company towards extraction over innovation and the…
I'm not so sure I buy the premise that engineers are really dismissing AI because it's still not good enough. At the very least, this framing does not get to the heart of why certain engineers dislike AI. Many of the…
Low mileage used cars don't come with a warranty, or probably have a more limited warranty if they're CPO. Leases can be better, but again they are usually better choices in high depreciation scenarios (like luxury…
Have you seen the prices of pre-owned Honda/Toyota sedans that are less than 5 years old? There are absolutely cars out there where trading in your new car after 3-4 years can make sense depending on the cost of the…
Those things also require more willpower than taking a medication. Willpower is generally determined by your particular psychology which is determined by genetics and environmental factors. People don't have a choice in…
"Real industry" also has quite a hard time getting things done these days. If you look around at the software landscape, you'll notice that "getting things done" is much easier for companies whose software interfaces…
I think the problem is false positives, not false negatives. The people you interact with during the interview process have all sorts of reasons to embellish the experience of working at their company.
The biggest jump was Opus 4.6. Since then they have gradually gotten better at finding issues in your reasoning, not hallucinating, and being rigorous with the code, but much much worse at explaining things and…
Yep. Simple answer is they want to IPO in the fall, and a new Haiku does literally nothing for them
That would actually be a reasonable thing to do with a model that is good at writing though. I'm sure if you used that method on Opus 4.6 you'd get a pretty decent result, because Claude models used to be quite good at…
I believe AI technical prose peaked with Opus 4.6. I still use it (mostly for that purpose) and I think it's legitimately great. I'm hoping they are able to reverse the trends that benchmaxxing and RLAIF have wrought on…
I assure you it won't be as bad as this (answer to FAQ: "Will OpenLogi support Logitech Flow?") > It's on the roadmap, at the far end: a cross-computer pointer and clipboard bridge is a very large feature. The half that…
The entire ecosystem of CC is designed to facilitate burning tokens. You have to ask the LLM to write a script for the app to tell you which folder you're working in and which branch you're on. There are commands that…
What output style have you found to actually fix Opus 5's grating prose then? I find it leans hard into its preferred grammatical structures and rote sayings no matter what I include in the output style.
> because it's really hard to learn anything by skimming code that's being pooped out by Claude. I think this is pretty dubious. It's true that you have to do more proactive learning with an agentic workflow, but it…
You're right—and it's worth calling out explicitly, because that completely changes the approach here.
Here is one data point for cost: https://artificialanalysis.ai/models?cost=intelligence-vs-co... Here is another data point for output token efficiency: https://artificialanalysis.ai/models?cost=intelligence-vs-co...
I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better…
> So I would rather share a match with the occasional cheater than run un-auditable ring-0 software on the same machine I use for anything private. The article makes an argument that anti-cheat is not worth the…
I'm basically restating what you said, but it's amazing to me that the vast majority of people you will meet, even educated people, are casual dualists and free-will libertarians. If they happen to acknowledge…
Because, as the OP said, working hours are powered by norms. There are salaried positions and companies and teams that certainly will make you work 6 days a week, or make you feel like a bad worker if you don't do…
[dead]
The real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?
But there is a way for even an aligned federal government to fight back against the slide into authoritarianism, even with an authoritarian president expanding the powers of the executive, and that is for the other…
> “In my trips to Wall Street,” Dyer told the panel, “one of my analyst friends took me to lunch one day and said, ‘Joe, you have to get iRobot out of the defense business. It’s killing your stock price.’ And I…
Sure, but the author is arguing that the outcome you're describing is tightly coupled to the perverse incentives that he describes in the article. Investors pushed the company towards extraction over innovation and the…
I'm not so sure I buy the premise that engineers are really dismissing AI because it's still not good enough. At the very least, this framing does not get to the heart of why certain engineers dislike AI. Many of the…
Low mileage used cars don't come with a warranty, or probably have a more limited warranty if they're CPO. Leases can be better, but again they are usually better choices in high depreciation scenarios (like luxury…
Have you seen the prices of pre-owned Honda/Toyota sedans that are less than 5 years old? There are absolutely cars out there where trading in your new car after 3-4 years can make sense depending on the cost of the…
Those things also require more willpower than taking a medication. Willpower is generally determined by your particular psychology which is determined by genetics and environmental factors. People don't have a choice in…
"Real industry" also has quite a hard time getting things done these days. If you look around at the software landscape, you'll notice that "getting things done" is much easier for companies whose software interfaces…
I think the problem is false positives, not false negatives. The people you interact with during the interview process have all sorts of reasons to embellish the experience of working at their company.