100 comments

[ 0.20 ms ] story [ 41.2 ms ] thread
After watermarking debacle and general repeated misanthropic behaviour, I don't see a single reason to keep using their products.

There are many alternatives, better, cheaper and more ethical. There's no way I'm going to support a family associated to Epstein with my or my company's money.

Oh yes. Claude still saves the day sometimes, but hopefully better alternatives pop up soon.

Recent case: had to plan a trip involving multiple bus switches. Gpt 5.6 Sol proposed a route that would bring me to a dead end, since it was sunday and a specific bus had a different route on weekends. Opus 5 correctly identified that and built a route that worked.

But yes, Darios wife trying to get funding from Epstein for a porn studio says a lot about the founder.

ok, i am not crazy
well, I wouldn't go that far .. but in this small, narrow case ... no.
You're saying he's crazy.
we're all a little crazy ... it's all relative!
ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task.

wish i was joking.

Aren't the model weights frozen?
I think there are other knobs that can be turned without retraining.
I've switched off claude this week; the last week has been significantly degraded in ability, many more screw-ups.
With the Opus models spouting more and more gibberish as version numbers increase, the joke about what "degraded performance" means basically makes itself
And we get our subscription usage cut in half tomorrow if I remember correctly? EDIT: By a third. Thx below.
This feels like a near daily occurrence.
This age: we made the thing that codes faster before we made the thing that does QA faster.
Nooooo I'm going to have to use my brain again and write 100% of my code like a caveman from December 2024.
What does that even mean :)
Claude is taking a watercooler break for now. Just like a human would.
Anthropic had really screwed up after 4.6. i don't know if they work to satisfy the ego of themselves or for releasing a better model for tasks.
There are whole sections of code work that 4.7+ can't do simply because it is both over fit and stubborn.

God save you if you have a company with narrow but correct technical tradeoffs, because you operate at scale.

Opus from 4.7 one will wreck your code and argue for hours with your engineers.

Certain parts of our company have had to mandate 4.6 and a training doc to explain why our current choice is both the cost efficient and performant one and shouldn't just be ripped out.

Newer models will re-litigate the same bad, known failed architectures over and over again.

Can you give some specific technical examples where 4.7+ are making the wrong architectural decisions?
This is exactly when I left claude and started using codex during April mid or so. It once argued with me and ran for 30 minutes with a half baked buggy fix.
Just recently went back to ChatGPT after abandoning it for Claude. I must say I was stunned at how good it had become and also how they introduced new product features that I really liked.

I wonder if from now on we have to switch providers every six months or so.

It was never good, never has been. Seeing you people goon over this model or that model is absolutely hilarious.
Clearly ego, you can always tell how full of themselves they are based on their media personalities going on the podcast circuit before product releases.
Agree. Opus 4.6 was the peak and after that they introduced the effort parameter and it was a downhill since then
Despite the years-long moaning on HN about AWS US East being a single point of failure, we've sold our souls to yet another unstable monolith.
Nobody forced you to sell your soul. You made a pact with the devil. We all know how this ends up
Must be a day ending in Y
Mondays are for GitHub, Tuesdays are for Anthropic
Wonder what Wednesday will be.
The AI apocalypse will definitely happen on a Monday. Remember, robots - unlike lazy humans - work weekends too!
Another week, another outage, another cache expiration of my prompts through no fault of my own.

At least OpenAI has the decency to reset after a serious outage.

Databricks on Azure went down too yesterday
What's the incentive to keep on improving the model beyond a point? 10 devs on a team will be cut to 2 devs, so that's 8 licenses lost. They have to increase the price many fold.
they unironically think that they can replace everyone in an organization
What I don't understand is, why not replace middle management, marketing, CTOs, CEOs and the like. Surely, LLMs are better at producing high quality looking slideware and vaporware than they are at producing software.

Heavy sarcasm here if it's not obvious. Of course I know why.

I mean, there a so many startups getting crazy funding that run without management, or without engineers... or at least with so many less people than before.. except: not really.

Is the whole "it's gonna replace people" even still on the table?

I don't really understand what you mean... are you saying that the entire industry is parroting a fantasy that they all know is fake. So, everyone is faking to get more funds?

Unfortunately, I don't think that's true.

You think AI will replace a significant portion of the workforce in the sector?

Do you see it already happening? I personally not. Lay offs are due to economic reasons, at least that's what I see. Why does Anthropic still had so many engineers?

I got the impression companies are already moving back from their initial excitement. Many went all-in AI "more is better", that's not the case anymore. Why restricting usage, it's much cheaper than paying an engineer.

While ancedata does not mean much, I have had horrible success with Claude lately. I have been using Claude to crosscheck some of the outputs from GPT and vice versa. It appears both Claude and GPT believe GPT's solutions are better (and so I do).

I still believe Claude has a better UI/UX in the web interface, but tolerating Anthropic's bullshit is not worth it.

To be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantization in a home setup. It feels and performs like a frontier model.

The only issue with relying on local models is when you need them to prompt other models, and you might need to offload or switch models constantly which adds significant overhead.

But when it all works, its truly awe inspiring.

Hopefully, a reset is coming.
It's been a while since the last reset. I think we're due one. Though I would prefer they just extend the +50% usage limit forever, it's been so long I can not imagine lossing a third of my current usage.
I developed a small plugin for claudeCode that allows you to directly see in the console whats the status of claude-code in general and the status for your current model check => https://github.com/moumine9/claude-status
Could’ve just used fewer tokens and redirected to the Codex signup page.

ba dum tsss

(sorry couldn’t help myself)