In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged or `Cyber`.
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
> Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
It can't be unlimited because you can spawn parallel streams.
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
Unlimited but account-level throttled tps (more parallel streams means more throttling across all) is OK IMO, as long as it isn't too crazy. The thing that makes subscriptions really suck is having to watch the quotas, because prompt cache maintenance.
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
It was really verbose and pedantic. I'm sure that made it more thorough. But compared to Fable (which it wasn't much cheaper than) where you could get the same rigour and more with a lot more concision, it was a tough sell. 5.5 is a lot cheaper and seems a lot better balanced.
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, given these constraints. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
> It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
In short, seems to describe vibe-coding to me?
What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
> Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
It sorta feels to me like extrapolating from "the internet has all the knowledge for free" to "we won't need tradespeople anymore"
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
The plumbing analogy doesn’t match software. You pay for a plumber once and you can always choose a new plumber. With software, it’s a monthly fee that will continue to increase over time. More features behind higher tiers. And probably taking and selling your data. So why wouldn’t a person try to build something custom for their needs? They have the ultimate feedback loop of actually using the product and telling AI the issue and having AI fix it. Most software people create are probably something they never used it their lives. But the best software comes from people building it that also use it. Even Shopify started when Tobi started a snowboarding store online and couldn’t suitable e-commerce software.
You’re right about people not watching plumbing videos and doing it themselves. But the equivalent example would be open-source software in the tech example. But instead of reading open-source code to see how different features were implemented, AI can go and dig into the code and figure it out.
You’re right, the monetization of software is fading and building actual moats is becoming difficult. If you have distribution, regardless of your app, you still have a long run way.
I was more pointing out that software gets enshittified. A plumber necessarily doesn’t and if the plumber does get worse, you can call a different one next time. Software, especially B2B, has switching costs and lock in. So you just have to put up with it.
Another point is that most software started with a few features and to get more market share and support more use cases, it became worse for the users using the early features. That’s why they try to build their own so it’s not bloated with features you will never use.
> I wonder how others are coping with it aside from denial and cynicism.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly no me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
> So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
One of the quotes which might help as well (I think I have this even in my HN profile): The only thing we know about the future is that it will surprise us.
Not even experts are much more likely to predict for what its worth than a coin toss in many cases (especially if they believe that only one theory/idea will mostly predict the future)
It's a blend of things and ideas and the sheer interconnectedness of them where a small pocket can grow large and then also shrink and taking into account all variables and factors is just simply impossible for a mind. I think that although we feel we are being more informed about the world, that in it of itself doesn't prevent things in the future from happening. It just makes us alert and sad and anxious about it.
Yet this life is one which shouldn't be lived with sorrow and anxiety. It is one of beauty and greatness. In many ways, we humanity have come so far from the past (Our medicine is something that not even the mightiest of kings could get) and yes, there are many problems in the world and some things feel as if they are staying just the same or getting worse real-time.
But even then, worrying about it could lead to nowhere other than a path of misery. Also these problems are complicated enough that its extremely hard for a single person to bring change (not that I wish to demotivate that person but rather seeing the system as a complex nature)
So to me, its also a form of inward action. I can work on myself to be better prepared for the world that comes next. In the same time, I think that the present for me as well is good. I have great family and friends and have many qualities that I am proud of and I wish to share that gratitude to the people who have helped me along the way (my family/friends/ Hackernews!.)
Within the hustle culture, there is no time to relax but it is within the time of relax that I believe some of the most fruitful actions can come. I believe it just makes my mind more productive being in a calmer state.
here's a quote from how to measure your life that I hope can help some people:
I genuinely believe that relationships with family and friends are one of the greatest sources of happiness in life. It sounds simple but like any important investment, it needs constant attention and care(...)
You'll be tempted to invest your resources elsewhere but if you don't nurture these relationships, they won't be there to support you in hardships or as one of the most important sources of happiness in your life.
So thank you hackernews and have a nice day and please, please try to say gratitude towards someone close to you (within these tough times) and try to keep a balance towards inward focus, sharing time with friends/family and also writing on hackernews (as is my past time nowadays), balance is necessary :-D
So once again, I hope that its a call to action to say gratitude towards anyone. Just send them a big message thanking them and make their day as well as yours memorable, have a nice day!
Also you are allowed to make mistakes (everyone makes them!) and even though I am saying (preaching?) these things, I have found myself sometimes failing to act on these things as well but I just think that these help in being more mindful about them hopefully and can help provide a perspective. I wish to adopt more of these things in my life myself as well hopefully :-D
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
Exactly. I'm saying that Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token.
> I want to retain some semblance of understanding
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase.
Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
I think going for more understanding is the way you need less understanding. The more solid your core understanding of your codebase is the less you need to know the details, the less missunderstandings the less iterations needed, the less mental capacity consumed
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.
A lot of old-school software engineering is about how to deal with this reality.
No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
Yes, it’s possible for a person to create software he doesn’t understand himself. In the old days this was pasting from Stack Overflow and changing things until it worked.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
I’m referring to competent engineers maintaining understanding over time of all the code they’ve produced. Long before agentic coding, codebases routinely grew beyond the comprehensive understanding of their own authors.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
I'm looking forward to the making of more software. I think there are probably people who have had really useful software ideas for a long time that they'd never be able to raise money for, but now for $20 a month, they can get up and running, serving their local and/or niche communities, without having to hire a team of engineers.
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.
This sounds... horrible? I mean, it's certainly a solution to the "wake up when this thing happens" problem, but... $SERVICE -> webhook -> $CHAT_APP -> MCP -> remote wakeup sounds both brittle and - as you said - the local code harness route is entirely unserved by something like this.
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
> The thing with watching CI in an agent loop is that it burns tons of tokens.
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
I guess folks' experiences with this stuff will vary wildly by what environment they work in. I use LLMs mostly at work, where I don't have any subscription plans, everything is billed per-token, and there's multiple coding harnesses with different token quotas available (and vastly different functionality). So the sharable CLI that works whether I'm in Claude Code (where tokens cost some outrageous amount) or Devin CLI (a horrible harness that also lacks any sort of scheduling system as far as I've ever figured out, but hey, there's GPT Luna and GLM available, at least) is a huge win.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
Yeah, I made this point above but LLMs just don't have a good sense of importance. They treat everything at the same level of importance and can spend considerable effort on things that just don't really matter.
I think that's what a future dev team is going to look like.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
You're thinking that the entire economy will collapse down to just 3-4 vendors?
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
> Another thing to think about is, what would it take for you to care less about the understanding
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
> what would it take for you to care less about the understanding
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Plan longer chains of work / higher level goals that can be broken down into multiple chains of work. This will allow you to automate more work units to be worked on.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
I think this is valid now, but not guaranteed to be valid forever. For engineers, there was a period where more checks, more tests, more auto code reviews improved results quite a bit. People were consuming tokens like crazy (including me). Then things improved via better effort/thinking levels, where you could see repeated code reviews plateaued, so now people don't really do that quite as much.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
> you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about.
If you have well-specified tasks, you can easily reach 10 simultaneous agents working on disjoint parts of the code in worktrees. That consumes tokens pretty quickly!
I created an orchestration skill for myself (using herdr but any persistent mechanism works). So I then only interact with a front session and it will triage and dispatch each request to the relevant spaces (each of them can have multiple worktrees of the same project), summarize movements and pending decisions for me all at once. I do not directly interact with a multiplexer or any dashboard.
Personally, it takes me longer to write the specifications than it takes the model to implement them (and it takes me much longer to review the resulting code, although maybe that makes me old-fashioned). Consequently I do not have enough tasks to run more than one agent at a time.
Exactly. I don't understand how so many developers seem to have a long tail of well written task specifications ready to submit to the LLM.
Who produces them?
Every new generation of model goes and cleans up the slop of its predecessor in my code bases, and it has turned out to be quite effective.
Last year I held off on implementing a few features knowing that a model like Opus 5.5 was around the corner. I'm now implementing them in a much more efficient and quality manner than I could have fall of 2025.
Indeed. Maybe people down voting me don't like to admit it but when models train on the entirety of human input you can assume they'd be better than the average human.
For my normal work I take ownership of the code, and end up with the exact code I want. I still have it go off and do a good amount of work a lot of the time, still queue up multiple tasks at the same time a lot of the time. Sometimes I throw it all away and re-prompt once it's time to commit to it, sometimes edit what it made, sometimes have it edit what it made etc.
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
I am actually going to go out on a limb and guess the opposite of this is true, and that on average vibecoders burn tokens less readily than veteran software engineers. I could list several reasons why I think this would be likely. No idea which of us is empirically right, though.
(With exceptions for what I can only call the "manic vibecoders" with like 10 simultaneous weird slopprojects they're spewing out at once. Generally with each project itself being something related to vibecoding. Steve Yegge being an example of a "manic vibecoder-actual programmer" hybrid.)
I’d think this matches my hypothesis. I’d say that I spend more tokens rewriting and fixing things, so that contributes more.
Also, I’d imagine the token-maxed user is a programmer that lives in chat. I’ll admit to having asked the LLM to move a method up/down in a file, and watched it burn tokens for a minute thinking and executing a menial task.
I mean, they're trading off the time to learn to program for LLM time. Which might make sense! Locally.
But, in my experience, the projects where I have a constant pulse on the core design and abstractions in the code end up moving much faster than the ones where I don't. And I've been working on one of each at work recently, so I have a decent point of comparison.
On browser based front ends it seems to be the case for me even though I still impose certain guidelines. On my C++ backends, no fucking way. Even the best models produce working but absolutely disastrous non scalable (performance wise and design wise) code unless watched over like a hen. Having said that - the value I get in either case is enormous.
I think it's not only a matter of token efficiency. If you don't know what you are doing development will eventually crawl to a halt invariably.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
Fixing code is not the same as fixing a fundamentally broken architecture. The latter requires understanding that the architecture is broken in the first place and that understanding comes from experience.
We’re at 99% for a lot of stuff today, you get a third nine from council reviews and labs have another one or two nines in the pipeline. At five nines your task length horizon extends far beyond the current frontier model release cadence. More out of distribution tasks lose a nine or two, still revolutionary. You can get an extra nine from a good set of skills around slicing and distributing work according to model capabilities.
Indeed, I've been coding for ~25 years (competitively and in open source for many of them) and I max out at least two subscriptions' worth of tokens every week across a half dozen active projects.
I still code "by hand" sometimes (mostly Ruby/Rails, C#, and random languages for code golf) but just for fun at this point. Serious projects started being 95-100% AI over a year ago.
That and some have bridged the gap with tooling. starter projects+ Strick typing + dead code detection, linting rules and today's models can get you pretty far, especially if you plan out some basic architectural patterns with your starter kit.
It's not perfect but any means but it helps manage ones sanity.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and very similarly
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they generally go for something that is so mainstream as to be superficially plausible in all cases and well-considered in none of them.
Does that mean you're doing work week by week that is well-scoped and planned for that week and you can't pick up the work for next week until that week happens? It just seems like mostly programmers work on things that extend out for a long time and you can kind of just throw more at the problem and pull it forward earlier
I have been using the “free” Ling Flash model on openrouter for a bit over a month now. My work is an aside project at home building a Rust binding for an open source Zig code base library. The result is nothing impressive but also not a total failure: I have a feature parity binding to use in Rust vs. Python/java/typescript.
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.
If I follow the principle "every single line of AI-generated code has to be reviewed and understood by me," I literally can't use up the $20 Claude subscription. I tried quite hard but only managed to do it once.
I do that and can use up the $20 subscription, but I haven't been able to hit the limits with the 5x subscription so far (even though I use Fable for making plans).
The other thing too is that I’m having a hard time reviewing these massive PRs that are being generated. So much so that I’m having it write me a book (also vibe coded) that I can read through to learn about all the stuff it did as part of the pr review
https://ai-lessons.oncanine.run/
This is boat I’m in too. I get a ton out if my pro subscription, and I don’t hit the limits, but my Microsoft buddy was just griping about how the company recently imposed $10000/month token budgets on his team and he blew through his quota on under a day. The mind boggles.
> Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
Not every user of Claude is a programmer. Or even exclusively a worker. Claude has uses beyond work. Something that many in HN struggle to understand.
> In our testing, it costs up to 30% less per task than its predecessor.
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally years behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
371 comments
[ 4.0 ms ] story [ 85.6 ms ] threadTerminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
This is bollocks. Their safeguards are shit.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
This is the way.
https://support.claude.com/en/articles/14604842-real-time-cy...
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
1 - https://bench.killswitch-lang.org
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
I use GLM directly from z.ai, they do not retain or train on your data accordingly to their TOS.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, given these constraints. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
You’re right about people not watching plumbing videos and doing it themselves. But the equivalent example would be open-source software in the tech example. But instead of reading open-source code to see how different features were implemented, AI can go and dig into the code and figure it out.
The fact that you think this is the way software is priced is telling.
The model being destroyed here is that every piece of software is something that needs to generate recurring revenue.
I was more pointing out that software gets enshittified. A plumber necessarily doesn’t and if the plumber does get worse, you can call a different one next time. Software, especially B2B, has switching costs and lock in. So you just have to put up with it.
Another point is that most software started with a few features and to get more market share and support more use cases, it became worse for the users using the early features. That’s why they try to build their own so it’s not bloated with features you will never use.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly no me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
One of the quotes which might help as well (I think I have this even in my HN profile): The only thing we know about the future is that it will surprise us.
Not even experts are much more likely to predict for what its worth than a coin toss in many cases (especially if they believe that only one theory/idea will mostly predict the future)
It's a blend of things and ideas and the sheer interconnectedness of them where a small pocket can grow large and then also shrink and taking into account all variables and factors is just simply impossible for a mind. I think that although we feel we are being more informed about the world, that in it of itself doesn't prevent things in the future from happening. It just makes us alert and sad and anxious about it.
Yet this life is one which shouldn't be lived with sorrow and anxiety. It is one of beauty and greatness. In many ways, we humanity have come so far from the past (Our medicine is something that not even the mightiest of kings could get) and yes, there are many problems in the world and some things feel as if they are staying just the same or getting worse real-time.
But even then, worrying about it could lead to nowhere other than a path of misery. Also these problems are complicated enough that its extremely hard for a single person to bring change (not that I wish to demotivate that person but rather seeing the system as a complex nature)
So to me, its also a form of inward action. I can work on myself to be better prepared for the world that comes next. In the same time, I think that the present for me as well is good. I have great family and friends and have many qualities that I am proud of and I wish to share that gratitude to the people who have helped me along the way (my family/friends/ Hackernews!.)
Within the hustle culture, there is no time to relax but it is within the time of relax that I believe some of the most fruitful actions can come. I believe it just makes my mind more productive being in a calmer state.
here's a quote from how to measure your life that I hope can help some people:
I genuinely believe that relationships with family and friends are one of the greatest sources of happiness in life. It sounds simple but like any important investment, it needs constant attention and care(...)
You'll be tempted to invest your resources elsewhere but if you don't nurture these relationships, they won't be there to support you in hardships or as one of the most important sources of happiness in your life.
So thank you hackernews and have a nice day and please, please try to say gratitude towards someone close to you (within these tough times) and try to keep a balance towards inward focus, sharing time with friends/family and also writing on hackernews (as is my past time nowadays), balance is necessary :-D
So once again, I hope that its a call to action to say gratitude towards anyone. Just send them a big message thanking them and make their day as well as yours memorable, have a nice day!
Also you are allowed to make mistakes (everyone makes them!) and even though I am saying (preaching?) these things, I have found myself sometimes failing to act on these things as well but I just think that these help in being more mindful about them hopefully and can help provide a perspective. I wish to adopt more of these things in my life myself as well hopefully :-D
[Pardon me for the long post]
Not sure how atheists are coping with the existential risks we are facing.
The proof of the pudding.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
> the economy is over
Hackernews' neuroticism remains undefeated
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A lot of old-school software engineering is about how to deal with this reality.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
> literally hand-write code they don’t understand even as they write it
I find this literally impossible. How can you even start typing anything without knowing what to type?
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
So accelerate the vibe coding of shit nobody wants or asked for, just to see some metric go up somewhere.
Are we still getting bonuses for the number of tokens we can burn?
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
You sure that wasn't just working at Microsoft?
Last year I held off on implementing a few features knowing that a model like Opus 5.5 was around the corner. I'm now implementing them in a much more efficient and quality manner than I could have fall of 2025.
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
(With exceptions for what I can only call the "manic vibecoders" with like 10 simultaneous weird slopprojects they're spewing out at once. Generally with each project itself being something related to vibecoding. Steve Yegge being an example of a "manic vibecoder-actual programmer" hybrid.)
Also, I’d imagine the token-maxed user is a programmer that lives in chat. I’ll admit to having asked the LLM to move a method up/down in a file, and watched it burn tokens for a minute thinking and executing a menial task.
But, in my experience, the projects where I have a constant pulse on the core design and abstractions in the code end up moving much faster than the ones where I don't. And I've been working on one of each at work recently, so I have a decent point of comparison.
On browser based front ends it seems to be the case for me even though I still impose certain guidelines. On my C++ backends, no fucking way. Even the best models produce working but absolutely disastrous non scalable (performance wise and design wise) code unless watched over like a hen. Having said that - the value I get in either case is enormous.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
I still code "by hand" sometimes (mostly Ruby/Rails, C#, and random languages for code golf) but just for fun at this point. Serious projects started being 95-100% AI over a year ago.
It's not perfect but any means but it helps manage ones sanity.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and very similarly
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they generally go for something that is so mainstream as to be superficially plausible in all cases and well-considered in none of them.
I mostly use Fable though, Opus only via sub-agents.
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.
Not every user of Claude is a programmer. Or even exclusively a worker. Claude has uses beyond work. Something that many in HN struggle to understand.
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally years behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
It's been out for an hour and you've already concluded this?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
Their new Ember-1 model is pretty good, fine-tune of Kimi3 with way less thinking
Does this make it an American model or is it still Chinese?
https://fireworks.ai/blog/ember-1
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
Anthropic made it that way, and I'd say the lower score is accurate.
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
That term is about hiding a system's design in order to secure something, rather than having secure design.
> it’s the first Sonnet model to launch with cyber safeguards
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.