363 comments

[ 0.27 ms ] story [ 34.9 ms ] thread
Luna saw a huge jump after the price cut and is one of the more competitive models at the new price on openrouter.

Maybe they want to see how much market they can grab with Sol?

This might help but there are already cheaper models with Sol's intelligence more or less, the most notable being Grok 4.6 at $6/m which makes it a tougher sell

Since when does Grok 4.6 have Sol 5.6's intelligence? I don't believe it.
I do have free sol and cursor ultra for 200 I prefer grok over sol, they are equally capable but grok is faster
I use Grok 4.6 every day; it's good, but it's not Opus 5 or Sol 5.6. The gap is closing, though.
I don't care how capable or how cheap Grok is, I refuse to financially support a company owned by a white supremacist that is actively working to disenfranchise me and millions of my fellow citizens.
Give it a rest dude.
Reposting because my other comment was incorrectly flagged. What im not allowed to speak out against my own disenfranchisemnet? Smfh

I don't care how capable or how cheap Grok is, I refuse to financially support a company owned by a white supremacist that is actively working to disenfranchise me and millions of my fellow citizens.

Do other people find 5.6 to be worse at most simple tasks and frequently over complicate things?

I asked it to write a user todo and it turned out a four page essay. I gave the same task to 5.4 and got the small list of checkboxes I expected.

It's your responsibility to set an appropriate level of Thinking. For simple tasks, I use the instant model. As an approximation, the choice is proportional to the amount of time I want it spending on the task. Also, you can always ask it to respond succinctly.
Yep. I have not yet had a single good experience with Sol or the 5.6 models on a variety of harnesses and configurations. It overthinks, overcomplicates and often makes my code into an unmaintainable sludge. It'll usually take 5+ turns of steering to get it in the right direction.
I've found it to be great for planning code changes (or new projects). I use the superpowers plug-in which I think guides the planning.

Then I switch models (to luna) before implementation. I find this combo nearly always does what I want.

I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above.

I like that skills are easy to try out, too.

I have the same setup you have, love it!
I stopped using superpowers because it wanted to turn every tiny bug fix into a $37MM DOD project. I got effective results but it took ages. I may try again - I need to find a good way to run different profiles in my harness so I can easily shut it off. The default planning workflow in OMP is pretty good though.

I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if the task is well defined and straightforward, but there are lots of models now that you can say that about.

I tell it not to use superpowers for small changes, for the same reason you say. (I often forget, though)

And yes, I sub in other coding models like DeepSeek. Mostly with good results.

I've found it's worse for simple tasks too, and I have to give it stricter guidelines, and sometimes it doesn't follow the same patterns I've grown to expect. I've found using 5.6 (sol) is good for diagnosing issues though, especially in terms of optimization of some given path
You would probably get better results with Luna for the real simple tasks, or Sol with low thinking effort.

I find that I get exactly the effort that I asked for, which is pretty nice. The other side of that coin is that these are the least lazy models I’ve used so far. They will go on elaborate tangents to complete the task when I want them to.

any good best practice for effort selection on claude. I always use high as default.
Opus 4.8 on Low is similar in price and quality to Sonnet 5 on High, but much faster.

But otherwise I don’t use Claude anymore.

The title looks to be misleading, since this price cut is limited to OpenRouter. It does not apply for the native OpenAI price listed at https://developers.openai.com/api/docs/models/gpt-5.6-sol
Which raises the question - who is subsidizing this, and why?
Possibly OAI? If you have OAI tokens you are a captive audience. If you have OpenRouter you are bidding on a free market.

OpenRouter attributes this promotion to OpenAI https://x.com/OpenRouter/status/2089416739398254662

Is this captive audience not going to switch providers for a 50% discount? Especially when the costs are simply swapping one URL for another?
Why though, to AB test/see the impact on a platform with multiple competitors?
OpenRouter is likely just leveraging Codex subscriptions.
Wouldn’t that be against TOS?
It's probably through a level of indirection.
Right? Should we switch from direct OpenAI API integration to OpenRouter?

What's the incentive here?

Open Responses API doesn't appear to support state management (yet)

Yes, and it's only for a month. This is an ad.
OpenRouter and the Vercel AI Gateway.

So yes, presumably a very small share of their total traffic.

[delayed]
They can’t decrypt the thinking traces.
You can train a LLM to inverse summarised thinking into thinking text. It’s not perfect, but it gets you maybe 80% of the quality?
The thinking traces are server-side, not exposed
I used over a billion tokens per day of gpt-5.6 sol xhigh starting last Wednesday through Sunday before reaching my reset limit. The $200 pro plan is still the best deal.
[dead]
Get a burner and use it? If you're spending $200/mo on something, $40 or whatever for a burner phone seems like a pretty cheap price.
You can get a phone number online for a few dollars.
Historically those are less useful because some of the verification systems require a real phone number and that your name is associated with the account, depending on what and how they verify. It's annoying, I use a google voice number as my primary, and it often gets rejected.
Good2Go is $5/mo for a real SIM with unlimited talk/text + 1GB data.
There are several tiers to these services, some are selling real us phone number verifications at about 0.5usd/text while others are selling virtual phone number verifications at much cheaper. From some limited experience with the former, there is rarely if ever any problems with rejections.
Oh no fuck that, business 101 is make sure that your checkout page works. There is plenty of competition in this sector, take your money elsewhere.
Use TextVerified, load up like $5 of credit and OAI verification is like $1.00. Then when your account is made, ensure 2FA/passkey is setup then you don't need to worry about the phone number.
Yes I filed a support ticket with them and explained that their system is broken and they just did not care. I explained how it was impossible for me to use it 3 times already as I've only made 2 chatgpt accounts EVER, and only recalling entering my phone number for one of the two chatgpt accounts. I told them that this issue locked me out of codex and chatgpt for work and they weren't willing to do anything about it. Totally useless support.

I ended up borrowing my gf's phone number just so I could get access for work. Ridiculous

I have mine churning like butter and I'm rarely hitting a billion tokens per day, what's your workflow look like?
Some people just do crazy stuff. For example this now ex yc guy who said he has agents constantly scanning Sf govt apis and forming dashboards just because
Pretty basic. The codex app with one conversation per project and several running simultaneously all hours. I’m going for prompt caching that way and it never gets lost even with compaction somehow. Each has a plan with milestones to keep up to date and a thin agents file. I check in on them in the Remote app. Use case is protocol and control reverse engineering of audio hardware. I think they must be identifying the heavy use agent sessions and cranking up their cache lives so it’s not a big deal for them.
A billion tokens a day is 11,000 tokens a second sustained. How many tokens per second are you getting off of GPT 5.6 Sol per project?
Often people are counting all tokens, including cached input tokens, for those more impressive "billions of tokens" quotes.
Ah, thanks, I'd missed that nuance in the other reply!
I'd like to see a benchmark on this specific topic: Reverse engineer the hardware protocol from a driver, or just migrate a driver from one OS to another.
I agree. I’ve chipped on it with each model since 5.2 but 5.6 sol is something else. When it first came out I’d get some refusals but they’ve since stopped.
5.6 has been a huge pivotal change in reverse engineering tasks for me too (largely extracting game assets from binary client files). Something I spent literally weeks on in January with claude models at the time was solved in about 30 minutes with 5.6 Sol at medium just yesterday. It's both extremely satisfying but at the same time also a little annoying how much time I had previously spent on it only for it to be solved so quickly now. I suspect improvements in AI will continue this happy but annoyed trend.
[delayed]
That's wild. I was gonna say 3-5 billion a month is more reasonable summed across all token types.
with ultracode, it goes fast. I can easily get to a billion on a busy day.
I spent $800 in a few hours when my sub maxed out because I was trying to get something done and had a long car ride to let it churn.

Their api pricing is absurdly expensive.

> Their api pricing is absurdly expensive.

I assume at this point that it subsidizes subscriptions.

yes, it absolutely does.

I've gotten more work done on a second chatgpt pro $100/mo subscription than I did with ~$150 of paying for usage through the app.

Massively subsidized. As soon as my Claude switches from subscription to overage I have to tap out quickly.
This is a big part of the reason I went local-only. Subscription limits are horrible for having a decent workflow.
Yeah, it sure was convenient that there was a RAM pricing crisis right when Apple was making local inference viable. All because of a promise that AI companies will buy more of it... with money they don't yet have, whereas Apple does have lots of money.
There's no universe in which just buying a second or a larger subscription isn't a billion times better and cheaper than any kind of comparable local workflow.

Privacy, experimenting with ML and "unorthodox" needs are currently the only acceptable reasons to do local.

Well I wouldn't say a billion times better. I've actually been having a surprising amount of success working with local models. And my investment has only been the equivalent of 4 months of a Max x20 subscription.

Experimentation and privacy are definitely advantages, but it's also quite a lot of fun.

Because API pricing is for corporations and subscriptions are for consumers.
Are we supposed to just shrug at the idea of businesses paying 10x more for raw materials than consumers? How long can this go on?
Hopefully for a long time this is like the beginnings of a vac startup, enjoy while you can before the enshitification comes
I mean this is kind of standard in a lot of industries? Businesses get reliability/support usually for the price increase. Just look at lots of industrial tech. Or business vs consumer 3D printers, or business vs consumer laptops etc....
I thought corporations were meant to be smart and get b2b / volume discounts - not pay 5-10x what the man in the street is paying.
Agreed I’ve seen what they’re paying at work for the API 0_o Try some add on credits next time if you can. I was curious at how far $20 would go (500 credits). Watched them go to zero over an hour and assumed it’d stop. It then ran for another 6 hours and completed the task despite the meter at 0. What a task means is very unclear but it’s definitely not pricing sol at $20 hr in tokens.
I’m exclusively using ultra and I run out in 3-4 days consistently. Those resets are great but I’ve noticed they like to cluster them at the start of the cycle, would be better if they spaced them out more.
A billion tokens per day?? Plausible estimates put the energy use at about 0.001 Wh/token, which means you're using 1000 kWh/day in electricity, just to generate slop. That's about the same as 50-100 houses. 300kg of CO2 per day - roughly the same as flying from London to New York every three days.

I think on average AI energy usage is not as big a deal as everyone is panicking about, but your usage is truly absurd and I don't know how you can live with that. It's immoral.

1. You do not know what they're using it for. 2. Get off your high horse please. 3. Immoral my ass.
> Immoral my ass.

Is any amount of tokenmaxxing moral?

Is any form of energy usage moral?
Only for plants, those weird bacteria that live near thermal vents, and maybe the tardigrade.
"slop"? come on. we're not in 2020 anymore, Dorothy.
I'm glad someone is voicing this. Overconsumption at that level is not defensible. However if they meant cached tokens so it's not that bad.
That’s for i/o tokens, mostly output. 90-98% is cache read usually, so you can divide electricity use by 10 at least.

As for co2, it depends on the provider, it could be way lower as well.

As for ethics, you don’t know what he works on, and how effectively - he might be saving 10x that much of co2 for the planet.

> just to generate slop

Do some people still deny you can do a shit ton of work with AI?

Ya, honestly those kinds of people are frustrating. It’s almost a religious unwillingness to use AI tools
Where do those estimates of 0.001 Wh/token come from?
I am also on the pro $200 plan, the limit is high enough to do what I want for a week, and binge run Ultra Fast last day to use the remaining credits.
Price wars did wonders for many businesses, like the bike sharing industry in China.

Overgrown datacenters or mounds of GPUs dumped into the harbour next ?

I would in such a scenario expect the GPUs to be dumped to industrial breakers who would send them to China for refurbishment and repackaging before being sold again on Amazon, AliExpress, and Taobao as last gen gaming cards from weird brands and specs.

This is what happened after the great crypto GPU dumping.

Honestly can't wait for that to happen, same with memory, drives, etc. There's going to be a massive amount of server pulls hitting the market.
The e-waste recyclers are pretty low on the pecking order, as the creditors will be first to strip these places for assets as Leopold Aschenbrenner discovered. =3
It's where the stuff ultimately goes since the creditors aren't interested in GPUs that no longer have much value. The context here is a crash, not just a basic bankruptcy. If it were that then absolutely it would be impounded by creditors (virtually or in reality) and then sold to the highest bidding data center.
>creditors aren't interested in GPUs that no longer have much value

Indeed, there is a point where the cost of disposal is higher than the expected market value of components. Thus, the asset turns into a liability if held too long.

I have seen factory liquidations, and everything goes... right down to the bolts in the floors. =3

Well one person can use at most one bicycle at a time.

One person can use as many GPUs as they want.

Mountains of GPUs next to the ET games in the landfill.
I'm loving this race to the bottom.
I'm not having that experience. So far each major model update has been at least slightly better than the last, in ways I've found useful. Can't say it's perfect, or able to do exactly what I want without a decent amount of instruction/implementation/docs, but it's been useful enough to keep paying for it.
(comment deleted)
Oh no the models are absolutely getting better, I'm just amazed that only 6th months ago I was using gpt-5.3-codex, and now I can use gpt-5.6-luna for similar results at like 1/15th the cost. Now 5.6-sol is being slashed by 50%? Amazing.
If they can cut the price of Sol by 50% and the price of Luna by 80%, then the original price might have carried a massive operating margin. They might still be serving the models at a profit after these price cuts, but we will never know.
OpenAI didn't cut the price of Sol by 50% like they did with Luna's 80%. Sol was unchanged. This is just a limited promo for OpenRouter non-BYOK.
I don’t think there’s a real answer for this. It depends on whatever number the accounting department wants to make up.

Do you include research costs? Of all models or only specific ones? What percent of the R&D budget do you allocate to model serving? What about data center capacity? Do you count future commitments? All the circular financing deals? Employee equity grants?

We have a simple definition for this: COGS.

We also have another solution for "whatever accounting decides": generally accepted accounting practices. It's far from perfect, but GAAP figures are what you should be looking at; not "adjusted GAAP" or whatever invention.

I always find it funny that Japanese pensioners are probably subsidizing my tokens.
What do you mean?
SoftBank is one of the largest investors in openAI having contributed more than 30B and pledged another 30.
Oh yeah. I forgot the geese are still at play here
I'm pretty sure tokens are priced to maximize revenue, not inference profit.
Or they have gotten new asics and can do now inference way cheaper
Why would they infer faster with new shoes?
It is a well-accepted fact new shoes make you faster. Current science suggests it is due to the lighter weight from lack of dirt, though there is a competing theory which says it's an optical illusion due to the fact the pure white streaks resemble speedforce.
After using Claude for a long time, I tested Sol 5.6 for the first time today. Love it, its an incredibly capable model and uses far fewer tokens/time thinking. Its what I imagine Fable would be if I haven't been downgraded on every conversation - even after completing the verification program. I think I may cancel my Claude subscription finally.
Fable is still the best there is. Sol close second but I find it gets way to stuck on details.

Also Opus 5 is fine if your codebase is simple.

Fable feels less cumbersome to work with, but it is SO DAMN ANNOYING with the refusals that I'm leaning more and more on Sol, and very much looking forward to GPT6. Just seems like Anthropic is trying their hardest to ruin their reputation and user experience.
Remember, Dario knows best
I’ve never had it refuse anything. Even vulnerability searching in my codebase.
[delayed]
Can I ask which country you're in? I have a theory that the safeguards differ depending on the user's country.

I'm in Australia, and Fable downgrades to Opus when testing for bugs in memory in a legacy C code base. If Fable starts taking initiative and writes a test case that involves writing to a null pointer, that's the end of the conversation.

That's interesting, thanks for sharing. Maybe the classifiers are really sensitive to anything that has to do with memory safety.
As one random example, today I had it hit a refusal loop when adding a country selector dropdown to a form, presumably because it contained a “bad” country name? I hit refusals at least 2-3 times per day, sometimes many more. The worst part is it is often right in the middle of a multi-stage task, so the only option is really to switch to Opus 5 and let it defecate its absurdly verbose comments all over the rest of the edits in the turn and hope it doesn’t go on one of its tangents, then have Sol do damage control. Oh and I got approved for their “cyber verification program” blessing, which comically does absolutely nothing for Fable.
I think Fable's dominance is overstated. It definitely has the lead, but quantifying what that lead actually is is really hard. I'm using GPT 5.6 Sol to do some shit that I personally would consider "crazy" - low level undocumented hardware driver alchemy, reverse engineering highly obfuscated code, even a bit of screwing around with a rendering engine in Vulkan, really just about the most complex tasks I can get any model to do, and it does great. For the more advanced stuff, it definitely needs the effort bumped. But even with the effort bumped, the token usage really doesn't seem to skyrocket too badly until at least you hit xhigh and max, which really only seem to be necessary if you are doing genuine crazy stuff, so it's not that bad. I did similar stuff with Fable. In fact, I went directly from an Anthropic subscription with Fable to an OpenAI subscription with Sol, more or less, and it really felt pretty seamless. If anything, I was thrilled to realize how much I actually preferred Codex CLI, to the point where I started using it at work too.

Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).

However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.

I used Sol to extract the remaining decryption keys from the Super Mario Maker 2 (Switch) game files. Someone had previously extracted all the keys from the original release, but not any of the new ones from updates. Not only did it succeed, but it helped me understand the data sufficiently to add support for “Super World” rendering to my level viewer (which I made back in 2021), eg the little widget at the top of https://www.smm2-viewer.com/players/B16-306-GVG

I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.

Fable is almost unusable for anything but super boring mainstream stuff. I was getting safeguard flagged so often I’ve significantly reduced my usage out of fear they will blacklist/ban me.

Some of the topics it’s flagged have been hard for me to understand what it seeing that can be remotely concerning in my requests.

The safety is really funny to me. I ask it a lot of extreme stuff and it goes through, but I ask it mundane stuff and hit the filters all the time.
It has learned a little too well how it works in human society.
I've gotten flagged for asking questions about tokens and tensors. That makes me believe it's not about safety, it's about protecting their turf. I cancelled my subscription - same fear about getting flagged too much leading to a ban.
OpenAI and Kimi are both pretty okay alternatives! I guess GLM 5.3 on Max reasoning as well but for more limited domains.
They said they also block usage of Claude models to build ML models.

Which is definitely protecting their turf, but also probably a little bit hiding their “RSI” abilities for competitive reasons. My theory is that a lot of “safety blocking” is actually WIP training of new business directions. Anthropic has started hiring biologists and has opened a preview of a “Claude code for bioinformatics”. I’m guessing they’re tweaking their bioinformatics market play, and block “bio safety” requests so competitors can’t learn about their training.

What's the distinction between "protecting their turf" and "competitive reasons"? I see them as the same, but I could be missing something.

That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.

My differentiation for this comment, was “preventing someone from using your product to build a competitor” vs “letting a competitor see your strengths/weakness to benchmark against you”.

Competition is competition and it’s two sides of the same coin.

I agree. It’s flagged me on discussing fast GEMM implementations for large regressions, a discussion on theoretical physics math, a discussion on designing a type of RAG system. I’m super baffled as to what the safety instructions actually are other than “advanced anything” and even then their definition of advanced is a joke, I’m an idiot and my questions are almost laughable.
Same here - my questions were sophomore level. I think it's notable that when I edited my question to say it was about Gemma 4, it answered without blocking. A cynic like myself would interpret that as evidence they don't care about sharing information if it involves their competitors.
Getting downgraded for asking "What is digestion" to Fable is where it's just ridiculous and clearly a limitation of the technologies involved.
The dangerous part isn't that a model refuses extreme requests. It's when mundane requests become unpredictable enough that you stop trusting the model.
I cancelled my Claude max subscription. Somehow every query I sent was flagged as bio or chem, even pure mathematics questions. Not going to waste money paying for a “max” subscription that won’t ever let me use the top tier model…

Sol is great and has never blocked a request, and generally gives great answers. Happily switched over to it now.

Have they made Sol do less unwanted autonomy than the previous Codex models did?
I was having it look at creating a driver for some old scanner and it actively looked up exactly where that gray area for my country was wrt decompilation.
Interesting. I've been wishing that old 'Stars!' game from the 90s would play easily on modern systems. I'd love it if we'd got to the point where I could point Codex at a folder with an ISO from my CD of the game and tell it to go reverse engineer it all for understanding of game mechanics, then go recreate it in a modern language capable of running cross-platform. Scarcely any need to improve on graphics, it could even be a PWA.

There's people that have tried to contact Jeff McBride and follow the IP trail but the IP is currently owned by a company that went defunct. Not sold, but no one is even bothering to register its LLC any more, it's simply dead.

Stars! was awesome, many fond memories of play-by-email games of that...
things that are alchemical are rarely alchemy. That is to say things are very fiddly but stick a room of monkeys on typewriters, a schizophrenic developer with HolyC and adderall or an LLM, persistence is the key to many of these things like drivers, extracting keys from vintage security domains, etc. Dropping into xdd to a human is a chore, not for an LLM.
Although I am not exactly sure what you mean, I am not really claiming it is doing anything I couldn't do - but yes, it does so with much less effort. For example, I can have it set up probes and tracing on Linux that I personally would have to consult documentation to do. It might not even have to consult the documentation due to having the information on-tap, but even if it does, it's nothing that would cause it any fatigue, it's just going to keep moving forward in a loop until it is satisfied that it meets the criteria. I could've done all of this alone - I really could have. I just would not have. Being able to do something 10 times faster or with 10 times less effort is, in some senses, sometimes more impactful than being able to do entirely new things you couldn't do before.
This is a really good point that I definitely failed to grasp when first hearing about these tools. At least for me, the best way to use these tools is as a way to free myself from having to spend time thinking about the things that aren't worthwhile so I can focus on the things that truly are. I've had times in my life spending hours reading documentation and googling random things to try to tease out the correct sequence of commands or the exact right shape of an API to be able to make things work to know that it doesn't make me more productive to do that myself rather than point an LLM at the thing and let it spit out the answer after a few minutes. Meanwhile, I can spend that time thinking about what comes next, or what the correct way to take that one-off output and abstract it to something that can be used meaningfully in more flexible ways.

The only obvious objection I can think of to this line of thinking (at least from a technical perspective) is "how does someone build up the knowledge to be able to use a tool effectively in that way if not by doing things by hand at first?" The honest answer that is "I don't know, but that's also pretty much exactly the type of thing my employers have never been paying me to solve in the first place". Even just a decade into my career, there have already been plenty of times in my career I've struggle to convince people that we should do stuff in a way that won't bite us in the ass a month or two down the line, and in the times I've managed to succeed, it's usually only by putting in more of my own time and effort to make the initial investment seem more palatable. Luckily right now I'm not in one of those times when I'm having to go full throttle to keep the lights on a few months from now, but I don't have enough fuel in reserves to work on a plan for when we need to build a new rocket in another ten years. Maybe ask me next month.

Terry A Davis wouldn't have touched LLMs. He'd probably claim they're demonic.
FYI I run it consistently in xhigh regardless of difficulty of the task at hand. I remember high being very fast, but I'd rather wait a bit more and get better output. AIs are insanely fast compared to me anyway, even on xhigh. Consumes more usage, but even at 100 EUR/m I don't hit limits.
To be fair, I actually do run xhigh as my default. However, for the first time in my experience of trying and using LLMs, with Sol.. sometimes I feel confident enough to set the effort level to "Low". I just had Sol prototype some AWS stuff on low earlier. Great result, did exactly what I wanted.
After hitting the session limit on my company's plan so many times with Claude when I was using it, I mostly keep Codex on "high" rather than "xhigh" as a way to leave the tokens for my more ambitious coworkers. It's possible that having it higher might end up with better output, but so far at least I've yet to see a way to get any model to do 100% of what I need up front without any need for me to make changes that end up being more tedious to do via interaction than by hand, and it doesn't feel worth spending a bunch more tokens trying to figure out how to better communicate to it up front how the dominoes get set up so they fall in place properly the next time.
How do you handle context limits? With more thinking tokens you fill it up earlier. Compaction degrades performance too. What's your strategy?
Initially I was planning heavily around context limits, but I've learned to just ignore it completely. Compaction is seamless for me. If details are lost in compaction, the model just re-reads what's needed. My conclusion is that at least for Sol, the summaries (which I've never seen) must be amazing. Every now and then a detail gets lost and I have to repeat it. I don't think there is performance degration, because the model is smart enough to re-read relevant files as needed.
And Mai-Code-1.1-Flash seems like a really good cooperative player to GPT 5.6 Sol. You get Sol to help you make a detailed plan, and Mai codes it up and you can get pretty decent code out the other end without too many tokens if you are careful.
Why wouldn’t you use Luna for that? It’s super cheap.
> Why wouldn’t you use Luna for that? It’s super cheap.

MAI also offers a ultra cheap version that's competitive with Luna.

So much so that the models look like they were designed by a product manager explicitly to eat away OpenAI's market share.

Vscode even pushed them quite hard onto users with the latest release, going to the extent of putting up a modal to convince users to try them out.

How detailed of a plan? Are you including code snippets or just behavior and letting the lesser model decide how to implement?
I feel like those examples are considered difficult because they're niche topics, but aren't actually all that difficult in a general sense. What I consider truly difficult are things like taking a ticket and implementing it in a preexisting codebase, using a clean and reasonable design that fits the existing style and makes sense to a human, and avoids the footguns I learned by working with the codebase for over a day.
If you said this in 2025 I would've 100% understood, but to be honest getting AI models to do a pretty good job on day-to-day ticket work has become so boring that we don't even bother using the top tier models and higher effort slots for that anymore. I personally wind up tweaking the results a lot and recursively having fresh agents review the diff, but that's just because I'm picky; in a lot of cases the first diff is actually pretty damn decent.

Compared to what I am doing locally, I feel like day-to-day work is absolutely nothing. Not only am I also working with existing codebases in my experimental prototyping, but I am also doing things vastly more complex with vastly harder constraints.

All non-trivial code terra has generated for me has had at least one serious bug in it. Typically caught by a review from myself or Sol.

But I wouldn't trust lower tier models for end to end solutions.

Personally I wouldn't want bots running autonomously on a repo, even if there were other bots cross checking them. But that having been said, I'd also say that my experience was similar with human code: it is rare to not find at least some issue worth at least pointing out. The only real difference is that the LLMs have vastly different holes than people do, making the real challenge trying to make sure you're covering them. In most cases for me the best solution seems to be just giving them a way to test and attempt to prove things out in a realistic environment. But clearly, we haven't really left the era of having humans in the loop. Fully vibe-coded codebases clearly suffer from a myriad of issues.
Multiple security holes. Issues marked fixed that aren't really. Giant holes left in solutions.

Sol over engineers now and then (hey please don't factor that function out into its own file....) but it doesn't do the same level of stupid terra does.

That said, plan with Sol, implement with terra, have Sol fix all the mistakes, then I go over the code and make recommendations for the architecture to fix Sol's foolishness.

My experience is that current models are pretty good at making functional changes without too many more bugs than a human would make, but are still bad at making good high quality changes. I have a degree in software engineering specifically, so perhaps I am overly sensitive to design issues.
This is true in some sense.

Getting the AI to output code that you like is difficult.

As an example, let's say in React you have a "useLocale()" hook.

The AI will happily pass down locale as a prop to 5 child components instead of just calling the hook in the component.

A review from another model did not flag such stylistic issues either.

I believe that the latest models are very good at functionally achieving the goal, but still have poor taste for UX or code quality.

The most productive use of AI for software development happens in an environment where you do not review the code but test the UX end to end.

I use the AGENTS.md to show it how i want the code to look like. Something like "when implementing hooks adhere to the guidelines in docs/react-hooks.md". And then react-hooks describes your heuristics and what you consider best practices. There is a clear difference in code quality for me when using codex with a well crafted AGENTS.md vs. without one, you can run the experiment yourself pretty easily. As I mentioned in another comment, I think Claude poisoned users to stop relying on their Claude.md files and new codex users might be surprised at how well it adheres to guidelines.
We use skills for similar purposes, with positive and negative examples.

I think it sometimes worked, for example for testing preferences, but sometimes it did not.

Could be a problem with the harness also.

In any case, I feel that it's a bit playing whac-a-mole with explicit rules for things that a more intelligent model should do by default.

AI-pilled obsession with "taste" is bordering on insanity

It's just vibes

It's vibes all the way down.

It's... really just vibes?

Always has been.

That's what they're always going to be, so not sure what would be "insane" about it. They literally feed on and emit natural language, and are put to work on informally defined, arbitrary tasks.

When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ). This is not easy.

This sounds about as unhinged as Google saying Material Design 3 is 30% more rebellious
(comment deleted)
Doesn't just sound like it, it is. That's life for you. That's what I'm pointing out to you. [0]

Just consider your own example. Do you think a less or more "rebellious look" is not something designers can actually ellicit? Less so in software design, sure, but in character design for example? Or general product design? Do you think e.g. Monster energy drinks are branded the way they are completely due to happenstance or something?

Except people don't usually put numbers to it, because they understand that that's hard to defend. You're the one who's describing such an idea, and wants such a thing to happen, classifying anything else as just vibes and unhingedness. You're handwaving the difficulty and fundamentally limited nature of that, assuming that it is some laziness or mental delusion that's preventing it instead. What I'm telling you is that you're wrong about that.

The guy above didn't put numbers to his vibe assessment, they just drew a comparison, exactly because they know that there's not much else they can earnestly offer. You're sulking at them not lying to you by overstating their rigor, and you're flipping the arrow as if this limitation was some sort of mistake, not a necessary and intrinsic property, which it absolutely is.

[0] In fancier and more mathematical terms: https://abeljansma.nl/2026/07/10/truth-is-not-a-direction.ht...

I think we’re saying the same thing?? It is literally just vibes and all this faffing over which model has better “taste” is pointless

Where you’re wrong is pretending there is any intellectual rigor to the discussion which justifies promoting from the domain of vibes to actual reasoned debate

It's easy to dismiss "taste" when you either have none or just fail to appreciate it, but nothing gives you an appreciation for the importance of taste like LLMs. There is no benchmark for taste, so while many things improve taste does not. Bad taste is, in fact, a huge component of what makes AI slop so sloppy.

But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.

Thinking the giant array of GPUs has “taste” is exactly the problem
You are confusing personification, which is purely a rhetorical device, with anthropomorphization. I am not anthropomorphizing GPUs. There is nothing particularly weird about personifying model weights, as we do with computer programs, cars, and all other manner of inanimate objects every day.

To be honest, in this case, I wasn't even personifying them, because I in fact didn't say that an array of GPUs has "taste", or in fact even that model weights did. I was saying I preferred GPT 5.6 Sol's outputs as a matter of taste. I can see why someone would confuse the two statements since in this case they're pretty much the same thing, but if you re-read what I said I was actually more careful than you're giving me credit.

Yeah at this point claude is overrated, overly expensive, weird writing style (elliptical), and the worst part is the aggressive guardrails that even normal convos get interrupted, meanwhile openAI is still I would say at the normal balance, if you ask something too obvious or direct it will stop you other than that, it work flawlessly, plus, I have yet to hit the limit despite heavily using it these past weeks.
[delayed]
You've replied incredulously to a similar stated experience in this thread already and proceeded to ignore the follow-up. Why are you again asking a question to which you have no intention to field an answer?
I left both comments 12 hours ago. The first reply to my other comment was 11 hours ago.
Sol is way too eager to hone in on small details and ends up with massive over-engineering. Fable does it too - to be fair - but noticeably less.

After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.

Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value, by far.

---

Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.

Both models suffer from doing-too-much though. I think it's really close and pricing cuts really spice things up for us consumers!

Interesting the use of Max and Ultra. I don’t doubt the complexity, but would someone use Max or Ultra on Typescript or Go, for example?

Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?

*The "Max" I referred to was the plan tier, not the effort level btw

For small tasks, you can just use something like low or medium effort and it can usually avoid mistakes; after all, the model will test the code anyways and can do some baseline level of iterating.

In regards to cost, we need to acknowledge how generous OpenAI was in the last couple months with Codex usage credits (no weekly limits) and usage resets. It afforded me many a dive with Codex! Yes it uses more tokens, but sometimes it's worth it -- just depends on what you're working on.

Finally, Ultra(code) isn't that bad when it comes to cached tokens. I think folks overstate the general token usage of ultra effort on both providers.

---

Both models are great at green-fielding a project when given detailed specs.

Both models overthink too liberally (imo) during these larger multi-shots. Sol overthinks more than Fable.

Both models are really smart and perform great for general knowledge and regular coding tasks.

(comment deleted)
Yeah, I very much agree on this. I think Sol and Fable code quality is on par. Maybe Fable is just a tiny bit better, but Sol compensates with its ability to work through things, while Fable, in my experience, generally tends to avoid solving problems that require many LOC.

However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).

One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.

I get a lot of mileage using Fable to spec, then Sol to review the spec, then Fable to plan, then Sol to review the plan, then Fable to implement, then Sol to review the implementation. It's a lot of steps, but a great boost in quality of output.
with what levels of effort? I only see you mention sol-5.6 on ultra, which I don't find is worth it at all. Sol-5.6 on medium is amazing.

https://winstonrc.github.io/ai-coding-agents-leaderboard/

Sol 5.6 Ultra doesn't seem worth it because its long-horizon task scaling is not that good.

Remove the one thing that differentiates these 2 SOTA models (long-horizon task scaling) and you're left with assessing the raw intelligence of both models.

Both models are really fucking smart. We should be intentional when discussing effort levels when it comes so SOTA models because currently, effort levels are the essential lever to evaluate task scaling.

I've never seen that leaderboard link but I think my sentiments reflect exactly the findings: Sol 5.6 is smart, Fable is smart, but Sol is more value for the end user (even more so when you lower effort levels because that doesn't degrade the model's raw intelligence/knowledge). Not to mention Codex resets, that's just the cherry on top!

But smart != capable and this is evident once you start assessing both models on long-horizon tasks with higher effort levels. While Fable is (imo) at least a little bit better, both are still very good. If you want to test the raw intelligence of said models, you should lower the effort level.. if you want to test the model's capabilities fully, you should increase the effort level.

Economics aside, Fable is the better model (imo), but there's no need for a binary stance here. Both models are very good yet there is a clear winner on the value front.

Sol has held stuff for a while to do the same sort of hazard checks I assume Fable is doing, but it always releases them. I think that's the better way to handle it rather than preventing me from seeing how far I can get generating schematics to use in Minecraft. Currently: a mostly normal voxel house.
Sol is my daily driver but there are still times I reach for Fable when Sol doesn’t cut it. Just yesterday for example, I was trying to build a self-modifying hot-reloaded agent harness in Elixir for fun and Sol just kept doing silly things like thin wrappers and unnecessary abstractions. Fable handled the task elegantly. Sol is really good as a reviewer for finding bugs due to its thoroughness however.
I too switched to OpenAI after I got sick of Anthropic's constant "safety" downgrades. Sol is definitely a breath of fresh air.

> even after completing the verification program

Was it easy to complete it?

I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.

Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.

I have a "strategy / life-coach" project, and was surprised at how much better Sol is than Fable on it, as I've found Fable to have the edge for most things for me so far. But Sol: questions were better, insight was better, it got the brief better.
Claude as a harness at all really spends too much time before giving user feedback

Its a crutch that is no longer competitive

you should check out the codex desktop app. people who've been using claude code for a long time will surely be surprised.
I was pleasantly surprised to find that the GPT models are much stricter in adhering to my AGENTS.md guidelines and heuristics than Claude.
sol is much better imho than Fable but i can understand if they will perform wildly different for different people with different levels of expertise aswell as different needs. I dislike fable myself it doesnt really work for me.

Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.

cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).

what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...

for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.

Okay I'm a bit late to this, but my side project is www.freepi.ai, it's free ad+training powered inference in a Pi Harness. So not sure if that something you'd be down to try but if you did and have feedback I'd love to try and make it work for you!
5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more/less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.
I cancelled my subscription recently and moved to Sol. So far - it has been a great experience. The only aspect where Fable/Claude is better I feel is doing some research from the web and summarising the facts.
I generally prefer Fable but in my experience Sol is a much better web researcher
It's the complete opposite for me. The model might be the worst model I have ever used when compared to other models in the class. You just can't get it not to just write the most enterprise over complex over engineered solutions for every little thing you ask it to do.

It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.

I don't even want to see the codebases this model is fucking up.

It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.

The final straw for Claude was its refusal to give me a list of the most recent rapes reported by the BBC and basic information about them (location, date, names, just things reported in mainstream media). It outright REFUSED to complete this task.

I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.

output token efficiency bruv. u never go wrong with it
Used claude since 7/2025. Switched to codex after fable got blocked. It was still 5.5 but I knew they had to come up with something. As soon as I switched, wow. It wasn't super intelligent, but it was stable. Every day it was the same performance. This consistency is definitely worth paying for.
I switched from claude to gpt when 5.6 came out for the same reasons. I don't understand why so many people are still using claude when GPT and open source are so much better.
Really?

I have witnessed 5.6 Sol Ultra edit line after line of literally empty lines ... for hours.

I wasn't literally watching it, I came back to a goal (that it started for itself without my approval!) that had done nothing but that for some reason.

It couldn't explain why it had started.

indeed, I'm switching from claude to codex
I also cancelled Claude recently. GPT 5.6 models are really good, and they have much better usage limits.
Fable is the only one that follows my instructions correctly, which I find quite important. It takes STYLE.md and SPEC.md as law, and code just like I would code, with the same mistakes and all.

I just can't get Opus (Opus 5 is dumb as a rock, to be fair) or Sol to do that, so I exclusively use Fable for personal work. When I reach my weekly limit, usually on the last day close to the reset, I just go back to coding by hand ¯\_(ツ)_/¯

Heck, it even does security reviews and fixes, as long as I don't ask it to "attack" the codebase. I'm planning on using Kimi or GLM for that part.

Does OpenRouter eat this cost to get their hands on a copy of the conversations people are using with the model?
No, this is OpenAI doing the discount, not Openrouter by themselves. OpenAI is crushing it with their 5.6 models, and they probably decided there was no better time to grab as much market share as possible.
I don’t understand this at all. They have never been profitable yet. How is this helping them? When it be more likely the case that not enough, people are using it as the prices they established already? So now they have to lower the prices?
They're prepping to IPO. They want top level metrics like usage they can use to pump investors, not nonsense like profitability.
You lower prices for marketshare. Fable became a mythological model to leadership because they were the first story of ai escaping and hacking another company. The it's so dangerous the public can't use it narrative is sticky so OpenAI is showing off its model so as many eyeballs as possible. We're in the samples in the supermarket phase.
But the discount is only available via OpenRouter.
OpenRouter offers 1% discount to save your conversations, explicitly opt-in.

Anything else they don't save it. Even if they tell you the model provider saves your data for training.

For context, Stripe has just acquired OpenRouter for >$7B.

I’d bet that explains this move!

Why would the potential acquisition have anything to do with this? They do discounts all the time on various models. Luna was 50% off last week..
Not sure as OpenAI models (Sol, Luna,..) are also discounted on the Vercel AI Gateway rn. My bet is on OpenAI trying to drive more enterprise customers to their models through API.
I saw this for Luna and then looked at the uptime and it said 85%. My interpretation is that this is just a gimmick where they serve the OpenAI flex tier at the same discount OpenAI provides for flex and then fall back to azure
Even at these prices, switching from subsidized subscriptions to the API just isn't worth it. Not even close.
Oh hey, that's cheaper than Kimi K3! Amusing to see a SOTA OpenAI model be cheaper than a Chinese open weight model.

Fwiw I love K3 and use it as a daily driver. I haven't tried Sol, as I dislike OpenAI.

Has anyone had mixed experience running Ultra with and without /goal? I come back to it after 8 hours to find it got stuck navel gazing imagined and Byzantine errors.
How can they do this? Are they subsidizing it out of pocket?
Absolutely not

I’ve used Claude exclusively for the past few months

Was excited when Sol came out a few weeks ago and loaded it up

I made the mistake of treating it as if it were Claude - I’d assumed they were close enough in ability and treated them that way

Well, turns out my instruction sets for Claude are 100% too complicated for Sol

Sol made the stupidest assumptions, constantly did things that it wasn’t asked to do and always approached code in what I considered a weird way - I had redo a lot of my prompts to get it anywhere close

Now, did it do good work?

Yes, on occasion. But with LLMs and coding, consistency is the name of the game. Constantly having to correct the LLM and constantly feeling paranoid that it won’t listen makes for an exhausting session

Maybe if you “came up” in the codex world you’re more fluent with it, but sticking with Claude for now

Why are you being downvoted, is this post an ad or something.
Probably because it's PEBCAK.
Hard to take offense from someone who uses pebkac seriously

Kudos to you though for being your authentic self so publicly

"rewrite unreal engine in rust, make no mistakes"
That’s exactly what I did! How’d you know?

Didn’t mean to make you upset sorry

At this point the models are “good enough” and whoever wins long term is gonna be whoever is the cheapest.

That’s why Chinese models are gaining traction and it’ll be the only way for OpenAI or Anthropic to keep up.

Reading the comments in this thread, i honestly dont get it. 5.6-sol has felt like a regression in capability. In fact, every model since 5.3-codex has been a regression from OpenAI. I just find 5.6-Sol over engineers problems, takes absolutely ages to solve basic problems....

At this point, I'm considering going back to cursor over codex due to the ability to get more control over what model I use since there is clearly a heap of user preference and having frontier providers constantly shift the goal post with "State of the Art" is complete non-sense.

It depends on what effort you're using etc. As an example [1] of what codex is capable of, here's hugo (written in golang) ported to TypeScript - and then a TypeScript to Rust transpiler which converts arbitrary TypeScript into Rust.

The TypeScript code which was transpiled into Rust (and is compatible with most hugo templates) runs faster than the original hugo.

[1]: https://github.com/tsoniclang/tsonic-examples/tree/main/rust...

The transpiler is still WIP, but the fact that it can do this says a lot of about how far LLMs have come.

I find all effort levels of sol are the same in terms of amount of hallucinated unnecessary changes. Luna is much better all round on xhigh but my point still stands, every release of these new models is not an upgrade, its re-learning how to work with it.

Its like rehiring an employee every few months then training them up. Its honestly tiring and cant stay like this.

Opus has the same problem too…

This sure looks like a race to the bottom to me, and I love it.

If Sol isn't the best model, it is up there...

You don't cut the price of the best model for no reason...

> This sure looks like a race to the bottom

Always has been. My prediction is that both OpenAI and Claude will go bust unless they deliver a killer product. And unlike scrappy startups, they have a pretty serious deadline because creditors will come a-knockin'.

There's little to no functional difference between Kimi, Qwen, Sol, Fable, etc. All flagship models are within like 1-5% of each other and the real moat will be what's always been the hard part: making a good product.

> All flagship models are within like 1-5% of each other

Don't know about that.

I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc.

Ran a complete Fable/max code review. Took over a month on a subscription. Now I've switched to OpenAI and am repeating the exact same review with Sol/max.

It's still not done yet but preliminary findings suggest Sol can only reproduce 70-90% of Fable's findings. So I think these models aren't as close as we've been led to believe.

Check the remainder for hallucination for sure.
check all for verifiability
I literally had Fable tell me that I've checked out a nonexistent upstream commit of LMDB - two days ago. Sol called its BS, of course.
Absolutely. I have an independent audit pass to verify claims.

My methodology consists of launching a 242 cell parallel code review matrix and committing all Fable/Sol max effort agent prompts and their full reports to a private orphan branch on the repository. This is the part that is taking me months to complete. This thing can kill my $100 subscription in about 12 hours.

When done, these raw findings will be semantically deduplicated and merged into a list of findings per model. This list will then be audited for hallucinated or otherwise made up findings. This will refine the list, and hallucination rate is its own data point. I'm also counting things like cybersecurity refusals and downgrades.

There is a massive difference even between Opus and Fable, same provider, before various harnesses and other optimizations come into play. Don't be deceived by rankings and benchmarks, try for yourself.
The problem is that most of the volume doesn't come from proprietary products, it comes from API use which has no stickiness.

Claude already has a killer product (claude.ai/chat is a Swiss army knife) but just relying on people typing stuff into chat is not enough to sustain the company.

The other strategy is entrenching yourself as the LLM of choice into existing products (like ChatGPT is on Apple products).

> All flagship models are within like 1-5% of each other

Depends on your use case. the Chinese models are not there yet.

Compute is the moat.

The Chinese models are cheap because no one is using them. But they can't actually afford (or have capacity) to serve enough people to kill the giants. This is evidenced by them all recently hiking prices or limiting usage.

It's possible that they build out in China at an unreal pace, China doesn't have concept of "community input" to drag down state projects, but then you are left giving your IP to China. Just ask western hardware businesses how well that goes.

Isn’t almost all of anthropic and OpenAI’s compute leased?

If compute is the moat, that doesn’t really make their position any less precarious.

No one is in a better position to more efficiently use it - that's what happens when you poach every top 0.01% engineer/researcher in AI.

They're guaranteed to get over whatever hump you think they're in unironically. Uber/Tesla have been in far worse situations and despite Elon being an idiot/liar you see how they performed when even the most bullish of investors called for their heads

I think I generally agree with you - OpenAI and anthropic will probably succeed here. If the ai bubble pops, they’ll come out on top.

I just think the whole “moat” discourse is silly. OpenAI and Anthropic’s success depends on the same thing every business’s success depends on: their customer base, and their continued delivery of services their customers want to pay for. Not their tech or their compute or anything else. They have no moat because moats aren’t a thing.

It's hard for me to imagine how you could define "killer product" to exclude something that takes you from $9B to $65B ARR in 8 months.

https://epoch.ai/data/ai-companies?view=graph&tab=revenue

It's not hard to sell a dollar for 50 cents. Imo, these businesses are pretty clearly not doing well financially (the revolving door of unvested C-levels is a good hint, the constant postponing of S-1s is another).
That graph looks a lot less impressive when not "annualized". Annualized revenue can be gamed in a number of ways, and is a big reason companies tend to compare YoY to investors once they're public.
Have you found an alternative to codex / Claude code? I’ve tried opencode but it’s been buggier / feels like an inferior product to both.
The less the merrier.
Competition is good for users.
If you believe https://artificialanalysis.ai/

This is basically undercutting KimiK3 and Grok 4.6 where previously utilised gad soke advantages but was a step more expensive

But who can take a benchmarking website seriously when they literally change their benchmarks to appease anthropic, like artificialanalysis.ai did the other day?
Is this pricing change only for openrouter? I don't see official OpenAI info about this.
I was wondering the same, and it clearly says : 50% off, aka a sale, not normal price cut.

I don't get this thread.... Really. Is it full of bots?

I doubt bots, the linked page is really boring to read and looks like a dashboard instead of easily consumed reading. I think people are therefore drawing conclusions from the headline more than usual since the link makes the real information sort of opaque.

However it's ignorant to think that there aren't bots on HN and especially for the very many motives people have for swaying public opinion.

They are really getting desperate. The bubble is coming closer to an end here.