Kinda same but I miss my 5.3 Codex. Thing lasted forever on my $20 subscription and with detailed prompts was able to pretty much implement everything I requested it to do with a acceptable quality.
Same. 5.5 got work done then 5.6 was also fine then 6 was maybe not quite as good. Now with 6.1 they are cutting usage and raising prices and introducing ultra fast mode, but things were good enough 5 months ago.
Astra is noticeably smarter than any OpenAI model before it. Sol 6.1 is very noticeably smarter than sol 6 even after half a day of using it (sol 6 was actually terra 6 and opus 5.5 has taken them by a total complete surprise)
Because you're probably using it for stuff like "edit this function"/"refactor this class". 5.5 to 5.6 Sol was a giant jump. 5.6 to 6.1 seems very large as well.
It's the same for most tasks. Where I do notice it is long agent runs, where agents take more steps and the performance difference definitely compounds over the iterations.
This has got to be a panic move from OpenAI, right? They’ve had some bad press lately because from their billing changes, and Anthropic have finally released a fast, relatively cheap Opus with improved written English.
What about those do you think are "shady"? Price discrimination in favor of small customers at the expense of large customers is somewhat common; businesses want customers to buy more of their products, but customers are not obliged to buy more if they like their current deals. Having limits on using finite resources seems even easier to justify.
Humans see differences in prices as unfair. The greatest example of this is price gouging during an emergency, but also look at the level of hate that scalpers get.
Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.
Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.
I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.
Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.
Shady as in not clearly communicated and change with random tweets being the only indicator, or a blog post if you're lucky.
x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".
That doesn't resolve them from providing a stable service for something they market. I'm not buying something under the counter, and it's not advertised to be an alternative for people not willing to pay the full price. From my perspective, the subscription is the full price and I expect stable, reliable service.
I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.
> We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day
Was it not essentially confirmed that all of the frontier labs have much stronger models internally that don't make sense to serve at scale yet, which they use for development and to train models that are able to be served at scale?
Still waiting for the moment where GPT either orders 4,000 pounds of raw beef or just deletes the whole OpenAI repo because it "thought the easiest way to remove all bugs is by deleting the whole repository"
It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be
be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
In my case it actually made a difference. I tested it on my own code review set with Gemini Flash. On low it found fewer bugs than on default around 90% versus 97% but was about three times faster and a lot cheaper. On high it actually found everything in the hardest case but took almost three minutes per call. so I think the difference is real you only see it if you run the same fixed cases a few times and not by feel. And yes I used Gemini rather than the GPT Sol that is mentioned here would be interesting to see the same tests on Sol.
I suspect it's partly because people didn't jump from GPT-3 to GPT-6.1 Sol and partly because SOTA models from the last few(?) months have been able to tackle most of the regular tasks. It means this new model isn't different in that regard from Opus 4.8, if your mental benchmark is that they both are capable of implementing something like a CRUD app.
I like the fast releases. 6.1 coming so fast after 6.0 means they found some improvement solid enough for a new rollout.
The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now
I’ve seen overwhelmingly that when a model is good people see it, and when it isn’t, they criticize.
I’ve seen nothing but ‘wow this is a huge step up’ from Opus 5.5. I felt this way about Opus 4.5, GPT-5.6 Sol/Luna, and to a lesser extent with Fable and Astra.
But Opus 5 was ass, and the entire gpt 6 line feels like OpenAI’s version of that.
They've been pretty capable for a very long time. I don't think the models are getting more capable so much as people are getting better at using them and more people are getting the opportunity to be impressed.
They probably figured they could increase their margins by releasing "6 Terra" as "6 Sol". After being surprised by the Opus 5.5 launch, they release the true "6 Sol" as "6.1 Sol" and with very aggressive pricing.
> the cache read discount rises from 90% to 95%. GPT-6.1 Sol’s overall blended price for agentic workloads is therefore slightly lower than GPT-6 Sol. This represents an additional price cut, following GPT-6 Sol’s original 50% discount from GPT-5.6 Sol
Wow, just wow. We are racing to the bottom with these prices.
50 comments
[ 3.0 ms ] story [ 14.6 ms ] threadThe results are less buggy, animations are much better.
It can work autonomously for hours and the result is decent most of the time.
That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.
Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.
Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.
I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.
Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.
x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".
Couponers have to work extra to find and maintain a good deal.
The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.
We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day
…but not for all models, which is pretty annoying.
Nothing, really. It's like oversampling your data set. You usually get a much better overall baseline performance if you use the default setting.
Every time a new model comes out, people come out in droves "oh I don't notice anything different".
People have been saying this about <currentModel-1> for 2 years now, and the entire state of AI has changed dramatically.
It cannot be that the next AI model isn't better, but also suddenly what they are capable of is on an entirely different level.
The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now
If the previous one was as good as worth retiring in 7 days, there's a good chance the new one is also not great.
"We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run"
Now I can finally build that half a thing I've had my eye on.
Still seems to me like google is in a good position with AI just based on their size and approach.
Wow, just wow. We are racing to the bottom with these prices.