Any previewers have hot takes? I've really preferred gpt-5.5 over Opus 4.8 for data analysis and scientific software work. It seems much more reliable. Fable is unusable for the type of work that I do (due to guardrails). Really looking forward to trying these new OpenAI models out.
> It's a damn good model. Not quite as "smart" as Fable, but it is incredibly capable. Fixed all the problems I had with GPT-5.5.
> It is incredibly determined. Will run for a day without even using a /goal. It understands subagents incredibly well and is great at orchestrating. It's super pleasant in use cases like OpenClaw and Hermes Agent. It knows iOS dev incredibly well.
> It has rough edges too, but FAR fewer than 5.5 did.
> For many things, gpt-5.6-sol will become my obvious defaults.
> It is better about [following instructions] than 5.5 was. Understands intent well and hammers until it gets there. Sometimes a bit too hard.
Also[^1]:
> gpt-5.6-sol is world leading in computer use. It made me use it 100x more. When we lost access to 5.6, I quickly started to go insane without it
I will stop here, sorry but I think we have limited time to listen to opinions and nowadays since they are abundant on social media we should give preference to the substantiated ones.
I’m bouncing back between Codex and Claude like a ping-pong ball. I much prefer the experience using Codex, less verbose and to-the-point I’ve found. But Fable, being as strong as it is, is a big draw for Claude right now. I’ll likely switch back to Codex if 5.6 Sol is comparable.
Damn this is exciting. I love that gpt models are much faster, efficient and cheaper than Claude models. They are so fast even on high/xhigh that I don’t find myself using the parallel agent setup anymore much since its cognitively less demanding to just follow along what the model is doing and most tasks it will complete in <5-<10mins anyway.
I'm most curious about whether OpenAI finally taught its models how to design interfaces. They have been behind the other labs in this area for what feels like ages.
I find codex way more usable. It’s not pretentiously verbose like Claude. It’s also responsive - I can see the progress easily and steer the conversation. With Claude, it might take 15 minutes and I would lose patience.
Coding with AI it feels like if you're not using the best model then you're possibly missing out - creating less capable, maintainable, just plain 'good' code. Why waste time using anything less than the best and cleaning up the mess later on. This is why I feel like local models and Chinese models aren't taking off (and Gemini/Grok) - they work, but they're plain just not as good as OpenAI/Anthropic. If you have the money then it doesn't make sense to code with anything else.
I’ve been using mostly deepseek v4, kimi k2.6, and gpt 5.3-codex
I sometimes chuck a few tokens to gpt 5.5 and opus 4.8 and they can sometimes solve a problem one of the other models couldn’t, but they’re not like 10x better or anything in my experience. More like 1.2x better
Earlier I predicted that Fable and Sol would be of similar capability, I think I will be wrong. Here is why: there is no indication that there are any classifiers like in Fable. I think OpenAI found out how to lobotomise the model without classifiers but the tradeoff is that it is a weaker model. I wonder how people feel about that. Would you like a highly intelligent jagged model with classifiers or slightly less intelligent smooth model without classifiers?
I know a few of my comments are related to this, but these new names are horrible. Why introduce ANOTHER layer of confusion and drop the mini, nano suffixes that people got used to?
How does this go through so many layers of management at a trillion dollar company without who has a say raising this? I simply can't believe how stupid the naming scheme from OpenAI was and continues to be even after they acknowledged it earlier.
I've been running a custom enterprise agent on 5.4 and it's been very good so far. I am looking forward to trying it with the monster model to see if we can approach some additional business cases.
I think if you are not seeing reasonable performance in your agent loops as of 5.5, it's likely there is a deficit with how the loop, prompt or tools interact with the environment.
Fable 5 for the planning, thinking, reasoning part, then GPT 5.5 to implement is an almost perfect combo, with Fable then reviewing GPT's code.
Codex CLI just seems faster at coding than Claude Code but Fable is just a level above intelligence wise, it's truly like taking to very very very smart human.
With GPT 5.6 though will be interesting to see if things flip, to have Codex speed (or faster) with Fable level intelligence is a game changer.
I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc.
In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?
28 comments
[ 3.6 ms ] story [ 76.9 ms ] thread> It's a damn good model. Not quite as "smart" as Fable, but it is incredibly capable. Fixed all the problems I had with GPT-5.5.
> It is incredibly determined. Will run for a day without even using a /goal. It understands subagents incredibly well and is great at orchestrating. It's super pleasant in use cases like OpenClaw and Hermes Agent. It knows iOS dev incredibly well.
> It has rough edges too, but FAR fewer than 5.5 did.
> For many things, gpt-5.6-sol will become my obvious defaults.
> It is better about [following instructions] than 5.5 was. Understands intent well and hammers until it gets there. Sometimes a bit too hard.
Also[^1]:
> gpt-5.6-sol is world leading in computer use. It made me use it 100x more. When we lost access to 5.6, I quickly started to go insane without it
[^0]: https://nitter.net/theo/status/2074708892341481755 [^1]: https://nitter.net/theo/status/2074720467395756499
I will stop here, sorry but I think we have limited time to listen to opinions and nowadays since they are abundant on social media we should give preference to the substantiated ones.
I sometimes chuck a few tokens to gpt 5.5 and opus 4.8 and they can sometimes solve a problem one of the other models couldn’t, but they’re not like 10x better or anything in my experience. More like 1.2x better
How does this go through so many layers of management at a trillion dollar company without who has a say raising this? I simply can't believe how stupid the naming scheme from OpenAI was and continues to be even after they acknowledged it earlier.
I think if you are not seeing reasonable performance in your agent loops as of 5.5, it's likely there is a deficit with how the loop, prompt or tools interact with the environment.
My quota is about to reset. Really can’t wait to use it.
Though its been just 3 days I started using.
Half way through the chat, GPT 5.6 Sol stops and does a safety verification, pretty annoying
Codex CLI just seems faster at coding than Claude Code but Fable is just a level above intelligence wise, it's truly like taking to very very very smart human.
With GPT 5.6 though will be interesting to see if things flip, to have Codex speed (or faster) with Fable level intelligence is a game changer.
In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?