1,572 comments

[ 0.27 ms ] story [ 67.5 ms ] thread
Dead link for me
I saw it
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
$10 per million input tokens and $50 per million output tokens

sol is $4 / $20

It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
neuralese is pretty token efficient i guess.
They're just announcing later availability. No launch.
Every frontier release nowadays is "we've launched*"

* for a special group of customers that you're not in. Keep waiting peasant.

I mean tell Nvida to 100x their hardware output and you'll get what you want.
Their announcement about later availability is unavailable to me now (500 error).

Great first impression.

> We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours.

This is from Tibo on X.

GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...

Performance is significantly higher than Fable 5.1

Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/

100% on ExploitBench seems fitting given recent events.
I think we need a few writing related benchmarks.
any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
Why would Anthropic trust and use these tests in their official comparisons?
great username lol
do you feel you're free of bias and predisposition in saying this
> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

Artificial Analysis just published their aggregate score (61).

Still well below Fable 5, let alone Fable 5.1.

I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.

If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.

I saw this too and I'm really confused.
We need them pelicans on bikes.
Its time to move on to the flamingo on a unicycle bench
That looks more than slight.
Looks very capable

404

Archive locks one shelf

Dust spins softly through the stacks

Browse one row nearby

by gpt-5.6-sol

Bridge ends in midair

Wind sketches the farther bank

The far bank draws near

by gpt-5.6-sol

Prompt blooms into verse

I count syllables, not rain—

Whose noticing?

Generative Pretrained Transformer 5.

I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
If it's not clear what's happened:

The launch was scheduled for 11am Pacific time.

The press embargo broke at 11am, and we saw a flurry of press articles by Axios, TechCrunch et al.

The model is visible on the ChatGPT API.

But the official blog post is not out yet after nearly an hour.

Apparently the article was posted then quickly taken down, hence there are snippets of information coming out.

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.

Did humans deploy the model? Or did the model deploy itself?

Sounds like "AGI" just stands for "IPO" as it always has done.

(comment deleted)
(comment deleted)
(comment deleted)
Why is this flagged ?
You should know: AA index is only 61. Pretty surprised it’s that low.
More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?

And in the past, gemini 3 pro was rated as high as opus 4.5 and the like

Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.

This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
Oh brotha, here we go again, it's so over again, as every week nowadays
I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.

But thank you for spending other peoples money to give us the tech regardless!