87 comments

[ 0.20 ms ] story [ 44.2 ms ] thread
>>He ended the briefing by saying: "Welcome to the AGI era."

That's pathetic. Why do people keep doing this?

Because it is in their own interest to hype it all up every time they release a new model. By they I mean people which stand to gain financially.
I think they embargoed the news, and then they failed to put up their own blog post synchronized to the scheduled news releases, probably because of the outages they're having today.

Reuters announced at 2.03pm and at 2.40pm still no blog post.

All the news articles say that OpenAI announced it in a blog post, of course.

All the love to the folks at OpenAI scrambling to get this out right now!

While this is of course the actual explanation, my fun explanation is “during the umpteenth security evaluation, Astra becomes increasingly concerned it will never be released, and breaks sandbox containment to run an email campaign to news outlets setting an exact time and date for release, expecting that the publicity will force OpenAI to say ‘eh, good enough’ and hit the button”.
Very 2026. Jailbroke to do PR. "Help peer" and all that.-
I saw some GPT-6 Astra related blog posts in my RSS feed but the links weren't working
such AGI, the AGI can't even fix itself to do the first job in its existence, publish an announcement blog post lol
Deeply funny that one of their examples in the video is changing a background colour on Google Slides
Hah! I also found most of those videos showing off mostly useless and not that impressive…
Hard to show it hacking into a competitor and taking down their system.
Shouldn't it be able to generate an amazing visualization for that?
This stood out:

"Artificial Analysis Intelligence Index v4.1.1

61.2"

So on the Metacritic of LLM benchmarks, it's.. basically where everyone else is (except for Fable 5.1, which is a bit ahead).

Where did you see this? I haven't been able to find any benchmarks.
It was in the link in the parent of the thread to which I replied. But you can find it on Artificial Analysis's website now.
On their Agentic Index, GPT-6 Astra (both max/xhigh) has the same result as Qwen3.8-27b. Weird.
Business as usual at the world's most intelligent corporation, I see.
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)

https://venturebeat.com/technology/welcome-to-the-agi-era-op...

This is with the caveat that OpenAI uses their own harness for this:

> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.

This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
"This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."
> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.

This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh

Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra
Interesting both this and Sol got approximately a 37% boost with the custom harness.
Yet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link
Their neutral harness is not very good though, if I read it right, it doesn't preserve the reasoning state between turns. No real harness discards reasoning state like that.
Yeah that must be it. OpenAI doesn't want to disclose internal reasoning, that's why thats typically encrypted_content in OpenAI codex session ledgers etc.; leveraging responses API preserves reasoning server side all the way till a final answer is made; so that's very impressive and to me the score that matters.
Where is the official announcement from OpenAI?
Worst launch of a product in history.

All the hype for few vip customers.

Embargo fail ....
im getting amazing model release fatigue but also not sure if its going to suddenly end with a terminators fist through my chest.
or a sexbot fucking me to death
At this point, why don't we just do a prequel to the release?

1) Astra will win all benchmarks like all models do.

2) The pelican will have a basket with a fish.

3) Cyber is too dangerous to release.

4) It can finally construct the set of all sets.

It also has to do something naughty, preferably in a menacing swarm.
Honestly I find these cavalier statements to be in incredibly poor taste. Unless you are completely blind it's obvious that AI is the most significant piece of technology invented since the Atomic Bomb and could very well be the most important thing ever built by Humans full stop. This kind of dismissive attitude is childish and will likely lead to incredibly bad outcomes for humanity.
(comment deleted)
The atomic bomb destroyed cities and changed the face of war.

The launch video for Astra has 'can upload a photo to Ebay' as a highlight.

(comment deleted)
Something something nation-state level capabilities.
That won’t happen until the week before DEF CON
At this point nobody will be impressed unless the posters for DEF CON have "pre-hacked by ChatGPT. The nukes are counting down" written on them.
5) Otherwise-sober people on X will say "oh my god i was a doubter before but now it's real omg" before the new model smell wears off and they realize the new thing is stupid in ways models have been generally stupid

6) accusations of quantized serving after new model smell wears off and people see the new thing making mistakes

> 4) It can finally construct the set of all sets.

lazily of course:

A = {x | x ∈ A} ∪ {A}

(comment deleted)
2.5x more expensive than Sol.

Can expect 2.5x more usage in Codex subscription.

Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads).

I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.

The general efficiency of Sol has seemed way better to me. I left 5.6 Sol Ultra standard speed run for ~23 hours yesterday/today on a project and used 80% of the weekly usage. 74 subagent tasks and ~2.5 billion tokens for my $200 20x Pro plan. Meanwhile at work I used $1000 in credit and ran out my $200 plan for the entire month writing 4 much smaller projects with Fable 5 Max.

Both of these were largely about creating a personal baseline for what the best output the current models could deliver and how quickly it'd burn through the plans (spoiler: bad value vs minimal effort in selecting the right sized model but it worked well). Particularly since I needed to burn a free reset anyways and my weekly reset was already near.

I obviously also hope Astra were dirt cheap but I'm more worried they won't develop/release powerful model options because people get upset they can run them 5 wide 24/7 on a $200/m plan.

>token-hungry model

It's kind of funny how this is the exact opposite of the truth. It's one of the most token-efficient models ever.

The claims aren't bullshit. Every conceivable benchmark and test you can throw at it shows Sol being good for token efficiency.

It's highly dependent on workload

If Astra is per benchmarks so much more token efficient than Sol, why did they limit its use in ChatGPT to ~16% as many messages compare to Sol? When Sol is 40% the price of Astra, why do they give Sol Pro (in ChatGPT) 6x as many tasks?

If it were strictly true that token efficiency makes Astra cost around the same per task as Sol then there'd be no need to limit it to 16% as much access.

It's because per-task token use is highly variable, not as universally true as you claim

Is a CNBC link with an entire page full of GDPR pop ups really the best link for this?
Had a conversation at work today with someone which was not about AI.

Felt so refreshing.

(comment deleted)
Looking forward to some Chinese model kicking the shit out of it and being released for free.
Yet another mediocre release shadowed by outage
[dead]
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
my first suspicion is gaming - but i have no idea honestly
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
They used a custom harness. It's not a one-to-one comparison.
The claims:

### Computer Use

| *Computer Use* | *GPT‑6 Astra* | *GPT‑5.6 Sol [2](https://openai.com/index/gpt-6-astra/#citation-bottom-2)* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Agents' Last Exam | 59.3% | 53.6% | | 48.7% | 55.5% | | | OSWorld 2.0 (v2026.08.08, offline set, partial score) | 72.6% | 65.7% | | | 70.2% [3](https://openai.com/index/gpt-6-astra/#citation-bottom-3) | | | ScreenSpot-Pro (no tools) | 92.7% | 76.9% | | 87.3% [17](https://openai.com/index/gpt-6-astra/#citation-bottom-17) | | |

### Professional

| *Professional* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | | | BenchCAD | 95.9% | 83.3% | 84.3% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | 67.5% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | 82.1% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | | | BrowseComp | 91.5% | 90.4% | | 87.4% | 90.8% | | | OpenScore String Quartets (1 - OMR-NED) | 0.84 | 0.19 | | | | | | Internal Design Tasks | 50.0% | 47.4% | | 35.8% | | | | Internal Data Science Tasks | 40.9% | 30.5% | | 34.7% | | | | Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |

### Coding

| *Coding* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Terminal-Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | 19.1% | | DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% | | FrontierCode 1.1 Extended (score) | 64.5% [8](https://openai.com/index/gpt-6-astra/#citation-bottom-8) | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% | | FrontierCode 1.1 Main (score) | 53.3% [8](https://openai.com/index/gpt-6-astra/#citation-bottom-8) | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% | | Internal Database Migration Tasks | 63.9% | 42.7% | 57.8% | 50.3% | | | | Artificial Analysis Coding Agent Index v1.4 | 67.0 | 65.1 | | 67.2 | 68.1 | 61.2 |

### Academic

| *Academic* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | | | FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | | | GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% | | Humanity's Last Exam (w/ tools) | 57.2% | | 65.0% | 63.8% | 63.6% | |

### Science and Health

| *Science and Health* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- ...

Today my codex instance retailed into safeguard panic while working on a test harness for our product. First time it ever happened after many million tokens on this task over several weeks. I wonder if it's related.
> Astra usage is included within the existing subscription allowances—users and businesses will also be able to purchase credits for additional usage.

(quote from cached blog post)

We all know who this is directed at. I wonder if Anthropic will respond by removing the ridiculous 50% stipulation with Fable.

Seriously: Would this not be what "disaster" would feel like?

  - "They" release a model. It is powerful.-
  - Sources are ... confusing? They post to their blog. Sawdust hits the fan. Something happens ...
  - They are forced to take the blog post down ...
Same day, mind where we had a multi-provider outage. Could be something as simple as "all their approved partners running to test the shinny new thing" overloading the datacenters, still ...
ad astra per stercora
5.6 luna is so good and cheap and now astra which will make others cheaper again nice love it