I think they embargoed the news, and then they failed to put up their own blog post synchronized to the scheduled news releases, probably because of the outages they're having today.
Reuters announced at 2.03pm and at 2.40pm still no blog post.
All the news articles say that OpenAI announced it in a blog post, of course.
All the love to the folks at OpenAI scrambling to get this out right now!
While this is of course the actual explanation, my fun explanation is “during the umpteenth security evaluation, Astra becomes increasingly concerned it will never be released, and breaks sandbox containment to run an email campaign to news outlets setting an exact time and date for release, expecting that the publicity will force OpenAI to say ‘eh, good enough’ and hit the button”.
This is with the caveat that OpenAI uses their own harness for this:
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.
This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh
Their neutral harness is not very good though, if I read it right, it doesn't preserve the reasoning state between turns. No real harness discards reasoning state like that.
Yeah that must be it. OpenAI doesn't want to disclose internal reasoning, that's why thats typically encrypted_content in OpenAI codex session ledgers etc.; leveraging responses API preserves reasoning server side all the way till a final answer is made; so that's very impressive and to me the score that matters.
Honestly I find these cavalier statements to be in incredibly poor taste. Unless you are completely blind it's obvious that AI is the most significant piece of technology invented since the Atomic Bomb and could very well be the most important thing ever built by Humans full stop. This kind of dismissive attitude is childish and will likely lead to incredibly bad outcomes for humanity.
5) Otherwise-sober people on X will say "oh my god i was a doubter before but now it's real omg" before the new model smell wears off and they realize the new thing is stupid in ways models have been generally stupid
6) accusations of quantized serving after new model smell wears off and people see the new thing making mistakes
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads).
I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
The general efficiency of Sol has seemed way better to me. I left 5.6 Sol Ultra standard speed run for ~23 hours yesterday/today on a project and used 80% of the weekly usage. 74 subagent tasks and ~2.5 billion tokens for my $200 20x Pro plan. Meanwhile at work I used $1000 in credit and ran out my $200 plan for the entire month writing 4 much smaller projects with Fable 5 Max.
Both of these were largely about creating a personal baseline for what the best output the current models could deliver and how quickly it'd burn through the plans (spoiler: bad value vs minimal effort in selecting the right sized model but it worked well). Particularly since I needed to burn a free reset anyways and my weekly reset was already near.
I obviously also hope Astra were dirt cheap but I'm more worried they won't develop/release powerful model options because people get upset they can run them 5 wide 24/7 on a $200/m plan.
If Astra is per benchmarks so much more token efficient than Sol, why did they limit its use in ChatGPT to ~16% as many messages compare to Sol? When Sol is 40% the price of Astra, why do they give Sol Pro (in ChatGPT) 6x as many tasks?
If it were strictly true that token efficiency makes Astra cost around the same per task as Sol then there'd be no need to limit it to 16% as much access.
It's because per-task token use is highly variable, not as universally true as you claim
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
Today my codex instance retailed into safeguard panic while working on a test harness for our product. First time it ever happened after many million tokens on this task over several weeks. I wonder if it's related.
Seriously: Would this not be what "disaster" would feel like?
- "They" release a model. It is powerful.-
- Sources are ... confusing? They post to their blog. Sawdust hits the fan. Something happens ...
- They are forced to take the blog post down ...
Same day, mind where we had a multi-provider outage. Could be something as simple as "all their approved partners running to test the shinny new thing" overloading the datacenters, still ...
87 comments
[ 0.20 ms ] story [ 44.2 ms ] threadThat's pathetic. Why do people keep doing this?
[1]https://theonion.com/amazing-new-hyperbolic-chamber-greatest...
Reuters announced at 2.03pm and at 2.40pm still no blog post.
All the news articles say that OpenAI announced it in a blog post, of course.
All the love to the folks at OpenAI scrambling to get this out right now!
"ChatGPT maker claims its ‘Astra’ could be considered ‘artificial general intelligence’" - https://www.ft.com/content/55ab40c0-59e2-4c0b-97c9-4f4f5a71a...
"Artificial Analysis Intelligence Index v4.1.1
61.2"
So on the Metacritic of LLM benchmarks, it's.. basically where everyone else is (except for Fable 5.1, which is a bit ahead).
https://venturebeat.com/technology/welcome-to-the-agi-era-op...
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.
This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh
All the hype for few vip customers.
1) Astra will win all benchmarks like all models do.
2) The pelican will have a basket with a fish.
3) Cyber is too dangerous to release.
4) It can finally construct the set of all sets.
The launch video for Astra has 'can upload a photo to Ebay' as a highlight.
6) accusations of quantized serving after new model smell wears off and people see the new thing making mistakes
lazily of course:
A = {x | x ∈ A} ∪ {A}
Can expect 2.5x more usage in Codex subscription.
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads).
I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
Both of these were largely about creating a personal baseline for what the best output the current models could deliver and how quickly it'd burn through the plans (spoiler: bad value vs minimal effort in selecting the right sized model but it worked well). Particularly since I needed to burn a free reset anyways and my weekly reset was already near.
I obviously also hope Astra were dirt cheap but I'm more worried they won't develop/release powerful model options because people get upset they can run them 5 wide 24/7 on a $200/m plan.
It's kind of funny how this is the exact opposite of the truth. It's one of the most token-efficient models ever.
The claims aren't bullshit. Every conceivable benchmark and test you can throw at it shows Sol being good for token efficiency.
If Astra is per benchmarks so much more token efficient than Sol, why did they limit its use in ChatGPT to ~16% as many messages compare to Sol? When Sol is 40% the price of Astra, why do they give Sol Pro (in ChatGPT) 6x as many tasks?
If it were strictly true that token efficiency makes Astra cost around the same per task as Sol then there'd be no need to limit it to 16% as much access.
It's because per-task token use is highly variable, not as universally true as you claim
Felt so refreshing.
### Computer Use
| *Computer Use* | *GPT‑6 Astra* | *GPT‑5.6 Sol [2](https://openai.com/index/gpt-6-astra/#citation-bottom-2)* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Agents' Last Exam | 59.3% | 53.6% | | 48.7% | 55.5% | | | OSWorld 2.0 (v2026.08.08, offline set, partial score) | 72.6% | 65.7% | | | 70.2% [3](https://openai.com/index/gpt-6-astra/#citation-bottom-3) | | | ScreenSpot-Pro (no tools) | 92.7% | 76.9% | | 87.3% [17](https://openai.com/index/gpt-6-astra/#citation-bottom-17) | | |
### Professional
| *Professional* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | | | BenchCAD | 95.9% | 83.3% | 84.3% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | 67.5% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | 82.1% [5](https://openai.com/index/gpt-6-astra/#citation-bottom-5) | | | BrowseComp | 91.5% | 90.4% | | 87.4% | 90.8% | | | OpenScore String Quartets (1 - OMR-NED) | 0.84 | 0.19 | | | | | | Internal Design Tasks | 50.0% | 47.4% | | 35.8% | | | | Internal Data Science Tasks | 40.9% | 30.5% | | 34.7% | | | | Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
### Coding
| *Coding* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Terminal-Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | 19.1% | | DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% | | FrontierCode 1.1 Extended (score) | 64.5% [8](https://openai.com/index/gpt-6-astra/#citation-bottom-8) | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% | | FrontierCode 1.1 Main (score) | 53.3% [8](https://openai.com/index/gpt-6-astra/#citation-bottom-8) | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% | | Internal Database Migration Tasks | 63.9% | 42.7% | 57.8% | 50.3% | | | | Artificial Analysis Coding Agent Index v1.4 | 67.0 | 65.1 | | 67.2 | 68.1 | 61.2 |
### Academic
| *Academic* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- | | Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | | | FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | | | GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% | | Humanity's Last Exam (w/ tools) | 57.2% | | 65.0% | 63.8% | 63.6% | |
### Science and Health
| *Science and Health* | *GPT‑6 Astra* | *GPT‑5.6 Sol* | *Claude Fable 5.1* | *Claude Fable 5* | *Claude Opus 5* | *Gemini 3.8 Flash* | | --- | --- | --- | --- | --- | --- | --- ...
(quote from cached blog post)
We all know who this is directed at. I wonder if Anthropic will respond by removing the ridiculous 50% stipulation with Fable.