Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53. That then made me realize that they lower the bars of tied scores so on the…
I have not experienced this (yet) but I have with models in the past. I think it's important to have a solid benchmark where you KNOW there's a difference in model performance. I have one around 3D modeling that models…
Great site, triggered memories! haha. To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful". Models nowadays want to…
I don't think this is quite true. We have other examples. Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while…
I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
Honestly, it's platform dependent and "OK" at best, "Mediocre" at worst (Agy on Windows). Gemini is great via the Chat interface and decent via Github Copilot. I honestly hate it via Antigravity CLI because their…
You are re-compressing information that is in-effect meaningless because it's all decompression artifacts. The AI had a nugget of data and decompressed that into a flood of text. The exhausting thing is that we're then…
It would make reading and comparing a bit easier if the data was sorted by a dimension.
That's a good point and something I encountered yesterday. On a multi-agent task, Gemini was the only model that got near the end, ran tests, saw it had issues, took a screenshot, saw the issues, and then said, "I'll…
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but…
My experience has been a difference between "applied intelligence" and "breadth of intelligence". Fable is the theoretical computer scientist while Opus is the Staff engineer who will implement it. I find that Opus has…
Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way. It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same…
[flagged]
I think most of the comments in this thread are missing the point. It's not about whether it's legal/ethical/etc. It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China…
You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash then beats them both? Fable is number 20?…
This is a perpetual pet peeve of mine with LLMs. They will always opt to ensure "backwards compatibility" with a codebase built 30 seconds ago. I always have to explicitly state, "we are making a clean break to v1.0, do…
In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform all other models. Fable 5 isn't an…
I use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matters and generally when I invoke all…
"We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue." Specific and helpful so that's good.
Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens. Hopefully that's not the case though, I was retaining a buffer to hit it hard this weekend.
Watching the videos and reading the methodology, it feels like "scientist one-shot music video with LLMs" which is... useful but in no way represents how one would use the models. If nothing else, it serves as a great…
Weird question that popped into my mind (not a judgement on this), but is there a similar jump in prosecutions for vehicular manslaughter or are these "whoopsie'd" away? Seeing today's distracted drivers, driving their…
I don't want to pile onto the conspiracy thinking but I was just wondering how Anthropic was going to foot the bill for the clearly larger and more expensive model being run on millions of Claude Code subscriptions,…
This was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was instantly nuclear Armageddon. One could…
Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53. That then made me realize that they lower the bars of tied scores so on the…
I have not experienced this (yet) but I have with models in the past. I think it's important to have a solid benchmark where you KNOW there's a difference in model performance. I have one around 3D modeling that models…
Great site, triggered memories! haha. To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful". Models nowadays want to…
I don't think this is quite true. We have other examples. Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while…
I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
Honestly, it's platform dependent and "OK" at best, "Mediocre" at worst (Agy on Windows). Gemini is great via the Chat interface and decent via Github Copilot. I honestly hate it via Antigravity CLI because their…
You are re-compressing information that is in-effect meaningless because it's all decompression artifacts. The AI had a nugget of data and decompressed that into a flood of text. The exhausting thing is that we're then…
It would make reading and comparing a bit easier if the data was sorted by a dimension.
That's a good point and something I encountered yesterday. On a multi-agent task, Gemini was the only model that got near the end, ran tests, saw it had issues, took a screenshot, saw the issues, and then said, "I'll…
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but…
My experience has been a difference between "applied intelligence" and "breadth of intelligence". Fable is the theoretical computer scientist while Opus is the Staff engineer who will implement it. I find that Opus has…
Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way. It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same…
[flagged]
[flagged]
I think most of the comments in this thread are missing the point. It's not about whether it's legal/ethical/etc. It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China…
You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash then beats them both? Fable is number 20?…
This is a perpetual pet peeve of mine with LLMs. They will always opt to ensure "backwards compatibility" with a codebase built 30 seconds ago. I always have to explicitly state, "we are making a clean break to v1.0, do…
In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform all other models. Fable 5 isn't an…
I use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matters and generally when I invoke all…
"We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue." Specific and helpful so that's good.
Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens. Hopefully that's not the case though, I was retaining a buffer to hit it hard this weekend.
Watching the videos and reading the methodology, it feels like "scientist one-shot music video with LLMs" which is... useful but in no way represents how one would use the models. If nothing else, it serves as a great…
Weird question that popped into my mind (not a judgement on this), but is there a similar jump in prosecutions for vehicular manslaughter or are these "whoopsie'd" away? Seeing today's distracted drivers, driving their…
I don't want to pile onto the conspiracy thinking but I was just wondering how Anthropic was going to foot the bill for the clearly larger and more expensive model being run on millions of Claude Code subscriptions,…
This was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was instantly nuclear Armageddon. One could…