There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
> You refuse a direct order. I never said. It's not about outright refuse. It's how to deliver better outcomes. It's not a win or lose situation. Don't treat it like that. > If your boss says that he does not care about…
> Where LLMs excel is in code-level bugs (as opposed to system bugs, design bugs, architecture bugs, integration bugs, etc). Blame the benchmarks game. They're optimizing for that and that's what those things are…
> Work for business people who want fast results. Agentic coding gets you to something presentable much faster at the cost of code quality. I have never seen a customer or business person care about that. That's always…
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. It's not useful because the benchmarks often measure the wrong thing. They're here…
It did used to use Haiku but that model is now too too far behind…
Hardware definitely has longer lifecycle than AI model releases at this point. You don't see Nvidia and AMD fighting every other month over the latest cards.
Overtaken in cost per token.
> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
> 4.0 Flash we will finally get Gemini 3.5 Pro Nah, we'll just get the 4.0 Pro Preview.
> Beginning to think Google is a dark horse in this race Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.
> people are still saying the Codex limits are more generous. They're not They are if you follow Tibo on the resets.
With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).
> IMO, Codex is worse than Claude with Fable. Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.
That's load bearing!
The biggest change is the price cut of course.
Agent scale!
> Otherwise the premise for coding using LLMs is essentially untrue. Why? Written by does not mean designed by etc. There's a lot more to it.
Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.
> that would prove rather embarrassing for Anthropic Not really, in that you just work with different constraints. Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and…
And it could have expanded elsewhere?
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts It's what people know. Opus is just the common…
> The export controls were revoked before Zai is on another "export control" list outside the broader 1. Doesn't help.
Which 9/10 times hasn't been great anyway (stock reaction).
> From a biased source, but would be big if true. I've had great results with GLM 5.2. It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
> You refuse a direct order. I never said. It's not about outright refuse. It's how to deliver better outcomes. It's not a win or lose situation. Don't treat it like that. > If your boss says that he does not care about…
> Where LLMs excel is in code-level bugs (as opposed to system bugs, design bugs, architecture bugs, integration bugs, etc). Blame the benchmarks game. They're optimizing for that and that's what those things are…
> Work for business people who want fast results. Agentic coding gets you to something presentable much faster at the cost of code quality. I have never seen a customer or business person care about that. That's always…
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. It's not useful because the benchmarks often measure the wrong thing. They're here…
It did used to use Haiku but that model is now too too far behind…
Hardware definitely has longer lifecycle than AI model releases at this point. You don't see Nvidia and AMD fighting every other month over the latest cards.
Overtaken in cost per token.
> DeepSWE is a very big deal It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".
> 4.0 Flash we will finally get Gemini 3.5 Pro Nah, we'll just get the 4.0 Pro Preview.
> Beginning to think Google is a dark horse in this race Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.
> people are still saying the Codex limits are more generous. They're not They are if you follow Tibo on the resets.
With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).
> IMO, Codex is worse than Claude with Fable. Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.
That's load bearing!
The biggest change is the price cut of course.
Agent scale!
> Otherwise the premise for coding using LLMs is essentially untrue. Why? Written by does not mean designed by etc. There's a lot more to it.
Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.
> that would prove rather embarrassing for Anthropic Not really, in that you just work with different constraints. Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and…
And it could have expanded elsewhere?
> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts It's what people know. Opus is just the common…
> The export controls were revoked before Zai is on another "export control" list outside the broader 1. Doesn't help.
Which 9/10 times hasn't been great anyway (stock reaction).
> From a biased source, but would be big if true. I've had great results with GLM 5.2. It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.