You need to validate what providers are actually serving. Add benchmarks, properly showcase what quantization they are serving on the model and KV cache, etc. Until that happens, your service is doomed to be shitty.
I've had Astra say that it is a Qwen model. They are all cross-trained and distill each other.
I wouldn't trust OpenCode Go with my lunch after the shit-sandwich they served us with horrible V4 quants and 0 transparency. Then blaming it on their partners.
Considering the fact that Google/Anthropic/OpenAI have WAY more compute and the race is this close, it's obvious that DeepSeek/GLM/Qwen teams are better or we're approaching a wall in terms of progress.
Garbage.
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
The only good take here.
Especially since they still serve Codex-Spark, which is dogshit.
But Bonsai is garbage.
As if our EU leaders aren't a complete joke as well. Pushing ChatControl, fascism and gambling everywhere. We're just as much of a joke.
This will degrade performance significantly. LLama.cpp has had this for a while and it tanks benchmark performance. I ran GPQA on GLM 5.2 using the llama implementation and it came back 19 points under the regular…
Or you can use parlor to chat with it directly https://github.com/fikrikarim/parlor/
Better than Opus 4.8 on complex tasks but tends to overthink. It found a bunch of bugs and architecture issues that only 5.6 Sol Max and Fable on my C++ projects.
It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
Yup. Smells like marketing.
[flagged]
They've been investing heavily into this over the last 2 releases, but it's just that SideFX is a really small company. I think they have 50 people or so right now. Karma is way more accurate than Cycles, but slow.…
And that's when you learn Houdini and open up a workflow from 2008 and it still works the same, despite SideFX pushing out an insane amount of features every release (look at their SneakPeaks).
Work more for my boss and landlord so the fiefdom can survive.
It also goes to show that Fable/Sol must be 4-5T in size.
Ling/Ring 1T-A50B and the new Inkling 975B-A41B deserve to be on that list.
Which would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.
A shame it has Marc Maron in it.
Please don't use that garbage. Just use the base Qwen models or Nex/Orinth, as those are the only properly post-trained finetunes. The Qwopus models are marketing.
"DeepSeek-V4-Flash will fit" At Q2, 2bit? Lobotomized to death.
You need to validate what providers are actually serving. Add benchmarks, properly showcase what quantization they are serving on the model and KV cache, etc. Until that happens, your service is doomed to be shitty.
I've had Astra say that it is a Qwen model. They are all cross-trained and distill each other.
I wouldn't trust OpenCode Go with my lunch after the shit-sandwich they served us with horrible V4 quants and 0 transparency. Then blaming it on their partners.
Considering the fact that Google/Anthropic/OpenAI have WAY more compute and the race is this close, it's obvious that DeepSeek/GLM/Qwen teams are better or we're approaching a wall in terms of progress.
Garbage.
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
The only good take here.
Especially since they still serve Codex-Spark, which is dogshit.
But Bonsai is garbage.
As if our EU leaders aren't a complete joke as well. Pushing ChatControl, fascism and gambling everywhere. We're just as much of a joke.
This will degrade performance significantly. LLama.cpp has had this for a while and it tanks benchmark performance. I ran GPQA on GLM 5.2 using the llama implementation and it came back 19 points under the regular…
Or you can use parlor to chat with it directly https://github.com/fikrikarim/parlor/
Better than Opus 4.8 on complex tasks but tends to overthink. It found a bunch of bugs and architecture issues that only 5.6 Sol Max and Fable on my C++ projects.
It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
Yup. Smells like marketing.
[flagged]
They've been investing heavily into this over the last 2 releases, but it's just that SideFX is a really small company. I think they have 50 people or so right now. Karma is way more accurate than Cycles, but slow.…
And that's when you learn Houdini and open up a workflow from 2008 and it still works the same, despite SideFX pushing out an insane amount of features every release (look at their SneakPeaks).
Work more for my boss and landlord so the fiefdom can survive.
It also goes to show that Fable/Sol must be 4-5T in size.
Ling/Ring 1T-A50B and the new Inkling 975B-A41B deserve to be on that list.
Which would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.
A shame it has Marc Maron in it.
Please don't use that garbage. Just use the base Qwen models or Nex/Orinth, as those are the only properly post-trained finetunes. The Qwopus models are marketing.
"DeepSeek-V4-Flash will fit" At Q2, 2bit? Lobotomized to death.