A flood of releases today, really difficult to make out for someone who does not use or test all these models on complex real world use cases as to how people decide which ones to use (besides price)
OpenAI and Anthropic need to just go ahead and give people access to the cyber models.
Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
I can run this at home. No guardrails, this is not shy of Sol and Fable, this crushes them in my book. It's not just about evals, but what I can do with the damn model.
>> This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results.
Agreed.
This release is the first time I'm able to employ a GLM model to write a substantive plan for a complex Clojure PR [1] with both Opus 5 and GPT-5.x playing supporting / reviewer roles.
Initial results are __very__ encouraging. GLM 5.3 -
- follows directions,
- digs into detail, and
- correlates well.
Still not confident about entrusting GLM with implementation - but IMHO, western labs are entirely cooked.
[1] 2K LoC PR in a 55K LoC Clojure + Clojurescript repo
I am in the process of creating my own Pi Coding Agent harness to leverage the power of Deepseek V4 Flash 0731 and other models (you can do that when you build your own harness! easily route opinions from other models whenever you're stuck, etc) and cancelling my Codex account next week.
No Hugging Face link yet. I wish they would release it under a true FOSS license.
Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
just imagine the world without these open weight models - we'd probably have to reverse mortgage our homes to pay for tokens to those trillion $ companies to have access to their models.
Feels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?
People familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this trend to continue? ByteDance is training a 10T-parameter model. Here, GLM 5.3 outperforms models 3-4x its size of roughly 700B, so parameter count doesn’t seem to be a direct correlation anymore.
I might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it?
It does feel like it respects both me and my time.
Thank you, Z.AI.
Amazing what difference it makes when the top of your org are actual university professors.
> Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.
I appreciate they don't just take the opportunity to self-glaze.
Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it.
For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.
The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
> Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
Data and content related to "biological weapons" already exist on the internet, in books, etc. The real issue is access to facilities and tools. There are models that help researchers, but they are not LLMs, rather they are models trained specifically on biological data (like AlphaFold).
Cybersecurity is basically used like a dog whistle pioneered by Anthropic to achieve regulatory capture. Otherwise, the widespread availability of good tooling for security analysis would eliminate more of these cyber threats, rather than gatekeeping them for a few private companies.
> Data and content related to "biological weapons" already exist on the internet, in books, etc. The real issue is access to facilities and tools.
No, I think tools are easy to come by (unlike in nuclear research), the real issue is the know-how to create biological weapons, which you can't easily get out of books, but much more easily out of an amoral LLM.
99 comments
[ 2.8 ms ] story [ 36.5 ms ] threadOtherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
Agreed.
This release is the first time I'm able to employ a GLM model to write a substantive plan for a complex Clojure PR [1] with both Opus 5 and GPT-5.x playing supporting / reviewer roles.
Initial results are __very__ encouraging. GLM 5.3 -
- follows directions,
- digs into detail, and
- correlates well.
Still not confident about entrusting GLM with implementation - but IMHO, western labs are entirely cooked.
[1] 2K LoC PR in a 55K LoC Clojure + Clojurescript repo
I am in the process of creating my own Pi Coding Agent harness to leverage the power of Deepseek V4 Flash 0731 and other models (you can do that when you build your own harness! easily route opinions from other models whenever you're stuck, etc) and cancelling my Codex account next week.
Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
Love this opening line. And wow, great results.
> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
It does feel like it respects both me and my time.
Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.
JieTang (Founder of Z.ai): It won't take that long
https://x.com/i/trending/2067626647050670400?lang=en
We just had amazing releases this past two months
kimi k3, glm5.3 qwen3.8 and now glm5.3
These open models are getting really good
I appreciate they don't just take the opportunity to self-glaze.
Original images: https://image.non.io/neonRamenDesigns.webp
GLM 5.3 build: https://html.non.io/neonRamenGLM5.3
Opus 5 build for comparison: https://html.non.io/neonRamen
For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.
The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
Data and content related to "biological weapons" already exist on the internet, in books, etc. The real issue is access to facilities and tools. There are models that help researchers, but they are not LLMs, rather they are models trained specifically on biological data (like AlphaFold).
Cybersecurity is basically used like a dog whistle pioneered by Anthropic to achieve regulatory capture. Otherwise, the widespread availability of good tooling for security analysis would eliminate more of these cyber threats, rather than gatekeeping them for a few private companies.
No, I think tools are easy to come by (unlike in nuclear research), the real issue is the know-how to create biological weapons, which you can't easily get out of books, but much more easily out of an amoral LLM.