> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. Build a better simulator to train them in (i.e. more expensive) that includes a…
> Before December 2025 they were still intelligent code autocomplete or Stack Overflow bots This is false. Coding agents have been usable since at least May of 2025. I can't speak to earlier than that as May last year…
Here's how it make senses: - reality they can't keep up their pace - that's bad for forward projections of their revenue - that's bad for their stock price - being 'forced' to slowdown by the government is less bad for…
The lack of access to brute force their training seems to be resulting in them training more efficiently too such that they are quickly catching up despite current restrictions. Between that and the chip manufacturing…
We've had software factories for decades. They're usually called compilers, linkers, toolchains, etc.
Who is this vendor that is consistently providing high quality inference for all families of open weight models at a competitive cost? That's a serious questions that I really interested in the answer too. I have 25…
It's more than just quantization. The middleware the provider is running matters a lot even to the point of exactly which version they are running due to defects being introduced / resolved. In my coding agent harness…
Take a computer that can run the biggest LLM available today. It can also run any smaller LLM as well. It can also run software that isn't an LLM at all. Brains and LLMs are not at all equivalent as LLMs lack a stateful…
They aren't grown/evolved from data, they are fit to the data. The fitting process can be fully deterministic although its fairly easy to screw things up such that it isn't deterministic, but that just a defect not some…
Simple cellular automata demonstrate emergent behavior. Emergent behavior is nothing new in computer science and is not remotely unique to LLMs.
This assumes the functionality of brains can be fully captured as a deterministic mathematical function, but the function of the brain may well depend on nondeterministic quantum states that can't be reduced to…
The problem with this line of thinking is that modern computers are nothing like the brain. LLMs don't stand on their own, they have to be run on these modern computers, but doing so does not change the physical…
Not at all. We all did it when I was child that age in the 80s.
Claude Code may simply be best used with Anthropic's models and quite bad with Kimi. An alternate solution is to remove Claude Code from the diagram if its so far off from the others that it causes scaling problems. I…
I was excited until I saw cost was only provided as the median. Your provider will bill you for all your tasks and one can get back to that total from the mean by multiplying by the number of tasks. This isn't possible…
We're done here.
I've been using Brave for years and its great. Having JS off by default (i.e. shields) is excellent for security. They provide MV2 versions of AdGuard and uBlock Origin that you can enable directly from the browser…
> Third you’re going to have to use TS for the frontend anyway. You can’t escape ts. This is obviously wrong. Sure you have to use Javascript for the web frontend that has interactivity without round-tripping to the…
That's AI. That's the most flattering thing I can think of to say about statements that are so confidently wrong.
Emergent behavior is not new in the realm of software. Cellular automata has emergent behavior that can get pretty wild. For example: https://en.wikipedia.org/wiki/Lenia There are huge differences between brains and…
They act like LLMs. They are unprecedented.
You are correct.
Maybe's its a problem with the hosting at novita.ai but I didn't got much useful out of this model as a coding agent.
I was hoping to see the reasoning_content mess get robustly fixed, but all we got was this doc change: https://github.com/vllm-project/vllm/pull/50624
That token dump is the funniest thing I've ever read here.
> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. Build a better simulator to train them in (i.e. more expensive) that includes a…
> Before December 2025 they were still intelligent code autocomplete or Stack Overflow bots This is false. Coding agents have been usable since at least May of 2025. I can't speak to earlier than that as May last year…
Here's how it make senses: - reality they can't keep up their pace - that's bad for forward projections of their revenue - that's bad for their stock price - being 'forced' to slowdown by the government is less bad for…
The lack of access to brute force their training seems to be resulting in them training more efficiently too such that they are quickly catching up despite current restrictions. Between that and the chip manufacturing…
We've had software factories for decades. They're usually called compilers, linkers, toolchains, etc.
Who is this vendor that is consistently providing high quality inference for all families of open weight models at a competitive cost? That's a serious questions that I really interested in the answer too. I have 25…
It's more than just quantization. The middleware the provider is running matters a lot even to the point of exactly which version they are running due to defects being introduced / resolved. In my coding agent harness…
Take a computer that can run the biggest LLM available today. It can also run any smaller LLM as well. It can also run software that isn't an LLM at all. Brains and LLMs are not at all equivalent as LLMs lack a stateful…
They aren't grown/evolved from data, they are fit to the data. The fitting process can be fully deterministic although its fairly easy to screw things up such that it isn't deterministic, but that just a defect not some…
Simple cellular automata demonstrate emergent behavior. Emergent behavior is nothing new in computer science and is not remotely unique to LLMs.
This assumes the functionality of brains can be fully captured as a deterministic mathematical function, but the function of the brain may well depend on nondeterministic quantum states that can't be reduced to…
The problem with this line of thinking is that modern computers are nothing like the brain. LLMs don't stand on their own, they have to be run on these modern computers, but doing so does not change the physical…
Not at all. We all did it when I was child that age in the 80s.
Claude Code may simply be best used with Anthropic's models and quite bad with Kimi. An alternate solution is to remove Claude Code from the diagram if its so far off from the others that it causes scaling problems. I…
I was excited until I saw cost was only provided as the median. Your provider will bill you for all your tasks and one can get back to that total from the mean by multiplying by the number of tasks. This isn't possible…
We're done here.
I've been using Brave for years and its great. Having JS off by default (i.e. shields) is excellent for security. They provide MV2 versions of AdGuard and uBlock Origin that you can enable directly from the browser…
> Third you’re going to have to use TS for the frontend anyway. You can’t escape ts. This is obviously wrong. Sure you have to use Javascript for the web frontend that has interactivity without round-tripping to the…
That's AI. That's the most flattering thing I can think of to say about statements that are so confidently wrong.
Emergent behavior is not new in the realm of software. Cellular automata has emergent behavior that can get pretty wild. For example: https://en.wikipedia.org/wiki/Lenia There are huge differences between brains and…
They act like LLMs. They are unprecedented.
You are correct.
Maybe's its a problem with the hosting at novita.ai but I didn't got much useful out of this model as a coding agent.
I was hoping to see the reasoning_content mess get robustly fixed, but all we got was this doc change: https://github.com/vllm-project/vllm/pull/50624
That token dump is the funniest thing I've ever read here.