16 comments

[ 11.8 ms ] story [ 861 ms ] thread
(2025)

As it's 9 months old and they just had a major model release

Another banger from Zhang et. al
If you want to believe that the success of Kimi is about distillation attacks, ignore this.
Old but relevant: if you read the recently-released Kimi K3 paper[0], you'll see that it's heavily based on Kimi Linear discussed here, scaling it up and adding a bunch more things (like native vision and RL improvements).

[0] https://arxiv.org/abs/2607.24653

I started creating internal models using it, then the Gated Deltanet 2 came out( https://arxiv.org/abs/2605.22791), and it seems like an evolution of it in expressiveness. And in our tests it is really better than.
Does any expert in the field know whether it is really the case that this intelligence we are seeing with frontier models is an "emerging" phenomena, only coming up when the architecture is scaled?

Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture?

It's unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won't beat quicksort if you put more processing to them, but in the modern LLM scene it seems people are in a race to scaling up, experimenting empirically and hoping the same algorithm/architecture comes to a solution.

My expertise lies in deep learning theory, and yes, the "intelligence" is coming primarily from scaling up, among other things. There are good reasons for this, but essentially it comes down to taking advantage of a narrow statistical trick, where a very well-crafted model/optimizer pair that has a strong implicit bias toward simplicity can exhibit progressively increasing performance with respect to model size. Marcus Hutter's lab has shown that you can phrase this in terms of Solomonoff induction, so this bias is truly universally effective. An effective bias can continue to improve performance with larger model sizes by taking advantage of the curse of dimensionality in a way not dissimilar to how more data generally gives you a better answer (indeed, there is a duality taking place here, but I digress).

To be clear, it is an extremely narrow model class that can do this; we just got "lucky" and worked our way to it. That's why we still teach general statistical principles which often forbid this sort of behavior as a rule of thumb.

I'm out of the scene these days, and my field is more RL than LLMs, but my take is that all of this is telling us that meaning (and intelligence) is in the medium. If an LLM is a function mapping an input to some stateful internal representation that in-turn then gets mapped to an output, a smaller model probably doesn't have the capacity to learn that internal mapping, at least not from scratch. Or it would take a very long time, in same way that a naive sorting algorithm would still get there, but it might be computationally prohibitive in practice.

The bigger model has a better chance of gaining a foothold in that internal representation space where inputs are mapped to meanings and outputs, and eventually, it optimizes to the point where most of the weights aren't doing much. It's not immediately clear how much expressivity is required by the network to learn that space, but so far the answer seems to be in the billions of parameters.

The more interesting question to me is to what extent we should expect the models to be invariant to data. For instance, if I learn a certain type of analysis, that skill shouldn't depend on the data I'm looking at-- it should be repeatable for any given data of the same type/class. I'm curious to what extent skills are embedded in the weights versus data and "facts". My hunch for why mathematical reasoning and programming resulted in large step changes in model performance across the board is because these are inherently skills that are widely repeatable for a large class of tasks. The ability to express programmatic logic is invariant to both the language and the task at hand. And to me, that's how you get to smaller models: by focusing on the skills.

Increasing raw model scale is one of the most consistent and reliable ways of increasing its intelligence.

In a way, training an AI is: using an algorithm to find, discover and refine other algorithms computationally. When we scale, we pour more raw inputs into finding the right algorithms for a given objective. Is it surprising, then, that we find a better algorithm for it?

As a very stupid analogy - for a given task, a small model's internal algorithm might, under its capacity and training signal constraints, top out at slightly above "bubble sort". While a larger one could dig deeper and get closer to "quicksort" internally. A naive, dirty algorithm got replaced by a more sophisticated algorithm that performs better.

Keep in mind: intelligence is very much not a binary. There's no "threshold" at which a model goes from "this is dumb statistics" to "this is actual intelligence". A 1B LLM and a 10T LLM both have some amount of intelligence. It's just that one would have intelligence that's so weak, underdeveloped, and overindexed on statistical regularities that it's very easy to dismiss it altogether. And the other might have enough of it to snipe unresolved conjectures with novel counterexamples in math. Makes it considerably harder to dismiss it outright. While the curve between the still two looks less like an abrupt jump and more like little increments building up to an avalanche.

The jump in something advanced and specific like "math abilities" can look quite sharp - but under it, there are far more generic capabilities that back it. They build up slowly to eventually enable that jump. A more advanced model makes less reasoning mistakes and recovers from reasoning mistakes more gracefully - two very generic capabilities - but once those capabilities improve enough, a whole new type of logic problem might fall to it.

It is _emerging_ in the sense that complex (sic non-linear) systems can exhibit unintended or unexpected behaviors.

It is _intelligent_ in the sense that it optimizes a thing that is hard for humans to not anthropomorphize.

Things like percolation theory, swarm theory, and others are similar topics in “complex systems”. Neural networks are interesting because they combine aspects of both complex systems and dynamical/adaptive systems (ie systems with a feedback loop).

A neural network at its very core is a function fitting algorithm. It stores matrices of parameters (ie weights and biases) such that every parameter (and combinations thereof) captures a relationship of your data in exactly the same way the slope and intercept are obtained through Linear Interpolation. All of the Regularization tricks applied just are attempts to incentivize a given parameter to not encode any trends too specifically (think an instance vs a type).

In this way you can ask yourself “how many aspects of your data are required to capture it adequately?” This is what scaling offers.

Mice have around 70 Million Neurons, humans around 50-100 billion.

Yes, you can't compare biological neurons to parameters in an LLM 1:1, neurons do much more, but still - It is plausible that higher intelligence just requires scale.

At least not even evolution over millions of years seems to have found a way to produce human-level intelligence in a smaller scale, even though it had a lot of incentives to do so (the human brain needs a lot of kcal).

Emergent phenomena = pass/fail grading of multi part problems. If you're just gonna count it as a fail you're not gonna see the progress until everything works, even if every single subproblem improves linearly.
>To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

This is just awesome.

Any one knows how this holds up on long context retrieval (needle in haystack , ruler) vs same size full attention model? efficiency gains look great but that usually where linear attention hybrids fall apart.
if non standard transformers like this take off are companies like etched forked?