I didn’t participate in the discussion yesterday because I find it implausible Fable was available long enough (3-4 weeks cumulative?) to get data and train and have it fundamentally affect.
But I don’t grok training enough to know that’s silly.
If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.
All frontier models have been trained without any regards for IP protection laws. I don't see how anyone can argue in good faith that distillation is not fair game and does not ultimately "benefit humanity™"
This data sort of disqualifies itself: unless Moonshot has a time machine, K3 should be more similar to Opus 4.5-4.8 than Fable 5.
Keep in mind, Anthropic started limiting access and introduced anti-distillation measures around 4.5-4.6 (?). So the majority of distillation should have happened on earlier models.
Maybe a better explanation is that they have access to the same training datasets? Which if private can again raise questions about theft, but on a very different level.
given that github is full of vibe coded repos, reddit is full of bot comments, and most blog posts are ai slop, isnt it guaranteed that anyone training on big public data sets will be closely tracking each other in text style?
Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2 weeks apart, and were unmistakably different. Results were wildly inconsistent run to run. It didn't stop the crowd believing its creator in that DeepSeek R1 was trained on o1-preview (which was obvious bullshit as well, they were as different as two models can be). It's amazing how you can put anything on the web and everybody will believe you without checking or even understanding of what they're looking at.
K3 was trained on Claude's outputs, though - it repeats Anthropic's prompt injections 1:1 in its reasoning, which you should know if you ever tinkered with both models long enough. Good for them.
15 comments
[ 1.4 ms ] story [ 25.5 ms ] threadThere is an optimal answer to any question. Something that maximizes utility and minimizes tokens.
But I don’t grok training enough to know that’s silly.
If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.
Keep in mind, Anthropic started limiting access and introduced anti-distillation measures around 4.5-4.6 (?). So the majority of distillation should have happened on earlier models.
Maybe a better explanation is that they have access to the same training datasets? Which if private can again raise questions about theft, but on a very different level.
K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?
Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?
So who's training on who's outputs?
https://news.ycombinator.com/item?id=48990086
https://x.com/stevibe/status/2026227392076018101
Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2 weeks apart, and were unmistakably different. Results were wildly inconsistent run to run. It didn't stop the crowd believing its creator in that DeepSeek R1 was trained on o1-preview (which was obvious bullshit as well, they were as different as two models can be). It's amazing how you can put anything on the web and everybody will believe you without checking or even understanding of what they're looking at.
K3 was trained on Claude's outputs, though - it repeats Anthropic's prompt injections 1:1 in its reasoning, which you should know if you ever tinkered with both models long enough. Good for them.