208 comments

[ 0.26 ms ] story [ 51.9 ms ] thread
Proof - they also claimed that China has an ASML UEV machine - crickets when ASML said it was impossible due to all the safeguards and assistance needed to operate one.

The current US administration is known to be collection of BS artists and liars.

how is it possible to distill fable only a month after its release? maybe they are confusing opus with fable.
I think they deserve, by Justice, to have their models pillaged and raped, just like they did to the internet. They didn't ask for permission when they took the entire of the internet, after all, and given their behaviour is nefarious, it's of Justice that they receive nefarious treatment by others, including chinese AI labs.

The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.

Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited.

How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?

I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies

I think the accusation implies Kimi has gained time travel capability (distilled from fable probably) to have enough time distilling fable. Given they can travel time now, I think it is fair to call them a threat to national security.
If distillation truly is the cheat code they act like it is, then all the US and EU AI labs have no excuse for not having Fable-level models already
Do you get a token trophy for a few (many) trillion tokens purchased in distilation?
(comment deleted)
Yeah if anything Kimi's ability to distill that quickly is a major technological breakthrough
"we have information" says a US Government official who almost certainly has had Anthropic and/or OpenAI on the phone spinning him stories.

See also, don't trust anyone in Trump's government who says "we have information".

"they distilled us" is fast becoming standard US FUD.

The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".

I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]

     It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...
Assuming they did then they surely paid for them, which makes it "not stealing". Am I also "stealing proprietary U.S. technology" by harvesting my Claude chats from my `.claude` directory and training a bunch of models on them?

That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.

This constant FUD spread by Anthropic is so tiring.

I wonder how they detect this kind of thing. Seems like this is going to be a perpetual issue until it stops being worth doing.

Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?

Is distillation something we have to live with or are there ways to prevent it?
If you build a device that can help build devices, you shouldn't be too surprised when people use it to build devices.
To the extent that we have to live with AI and LLM, I believe we have to live with distilling also.
(comment deleted)
If web scraping is legal, so is distilling.
I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?
It’s the same as the patent argument. If everyone could freely copy everything then yes in the short term prices would drop and consumers would benefit, but over the long term it would discourage investment into new technology because a return would be impossible.
“If you can’t compete with them, get them banned”

- US AI companies

Anthropic should think hard about all their fear mongering. It will only end up backfiring on them and everyone else involved.

They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165

or other similar "AI" startups https://x.com/envconfig/status/2079613455296827402

Didn't they just pay a fine for stealing all those books?
You wouldn't distill a car.
wearing a lab coat and safety glasses and mixing cool looking liquids in a beaker Oh, I certainly would.
Poor AI labs... All they hard earned training, done via scraping a lot of people works for free, now being scraped through payed subscriptions...
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.

So here robbers are blaming robbers?

These claims are just pointless, everytime

If they did this in the US they would almost certainly be sued.

Meta, for examples, doesn’t want employees to use Claude Code due to distillation risk.

(comment deleted)
honestly if they did what he said they did, it seems like it would be cheaper just to train your own model from the get go
Does this matter? Distillation is not illegal by every definition of the word.

There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.

And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.

It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.

The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".

Legality or lack thereof is one of many data points about whether a thing is noteworthy.

Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.

As consumers this is probably a positive development.

It does matter in that these LLM companies need to be run into the ground, and every embarrassing clod working for them run out of town.

It's showing that 'distillation' is a viable way to reclaim all of what they stole and hoard, and with enough luck their debts will come due in time for them to feel it.

I agree distillation isn't illegal; I also think Moonshot/Kimi is very impressive. But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic. If you can only play catchup (however quickly you do that), then you're never going to be at the frontier - I think that's why distillation matters.
> The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens

Note that Chinese companies are free to rent from GB300 clouds internationally. There are large datacenter hubs in Singapore and Malaysia serving chinese and other customers.

Though there is also reported [1] significant smuggling of Nvidia chips into China as well.

[1] https://epoch.ai/publications/chip-smuggling

Boy I tell you, I am having an awful hard time summoning pity for the organizations that have themselves distilled all of humanity's knowledge into mysterious labor-market-masticating black boxes.
It matters because everyone imagines the inevitable "closing of the gap" between closed and open source, but the rate at which open source catches up with closed source seems to depend on being able to train on and distill the outputs of open source models. As long as performance of open source models is at least partially dependent on frontier-model outputs, then that gap will remain in place by definition.

>Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

If the distillation is irrelevant to why it is competitive, why do they do it then? Obviously is helps improve their benchmarks/performance to some degree, otherwise they wouldn't need to do it.

I wouldn't be surprised at all if US labs are also distilling Chinese models, except we'd never know since they can simply self-host them
> Does this matter? Distillation is not illegal by every definition of the word.

Correct, but it at least helps answer the question of "how do they make such good models for a fraction of the price???" The answer is someone else spends the untold billions and Chinese labs do a little tweaking.

Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled.

Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.

It is incredibly important to whether the US can maintain its AI lead. If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.

US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other countries can independently train competitive models, international cooperation is required.

Edit: It is also important for the business model. Companies won't be able to justify tremendous training costs if competitors can replicate their product much more cheaply via distillation.

It matters because it means that lab could not train that model without distilling another frontier model, and their progress would slow once they get properly cut-off. If I funded that lab, I would want to know that.
It certainly matters as familiar sounding words to their stock holders to please not drop the valuation.

Because what they want them to think is "the AI factory has unique proprietary technology that cannot be replicated"

What they don't want them to think is "it's relatively easy once you know the basics to bootstrap to near SOTA and so the commercial case for selling inference has an extremely short profitability horizon with little if any brand loyalty or lock in".

> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

They claim it because Anthropic are planning to push for protectionism. They just doubled their political spending to $40 million for the midterms to "push for AI regulation" Gee, I wonder what it is they are lobbying for. Certainly won't be OFAC sanctions right? ICTS import controls?

US GOV, under lobbying pressure from Anthropic and OpenAI are going to go full protectionism and restrict Chinese models, I'd almost be willing to bet money on it. They can't really enforce for individuals, but they can definitely tell US based hpyerscalers they can't host them, make it illegal to host the weights, and government procurement restrictions.

> Distillation is not illegal by every definition of the word.

I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.

I don't remember the details but I believe one of these instances was when one company trained its AI on the knowledge base of another company and turned it into a competing product. Fair use was denied because of that direct competition. Distilling a LLM to make a competing LLM looks kind of like this, or maybe not, I don't know.

It would make sense for distillation to be legal in every way, LLMs are built on a broad interpretation of fair use, but sometimes, law is weird.

> The only argument they have here is that they use GB300 GPU's

I don't follow events closely, but the US has constantly flipflopped on what sort of GPUs the Chinese are allowed to have, not in small part because much of the AI boom's valuation is based on demand for US-made hardware, for which the Chinese have inexhaustible and well-financed demand.

So even this feels a bit hypocritical to me, but my understanding is that Chinese native AI hardware is getting good enough that labs dont feel a huge disadvantage by being forced to buy at home, even if they'd have preferred to buy US chips.

Which is a situation that was manufactured by the constant thread of having their access to advanced GPUs revoked.