54 comments

[ 3.1 ms ] story [ 52.0 ms ] thread
> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.

I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.

YC does better if its startups get open weight frontier benefits. Garry’s just advocating for his book, which is his job. Consider how much capital YC portfolio companies would have to burn until liquidity if they have to pay OpenAI and Anthropic, versus relying on open weight frontier capabilities.
If someone proposes the right thing for selfish reasons, do we call that bad? Or do we call it proper incentive alignment?
It's an observation, how you feel about it is a personal opinion. I have no opinion on whether it is "bad vs "good." It just is.
We used to call it enlightened self-interest.
(comment deleted)
I reflected on this myself recently. Model distillation seems to be at least as fair a use as distilling a book.
More than fair if you consider that the tokens are paid for.
And LLM output can't be copyrighted, even.
Also, as said elsewhere: "Lab" is rich here, for outfits that, facing these giant, energy swallowing black boxes have really no clue what's going on inside.-

The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.-

Of course they are entitled to kill off a few mice, or pillage the commons to forward their "lab" work.-

We also associate laboratories with evil scientists and Frankenstein and the like. I can just hear Boris Karloff (er Bobby Picket) uttering “I was working in the lab late one night. When my eyes beheld an eerie sight… … … …the monster mash”. If anything, I associate _uncertainty_ with labs. The result is never known up front, they’re a place of discovery.

But I get your meaning. What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?

“We don’t know what’s going on” is essentially marketing. Sure we don’t _know_ but we have intuitions about why, where, and how to make certain changes…
(comment deleted)
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.

I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.

I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.

why can't I use the tokens i paid for anyway?
I'm sure they put some BS in their TOS
I'm also certain that they violated countless ToS when they scrapped the internet for training purpose.
Ethically sure but that doesn't mean taking from them is nothing "illicit".
(comment deleted)
Abolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain.

LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.

One of the most infamous legal challenges to sampled music was MARRS "Pump Up the Volume" in the 1980s, and that was preceded by other famous cases. Not sure why you think that started in the 1990s.
This. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.

Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.

This could mean every potential serious customer would have no option but to seek alternatives to these online services.

> OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.

I think this is being misunderstood. Codex has a toggle to allow your prompts to be included in training data. They’re saying they can’t be sure if the person had it on or off while using Codex to discuss the work.

They’re not saying that some prompts are mysteriously jumping into training data.

Also, there is a large market for AI services which don’t retain anything under any circumstances for enterprise customers.

All correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM.

Figure we’ll have to reckon with this next year in any case, guess we’ll see.

If the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs.

Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.

There is nothing illegal about training on traces from frontier models.

However the frontier labs don’t have to serve customers who are farming the service for distillation purposes. That’s their choice and they’re free to make it if they detect distillation happening.

I see it as analogous to companies building fiber in the public ROW during the last big infrastructure bubble. Under the Telecoms Act, these companies had to allow competitors to use their fiber at a fair price.

Similarly, AI companies should be required to allow distillation at a fair price. Fair Use doesn’t make sense as a social contract if it only cuts one way!

But if they tried to set a fair price they would have to report how much money they are losing on each token sold. This might be bad for the real business of ai firms, hoovering up as much capital as they can
There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
Agreed! Allow US companies to innovate by creating an ecosystem of smaller, more efficient open weight models and it will be a net benefit for everyone. Preventing consumers from developing competing products should be litigated as anti-competitive behavior.

Distillation is a good thing.

https://youtu.be/ZIaOBAjvc38

Garry Tan and Sam Altman recently did this interview together. They seemed pretty friendly with each other during it. Wonder what Sam Altman would say about Tan advocating for OpenAI’s models to be distilled.

Then again this is the same OpenAI that has gotten into legal trouble recently regarding Apple’s IP so who knows

Like Gates saying there should be UBI, or Musk saying... well, whatever.

They know it won't happen, so arguing for it is 'effectively free' and purely personal marketing.

A bullshit game played by politicians and wannabes.

If it were so easy why aren't the frontier labs doing it themselves?
Garry also goes to Thiels silicon valley church.
Ok, but how do the economics of this work? Based on its settlement, Anthropic paid an average of $3000 per work they scanned based on their settlement (https://tech-insider.org/au/anthropic-copyright-settlement-2...). They and OpenAI pay billions per year for a mix of experts and normal people to label or create data. Why would they continue doing this if the value of this is immediately copied by open models? If your goal is to end the economics of generating and buying data for AI (and I recognize for some people this is really the goal) then sure, but if you want AI for various subfields of interest to continue improving then it's not workable.

Back when people made arguments for software privacy, the argument was usually "big business will still pay and consumers wouldn't have paid anyways so it's ok for us to pirate" - I actually think that was fine for business software but terrible for indie games, whose market was 0% businesses.

But in the AI case, it's not like they get to keep some of the value of their investment - it all gets cloned into models that businesses and consumers alike are happy to use. If someone knows how labs could continue to fund data creation and acquisition in this model, please do share!

Surely if you hoover up every book in existence to feed into an ai model you must be extracting more than 1.5B in value. If not then it’s not a viable business.
This is all based on the delusion that Chinese labs are mindlessly distilling the frontier.

I would love for a US lab to be at or near the frontier with an open weight model, but it’s going to take some serious elbow grease, and yes some distillation (which btw OAI, anthropic et al, also use distillation of other’s outputs in their training)

I think OpenAI and Anthropic will go bust, or at least be scrapped for parts in the next 5 years or so. It's clear that the extreme cost used up for training is impossible to recoup, as inference is already being subsidized.

It's also clear that, as Tan indicates, open-weight models will be (and basically already are) just as good as frontier models. It's all about the harness, baby. We will have two main forks in the road, and two new industries created:

    - AI hardware (NVidia/Cerebras/etc.), the equivalent of Intel/AMD
    - AI software (harnesses, assistants, etc.) the equivalent of Microsoft/Apple
We already saw a glimmer of this with popularity of OpenClaw—the problem is that it's janky, hard to set up, inconsistent, and very hacker-esque. Imo "AI labs" will be a dying breed because there's no real money in the actual models if they get commodetized, which they already kind of are.
If harness is all that matters, a co-developed harness + model stack + large compute availability advantage + massive distribution advantage with data for post training will win the market.
Society as a whole has paid into this technology: through the theft of its intellectual property, through having to deal with the pillaging of so many commons (digital or otherwise) by it, through skyrocketing energy and computing device prices, and even just through ordinary investment. Democratize the technology! At the very least, don't step in legally to prevent this from happening.
Controlling what users and customers do with API calls to closed weight models feels constraining, and there’s a role government can play here to normalize the fact that access to intelligence that was trained on broad public access data should itself also be more a form of a public good than something locked away behind restrictive terms of service

I do not agree with this man all that often, but that is very concisely put.

> To him, the true AI doomer scenario is for all the immense power of frontier AI to wind up in the hands of a single powerful, proprietary provider. “The nightmare scenario, the doomer scenario for AI is that there’s just one company,” he said. “It has the best access to capital. It has the best AI researchers. It runs away with it and suddenly there’s one company that’s monolithic. And that would be bad.

Well yes, as I think I said in a previous comment, on the current trajectory OpenAI and Anthropic will really stop releasing models due to distillation and regulatory pressures. Then, they would eat all knowledge work themselves, which would be the end of YC, among many other changes.

Frontier labs trained their models on the entirety of human knowledge and didn't ask permission. It's a "want" or "should" it's a moral imperative to distill their models.
Given the short-term pragmatic, conflicted way that AI tech adoption is happening... won't encouraging distillation effectively taint the entire space of open weights models, with the undisclosed biases of a few models that are under the influence of parties (certain billionaires and politicians) known for aggression and duplicity, and not for admirable ethics?

Following news of companies and companies increasingly moving to open weights models.

As AI gets more central to society, we really need to know how the weights were determined.

Open weights isn't just "free as in beer"; it can be "free as in the mystery drug that creepy guy chatting you up at the bar offered you". And maybe even he doesn't even know everything that went into the tablets, since he too was being worked, by an organ-theft ring who will be harvesting both of you tonight.

That's an analogy to get your attention. Your LLM probably isn't going to steal your organs. But in the current environment, it does and will have ideological biases determined by those with direct and indirect influence over it. And there will be a massive market for commercial influence biases (look at how previous generations of adtech invaded almost all technology companies). And there's incentive for military and spying capabilities to be buried in the models, perhaps as long-term sleepers. Maybe some organized crime trojans, too, depending which model you pick up.

In this low-trust environment of the current real world, we need genuine open source models, not closed "open weights", and not mindlessly distilling black boxes gifted by sketchy powerful interests.

(comment deleted)