Every author ever has learned from, been inspired by and influenced by the books they've read before. Same for every poet, painter, musician, sculptor. So much so that it's a common interview question to ask such people about their influences.
But somehow, if it's AI instead of a human, the same behavior is supposed to be copyright infringement in the eyes of the greedy copyright cartel? Preposterous.
Computers aren't people and acting like a computer doing something a human can do means it has all the same rights as a human is not established law anywhere in this solar system.
Computers "act" on behalf of the human(s) who own them. Are we gatekeeping how humans can learn? You can learn with your own brain and maybe a pencil and paper notebook, but your computer cannot help you learn?
There are services that are provided based on implicit assumption that their consumers are human and they are priced accordingly.
You can't go to theater and instruct your computer(a.k.a. camera) to consume & learn or summarize the movie/performance on your behalf. If that's possible and allowed, I'm sure the admission fee will be drastically different.
Google will not let you instruct your computer to consume all the search results for a keyword and filter out good results & summarize them (you can technically still do that but Google will try their best to make sure you don't). If you do, they'll throw you into a hell that is infinite captcha loop. Yes, they are hypocritical to do that but that is a different issue.
A human and a computer are not comparable either in capabilities or their lifetimes. Let's not mix them up.
> is not established law anywhere in this solar system
Fair Use is the main doctrine relevant in this situation, and does not inherently distinguish between manual and automated creation. Use of traditional algorithms, like the caching/thumbnailing/snippeting done by search engines, is already covered by it.
There are cases where "but it's not human" makes sense because the relevant law does make that distinction (e.g: registering a work for copyright in the US depends on human authorship), but I think it's a fairly weak point in this case. You could instead argue that it's not solely "doing something a human can do" and how that impacts on the Fair Use factors.
Our legal system and government is created for humans, computer code is not human. Tech people have a tendency to create abstractions of the real world and then demand the real world respect the abstractions they create - that’s the preposterous part.
LLMs are not AI. LLMs have nothing to do with how the human brain works, they have at most, a passing resemblance to how a brain could work, with the tech we have available. They aren’t human. They are code, and the humans who create these programs cannot hide behind “yerm acktchually it’s a brain so your laws don’t apply”
It exists so that you are motivated to create, this makes people less motivated to create. Because, they put in work, you take it for free, repackage it and benefit from it.
I can't speak for LLMs, but I'm pretty sure that by this point I've seen more AI created fantasy elf images than I've seen non-AI ones. AI use has definitely exploded the number of fantasy images that I enjoy.
Obviously the input data is also useful, so we should protect the production of it.
Want proof that it’s useful?
> “it would be impossible to train today’s leading AI models without using copyrighted materials“ - OpenAI
Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
> Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
1) Infringement is not theft and is usually handled differently. The (poorly-named imho) NET act, which criminalized some non-commercial infringement, had to do so explicitly.
2) Sometimes exclusive/monopoly rights are not in the public interest, and compulsory licenses are desirable.
I don’t see what either of these points add to the conversation. I didn’t accuse them of theft, nor did I claim all exclusivity is always in the public interest.
Advancement of. The input data might be useful, but that doesn't mean it needs to enjoy copyright protection in every single avenue. That's why we have fair use. I can't really imagine an LLM not being transformative.
>Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
Did you get permission from OpenAI to share that quote or are you rent-seeking/stealing?
I think you'd agree that there are more reasons than just that to not always enter into agreements.
Did my use of that quote harm the incentive of OpenAI to produce it in the first place? No of course not.
Contrast with the entire value prop of LLMs being that you can utilize the knowledge in them without paying any time or credit toward the source, which does indeed destroy all commercial and most non-commercial incentives of producing and sharing such information.
The entire point of copyright law is to protect the incentives to produce work. The models are certainly doing something of value that also should be protected (and surely they’ll utilize law to do so), but ultimately a system that sucks up all prior works and obviates the need to view/buy them and destroys almost all incentive to produce new ones will not stand.
Of course this will boil down to model creators saying it doesn’t create incentive (except for “possibly capturing all future value of the light cone,” when talking amongst themselves) and original creators claiming that it does.
If they get to be treated as humans, then the far more important issue is how we immediately imprison all of these company executives for the crime of slavery.
Doesn’t this make sense? If your actions of copying content disrupt or prevent the content owner’s commercial practice, then it does seem like infringement.
> Many other news publishers, including the Financial Times, the Associated Press and Axel Springer, have instead opted to strike paid deals with AI companies for millions of dollars annually, undermining the Times' argument that it should be compensated billions of dollars in damages
I would see this as strengthening, rather than undermining, the argument. The AI companies in those agreements are paying for something, which implies that the work both has value, and should be paid for.
24 comments
[ 3.1 ms ] story [ 60.4 ms ] threadEvery author ever has learned from, been inspired by and influenced by the books they've read before. Same for every poet, painter, musician, sculptor. So much so that it's a common interview question to ask such people about their influences.
But somehow, if it's AI instead of a human, the same behavior is supposed to be copyright infringement in the eyes of the greedy copyright cartel? Preposterous.
You can't go to theater and instruct your computer(a.k.a. camera) to consume & learn or summarize the movie/performance on your behalf. If that's possible and allowed, I'm sure the admission fee will be drastically different.
Google will not let you instruct your computer to consume all the search results for a keyword and filter out good results & summarize them (you can technically still do that but Google will try their best to make sure you don't). If you do, they'll throw you into a hell that is infinite captcha loop. Yes, they are hypocritical to do that but that is a different issue.
A human and a computer are not comparable either in capabilities or their lifetimes. Let's not mix them up.
Fair Use is the main doctrine relevant in this situation, and does not inherently distinguish between manual and automated creation. Use of traditional algorithms, like the caching/thumbnailing/snippeting done by search engines, is already covered by it.
There are cases where "but it's not human" makes sense because the relevant law does make that distinction (e.g: registering a work for copyright in the US depends on human authorship), but I think it's a fairly weak point in this case. You could instead argue that it's not solely "doing something a human can do" and how that impacts on the Fair Use factors.
LLMs are not AI. LLMs have nothing to do with how the human brain works, they have at most, a passing resemblance to how a brain could work, with the tech we have available. They aren’t human. They are code, and the humans who create these programs cannot hide behind “yerm acktchually it’s a brain so your laws don’t apply”
On that note, I would like someone to point out which bits in the model file contain the copy that's infringing.
Copyright is not some kind of divine law handed down by god.
It is exactly what copyright is for.
Want proof that it’s useful?
> “it would be impossible to train today’s leading AI models without using copyrighted materials“ - OpenAI
Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
1) Infringement is not theft and is usually handled differently. The (poorly-named imho) NET act, which criminalized some non-commercial infringement, had to do so explicitly.
2) Sometimes exclusive/monopoly rights are not in the public interest, and compulsory licenses are desirable.
https://en.wikipedia.org/wiki/Compulsory_license
2) Supports "rent-seeking" by providing an example of legislative counterbalance.
>Being unable to survive if you had to enter into mutually consensual agreements with your suppliers is a pretty good sign that you’re rent-seeking or stealing.
Did you get permission from OpenAI to share that quote or are you rent-seeking/stealing?
I think you'd agree that there are more reasons than just that to not always enter into agreements.
Contrast with the entire value prop of LLMs being that you can utilize the knowledge in them without paying any time or credit toward the source, which does indeed destroy all commercial and most non-commercial incentives of producing and sharing such information.
The entire point of copyright law is to protect the incentives to produce work. The models are certainly doing something of value that also should be protected (and surely they’ll utilize law to do so), but ultimately a system that sucks up all prior works and obviates the need to view/buy them and destroys almost all incentive to produce new ones will not stand.
Of course this will boil down to model creators saying it doesn’t create incentive (except for “possibly capturing all future value of the light cone,” when talking amongst themselves) and original creators claiming that it does.
I would see this as strengthening, rather than undermining, the argument. The AI companies in those agreements are paying for something, which implies that the work both has value, and should be paid for.
I think it’s time to start managing our expectations around GPT-5