Interesting. In addition to Anthropic's watermark use to prevent model collapse, we can definitively call something copyrighted or not copyrighted. This is a boon to everyone who consumes culture.
Lmao. No. Content that is wholesale generated by AI is also not subject to copyright in the US. It relies on the honor system and you can always modify the material juuuust enough that you can claim the copyright. In practical terms, this changes nothing.
Judges are not at all stupid enough to fall for that. You would get laughed out of court for that. Judges are not computers, and they can not be tricked with these kinds of technicalities.
In fact, the defense of "I wrote the prompts that led to the code that the LLM wrote" would be a much better defense.
My bet would be to build anything from multiple independently generated pieces of AI code. Thus the software build from AI generated blocks by human would at least have protection on basis of the structure and work put into that.
That was a terrible ruling. It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams), there isn't even "a button".
After this ruling, I had an idea for a photo where I set up the entire set, camera, etc, but the photo entailed the model clicking the shutter while we were both in the shot. Just in case, I had the model write a quick note ceding the rights.
It's fine, I guess. How does it work in cinema? A director who is the creator of the project must have to get rights from every camera/mic operator.
In cinema it works the same as in software engineering. There is an exception to normal copyright law where works made in the course of employment are treated as if the author is the employer rather than actual author of the work. Its known in copyright law as a "work for hire" https://en.wikipedia.org/wiki/Work_for_hire
> In the United States, United Kingdom, and several other jurisdictions, if a work is created by an employee as part of their job duties, the employer is considered the legal author or first owner of copyright.
> It is an exception to the general rule that the person who actually creates a work is the legally-recognized author of that work.
Not paying anyone would be a violation of minimum wage laws. Or it would mean they are truly equal partners in the project who have no "employer" they are working under. In which case it does seem most fair that they would all own the copyrights to their contributions absent some other agreement to assign copyright.
>It should be about who got the camera there and set it up, not who pushed the button. With automated recording (dashcams), there isn't even "a button".
I'm actually having a bit of trouble thinking of what sufficient societal good there is/would be in granting copyright on raw dashcam or security camera or the like footage? None of those purely mechanical automated systems need a subsidy or encouragement to generate more. Certainly someone can use that sort of thing in the creation of a copyrighted work but what would be the issue with the underlying material in that case being unprotected?
I am in favor of copyright reform so eventually it becomes something like two years automatic with a one time two year extension possible an I agree with you because of one key word — raw. Once the footage is used in a package, be it a movie, a news report, or a music video, that package can be copyrighted. The problem is there is no way the owner of a dashcam can prohibit commercial use of a video they own if it cannot be copyrighted in the US. In a way, we have conflated copyrights with other so called intellectual property (details). Maybe the solution is to strengthen privacy rights somehow? I'm not sure.
Honestly, I think we just need fo go back to life+75 or something that covers the next generation.
Copyright is a fairly basic protection on Work, and wanting to alter it to be based on percived value is catastrophic- On one end you'll deny copyright to security camera footage- nevermind it's of a meteorite and everybody is clamoring for a copy. Then, on the other hand you'll have micky mouse where it's such a big thing and Walt put so much effort into it that the copyright is perpetual.
We need patent reform far more than any complaint against Copyright that I've seen- and far too many people conflate the two. Lets not even talk about trademarks.
Thinking more about monkeys, I think that case was made relatively easy because there was a normal living recipient for copyright, but that wasn't eligible for such. I'm fairly certain that dashcam footage, which is 100% automated, as you mention, is copyrighted by default. It's a fairly common source of footage of novel events like a meteorite, rocket impact, or whatever else. And that footage is licensed to media companies.
In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
Theoretically, in the US dashcam, CCTV, and similar automated recordings are not copyrightable. In the US you can only copyright creative expression. However, the question hasn't been answered in court, yet, AFAIU.
That doesn't stop people from treating automated recordings as if they were copyrighted, though. Especially for media companies, licensing is standard operating procedure. Even if a media company's lawyers believe something isn't actually copyrightable, if the licensor and licensee believe it is, and especially if distribution outlets (e.g. YouTube) treat it as such, then it all works out.
In some other countries automated footage is copyrightable, AFAIU, but it varies.
> In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
In the US the initial burden is on the person claiming a copyright violate. In court that is met by simply showing a certificate from the Copyright Office; in fact, it's actually a requirement--you can't sue in court to enforce a copyright claim with a certificate. But in getting the certificate you would be telling the copyright office it was made by a human.
Dashcams and security footage definitely don’t have copyright protection. Are you seriously thinking that the construction worker who set up the camera would then have copyright over everything it records?
Really? Regardless of the debate on what should and shouldn't be copyrighted, in the current system there is no way to set up a continuous camera with intention to create something and not be afforded the same rights as if I took a snapshot?
Yes the security guard’s footage would have copyright protection. However, a security guard is also typically paid and under contract. It’s a commissioned work.
Btw, this is the same agreement when hiring a software engineer.
I think the issue you then run into is imagine another human came, picked up the camera, and used it to take a picture of whatever. It'd be fairly self evident that that the photographer would own the copyright to that work, but in your take - it'd be the camera owner.
So you're more of trying to create a special rule where if the normal recipient of a copyright would be invalid, then it slides to the 'nearest' most appropriate individual, but that seems extremely fragile and difficult to define.
If copyright does not protect AI-generated content, then AI should not be restricted from generating content that falls under copyright protection; yet, the reality is otherwise.
I'm not sure that logic holds. If I create a frame for frame copy of Star Wars then I cannot copyright it. However, that doesn't mean that I can still distribute my copy.
I feel like this question is typed wrong but the same reason someone wants copyright on human work to make money by prevent other people from stealing your work
As software, creative works, science, etc become more and more contributed via AI does that mean all future works will not be copyrighted or patented? Under our current copyright culture and laws obviously not but that does open the question to how much does a human have to contribute and what evidence is required to show that a human contributed enough so that it can be copyrighted and patented. Some time in the future (may be many years) science may become to complicated for humans to understand fully with AI systems researching themselves. Current patent laws in many places including US says inventions created entirely by artificial intelligence cannot be patented. This future may be coming. What will the new copyright and patent laws look like in the future? Do we need copyrights and patents anymore?
Take note of the qualifier entirely. If you're working with an agent steering it to produce the results you want, it would be an entirely different story.
> Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.
which reads like you actually have to contribute by modifying the output from the system, i.e. prompts/inputs to the system do not count as a human contribution.
if you modify enough of the output as to make it unrecognizable as the machine output -- that's a new work created by a human. i.e. you don't have to change every line.
same thing applies to sampling in music. if i change enough of a sample as to make it unrecognizable, then i don't need to worry about registering/licensing the sample.
This is less relevant for logos & creative works, but things that enable companies to do production will go back to being closely guarded and sharded secrets, which is what the patent system was trying to resolve (though we can certainly argue the success of it).
If companies fail to protect their investments in generating IP, they will stop investing in generating it.
And unless IP generation costs (all in, including the humans telling them what to generate) fall close to zero, it will be bad for the world if companies cannot recoup investments in generating new IP.
We would expect this to hit those industries relying on IP protections the most, e.g. pharma.
"Content that is entirely generated by artificial intelligence is not protected by copyright."
If that is faithful reading of the law, that makes sense. I know a number of people who use AI, but none of them (that are making anything actually useful) have the output "entirely generated" (aside from some POC tests that never see the light of day).
I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
> I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
I don't think you're stretching your imagination enough. What if I used an AI like it were a fancy auto-completing dictation machine and I used it to write my great American novel? AI may have "entirely generated" all the text, but what if I micromanaged the shit out it?
I can imagine the difference between someone who fires off a lazy 5 minute prompt, and someone who labors for months and months to get exactly the results they want.
It’s almost like laboring for months and months would have draft revisions, dozens of changes, and many other interactions they could prove as theirs defeating your entire argument.
The current trend (since months now) is to have everything done through agentic loops. Meaning humans are only here to provide the initial prompt and do a few back and forth during the implementation.
> As software, creative works, science, etc become more and more contributed via AI does that mean all future works will not be copyrighted or patented?
No, just the stuff from people who think whatever they prompted from an AI is a contribution to creative works, science, etc.
Thank god! A victory for everyone who believes in the freedom of information, regardless of what you think about AI.
Copyright didn't always exist, nor should it continue to. Hell; it must not.
I think the words (read: hilarious 1.25pp pamphlet) of Aaron Swartz on the topic are just too poignant to ignore, given the paths of Reddit (corrupted yet democratic), IP law (malignant yet showing cracks), and government survellience have taken in the Trump era. Despite the dated context... he really says it best:
If the AI is treated as an agent and not merely as a tool, then this should also apply on the training side.
AI agents are effectively consuming people's output without permission. For code, the MIT license only gives rights to "any person obtaining a copy of this software".
So the rights are given to a 'person', and the rights pertain specifically to a person who performed the act of 'obtaining a copy'.
Scraping the code in-situ from millions of repos automatically and systematically, with no intent to use the software, does not involve a 'person obtaining a copy.' The subject at hand is an 'AI agent scraping the code'; that's not the same subject and the MIT license says nothing about this case.
This has already also been implied in the US. Courts upheld the Copyright Office's stance on human creation in the context of AI image generation. There's no reason to expect something that fundamental to copyright to be any different for other media, such as source code.
It's reasonable to expect this to hold for all Berne Convention countries.
All licenses are unenforceable if you don't hold the copyright, since you don't have a copyright to license. And attempting to do so would probably be perjury.
Im not sure it’s really a copyright issue here, and not patent or trademark related, or something else. Copyright is really only one aspect of IP laws and they all have their own legal nuances
I don't see an easy test here. Worse, I see the beginnings of a test that is technical and very likely to not match the future of how we will interact with these things. We have to start somewhere but I'm not sure 'Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.' is the right place to start. I think we need more examples of what does constitute 'human-centric' and work from there. I also don't think that any system that says 'you didn't do enough work so it isn't human-centric' works. Any system like that will require a reexamining of copyright in general. There is a lot of low work copyrighted material out there. Just because 'AI' didn't build it doesn't mean the same tests shouldn't apply.
Some examples of why I think this is really hard: Say I build a story generation system. I work hard on building an agent swarm of actors, critics, editors, researchers. I craft into the various agents concepts of story arcs, outlining techniques, character development. I build a huge well thought out process for how to agentically write an actually good story, so long as you give it a title. Heck, I even design and train my own custom LLM with original layer ideas and novel training techniques to use on this system. After all that I then take that final step and give it a title. Do I have no claim to that? I probably put more work and creativity into it than an author would have a book. What if I then gave it 500 titles? 5,000? Would my claim degrade the more titles I fed it? Is it a percentage of work question? What is the core concept here that defines 'human-centric'? What is the cut-off here?
Let's go even further. I don't prompt. I live in a world with unlimited context models. I have a conversation about the book I want it to write. During that process I reject some ideas and accept others. I didn't give it a 'system prompt' but essentially all I did was prompt it and select versions I liked. Is that not human centric? How about if I asked it for advice and it did some editing work on my story? Did that make it not human centric even though the starting text was mine? What if that starting text was 99% replaced with a version 10x as verbose. Defining based on how you interacted with the model (prompted and selected) just seems way to weak to be a clear test.
Information wants to be free. We should simply dispose of copyright. With LLMs, waters that were already muddy are now a thick slurry. And it's only going to get worse.
It's an antiquated mechanism and is far more abused than it is actually used at this point.
Actually, in Germany, copyright remains with the author and their heirs until seventy years after the author's death; it cannot be sold or given away, even if someone paid them doing it. Instead, there is "Verwertungsrecht" or "Nutzungsrecht" ("License", "Usage right"). In the USA, copyright can be sold completely to another legal entity.
Is it fair to say, then, that artists who have trained on the work of prior artists should also not be able to protect their works unless they have directly received consent and have compensated all the artists whose works they have viewed and potentially learned from over their lifetime?
There is a dramatic difference between learning from or being inspired by an artist and incorporating it wholesale into an LLM
Artists can use human judgement and knowledge of social norms to decide what art is or is not appropriate to utilize. (e.g. a painting in the museum versus someone's tribute to a dead loved one on DeviantArt)
There's a big difference both in the amount of learning and the amount of output between an AI and a human.
There's a big difference between an artist and a megacorporation.
Andy Warhol did in fact compensate the artists whose work he incorporated into his art. So there is definitely precident.
> I made minor modifications post-generation (open question, probably enough)
I don't think it would be enough. Compare for example the case where the US copyright office ruled that assembling AI-generated images and human-written text into a comic book, only the human-made elements themselves (text, arrangement) got copyright protection, but not the images.
So this means you can not use AI for the majority of open source projects since GPL, MIT, BSD, etc are all copyright declarations and they'd be being made for code which you can not copyright.
You can't enforce those licenses against AI-generated parts, because enforcement relies on the recipient having no other way to avoid copyright infringement.
However, there's no issue with including non-copyrightable code in otherwise copyrighted projects. There's already plenty of non-copyrightable code like auto-generated boilerplate.
>no issue with including non-copyrightable code in otherwise copyrighted projects
That is not how copyright/trademark/contract laws work, and isomorphic plagiarism is not a long-term business model. People also loved Napster at first too. Good luck =3
No, GPL is a contaminating license contract, so "AI" slop means you are probably in GPL violation by including isomorphic plagiarized uncopyrightable code. =3
I find it hard to find correct terminology in this case. AI generated content is copy right wise in state of nothingness. It simply does not have copy right status like other material can have. As such maybe best I can formulate is that you can't enforce license violations against ai generated material as you do not have ownership.
So no license is enforceable with code written by AI.
It is actually pretty easy when the area is very specific like 1 guys Perl library, and Claude ports it to Python for a group unaware of what happened.
All models know what Disney Micky Mouse looks like too. =3
We know for a fact that Bun port to Rust was done entirely by Claude! I’m wondering if that could be used against Anthropic in some ways (not that I want to, just curious what would be the angle)
wasnt there a case a while back, where GPL or LGPL code made its way into MIT licensed software via LLM? And they were forced to remove the copyleft code. I dont remember the details though.
Usually what I have seen is someone writes a Perl library, Claude mostly ports it to Python with isomorphic plagiarism, and a bunch of Brogrammers proclaim "AI" magic is real.
The issue is most GPL license fall under contract law, and scraped code can't legally have assigned "copy" rights on an "AI" vector search compaction output.
Indeed, but people will continue to think vector search compaction similarity absolves folks of Trademark and Copyright liability.
As the dark specter of Disney Mickey Mouse looms over every LLM model involved in isomorphic and character plagiarism. Yes, even motion capture is considered a performance act in the guilds, so video reskinning an unlicensed performances act people make are also a liability.
It would sure save a lot of money if you don't get caught, so people are gonna try it for sure. =3
> I mean the copyright has to belong to somebody right?
Why would it? It’s generated by blending together ~ every bit of content on the internet and in books that they could steal. Why would the operator of the blending machine suddenly get copyright?
Does a gambler own the copyright on the symbols generated by a slot machine?
There is the potential for a solid piece of artwork here, ie. at what point does a string of words become an artistic product.
Imagine a slot machine in a gallery that produces a 4 word sentence. Visitors press the button, and the artist copies the 4 words onto a piece of paper and puts it on the wall.
EDIT: my point being that causing a machine to generate words doesn’t give you copyright over the words but if you then do something with the words they become your work.
Lots of wild guesses about mixed human-AI copyright in here. Last time I read the copyright office’s statements in this, their stance was basically: The human owns exactly what they contribute. The rest is public domain.
Yes, that is vague. I think the examples were like:
If you paint a symbol and use an AI filter over that to stylize it, you own the symbol aspect of the image but not the stylized final result.
You can own a book of AI images as a curated collection. But, not the individual images.
Seems about right. Copyright is supposed to literally prevent outright copying. The output from an LLM is not a creative work of the prompter. It's genuinely the opposite, I use them when I don't care about something but have to do it anyway for whatever reason. It makes more time for me to do the things I like working on.
It feels shitty that they have been trained on the life sums of all of our work and online presences with absolutely no credit given... But then again, I'm not sure I'd want to know what parts of the weights were from me and which weren't.
I feel like people will be upset by this take... But you don't get credit for the creativity of an idea alone. Every asshole has ideas. It's called a work of art for a reason, emphasis on WORK.
I’m trying to imaging what giving credit would look like… “Credit: Everyone who has made anything freely observable on the internet.” Would folks be less upset if that was included?
Frankly, I should have been putting something similar attached to any drawings or paintings I make physically. My entire life I’ve been training my brain on countless works of others. And, I’ve never asked one of them for permission or given anyone any credit.
This post is glossing over most of the nuance in EU law.
The AI system must function merely as a tool or instrument (like a camera or Photoshop) guided by the human, rather than acting as the creator itself. The line may get a bit fuzzy case-by-case, but effectively the human must be the creative one, not the AI.
This is not unprecedented. Machine generated technical data, sensor outputs, automated surveilance photography, monkey selfies, purely algorithmic or generative music and such were already disqualified long before AI came along.
Yes it does because there's is still direct control from a human. When working with LLMs, even interactive prompting does not provide that level of control.
Funnily I think copyright protection for LLM model is 15 years. As to me it just looks like very fancy database and that is the protection for those...
Not a layer but a bit nerdy on the law but specially on the copyright side.
Copyright is a framework to protect creative work.
One could easily argue that an LLM model is just weights automatically produced from data and there is no creative work involved.
The counterpart would argue that the body of training is itself a curated body that is reflected in the model and that the pipeline training is a creative process required to obtain the final model, like some artist work; But it feels like a weak argument to me.
But then again, let's stop pretending law in US has any meaning other than what the billionaires or the dictator wants.
Interesting thought experiment is to consider an author who writes a novel in English, but wants to translate it into German.
They have a copyright on the original, and if they hire a human, the human would have a copyright on the translation (which would generally be licensed or transferred back to the author in some way).
If they use an AI for the translation, by the logic here, the translation wouldn't have its own independent copyright, but (based on other long established principles of copyright) it would still be a derived work of the original, so even if this decision holds it would not be legal to make unauthorised AI translations, pirate authorised AI translations, make further translations into other languages (or back to English), etc.
Which seems fairly reasonable! But consider:
If you start with, say, a 90,000 word novel, and ask for a translated novel, you (presumably) have sufficient rights to stop someone making unauthorised copies of the AI translated version.
If you start with a 300 word prompt, and ask for a logo, you (apparently) do not have sufficient rights to stop someone from using it without authorisation.
So some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point between these two extremes, and I think it's interesting to wonder what the boundaries are. Like, in theory you could graph input size versus output complexity, and sketch a frontier between "the author's protected expression survives in the output" and "the author's protected expression does not survive in the output". And I don't have the slightest idea what I think a fair frontier would look like.
Input size and output complexity are not necessarily the only two dimensions involved. There’s no reason to think that such graph would show a continuous frontier.
Suppose that an artist trains an AI model purely on their own works, and then the AI generates something similar to that artist's work. I would say that the artist must be allowed to assert copyright over that. The artist has copyright over all the training data, and the output of the AI is a derived work of that.
What do you think should happen if an artist views a lot of arts over their lifetime and then themself generates something similar to another artist's work? I would argue that collectively humans have been standing on the backs of other humans all throughout history. An artist today does not simply come out of the womb knowing how to make great art - they learn it by observing other art, learning about art and practicing (not altogether dissimilar to model training) over many, many hours.
151 comments
[ 0.22 ms ] story [ 19.3 ms ] threadIn fact, the defense of "I wrote the prompts that led to the code that the LLM wrote" would be a much better defense.
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
What if this applied in a photography class? The instructor owns the equipment and helped “set up” the photo. Does the instructor own the copyright?
It's fine, I guess. How does it work in cinema? A director who is the creator of the project must have to get rights from every camera/mic operator.
No, because they're already paid to do that job. And a director is paid too - he doesn't have any copyright for his work. The employer does.
> In the United States, United Kingdom, and several other jurisdictions, if a work is created by an employee as part of their job duties, the employer is considered the legal author or first owner of copyright.
> It is an exception to the general rule that the person who actually creates a work is the legally-recognized author of that work.
If employment is an exception, I wonder what would happen if an animal that you owned pressed the shutter on a camera.
I'm actually having a bit of trouble thinking of what sufficient societal good there is/would be in granting copyright on raw dashcam or security camera or the like footage? None of those purely mechanical automated systems need a subsidy or encouragement to generate more. Certainly someone can use that sort of thing in the creation of a copyrighted work but what would be the issue with the underlying material in that case being unprotected?
For what it’s worth, dashcam footage can absolutely be copyrightable.
The ruling is more about “only humans can get copyright protection”, not so much anything about whether a button is pressed or not.
Details https://ftp5.gwdg.de/pub/gnu/www/philosophy/words-to-avoid.h...
Copyright is a fairly basic protection on Work, and wanting to alter it to be based on percived value is catastrophic- On one end you'll deny copyright to security camera footage- nevermind it's of a meteorite and everybody is clamoring for a copy. Then, on the other hand you'll have micky mouse where it's such a big thing and Walt put so much effort into it that the copyright is perpetual.
We need patent reform far more than any complaint against Copyright that I've seen- and far too many people conflate the two. Lets not even talk about trademarks.
In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
That doesn't stop people from treating automated recordings as if they were copyrighted, though. Especially for media companies, licensing is standard operating procedure. Even if a media company's lawyers believe something isn't actually copyrightable, if the licensor and licensee believe it is, and especially if distribution outlets (e.g. YouTube) treat it as such, then it all works out.
In some other countries automated footage is copyrightable, AFAIU, but it varies.
> In any case, any country that decides neural network generated media isn't copyrighted then faces a major, and probably impossible, problem of trying to prove it wasn't made by a human.
In the US the initial burden is on the person claiming a copyright violate. In court that is met by simply showing a certificate from the Copyright Office; in fact, it's actually a requirement--you can't sue in court to enforce a copyright claim with a certificate. But in getting the certificate you would be telling the copyright office it was made by a human.
Registering a work provides additional benefits.
That’s it.
Btw, this is the same agreement when hiring a software engineer.
So you're more of trying to create a special rule where if the normal recipient of a copyright would be invalid, then it slides to the 'nearest' most appropriate individual, but that seems extremely fragile and difficult to define.
Take note of the qualifier entirely. If you're working with an agent steering it to produce the results you want, it would be an entirely different story.
> Neither mere prompting nor the selection between several AI suggestions is sufficient as a human creative contribution.
which reads like you actually have to contribute by modifying the output from the system, i.e. prompts/inputs to the system do not count as a human contribution.
same thing applies to sampling in music. if i change enough of a sample as to make it unrecognizable, then i don't need to worry about registering/licensing the sample.
> Munich Local Court has held that
Local Court in Germany is the equivalent of the Magistrates Court in the UK or Australia.
Definitely not the arbiter of ultimate truth.
If companies fail to protect their investments in generating IP, they will stop investing in generating it.
And unless IP generation costs (all in, including the humans telling them what to generate) fall close to zero, it will be bad for the world if companies cannot recoup investments in generating new IP.
We would expect this to hit those industries relying on IP protections the most, e.g. pharma.
If that is faithful reading of the law, that makes sense. I know a number of people who use AI, but none of them (that are making anything actually useful) have the output "entirely generated" (aside from some POC tests that never see the light of day).
I have a hard time believing anything of value, anything worth copyrighting, could be entirely generated by AI.
I don't think you're stretching your imagination enough. What if I used an AI like it were a fancy auto-completing dictation machine and I used it to write my great American novel? AI may have "entirely generated" all the text, but what if I micromanaged the shit out it?
I can imagine the difference between someone who fires off a lazy 5 minute prompt, and someone who labors for months and months to get exactly the results they want.
However, I just saw a product yesterday that was released and being built as you described. Interesting times.
No, just the stuff from people who think whatever they prompted from an AI is a contribution to creative works, science, etc.
2) If one applies a copyright message to AI generated output, is that fraudulent?
Copyright didn't always exist, nor should it continue to. Hell; it must not.
I think the words (read: hilarious 1.25pp pamphlet) of Aaron Swartz on the topic are just too poignant to ignore, given the paths of Reddit (corrupted yet democratic), IP law (malignant yet showing cracks), and government survellience have taken in the Trump era. Despite the dated context... he really says it best:
https://ia800101.us.archive.org/1/items/GuerillaOpenAccessMa...
AI agents are effectively consuming people's output without permission. For code, the MIT license only gives rights to "any person obtaining a copy of this software".
So the rights are given to a 'person', and the rights pertain specifically to a person who performed the act of 'obtaining a copy'.
Scraping the code in-situ from millions of repos automatically and systematically, with no intent to use the software, does not involve a 'person obtaining a copy.' The subject at hand is an 'AI agent scraping the code'; that's not the same subject and the MIT license says nothing about this case.
It's reasonable to expect this to hold for all Berne Convention countries.
All licenses are unenforceable if you don't hold the copyright, since you don't have a copyright to license. And attempting to do so would probably be perjury.
Does it enable decompilation remasters of classic games?
It feels like AI is a cleanroom laundromat
That sounds similar to how LLMs have been trained on GPL code and yet output is not (yet) considered derivative work
Some examples of why I think this is really hard: Say I build a story generation system. I work hard on building an agent swarm of actors, critics, editors, researchers. I craft into the various agents concepts of story arcs, outlining techniques, character development. I build a huge well thought out process for how to agentically write an actually good story, so long as you give it a title. Heck, I even design and train my own custom LLM with original layer ideas and novel training techniques to use on this system. After all that I then take that final step and give it a title. Do I have no claim to that? I probably put more work and creativity into it than an author would have a book. What if I then gave it 500 titles? 5,000? Would my claim degrade the more titles I fed it? Is it a percentage of work question? What is the core concept here that defines 'human-centric'? What is the cut-off here?
Let's go even further. I don't prompt. I live in a world with unlimited context models. I have a conversation about the book I want it to write. During that process I reject some ideas and accept others. I didn't give it a 'system prompt' but essentially all I did was prompt it and select versions I liked. Is that not human centric? How about if I asked it for advice and it did some editing work on my story? Did that make it not human centric even though the starting text was mine? What if that starting text was 99% replaced with a version 10x as verbose. Defining based on how you interacted with the model (prompted and selected) just seems way to weak to be a clear test.
It's an antiquated mechanism and is far more abused than it is actually used at this point.
If we go back 10 years and your friend says “I have an idea for an app, here it is,” and you build it, you own the copyright because you wrote it.
You give an idea to the pile of math calculated of the stolen work of humanity, the math owns it (which it can’t, so no one owns it).
No matter how detailed of a conversation you have with a friend, I don’t think they have justification to claim copyright over code written by you.
Based on your logic that should not qualify, but it currently clearly does: https://quayola.com/selected-unfinished-sculptures/
Alternatively, they could remove the art from the LLM and divide their past profits among the artists.
But it just begs the actual question of how much human contribution there needs to be:
- I wrote the prompt (not enough)
- I wrote many prompts and iteratively refined them using distinctly human skill (open question, but loosely seems still not enough?)
- I made minor modifications post-generation (open question)
- I made equal or more contribution to the final result (this better clearly have copyright protection or we are in real trouble)
Artist are human
There is a dramatic difference between learning from or being inspired by an artist and incorporating it wholesale into an LLM
Artists can use human judgement and knowledge of social norms to decide what art is or is not appropriate to utilize. (e.g. a painting in the museum versus someone's tribute to a dead loved one on DeviantArt)
There's a big difference both in the amount of learning and the amount of output between an AI and a human.
There's a big difference between an artist and a megacorporation.
Andy Warhol did in fact compensate the artists whose work he incorporated into his art. So there is definitely precident.
I don't think it would be enough. Compare for example the case where the US copyright office ruled that assembling AI-generated images and human-written text into a comic book, only the human-made elements themselves (text, arrangement) got copyright protection, but not the images.
80% of the code should be hand written by you. 20% can be allocated by AI for corrections or suggestions.
AI generated code should never be copyrighted otherwise. It's objectively common sense.
However, there's no issue with including non-copyrightable code in otherwise copyrighted projects. There's already plenty of non-copyrightable code like auto-generated boilerplate.
That is not how copyright/trademark/contract laws work, and isomorphic plagiarism is not a long-term business model. People also loved Napster at first too. Good luck =3
https://www.youtube.com/watch?v=YhgYMH6n004
So no license is enforceable with code written by AI.
Perhaps we'll have new iterations of FOSS licenses to adjust to legal declarations.
All models know what Disney Micky Mouse looks like too. =3
Add a subtle bug or odd behavior to your code, see if someone else’s code repros it (presumably via LLM regurgitation), sue.
The issue is most GPL license fall under contract law, and scraped code can't legally have assigned "copy" rights on an "AI" vector search compaction output.
https://www.youtube.com/watch?v=YhgYMH6n004
Indeed, these rules obviously don't apply in places like India, Russia, Iran, and China. =3
As the dark specter of Disney Mickey Mouse looms over every LLM model involved in isomorphic and character plagiarism. Yes, even motion capture is considered a performance act in the guilds, so video reskinning an unlicensed performances act people make are also a liability.
It would sure save a lot of money if you don't get caught, so people are gonna try it for sure. =3
I mean the copyright has to belong to somebody right?
Why would it? It’s generated by blending together ~ every bit of content on the internet and in books that they could steal. Why would the operator of the blending machine suddenly get copyright?
Does a gambler own the copyright on the symbols generated by a slot machine?
Imagine a slot machine in a gallery that produces a 4 word sentence. Visitors press the button, and the artist copies the 4 words onto a piece of paper and puts it on the wall.
EDIT: my point being that causing a machine to generate words doesn’t give you copyright over the words but if you then do something with the words they become your work.
Yes, that is vague. I think the examples were like:
If you paint a symbol and use an AI filter over that to stylize it, you own the symbol aspect of the image but not the stylized final result.
You can own a book of AI images as a curated collection. But, not the individual images.
It feels shitty that they have been trained on the life sums of all of our work and online presences with absolutely no credit given... But then again, I'm not sure I'd want to know what parts of the weights were from me and which weren't.
Frankly, I should have been putting something similar attached to any drawings or paintings I make physically. My entire life I’ve been training my brain on countless works of others. And, I’ve never asked one of them for permission or given anyone any credit.
Say I would prompt Unix like system in few prompts or started agent chain to make it. That OS wouldn't have a protection.
The AI system must function merely as a tool or instrument (like a camera or Photoshop) guided by the human, rather than acting as the creator itself. The line may get a bit fuzzy case-by-case, but effectively the human must be the creative one, not the AI.
This is not unprecedented. Machine generated technical data, sensor outputs, automated surveilance photography, monkey selfies, purely algorithmic or generative music and such were already disqualified long before AI came along.
But then again, let's stop pretending law in US has any meaning other than what the billionaires or the dictator wants.
They have a copyright on the original, and if they hire a human, the human would have a copyright on the translation (which would generally be licensed or transferred back to the author in some way).
If they use an AI for the translation, by the logic here, the translation wouldn't have its own independent copyright, but (based on other long established principles of copyright) it would still be a derived work of the original, so even if this decision holds it would not be legal to make unauthorised AI translations, pirate authorised AI translations, make further translations into other languages (or back to English), etc.
Which seems fairly reasonable! But consider:
If you start with, say, a 90,000 word novel, and ask for a translated novel, you (presumably) have sufficient rights to stop someone making unauthorised copies of the AI translated version.
If you start with a 300 word prompt, and ask for a logo, you (apparently) do not have sufficient rights to stop someone from using it without authorisation.
So some combination of the input (0.3k vs 90k) and the output (logo versus novel) crosses an inflection point between these two extremes, and I think it's interesting to wonder what the boundaries are. Like, in theory you could graph input size versus output complexity, and sketch a frontier between "the author's protected expression survives in the output" and "the author's protected expression does not survive in the output". And I don't have the slightest idea what I think a fair frontier would look like.
https://fmhy.net