> "it is ok to strip out my copyright string from my code"? How did they end up on that side?
The answer is, they didn't end up on that side. Copyright infringement does not involve merely using a copyrighted work. Copyright infringement involves copying part of a work's copyrightable expression into another Thing (for lack of a better word). In the US, if no part of the Thing is substantially similar [1] to any part of the original work's expression, then the Thing does not infringe on the original work's copyright. In such cases, there is no obligation to add any copyright management information (CMI) for the original work, and it makes no sense to argue that the CMI was "removed" from the Thing. Not every LLM output contains substantially similar to any particular copyrightable expression in the training set.
If they don't rely on the original works why do they incorporate them in the training corpus? But they do, so it is a derivative work.
When the industry favored copyright, it made sure to do clean room implementations of software by competitors. Programmers who had even read a single line were disqualified.
AI reads everything, so it is not a clean room implementation. The EFF knows this of course and still supports the industry (which is now of the side of theft, unless it is distilling).
Neither my words nor EFF's words suggested anything like that.
> so it is a derivative work.
Just because one work is derived from or relies on another does not implicate copyright. Copyright is not use-right or rely-right (nor should copyright be expanded to be them). (Contracts such as EULAs can go beyond the scope of copyright and may include use-restrictions.) If there is no substantial similarity (including obfuscated or mangled similarity) between the derivative work and the original work, then the derivative work does not infringe copyright.
> AI reads everything, so it is not a clean room implementation.
Very true, but substantial similarity matters. If there is no substantial similarity between the output and the original work, then an output is not an "implementation" of the original work. When I say output or Thing, I mean the output of an LLM or a human, not the LLM itself. If (if) a particular LLM itself infringes copyright, not every output of the LLM necessarily infringes copyright. If a particular output of an LLM infringes copyright, the LLM itself does not necessarily infringe copyright. (Unless someone intends to demonstrate that the particular LLM might as well be incapable of producing non-infringing output.)
There's no guarantee that a non-clean-room implementation always constitutes copyright infringement, especially considering that for software in particular the functional aspects are not always separatable from the creative expression. Theoretically, both clean-room and non-clean-room implementations of a very optimized program designed for non-entertainment purposes would be unavoidably substantially similar to the original work.
5 comments
[ 0.23 ms ] story [ 12.9 ms ] thread"The U.S. Court of Appeals for the Ninth Circuit handed internet users and programmers a big win today ..."
I am a programmer and I am not represented by the devious EFF liars. You support stealing my code.
The answer is, they didn't end up on that side. Copyright infringement does not involve merely using a copyrighted work. Copyright infringement involves copying part of a work's copyrightable expression into another Thing (for lack of a better word). In the US, if no part of the Thing is substantially similar [1] to any part of the original work's expression, then the Thing does not infringe on the original work's copyright. In such cases, there is no obligation to add any copyright management information (CMI) for the original work, and it makes no sense to argue that the CMI was "removed" from the Thing. Not every LLM output contains substantially similar to any particular copyrightable expression in the training set.
[1] https://en.wikipedia.org/wiki/Substantial_similarity
When the industry favored copyright, it made sure to do clean room implementations of software by competitors. Programmers who had even read a single line were disqualified.
AI reads everything, so it is not a clean room implementation. The EFF knows this of course and still supports the industry (which is now of the side of theft, unless it is distilling).
Neither my words nor EFF's words suggested anything like that.
> so it is a derivative work.
Just because one work is derived from or relies on another does not implicate copyright. Copyright is not use-right or rely-right (nor should copyright be expanded to be them). (Contracts such as EULAs can go beyond the scope of copyright and may include use-restrictions.) If there is no substantial similarity (including obfuscated or mangled similarity) between the derivative work and the original work, then the derivative work does not infringe copyright.
> AI reads everything, so it is not a clean room implementation.
Very true, but substantial similarity matters. If there is no substantial similarity between the output and the original work, then an output is not an "implementation" of the original work. When I say output or Thing, I mean the output of an LLM or a human, not the LLM itself. If (if) a particular LLM itself infringes copyright, not every output of the LLM necessarily infringes copyright. If a particular output of an LLM infringes copyright, the LLM itself does not necessarily infringe copyright. (Unless someone intends to demonstrate that the particular LLM might as well be incapable of producing non-infringing output.)
There's no guarantee that a non-clean-room implementation always constitutes copyright infringement, especially considering that for software in particular the functional aspects are not always separatable from the creative expression. Theoretically, both clean-room and non-clean-room implementations of a very optimized program designed for non-entertainment purposes would be unavoidably substantially similar to the original work.