It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free.
I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
The argument was "It's the robbery of all of our culture to sell it back to us at a mark-up".
Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.
Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.
I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
> companies spent a long time telling us downloading single songs via Napster was the worst thing ever
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?
No, for obvious reasons. "acknowledging" implies it's a self-evident truth that people are just refusing to see as such, when it just isn't. Completely different companies - in fact one is the Recording Industry Association of America, an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
Moreover, it's the people/companies from RIAA and adjacent circles (news publishing) that are stoking the "AI training is theft" arguments that people here are so breathlessly repeating - in a twist of irony, it's the people that suddenly decided to align themselves with media/news conglomerates, the same ones that were considered scum of the Earth before ChatGPT debuted.
> an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.
This is the most unintentionally hilarious misunderstanding of what the RIAA does, and the power relationship between artists and publishers I've read in years. In practice the RIAA exists to maintain the copyright monopoly of a few major labels. Rent seeking from the non-artist owned catalogues of the enormous majority of musicians who never 'recoup' their initial record deal.
> While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
It's far more seductive (since it's the default) to assume class relations don't exist, and wealth distribution is meritocratic. There may not be 'goodies and baddies', but there absolutely are rentiers and workers, billionaires and plebs.
All the major labels are public companies. Which means it's the very same people - the investment class, who claim ownership and extract wealth from say Warner and Open AI (should it make any money - obviously the whole house of cards could come down first).
> Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves.
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
> These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
> I don’t believe any of these companies have paid for all the books they have trained on.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
> Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
The multiplication comparison makes no sense because humans don’t learn by multiplying numbers to change weights. It’s like comparing the lubrication oil consumption of a car to the cooking oil consumption of a human to compare the carrying capacity. That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn. Otherwise, a human would learn much less in their whole life than a GPU does in one second.
> That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn.
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
> The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard.
Yeah, which is why it's only used to compare cars etc. among each other.
Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
> It was a marketing thing at the time to make steam engines look good.
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
> Except in your own examples it can easily be shown that it is not handled differently.
I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.
Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
> A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
> I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.
I was talking about cars vs horses. HP is Horse-power not human-power.
> Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
Cars are not free to roam the road. I don't know why it is so hard for you to understand that we use metrics like HP/CC etc to equalize with humans so that automobiles can be taxed just like humans. Without metrics like HP/CC etc there is no way to tax cars. Get it?
> And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
Come on you are clutching at straws here. It is not about "making sense". It is about using a metric to equalize unequal entities. When I am already saying they are unequal and have no direct relation to each other and any relation can only be arrived at indirectly. FLOPS is just an example I gave. I am not saying we should literally go with the FLOPS example itself. But we can use any metric and equalize it with human work. That's all I am getting it. It is the same argument as Horse-power.
> So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs, if the way they learn is fundamentally different?
Because that is the only metric we can use to measure how quickly GPUs can process arithmetic (you can label it "training" or "learning" or whatever name you want). There is no other metric that is deterministic and comparable to something humans do (which is also process arithmetic).
> Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?
It is fundamentally math. Every physical law in the Universe is expressed and backed by math. So on a fundamental level, yes "reading" is essentially maths only.
Do you not think that stealing requires the original owner to lose access?
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
> EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
> News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
> A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well.
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.
Break up the multi trillion, multifaceted companies into the their separate facets. These companies all thrived on far significant smaller portfolios in the past.
Cap the size a company can grow.
Regulate the amount of compute they are allowed to use.
Legally mandating that companies open the weights of, say, 18-month-old models might change their minds a little.
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
When a human learns, they carry that forward into their future ventures. They might use their learning to recreate the original work and profit from it without attribution. In the West we view that badly and have legislation to protect from some abuses. But the same person might later collaborate with the original author, or make a derrivaive work that improves on the original (a la most science).
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two:
- either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing
- or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
A simple law stating that laundering data through an LLM does not constitute "fair use", that the outputs can be subject to copyright claims of the original authors, and that the outputs themselves do not qualify for copyright protections would go a long way.
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
Yes, the whole IT sector is built on it. Robbing jobs, money, power, opportunities from billions of people and delegating them in to shitty jobs.
Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
There is no such thing as PROGRESS for progress' sake.
And FYI, agriculture is a wonderful invention.
Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.
Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.
I won't even go into climate change.
And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.
Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.
* * *
And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.
Yet our world can feed billions of people. More people are alive today than more people ever lived in this floating rock. Most people are living better than kings from 200 years ago.
We invent. We make progress. We make course correction. We end up in a better place.
Climate change? Lmao. We were told by 2000 20% of our country will be under water... It's 2026. Not under water yet.
What amount of pesky human intervention (positive or negative) will affect earth waking up from it's cold climate? Not luddites and their cow farts are global warming.
The real fight with "climate change" will be done with planetary scale technology derived from massive amounts of technical progress and energy production.
Your argument is as good as I should throw trash in the road or plastic in the sea. I'm just one individual. It doesn't matter. I'm not making the whole society do it!!
I fully agree with you. It is survival of the fittest in the face of progress. Competition is essential. That's how society grows. Humans prosper. May the odds be in your favor too.
Next time can you please start with stuff like this: "Climate change? Lmao." so that everyone reasonable can ignore you? The science beihnd climate change is accepted by freaking BP and Esso.
Does more CO2/Greenhouse gasses etc raise temp? Yeah duh.
Does not using those bad refrigerant gasses is good? Yeah duh!
EVs are better for environment than ICE? Oh yeah!
But climate is not just a first order effect. It has more orders and complications than we can calculate. Our models are horribly incorrect, predictions further than a few years are always wrong. But that's science. We do our best and try to correct asap.
What's wrong is zealots who spreads doom and gloom about this. Makes predictions like my country will go under water by year 2000. Do you know how frightening it was for a science interested kid to read that? Kilimanjaro snow will be gone within decade. One extra decade gone by already, why is there still snow? 96 months to climate/ecosystem collapse? Double 96 months gone by, where's the collapse?
EU shot its own economy to pieces chasing these garbage. Wealth creation is the answer. Look at China. Didn't give one shit about it, built up wealth and power. Now they have the resources and they are rapidly cleaning up. In 50 years, China will be more green than US/EU while Germany killed the cheap nuclear power that was the backbone of their economy.
The REAL climate change is not tiny amount of co2 or whatever humans are doing; it is earth and sun. Orbital geometry, solar output, ice sheets, oceans, volcanoes and tectonics.
You can remove 100% of "environmental damage" from earth by humans from start of humanity, and it will not make a dent if/when earth/sun decides it is time to warm up. And we know for a fact that earth goes thru high temperature times and right now it is in an ice age.
What we need is even faster pace of generating wealth, power, and technology that helps us to clean up the air and water. Not be regressive out of premature fear.
Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.
You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…
Social conventions are regressive and are followed blindly by luddites. Real progress comes from doing what is right and breaking idiotic conventions.
Pirating is fine.
Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.
Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.
When society becomes a free-for-all of people with haves making it impossible for the people with have little to have anything at all, it begins to collapse. People have no incentive to treat anyone else with dignity and respect. They no longer take care of one another.
Competition does not have to mean kill or be killed. We can still be decent, helpful, charitable while competing.
You are just bullshitting without any understanding of "people who have little". I'm one of those who came from "have little". As a consequence, almost everyone I knew growing up, were also "have little/have nothing" people. Among 100s of people in my super extended family (tree), among 100 of my "have little" classmates from school, it was always clear from early age who will "succeed." The ones who had grit, dedication, focus, and a shine in their eyes.
The ones who complained about wealth and society, are still in the have little group. The ones who didn't care, heads down worked hard, are living in luxury.
Taking care of one another is mostly a privilege afforded in society with wealth that leads to high trust situation. I grew up in a place without it; so I know how valuable (and desirable) "taking care of one another" is.
Current western generation (esp far left) has no idea how good they have it and are actively destroying wealth that will lead them to the chaos they have no idea about.
I don't want that "way down" just like you. But unlike you, I don't believe in fairy tales. Reality is capitalism works, wealth creation works, we all end up in better places (not equal...but overall better). Much better than road to hell is paved with good intentions.
It is shit compared to what it could have been. It is diverse in language (which is a negative), food (yet mostly boring) yet common in one thing: regulate to suffocate. Kill (most) innovation. Live in borrowed times. Spend money that we don't have. Fuck future generations even more. Then wonder why far right is flapping their wings.
It should be no surprise that it's a movement of lying thieves and scammers, Effective Altruism, that's behind close-source AI in the US.
Now I'm not sure justice won't come: SBF is behind bars for 25 years. He could turn out to not the be last scum from the EA movement to end behind bars.
As to open-weights models: at least it's not sold back at a mark-up and anyone can run them.
IMO if they didn’t have proper licensing to train on the data the model should not be copyrightable.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owns anything they invent or create. [0] It was a long time ago and I remember feeling sad watching the story.
It is this one line, Article 1 Section 8 Clause 8, that separated the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Right clause a communist revolution, Violent or not. That is the one thing the communists would change to make the United States a communist country.
In a communist society there is no profit (or incentive for), thus no need for copyright laws.
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
Well it's worth reading the linked wiki article section, which includes a link to another article "Copyright law of the Soviet Union", flatly contradicting you unless you maintain that the USSR wasn't truly communist.
In the United State, the individual or corporation owns the invention. In the Soviet Union the state automatically owned the invention. That clause is what ensures private ownership.
The clause is what ensures profits from market sales or licensing of ideas go to the creator.
Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.
In the Soviet Union the state owned all media (de facto) and you had to get permission from a party official before you published or copied anything. It's not the same thing as a free for all.
It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.
In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you had bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine.
> This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn.
Learning isn’t stealing, but if you’d set up a forum or hotline where paid staff would answer questions and write essay directly based on NYT content without a license for that content, that would probably be deemed illegal.
I would highly doubt it, because news outlets and blogs rehash each others' reporting all the time. How many times have you seen an article start with "Today <XYZ publication> broke the news that..."?
Since this is a copyright fight, rights extend only to verbatim copies of the full work or significant portions thereof. Abstract things like facts, ideas, concepts, themes, and patterns are explicitly not protected, and rightfully so. Yet those abstract things are what get repeated and distributed, and are what get encoded into model weights.
This is probably plaintiffs' biggest challenge because it has been very hard to get models to regurgitate entire works except for a very small handful of extremely popular works (and now there are guardrails against even that.)
Wiki already democratized it just fine and was legitimately free for people who know how to read.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
There are no viable free versions with sufficient computing power. The do-it-yourself AI is dangled as a carrot in front of users to camouflage the lock-in and rent seeking by Big AI. Go prove something new like Navier Stokes (replicating N-S itself no longer counts due to scraping and plagiarism) on your Mac Pro!
Paid influencers who perpetuate the open narrative are a whole new industry.
Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".
There is no lock in. Anyone with sufficient budget can download the weights for GLM 5.3. On OpenRouter, I count 29 different providers for this model [1].
In terms of purely local LLMs, one can run GLM 5.3 flash on a beefy workstation.
Ah, this remembers me when I was also young and naive.
Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..
Yes? I own a lot of books that were public domain when published (as reprints). I could read them on Gutenberg, but I'm paying for the nice paper formatting.
Our "culture" has long been the province of corporations. In prior epochs it was still the product of patronage and power.
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
Flourishing life? Where is your data leading to the conclusion that AI models are leading to the world populace leading a ‘flourishing life’? Are taxes on revenue of companies like OpenAI somehow being collected and turned into a UBI and I just didn’t hear about it?
I mean I'm producing more and better art and software each month than I used to do in a year. I feel like I'm flourishing because of AI. Tell me I'm wrong?
You’re not wrong. Me too. I spend most of my waking hours using AI for new product development.
But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?
I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.
Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.
I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?
Considering that we don't seem able to appropriately tax the biggest corporations, why do you think we'll suddenly be able to redistribute LLM gains?
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
These doomsday theories sound solid in practice but in order for someone to accumulate a lot of wealth there have to be consumers to chip into this. If we are all slaves, where is that wealth going to come from?!
Not necessarily as there are methods to redistribute wealth without the consumers having any choice.
Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.
Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.
There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.
> At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
We are not deluding ourselves that we will maintain our standard of living. If the boss’ dream of replacing all workers by AI that does their job poorly but cheaply will be realized. We will go to the gutter with the rest of the economically worthless people.
There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.
It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.
As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.
Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.
You assume there will still be free elections and democracy. There is also the possibility that the form of government will change into something else.
The value will tend towards infinity [1] because we can keep building more data centers. This will allow us to apply tokens to solving every problem we can conceive.
The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.
If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.
Industrialization led to iPhone, PlayStation, Spotify, and Waymo.
AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...
AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.
[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.
Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.
Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.
You think people are going to stop watching celebrities and influencers just because there are AIs and robots?
Do you think people will stop dating, trying to impress mates, buying luxury, etc.? Neither candlelight dinner to a robot no a Dior manufactured by robots have the same allure. There's a huge industry around this.
Sports aren't going to go away. Huge industry.
People aren't going to stop making art. I know a ton of artists who have embraced AI that are doing even bolder work using the tools. (I was a filmmaker pre-AI, and I know a lot of people in this field.)
People aren't going to stop traveling. And consuming. And eating human food and consuming human experiences.
There are going to be all new kinds of businesses and opportunities that spring up. OpenAI and Anthropic are not going to be the only two employees. They won't be staffed by only agents.
It may be the case that the potential output of a worker engaging in the productive process goes arbitrarily high. But the that won't matter, if they're not permitted to. Given a choice between involving a human worker who will demand compensation, and a fully general robot, which will the owner class choose?
The vast majority of human intelligence is already squandered: millions of potential geniuses in impoverished places, suppressed by lack of opportunity to flourish. It's not about the quality or quantity of intelligence. It's about who controls it.
The fruits of the industrial revolution didn't end up in the hands of the worker by divine grace, or by some natural law. They were won by the hard struggles of the labour movements, by leveraging their indispensibility to the process of production.
If we want the utopia you imagine to be accessible to ordinary people, workers must cease being so eager to build their replacements, and be prepared to collectively struggle for their share. But if we are lulled to complacency by the notion that this will be a passive process, that struggle will not be necessary, then prospects are grim.
In the 50's there were predictions that in 10 years no-one will have to work again because of the advances made in automation, like the washing machine for example. Any predictions of less work this time around are a complete joke.
For thousands of years humans have dreamed of reaching the stars.
Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.
On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.
You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?
That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.
US has been procrastinating reparations for slavery. LLM redistribution can't happen until the capitalism machines recursively solve "sins of our fathers".
> a world in which the labor required of each human to lead a flourishing life approaches zero.
The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.
Just focusing on redistribution, The frontier labs are based out of the US, so how will you redistribute to people who are not in the USA? This assumes you don’t have a riot over the very idea of transferring profits from US firms to other countries.
Does this mean that all countries will have to force labs to create subsidiaries within their borders, that are then taxed?
Also, this isn’t foreclosing a past that didnt exist. What has occurred is absurdist levels of theft, in a political and business environment that doesn’t have to worry about the wronged individuals being able to have their grievances heard.
>a world in which the labor required of each human to lead a flourishing life approaches zero.
Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?
Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.
At the very least we should foribly confiscate the models & make them available free-for-all as open-weights downloads. Failing that, bring back the guillotine.
So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.
Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.
The more fundamental objection seems to just to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
> Would that resolve the injury caused by the intellectual osmosis?
Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.
What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.
Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
Musicians and writers build new works on top off millennia of literary and musical history, then copyright their works and sell it back to us. Is this so bad? If Taylor Swift, consciously or subconsciously, gets an harmonic idea from a 1970's song and a fragment of a melody from some 1990's song... is that theft?
I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.
But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \
It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.
I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
The problem is not Taylor Swift is being "inspired" from other artists. We have tons of examples this throughout history.
The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.
This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...
We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.
Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
Ok, thought experiment: at what point in the past 1000 years do you think the balance of collective versus individual benefit from creative works was at its most fair?
> getting things were free was part of sending a message, not monetary gain.
As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.
We’re about the same age, probably. On this side of the pond, all of this stuff was rooted in inability to get something legally first, and if it was possible to get, it was prohibitively expensive to buy for us. So we got it for “free”. The thing is we were not using the software to earn money, so companies don’t care (this what I heard from a couple of companies directly). Music was in a similar position.
It was a righteous moral stand because we wanted to get things relatively affordable for us.
I for one prefer to buy my software and music nowadays, because it’s affordable and I can get it at the quality I want.
For your other question, the answer is probably 1800s, because the current model was not entrenched everywhere and creators had sane rights for what they created. So creators had to get what they made out to masses to show it, but lost their rights in relatively shorter times, so things were free to use for everyone. So you can’t excessively milk something till proverbial eternity.
On this side of the pond, wow were the college kids excited when they discovered they could reframe piracy as a political and countercultural statement, despite things being available paid.
At the time I pirated a lot of stuff. If we're of the same age, you may well have played video games I cracked in your teens. And, TBH, for me it was mostly collector mentality (I have have all the games!) and a bit of poverty (I can have the games I'd like to buy but can't afford), and about zero politics.
That said, I'm generally with you on 1800's. It's astounding how fast things changed. IIRC it wasn't until 1890 or so that international copyright was even a thing; you published in your home country and publishers in other countries just copied and published without permission or payment.
And then just 100 years later, copyright was essentially perpetual and global.
> It's the robbery of all of our culture to sell it back to us at a mark-up
Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
Imagine someone asks you how to do something at work, you tell them. Then they create a huge packet filled with bullshit about how it got done and now they are your boss.
It's a shitty move, but ultimately, in between the bullshit narrative, they also did the thing - not you - so the promotion rightfully belongs to them.
Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
Of course you tell them, because you're nice person, and not jealous of someone else succeeding in a thing you aren't even pursuing? You wouldn't want to act like "the dog in the manger" from childhood stories.
I bet you do, you probably don't recognize it because it has different names in different places (for me, it's known as "pies ogrodnika" - "the dog of the gardener"). It's an old tale.
Isn't that the same point piracy advocates have been making for a while? If I watch a pirated film, I didn't really consume a physical resource. Nothing physical is lost. Therefore it's not theft?
>Learning isn’t stealing.
This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid.
> Isn't that the same point piracy advocates have been making for a while?
No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.
Learning isn’t copyright infringement.
Learning isn’t theft.
Copyright infringement isn’t theft.
These are all different statements, and all are true.
> There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.
“Gaining exposure to” isn’t theft either.
Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.
There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.
If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.
If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.
If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.
If people stop making art because they can't outcompete generated art, that is also lost to LLMs.
So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.
> If $X dies because ..., the entirety of its knowledge is now only availabe through LLMs that trained on it.
And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.
The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.
> It's the robbery of all of our culture to sell it back to us at a mark-up.
Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.
And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:
Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.
To use another Slashdot cliche: "But you knew that already."
AI didn't steal any culture. It just made mediocre culture more accessible. Turns out, most people do like mediocre culture. Previously, public TV channels could at least pretend that people enjoy educational content or classical music. AI exposed that lie. And now we're shocked that the emperor is naked.
If we ignore the massive amount of books being destroyed, the outages and increased usage bills for once free websites (some resulting in closure), the increased difficulty to access public information such as Reddit and Twitter, and lastly ignore the amount of conversations being taken away from the public in favor of LLMs, I would agree.
Most of these things are bad. None of them is robbery, though. Some are of questionable legality, but most are not.
People need to face reality. The old Internet was fragile and couldn't have lasted long. It was dying slowly due to closed social media and LLMs are accelerating that death.
I'd rather it die quickly and be replaced by something more robust than decay slowly. I was sick of how it was before LLMs even came among.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.
Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.
They work very hard to protect their own IP so they obviously believe that IP is worth protecting.
Yep: the problem with IP is exactly the "Property". The texts are economically (even if only for leisure) valuable, so the producer deserves a fair comoensation unless he has explicitly dismissed it.
So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
It hasn't been literally stolen but indirectly it has been.
I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.
Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.
AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
Someone once told me, any illegal thing can become legal if you add extra steps. This is not _legally_ true, but it is _practically_ true. Evading a tariff is illegal, until you start using a proxy-country. Firing an employee, and not giving them severance is illegal, until you find a way to make their job so miserable, that they quit on their own. Really, if you can turn signal into _noise_, it becomes difficult and expensive for the enforcement organs to actually enforce their own rules.
It is unclear to me just what percentage of tech-companies, are in the business of adding extra steps to an illegal process, of turning signal into noise. For example, the vapes made by Juul are a kind of hack around public health laws. The Amazon marketplace shields merchants who sell counterfeit goods. Uber and Lyft bypassed the local laws that applied to taxi services and the medallion system. Airbnb did something similar with the laws governing hotels, and delineating who owns and who rents.
It is a special kind of disappointment, to be someone who loves technology, and to have to work in the technology industry -- because it appears to be run by people who hate everything that is not money.
I'm sad to hear your business model didn't survive the march of technology, but at the same time, this is what's been happening to everyone ultimately. And rightfully, no one is entitled to a business model working in perpetuity.
I know this is a real harm to individuals, and bringing this up is not some kind of anti-technological thinking (the luddites were OG with that, and I always argue in favor of them - they had a very valid point).
I also realize that in a few years, I very well may be in the same position here. So will most of us.
But that's a different argument than GP was making, different one from what I replied to. That was about losing culture, and losing a business model is not losing culture.
Also I was with you all the way until the last paragraph:
> AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet.
AI is doing exactly what users want it to - what I too use it for - it bypasses the spam and scams that stand between the user and the solution to their problem. Good riddance, in this.
> Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
And that is just bullshit. AI is not "taking everyone's content and selling it back to users", and the vendors are definitely not "keeping all of the profits" - on the contrary, they haven't even begun to figure out how to monetize almost any of the value their inference services provide, because they have zero visibility into how much any stream of tokens ends up being worth for their customers. They cannot tell whether the plumbing advice the LLM generated unclogged my toilet, or prevent a restaurant from having to close for the day; they cannot tell if the code their agent wrote won me a beer from a friend, or unblocked $2 000 000 dollar opportunity for my business. They capture none of the value from this either way.
> Except, of course, no one has actually been robbed
No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.
Yeah, this whole "we're being attacked by China, they're distilling our models!" thing is completely and utterly absurd. Insulting, even.
Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.
IP is obviously not "real" like physical property is real. Moreso, IP is a restriction on free speech as it walls off some expression as "Copyrighted" so you can't legally express it yourself without paying someone for the privilege.
That's not to say that IP is useless, but treating IP infringement as "stealing" was always a problematic shorthand. Previously, calling IP infingement "stealing" was the domain of corporate interest groups like the RIAA or MPAA, but with AI the winds have turned and supposed anti-corporate lefties are keen to treat IP infringement as theft to attack AI companies. There's very little intellectual consistency on either side of aisle here.
There are a wide variety of middle-ground views between them. Copyright law itself is one of them: it provides exceptions for transformative "fair uses", and there are some courts that have ruled that AI is one of them. Another might be Harbinger taxes on IP.
I think the blend of the two is that value added to any information, such as the first time its distilled or positioned in a certain way, that's IP and should be real.
When an AI company takes that type of value, and doesnt provide value back to the person who made it, thats theft.
Information / facts are real and free. But those who helped us get here should choose how to license their work, and an AI company picking it up and re-selling it is clearly unethical.
> But those who helped us get here should choose how to license their work
I can reasonably wrap my head around the idea that an author should be compensated for their work, but where in the social contract does this power to control how that work is used come from?
It's so destructive and tangles up courts and makes contracts complicated and we lose the original versions of the work because e.g., they have to change a background song due to complicated licensing.
I can respect protecting copy rights, but it should never be conditional; if you choose to make a work available to the world, you have a legal right to defend the unauthorized copy of it but you should not have a right to say how it gets used.
Intellectual property is a legal fiction, not an empirical fact about the universe. It's as real as societies want it to be, and that view can change over time.
It’s real but whether you can make a profit or not is dependent on whether you can protect the information. It has been this way since civilization as we know it started, I believe.
If you are freely posting sentences like this was on the internet you are giving away your IP for free. Anyone can read the sentence. If someone can make money off of it then that’s just markets at work. I can’t make money off of what I write here, for example.
But I also don’t think pirating a movie is theft either. You haven’t proven to me you’ve lost money. Maybe wouldn’t have watched it anyway.
That would be on what the AI model generates, not what it is trained on.
The latter is where the contention is, and it's a valid argument. So much so that some companies are not using stolen information to build their models.
IBM for example indemnifies its models for its customers and has detailed information on where the sources came from to train them.
> The latter is where the contention is, and it's a valid argument
It has been tried in court several ways already. Remember the lawsuit that forced Anthropic to use physical books? They tried to argue that the books couldn’t be trained on at all. It failed.
I'm not sure I understand this argument. If I go to the library every day for 10 years and learn everything there is to know about subject x I shouldn't be able to sell my skills to the world about it later because I didn't give the creators of the books I read any money?
Arguably LLM companies could have made large-scale deals with libraries and got the exact same knowledge (much, much more slowly). I wonder if people would have the same issues then? My guess is probably. Goes back to the meme that if libraries were proposed today there's no way they would ever be allowed.
it's of wrong scale, you cannot claim derivative work when you literally ingressed sum of knowledge while destroying it in the process so your "competitors" cannot do the same
Objectively, IP exists and is acted upon, so IP is real, obviously, but IMO it is bad, and should not exist.
People tend to forget that IP was not originally about digital distribution at all (copyright), it was about giving inventors exclusivity periods to profit without competition (patents).
It was a misguided attempt to stop sometimes literal theft of designs from rival inventors, by tying the design to the person instead of whoever possessed the schematic.
It's also a regime that in its modern incarnation protects businesses, not artists.
AI scraping for the goal of making a commercial LLM service is different from, say, a commercial file sharing platform.
The first difference is that LLM training is highly transformative. Let's say your LLM ingests the Harry Potter novels during its training. What you get at the other end is not the Harry Potter novels, it is a LLM that can talk to you about Harry Potter, it is not the same thing, and going from one to the other requires a significant amount of work, very expensive work in this case.
Not only that but there is no direct competition. People won't stop buying the Harry Potter novels because a LLM trained on it exists. If you want to read the books, you buy the books, you don't ask a LLM about it. A file sharing service on the other hand competes directly against the official channels, if you want to read the books, you can download it from this service instead of buying it on the official channels.
So, about how free you should be to get these weights from the AI companies. If you just share a 1:1 copy of the weights, that's the "file sharing" situation, not transformative, you took their work, didn't do any of your own. Usually considered unacceptable by IP laws.
Distillation is a more interesting case, you are using a LLM to train your own, it is transformative work, but you may also be competing directly against the LLM you are distilling. So, in a sense it is worse than scraping, but still, despite how much the likes of OpenAI and Anthropic are complaining, it seems to be legal.
So it is somewhat consistent: 1:1 copy and distribution is not allowed, be it source material or LLM weights, and training is, be it source material or another LLM (as in distillation).
> The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.
This is how IP has always worked. It has never protected the little guy. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.
And by constructing that false dichotomy, you're failing to see the actual truth, which is that Intellectual Property is just a legal construct to turn ideas into Capital, with the goal of incentivizing its creation.
Nothing more and nothing less, and all of the normal debates about what should remain Capital and what should be The Commons apply.
IP could be real and LLM training could be considered fair use. IP protects published information, so even abolishing all IP would not help us since the LLM weights are secret.
Nothing short of a global revolution can fix this. Proprietary LLMs should be illegal. Either we achieve post scarcity within this generation or it's literally over.
Content creators are taking steps to prevent bot from walking away without compensation for their visit as if they’d been stolen from. At least Google, in the beginning sent traffic back that could potentially be monetized by the creator.
Perhaps some. But I'd wager orders of magnitude more are using LLMs to brainstorm if not outright generate content for them, which they are then reframing as their own creation.
> orders of magnitude more are using LLMs to brainstorm if not outright generate content for them, which they are then reframing as their own creation
I doubt it, those are the same group of people anyway.
Orders of magnitude more people are using LLMs to brainstorm or generate content for them, which they are then using to solve their own problems and carry on with their life, increasingly solving many more problems for themselves this way, and not publishing any of it.
The world isn't made of "content creators" flinging crap around in hopes of monetizing it. Most people have actual jobs.
I don't know, and I don't care. Scaled proportionally to actual contribution, it'll be like $0.000001 / year, and I'm not going to begrudge the AI companies for not paying me a fractional cent, and I'm definitely not going to cry foul and demand stopping progress over that fractional cent they plausibly owe me.
Not the least because I'm already getting many orders of magnitude more value each day from them providing me inference as a service.
It is robbery because you're not paying the person that produced the knowledge, you're paying the person that took the knowledge and put a chatbot interface on it.
Many many people didn't give permission for the data to be ingested and used this way.
We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.
We have free plans and open-weight models you can host yourself, or pay someone for hosting, or host yourself and sell others to recoup the costs.
All of these lag SOTA by ~6 months on average, so yes, indeed, the thing you could only pay for half a year ago, is now free for entire world forever, and it only keeps getting better with time.
The past is still there, but our culture, being a live thing, is pretty much being affected by AI. We now have an artificial intelligence in this loop of exchange of ideas. Slopifying everything in its way.
That is a different argument entirely, though. Not the one GP presented.
On this, I have no concrete views. Sloppifying is real, and AI is short-circuiting "this loop of exchange of ideas", for sure. On the other hand, we had several precedents on this within last 700 years - the printing press, the radio, and the Internet to name the largest three - and culture turned out fine. Different, but fine. AI feels more than this, though, so I'm not putting that much stock in argument from history here.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form.
But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of denying it.
[1] Fuck you, Claude, for making me wince at that.
The problem is the hypocrisy. They scrape, but don't want others to scrape them. If they act like they've been stolen from when they get scraped, that is a confession of guilt to when they did it themselves. They should be punished according to the same rules they wish to impose on others.
I really agree with this critique! It's very hard (without hiding behind ToS) to claim that creating the models was A-OK, but distilling them is some sort of ethical breach or attack.
> the culture has not been stolen - it's still there
It's there in much diminished market, one that has been saturated overnight on a scale previously unfathomable. Their art is still there but it has been used against them to obliterate the playing field.
> nor are the people involved selling it back in any form.
Aren't all of these AI companies selling subscription models for people to create derivative art?
> And let's not forget what we got back for this
Get back for what, exactly? Didn't you just argue that no theft or sellback has occurred?
"It's still there" only in a technical sense because it's being buried by the automatic slop machines.
It used to be that, no matter what you want to do, there's a video on youtube of some guy who knows what he's doing showing you exactly how to do it. It could be, like, compressing a rear disc brake cylinder on a car or matching fiberglass gel coat colors, or carving a statue out of marble.
Theoretically those videos are still in there somewhere, all 15 years old at this point, but you'll not find them with youtube search instead you'll find a million AI slop videos no matter what you query.
> There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
There's a way that many people fear this is true: what you're getting for "free" has some external cost that you aren't taking into account.
Cost 1. The environment. Tech companies were once significantly interested in efficiency, their carbon footprint, etc. With the advent of LLMs, this was thrown out the window. Concerned people are now thinking about water footprint, heat footprint, noise footprint, and probably more. These things are hard to put a dollar figure on in a short comment, but the sustanability of a liveable climate for the billions in this planet is literally priceless.
Cost 2. Employment. We face a significant cull of employability -- graphic artists, programmers, mathematicians, paralegals, and more are rightly fearful for the end of their career. If we're getting intelligence for "free" in 2026 dollars, but we can't get jobs in 2030, will all that content still be affordable in 2035?
Cost 3. Infinite investor dollars. The AI companies are burning cash at a historically unprecedented rate. This has attracted a lot of attention from traditional investors, like pensions, banks, etc. If all of this ends up in a free product, that doesn't sound like it'll return those investments. And that can crash the economy -- I ask again about real affordability in 2035.
This has another side effect in that it breaks the power that consumers have in the market to have a say in how resources get allocated, by voting with their wallet.
The entire AI buildout is non-consensual. Individuals have zero say. We could all refuse to buy AI subscriptions and it would not matter because businesses will still buy, and, investors have decided that we are moving forward with this no matter what, seemingly whether anyone is actually buying or not.
AI isn't necessarily unique here, but it is one of the biggest new examples of wealth inequality and how society at large are no longer the ones who get to decide how resources are allocated under a capitalist system, especially when the investor dollars behind it amounts to the GDP of a small nation.
And this is fuel for climate change denial. See? Are the wealthy and powerful people acting like they are concerned about the climate? No. They are doubling down on energy use and consumption. It's all a fraud!
The idiocy of those arguments around AI and climate are indeed a fuel for climate change denial, because the only consistent views are that either the AI is not an environmental disaster, or the whole climate change thing must be overblown. The numbers don't add up, but if you're so blind with hatred towards "AI companies" while still desiring to have a remotely consistent worldview, it's seriousness of climate situation that has to go.
Item 2 is the big one. Because not only is AI about to destroy jobs wholesale; it's doing so in an environment that abhors wealth and income redistribution that could blunt the harms.
1. is mostly bullshit that's used by anti-AI crowd to gather further support. Data center companies did not stop caring about footprint, they're still pursuing that for the same reason they did it before: it aligns with overall lowering of costs. AI added an extra incentive to improve efficiency because the demand far outstrips supply of compute.
For 1., also the general argument stands: yes, these data centers use energy, just like everything else humans do. What matters is the value it provides to people, and in case of AI, it's one of the most objectively useful expenditure of watts on compute (your point 2. notwithstanding).
I heartily disagree on 1, and you've taken a cheap out on that point. What I said was, the environmental externalities are not in your cost model, and you've said nothing to even acknowledge that. Your point about efficiency is particularly obtuse: efficiency doesn't mean burning less fuel, and getting more compute per Joule. Not reducing the overall carbon footprint, but accelerating the burn rate as fast as we can build.
I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
I'm saying that environmental externalities are hugely overblown[0] and to a large extent are part of the cost model. Datacenters are not like the legacy industries, they have to pay for the land and energy they use.
> Your point about efficiency is particularly obtuse: efficiency doesn't mean burning less fuel, and getting more compute per Joule. Not reducing the overall carbon footprint, but accelerating the burn rate as fast as we can build.
Efficiency does mean getting more compute per Joule. They are doing that. They are also building out as fast as they can, because both of those address the same problem: demand for compute outstripping supply for it. They can't just pursue the build-out strategy, because hardware production is now supply-constrained too - that's where the "RAMpocalypse" came from.
Yes, this is a huge build-out, and little to none of that is 100% carbon-neutral, so ecological footprint adds up. But that's normal and expected and a kind of tradeoff humanity has been making forever: building new stuff, be it hospitals or airports or data centers, has an environmental footprint, and we only hope what we get from it is more important to us on the margin, and that we can offset the environmental costs some other way.
> I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
I'm not talking about financial "demand". I'm talking about real demand, for compute. There's nothing to doubt here, it's pretty clear that all major AI players are constantly running at capacity - it's obvious in the very structure and limits they place on even the highest tiers. Unless you believe half the data centers are just spinning busy-loops and burning energy to inflate stock prices and generate a fake reason to build out more compute capacity - if yes, then I don't know what to tell you.
--
[0] - Like the water story, where the reported datacenter usage numbers look big in isolation, but once you relate them to "how much water is there" and "how much is used by other industries", it turns out to be a nothingburger. The only numbers that survive are those showing the story is really about those other industries having a competitor for cheap water now, and having to pay a bit more than they're used to.
> reported datacenter usage numbers look big in isolation, but once you relate them to "how much water is there" and "how much is used by other industries", it turns out to be a nothingburger.
Water scarcity is a problem around the world; even my rainy hometown of Vancouver had steep water restrictions this year. The question "how much water is there" is incredibly disingenuous, to the point of feeling like a motte and bailey: yes, the universe, or the earth contains vast oceans of the stuff but datacenters are using drinking water which is quite scarce. Just because "other industries" use a lot of water doesn't make it okay -- I level exactly the same criticism at other industries which waste water at such an egregious degree.
To be clear: if data centers only charged up a pool of water to use as coolant, that's fine. But when the weather is warm, they run millions of gallons of tapwater down the drain just to cool off. And heat spikes are happening with increasing frequency, severity and duration.
We as a species need to conserve water and emit less CO₂, full stop. The tech industry doesn't get a pass just because other industries are dirty.
>real problems of real people, including individuals and non-profits
Again with the weasel words. Real people, including individuals? Pretty clear this is written from a pro-business perspective and can be safely dismissed.
>"robbery of all of our culture to sell it back to us at a mark-up".
Well, it is that. Why is it criminal for me to steal a big-business-movie for personal use, but they can just take all my blog posts and sell their deritives to others?
> Except, of course, no one has actually been robbed, the culture has not been stolen
I respectfully request that you get all the way outta here with this take.
Artists have caught AI generating literally their own work, for free, at scale. There have been tons of articles on this website about the AI hyperscalers slurping up books, copyrighted works, etc through legal and questionable ways. Even TFA says that the NYT suffered absolutely devastating CTR drops.
If you create a blog post about something super esoteric, it will guaranteed end up in the training sets for every big-lab frontier model within 24 hours (probably less) and probably show up in Google's AI Summaries at around the same time. No audience for you!
I have artist friends - asking AI to draw something similar to subjects they've drawn ends up making, line for line, the exact image they created, hallucinated alongside a couple others. They spent decades getting good at their craft. They spent so many years getting paying patrons and customers to commission them.
You'll never agree but I think you underestimate how many people consider all that the AI companies have done to be the wealthy stealing and selling things back.
> nor are the people involved selling it back in any form
I am convinced that the longer one works in AI the less one has any grasp on reality
> too cheap to meter
Its the most expensive buildout in human history - you can't just split half the cost. The spending is the only significant growth in the US economy. Unless you think that all the datacenters and all that capex are for training?
I have horse & wagon friends - they spent decades understanding the ins and outs of roads, some of them also creatively made up their own routes and roads.
However ordering a car with Uber through your phone just gets people to the destination faster, cheaper, more comfortably and reliably using those same roads.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there
It's actually not there. See how many websites have now closed doors or ceased to operate because of constantly being hammered and bombarded by robotic scrapers. For others, it made websites that turned a small profit into unsustainable money pits.
Sure, their content might now be ingested into an LLM training set sitting somewhere on proprietary servers. But the site itself (the origin of truth) now does not exist.
So no. It actually isn't there. Resting on this falsehood, the rest of your retort makes not much sense.
reified intelligence on a chip, almost too cheap to meter,
“Reified intelligence” is stretching the truth. It is certainly closer to intelligence than anything we’ve come up with before, it can certainly match or beat actual intelligence in some domains, but just as certainly it clearly lacks many features and capabilities of intelligence.
“On a chip” is also stretching the truth. The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them. Individuals who have free usage for a project commonly report that, had they paid, it would cost five or six figures. Companies have seen enormous bills, some approaching the size of their payroll. And all of this is true for tokens that are dramatically subsidized, by one of the most intense and largest concentrations of capital in history. It is the polar opposite - “almost too expensive to even do, and absolutely must be carefully metered”.
In summary, let us indeed not forget what we got back for this: extraordinary distortions of reality evenly intermingled with bald-faced lies.
> It is certainly closer to intelligence than anything we’ve come up with before, it can certainly match or beat actual intelligence in some domains, but just as certainly it clearly lacks many features and capabilities of intelligence.
It's not complete or that well-rounded. But it's something that was the domain of speculative science fiction only 5 years ago, and it's rounded enough to be applicable to ~everything to some degree.
> The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
By "on a chip" I meant more "in silica" than literally on a single chip" - though this actually is* true, but those chips aren't cheap.
> And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them.
You are looking at power users that use LLMs in agentic coding sessions. Most people just run off free tier of ChatGPT, which is free for them. There are equivalent open-weight models at this level, and while hardware to run one for yourself is expensive even for most westerners, the marginal inference cost is literally dirt cheap, which is why you can get that for near-free from smaller inference providers - or pony up some money, rent a bunch of compute with friends, and become an inference provider yourself.
(It's only a tough market because the major vendors are giving out better models than you can run for ~same or lower price than you can offer. Which either way is too cheap to meter in terms of solving useful problem for real people. Again, developers are a special case of power users, as usual.)
> because of how many real problems of real people, including individuals and non-profits, it addresses.
Not problems, laziness. The way I see students and colleagues use it is to get their work done with less effort. That's its selling point. There are not many real problems LLMs address.
> no one has actually been robbed
In our society, people get paid for work. If you think society is wrong, fine, but you must state so first. Under common assumptions, all that data has been produced through work, and that work represents value. Taking it for free is therefore theft.
> reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West.
I think this is quite the pollyannaish perspective and very much inline with those that think if you can take, then take and only apologize when caught.
We won't get anywhere in these discussion if good chunk of participants cannot admit to the trivially observable facts about the actual objective reality in which they live in.
It's become a bit clichéd to say it but I think it's different this time.
The kind of "piracy" you're talking about wasn't depriving the authors of anything because you could always make the argument you weren't going to pay for it anyway. If, on the other hand, you were making copies and charging people for them, you could definitely say you were depriving the legitimate authors of that revenue. AI companies are very much doing the latter, not the former.
The other part of it is it's not just copying. Previously, if I decided to make a copy of a work without paying, I'm only copying the work, not the author's whole writing style. Now the AI companies are depriving authors of revenue from works they haven't even made yet.
If they were able to create a model de novo then they could truly claim it hasn't just been lifted from existing culture.
Every time AI is used it is a theft from artists, programmers, writers, and all creative workers whose work was used to train it without permission or compensation and are now unable to find work or buyers. It is absolutely theft and you're being intentionally disingenuous to suggest otherwise.
“It is said that at the heart of every great fortune there is a great crime“. That quote is from a fiction author - you’re essentially quoting Spiderman “with great power comes great responsibility”.
Other than the rare books than have been ruined, all the same knowledge is still out there though. So it’s not robbed in the sense of a bank heist. Maybe in the sense of pirating a movie.
The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
Nobody has made much money from AI yet unless you count the shovel sellers (Nvidia). Not sure closed models would make any money ever and eventually the benefits should flow to everyone.
"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown, 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
I wasn't aware of the "sweat of the brow" doctrine until you mentioned it. But you seem to be using it incorrectly. The doctrine only states that creativity or originality isn't required to make a work copyright-able.
Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.
The parent comment I replied to is concerned with "life's work got appropriated without consideration, compensation or consent". To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions. Today in most jurisdictions copyright laws do not care the slightest about an LLM ingesting databases -- phone directories, sport fixtures and results, someone's life work measuring the dimensions of frogs, etc. 100% of the original factual data could be learned by the LLM, and 100% could be output all at once.
^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.
> To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions
No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" can extend beyond collecting facts. Originality and creativity are also work.
The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.
However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.
> LLMs are sufficiently transformative and do not hamper the potential market of authors of training material
And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.
We know what is going inside black holes in other galaxies, we know the details of Israel nuclear program..., the Windows source code got leaked, we got the NSA tools and full details and locations of the Echelon architecture... Phds in Maths warns us daily about the terrible secrets of the evil AI inside their labs. How, their numeric matrices and gradient descent Python scripts, are about to kill 10% of us all, I guess the sick or genetically less interesting ones...and use the rest, as some meat/metal hive drones part of some Borg collective...
There are ONLY TWO Stories and their details, that we collectively will never see.
1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...
So why is there so much outrage now when Google and other search engines did it 30 years ago? They literally said their objective was to collect and index all the world's knowledge, and when they scraped the internet clean a thousand times over they invested in scanning and digitizing everything that wasn't on the internet yet.
AI is better at repackaging it back to the end user but ultimately I'm arguing it's the same thing.
(caveat: yes I know there were plenty of people that objected to Google et al indexing everything; famously, Gmail was scary to a lot of people because they read your email to give you ads)
It's not the same at all though. Google indexes the original content, hosted at the original location. It just gives you a way to find it. LLMs ingest the original data, throw it away, and spit it back out in whatever form they like back to the user. And since it's gotten so popular, many peoples' interactions with the content are no longer in the original form at all; their site traffic / book sales diminished.
Google didn't immediately create $20 and $200 per month paid tiers. The free thing lasted a long time, and the initial monetization with ads, when it eventually came, was not immediately shoved in the user's face. the enshittification and bean counting came way later compared to the monetization blitz that occured with the AI companies.
Because they aggregated links to that information rather than copy it outright. It’s like the difference between an encyclopedia and its table of contents.
Google (and other search engines) gave you an easy and reliable way to opt-out of having your content indexed and linked to. People did in fact object pretty loudly every time Google attempted to surface the information on google.com rather than sending traffic to the source.
> It's the robbery of all of our culture to sell it back to us at a mark-up.
At a mark-up would mean it’s more expensive. The outrage is that they’re taking knowledge that was expensive to access because you had to hire experts or otherwise pay a lot of money for it and making it accessible to anyone who signs up for the ChatGPT free tier.
Calling it “robbery” is also specious as no knowledge was taken away from anyone. The content in the training sets was out there in the world one way or another. It still is!
I’m really perplexed by this sudden swing toward the idea that knowledge is something that we should encourage or incentivize to keep locked away or that other people should be forced to pay for use of knowledge. Roll back the clock a few years and tech sites would be almost unanimous about knowledge being free and unrestricted for the benefit of humanity. I’m keeping knowledge separate from actual direct rote duplication of content.
Now we have this amazing era where I can download models to my computer, run them locally, and have enormous amounts of derived knowledge at my fingertips for the cost of some compute cycles. Except now it’s a “crime against humanity”?
I don't know about "crimes against humanity", seems to diminish a bunch of actual crimes against humanity. And Microsoft is one to complain! Talk about a glass house. Just checking the annual report for 2025, the median employee compensation was 200k per year, but the total dividends divided by number of employees was 100k, meaning that each employee gets only about 66% of the value they've generated. Isn't that theft?
In any case, Microsoft has stolen 25 billion from its employees in 2025, and OpenAI has got 13 billion in revenue from "stolen" content in the same period, so that'd make them about equally bad villains, except OpenAI has mostly stolen from other companies.
People are getting lost in the weeds and ignoring the simple, fundamental fact that these companies are making fortunes off of labor that they didn't attempt to compensate the laborers for.
"Property" and "IP" discussions are distractions; no amount of it can rationally get us around the utter unfairness of what occurred and the way it will warp our economy at a basic level if not addressed.
It's not even enough to make the weights and models free; access should be free, and everyone who hitched their horse to this wagon should be on the hook for keeping the systems running, on their dollar. They took ownership of a venture that is short one (1) "Humanity's entire cultural corpus", and the only question is if we're going to issue a margin call.
first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft.
so, really, a lot will be coming out of it, we like it or not.
Well… Marx would say “theft of labor” is fundamental for the capitalist system to operate. The sin is when you “steal” from the capitalists. That’s why you can’t distill from a commercial model.
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.
We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).
I suppose it's theft in so far as you're deriving value from something that others labored for. I think that sticks in peoples throats. You can go to the library and do that already, gain knowledge, start a business or whatever. But the scale feels impersonal and monstrous in comparison.
The same point was made by many in the music industry about huge scale Internet music piracy vs people dubbing CDs onto mixtapes.
Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.
Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.
The original might no longer be there as the site might have already shut down due to AI scrapper bot overload. Or the original was a book Anthropic scanned and then shredded. Or the artist stopped publishing their works or doing art all together after all their creations were ingested into the model blob without their permission.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft.
The copy part was a recognized right, then taken away.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
My argument is if AI companies are ignoring copyright law and looking at all training data as commons, then we should look at LLM output as something that is not protected by copyright law.
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
I wanted to understand your background first because what you initially said is a very oversimplified view of LLM mechanics, and generally not quite the way how software is in practice written. Since you didn't answer that question, I will assume that you're not a SWE by a call. To give you an example of what I am trying to convey is: imagine a data-intensive workload hitting your storage/database/kernel implementation, and it's painfully slow, your customers are not happy. Then you as engineer sit down, spend days profiling and understanding the code, researching about existing algorithmic solutions to the same or similar issues found in the wild, you read some open-source implementations of viable approaches, you ditch some, some you take, you also read books, articles, other peoples experiences etc. And finally you end up, let's put it bluntly, with some sharded data structure by which you solve the bottleneck. It's not novel, the technique is so common and is already implemented across many many different products in slightly different flavors so I am wondering why do you think this is not a copyright breach but the LLM, which does more or less the same thing, is?
In my experience many execs know what they're talking about.
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
I think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed.
This sounds rather obvious, but I feel people forget this far too often.
Oh, I think MS execs (and all other execs, top-level politicians, and pundits for that matter) know what they're talking about, but they'll tell whatever lies they need to enrich and empower themselves.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
Under “conduct requirements” imposed by the CMA in June, UK websites are able to activate an opt-out to stop Google from scraping their content to power search features such as AI overviews - very similar conceptually to the news law passed in Australia.
I remember when this was the prevailing thinking online until about 2024. But that's when everyone was trying to justify their own piracy of GTA or whatever.
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
> That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.
I think the difference is that people are investing significant amounts of time, effort, and money into transforming their work into something useful, so naturally they would like some return on that investment. Giving away that work for free is not a particularly good business model if those people expect to be compensated for the value they create.
Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.
I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.
IP does not protect the little guy. This is nothing new. Self publish a book and then someone uploads a PDF of it? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Draw a picture and then people start putting it on t-shirts and posters without paying you? Better hope your real passion is filing takedown requests instead of making art.
AI training is the clearest example yet that companies are allowed to get away with what is treated as a serious crime only when an individual does it. There are many more examples of this, of course, but this one seems to be the most stark and obvious.
I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
> I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
Nope. You are making one or both of these mistakes. (1) Overlooking that the same logic can lead to different outcomes depending on the premises to which the logic is applied. (2) Overlooking that real life is analog, not digital, and so thinking the premises are the same when they are not.
The targeted outcome is the redistribution of wealth away from the capital class to the working class. Capitalist exploitation of labor was analog to begin with, and the same premise does apply: the capital class absorbs the fruits of labor, training, and education that is performed by the masses in order to enrich themselves. The capital class owes a tremendous debt to society and if they don't plan on paying we should plan to seize it.
The exploitation of labor has has a tremendous negative impact on the environment and the health of humans. Recall that it took dragging the factory bosses from their and beating them to death to get an 8 hour work day, a weekend, and restrictions on child labor.
We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.
Pretty sure we had to sanction/assassinate/goad into self-destructive military campaigns the Red Terror to beat it. And even then, the major survivor still beat us to cyberpunk dystopia (the cool one with hologram skyscrapers, not the uncool one with decaying suburbs).
There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.
The key problem is that IP is either proprietary to the creator or it is a commons type of situation.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
Taking the original of a painting from your house is stealing. Copying is only potentially violating the government-granted limited-time exclusivity that allows you to decide who can copy your work.
BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.
It's telling that you need to fall back to non-creative information in your argument.
Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.
Extrapolated further, this is a collapse of creative works as a profession.
Yikes. That’s some deep entitlement. Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.
If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet
> In the past, but today fewer people are getting paid less this way.
There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.
> Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
Uploading content online and getting a cut of ad revenue fits this criteria, no?
The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.
> There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.
And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.
IIUC, the question at hand is: does training require a special, separate, license or can you legally acquire a work and then use it for training?
I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.
> does training require a special, separate, license or can you legally acquire a work and then use it for training?
No. But it's not about current precedence or legality because the legal framework for accurately (according to general moral and societal acceptance) is decades behind where it needs to be. The courts will decide over the next few years.
> It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?
By your argument, I should be able to directly copy a book and sell it myself. The work was already done! What's more, they were fed by everyone's creativity, so of course I should be able to sell an exact copy.
In reality, the short-sighted greed is allowing widespread theft of intellectual property; do you think the number of writers would increase or decrease if there were no protections against content theft?
If you have such "immense gratitude", pay for the work.
Most likely increase, despite your intuition. This argument has been debunked so many times both intellectually and empirically. For one, read "Against Intellectual Property" by Stephan Kinsella.
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
That's what this boils down to: What are our rights?
Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.
If we take away that right, how would the Internet even work? It wouldn't.
Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.
If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!
People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.
My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.
But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
> People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
And profiting on it on a scale that’s hard to fathom. Someone who spent effort creating a great resource or doing some research and maybe got some income via donations, ads, whatever. Now that information from their resource is distilled into a big model. The original author is screwed, the model provider makes money through the effort of everyone else. It worked well for everyone before, because there was recognition, prestige, a sense of doing good for people, even a chance for some income. That’s completely eliminated with AI.
It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.
Indeed. If I take somebody's work, transform it somewhat and sell it, I should pay royalties (unless the author explicitly allowed me to do so). That is exactly the case.
Sadly, they're just following in the footsteps of every large media/publishing/music conglomerate that already screwed over the vast majority of artists/musicians/writers. For every Taylor Swift striking it rich, there's 999,999 who can't even pay their bills with what their copyright gets them.
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
570 comments
[ 0.20 ms ] story [ 85.7 ms ] threadIt is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
https://news.ycombinator.com/leaders
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Any argument that writers and artists lose from these existing, would remain unchanged.
Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.
Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.
I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
For instance a image/video generating model.
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
This is the most unintentionally hilarious misunderstanding of what the RIAA does, and the power relationship between artists and publishers I've read in years. In practice the RIAA exists to maintain the copyright monopoly of a few major labels. Rent seeking from the non-artist owned catalogues of the enormous majority of musicians who never 'recoup' their initial record deal.
> While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
It's far more seductive (since it's the default) to assume class relations don't exist, and wealth distribution is meritocratic. There may not be 'goodies and baddies', but there absolutely are rentiers and workers, billionaires and plebs.
All the major labels are public companies. Which means it's the very same people - the investment class, who claim ownership and extract wealth from say Warner and Open AI (should it make any money - obviously the whole house of cards could come down first).
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
1) AI companies all get sued out of existence.
2) AI companies can train networks with piracy, those networks don't fall under copyright, so give me a copy to do what I want with.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
Yeah, which is why it's only used to compare cars etc. among each other. Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.
Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
> A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
I was talking about cars vs horses. HP is Horse-power not human-power.
> Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
Cars are not free to roam the road. I don't know why it is so hard for you to understand that we use metrics like HP/CC etc to equalize with humans so that automobiles can be taxed just like humans. Without metrics like HP/CC etc there is no way to tax cars. Get it?
> And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
Come on you are clutching at straws here. It is not about "making sense". It is about using a metric to equalize unequal entities. When I am already saying they are unequal and have no direct relation to each other and any relation can only be arrived at indirectly. FLOPS is just an example I gave. I am not saying we should literally go with the FLOPS example itself. But we can use any metric and equalize it with human work. That's all I am getting it. It is the same argument as Horse-power.
> So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs, if the way they learn is fundamentally different?
Because that is the only metric we can use to measure how quickly GPUs can process arithmetic (you can label it "training" or "learning" or whatever name you want). There is no other metric that is deterministic and comparable to something humans do (which is also process arithmetic).
> Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?
It is fundamentally math. Every physical law in the Universe is expressed and backed by math. So on a fundamental level, yes "reading" is essentially maths only.
I went to the library. Didn't pay a cent.
Now what?
> Didn't pay a cent.
Taxpayers did pay on your behalf by funding the Library via the Government (if Library is public).
Nothing is free. Except ofcourse stealing, which is free.
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
That does not follow in any reasonable way.
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
I mean it's already happened, right?
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
There is no such thing as PROGRESS for progress' sake.
And FYI, agriculture is a wonderful invention.
Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.
Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.
I won't even go into climate change.
And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.
> "ripping off", "uprooting", "leeching (read: creating)" wealth
Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.
* * *
And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.
May the odds be ever in your favor.
We invent. We make progress. We make course correction. We end up in a better place.
Climate change? Lmao. We were told by 2000 20% of our country will be under water... It's 2026. Not under water yet.
What amount of pesky human intervention (positive or negative) will affect earth waking up from it's cold climate? Not luddites and their cow farts are global warming.
The real fight with "climate change" will be done with planetary scale technology derived from massive amounts of technical progress and energy production.
Your argument is as good as I should throw trash in the road or plastic in the sea. I'm just one individual. It doesn't matter. I'm not making the whole society do it!!
I fully agree with you. It is survival of the fittest in the face of progress. Competition is essential. That's how society grows. Humans prosper. May the odds be in your favor too.
Does more CO2/Greenhouse gasses etc raise temp? Yeah duh.
Does not using those bad refrigerant gasses is good? Yeah duh!
EVs are better for environment than ICE? Oh yeah!
But climate is not just a first order effect. It has more orders and complications than we can calculate. Our models are horribly incorrect, predictions further than a few years are always wrong. But that's science. We do our best and try to correct asap.
What's wrong is zealots who spreads doom and gloom about this. Makes predictions like my country will go under water by year 2000. Do you know how frightening it was for a science interested kid to read that? Kilimanjaro snow will be gone within decade. One extra decade gone by already, why is there still snow? 96 months to climate/ecosystem collapse? Double 96 months gone by, where's the collapse?
EU shot its own economy to pieces chasing these garbage. Wealth creation is the answer. Look at China. Didn't give one shit about it, built up wealth and power. Now they have the resources and they are rapidly cleaning up. In 50 years, China will be more green than US/EU while Germany killed the cheap nuclear power that was the backbone of their economy.
The REAL climate change is not tiny amount of co2 or whatever humans are doing; it is earth and sun. Orbital geometry, solar output, ice sheets, oceans, volcanoes and tectonics.
You can remove 100% of "environmental damage" from earth by humans from start of humanity, and it will not make a dent if/when earth/sun decides it is time to warm up. And we know for a fact that earth goes thru high temperature times and right now it is in an ice age.
What we need is even faster pace of generating wealth, power, and technology that helps us to clean up the air and water. Not be regressive out of premature fear.
We have social conventions regulating this.
Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.
You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…
Pirating is fine.
Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.
Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.
See you on the way down.
You are just bullshitting without any understanding of "people who have little". I'm one of those who came from "have little". As a consequence, almost everyone I knew growing up, were also "have little/have nothing" people. Among 100s of people in my super extended family (tree), among 100 of my "have little" classmates from school, it was always clear from early age who will "succeed." The ones who had grit, dedication, focus, and a shine in their eyes.
The ones who complained about wealth and society, are still in the have little group. The ones who didn't care, heads down worked hard, are living in luxury.
Taking care of one another is mostly a privilege afforded in society with wealth that leads to high trust situation. I grew up in a place without it; so I know how valuable (and desirable) "taking care of one another" is.
Current western generation (esp far left) has no idea how good they have it and are actively destroying wealth that will lead them to the chaos they have no idea about.
I don't want that "way down" just like you. But unlike you, I don't believe in fairy tales. Reality is capitalism works, wealth creation works, we all end up in better places (not equal...but overall better). Much better than road to hell is paved with good intentions.
It is shit compared to what it could have been. It is diverse in language (which is a negative), food (yet mostly boring) yet common in one thing: regulate to suffocate. Kill (most) innovation. Live in borrowed times. Spend money that we don't have. Fuck future generations even more. Then wonder why far right is flapping their wings.
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
JFC no, sit down.
You sound like one of those mentally unstable vegans who loves their avocado. How cute the animal has to be for them to care...?
It should be no surprise that it's a movement of lying thieves and scammers, Effective Altruism, that's behind close-source AI in the US.
Now I'm not sure justice won't come: SBF is behind bars for 25 years. He could turn out to not the be last scum from the EA movement to end behind bars.
As to open-weights models: at least it's not sold back at a mark-up and anyone can run them.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
lol at this edgy 5th grade statement. So ridiculous.
It is this one line, Article 1 Section 8 Clause 8, that separated the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Right clause a communist revolution, Violent or not. That is the one thing the communists would change to make the United States a communist country.
[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
Didn’t they themselves say they were a socialist society on the path to communism?
Does anyone think the USSR was communist?
The clause is what ensures profits from market sales or licensing of ideas go to the creator.
Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.
It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.
In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you had bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine.
The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.
Incorrect. You overlooked consideration.
Since this is a copyright fight, rights extend only to verbatim copies of the full work or significant portions thereof. Abstract things like facts, ideas, concepts, themes, and patterns are explicitly not protected, and rightfully so. Yet those abstract things are what get repeated and distributed, and are what get encoded into model weights.
This is probably plaintiffs' biggest challenge because it has been very hard to get models to regurgitate entire works except for a very small handful of extremely popular works (and now there are guardrails against even that.)
The 'sell it back to us' argument falls short in my view.
Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.
The comment here seems incredibly pessimistic and quite dramatical.
There will be new jobs coming from this.
The bountiful abundance of intelligence is truly the best thing that has happened this decade.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
Paid influencers who perpetuate the open narrative are a whole new industry.
Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".
In terms of purely local LLMs, one can run GLM 5.3 flash on a beefy workstation.
[1] https://openrouter.ai/z-ai/glm-5.3
Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?
I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.
Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.
I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
Don’t believe me? Try it. I have.
Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.
Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.
There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.
It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.
As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.
Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.
Gave me a chuckle. I think this is my personal ‘best sentence of the year’.
Great way to summarize how we will likely look back on 2026.
People can't help themselves when they don't believe they're being exploited.
The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.
If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.
Industrialization led to iPhone, PlayStation, Spotify, and Waymo.
AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...
AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.
[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.
[2] Until the robots leave earth, anyway
Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.
Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.
You need to show how having X more data centers is somehow going to translate into affordable, highly advanced robots in the immediate future.
There are hard problems about robotics we don't yet know how to solve. Not to mention, who the fuck is going to buy them if AI takes their jobs?
Do you think people will stop dating, trying to impress mates, buying luxury, etc.? Neither candlelight dinner to a robot no a Dior manufactured by robots have the same allure. There's a huge industry around this.
Sports aren't going to go away. Huge industry.
People aren't going to stop making art. I know a ton of artists who have embraced AI that are doing even bolder work using the tools. (I was a filmmaker pre-AI, and I know a lot of people in this field.)
People aren't going to stop traveling. And consuming. And eating human food and consuming human experiences.
There are going to be all new kinds of businesses and opportunities that spring up. OpenAI and Anthropic are not going to be the only two employees. They won't be staffed by only agents.
It may be the case that the potential output of a worker engaging in the productive process goes arbitrarily high. But the that won't matter, if they're not permitted to. Given a choice between involving a human worker who will demand compensation, and a fully general robot, which will the owner class choose?
The vast majority of human intelligence is already squandered: millions of potential geniuses in impoverished places, suppressed by lack of opportunity to flourish. It's not about the quality or quantity of intelligence. It's about who controls it.
The fruits of the industrial revolution didn't end up in the hands of the worker by divine grace, or by some natural law. They were won by the hard struggles of the labour movements, by leveraging their indispensibility to the process of production.
If we want the utopia you imagine to be accessible to ordinary people, workers must cease being so eager to build their replacements, and be prepared to collectively struggle for their share. But if we are lulled to complacency by the notion that this will be a passive process, that struggle will not be necessary, then prospects are grim.
Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.
On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.
You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?
The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.
Does this mean that all countries will have to force labs to create subsidiaries within their borders, that are then taxed?
Also, this isn’t foreclosing a past that didnt exist. What has occurred is absurdist levels of theft, in a political and business environment that doesn’t have to worry about the wronged individuals being able to have their grievances heard.
Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?
Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.
Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.
Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.
The more fundamental objection seems to just to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.
What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.
Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.
But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \
It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.
I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.
This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...
We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.
Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
> getting things were free was part of sending a message, not monetary gain.
As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.
It was a righteous moral stand because we wanted to get things relatively affordable for us.
I for one prefer to buy my software and music nowadays, because it’s affordable and I can get it at the quality I want.
For your other question, the answer is probably 1800s, because the current model was not entrenched everywhere and creators had sane rights for what they created. So creators had to get what they made out to masses to show it, but lost their rights in relatively shorter times, so things were free to use for everyone. So you can’t excessively milk something till proverbial eternity.
At the time I pirated a lot of stuff. If we're of the same age, you may well have played video games I cracked in your teens. And, TBH, for me it was mostly collector mentality (I have have all the games!) and a bit of poverty (I can have the games I'd like to buy but can't afford), and about zero politics.
That said, I'm generally with you on 1800's. It's astounding how fast things changed. IIRC it wasn't until 1890 or so that international copyright was even a thing; you published in your home country and publishers in other countries just copied and published without permission or payment.
And then just 100 years later, copyright was essentially perpetual and global.
therefore you have to work for me for free
However if I put a gun to their head and demand to live in their house or give me food for free they get mad.
Its lovely how everyone deserves to get paid except writers / musicians cause we are only fucking hippies when it comes to that.
I don't know what to think about that though.
Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
Now what do I do next time someone comes and asks how to do something?
https://en.wikipedia.org/wiki/The_Dog_in_the_Manger
>Learning isn’t stealing.
This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid.
No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.
Learning isn’t copyright infringement.
Learning isn’t theft.
Copyright infringement isn’t theft.
These are all different statements, and all are true.
> There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.
“Gaining exposure to” isn’t theft either.
Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.
There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.
If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.
If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.
If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.
If people stop making art because they can't outcompete generated art, that is also lost to LLMs.
So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.
And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.
The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.
Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.
Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.
And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:
Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.
To use another Slashdot cliche: "But you knew that already."
Citation needed
People need to face reality. The old Internet was fragile and couldn't have lasted long. It was dying slowly due to closed social media and LLMs are accelerating that death.
I'd rather it die quickly and be replaced by something more robust than decay slowly. I was sick of how it was before LLMs even came among.
This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.
Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.
They work very hard to protect their own IP so they obviously believe that IP is worth protecting.
So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.
It hasn't been literally stolen but indirectly it has been.
I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.
Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.
AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
It is unclear to me just what percentage of tech-companies, are in the business of adding extra steps to an illegal process, of turning signal into noise. For example, the vapes made by Juul are a kind of hack around public health laws. The Amazon marketplace shields merchants who sell counterfeit goods. Uber and Lyft bypassed the local laws that applied to taxi services and the medallion system. Airbnb did something similar with the laws governing hotels, and delineating who owns and who rents.
It is a special kind of disappointment, to be someone who loves technology, and to have to work in the technology industry -- because it appears to be run by people who hate everything that is not money.
I know this is a real harm to individuals, and bringing this up is not some kind of anti-technological thinking (the luddites were OG with that, and I always argue in favor of them - they had a very valid point).
I also realize that in a few years, I very well may be in the same position here. So will most of us.
But that's a different argument than GP was making, different one from what I replied to. That was about losing culture, and losing a business model is not losing culture.
Also I was with you all the way until the last paragraph:
> AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet.
AI is doing exactly what users want it to - what I too use it for - it bypasses the spam and scams that stand between the user and the solution to their problem. Good riddance, in this.
> Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
And that is just bullshit. AI is not "taking everyone's content and selling it back to users", and the vendors are definitely not "keeping all of the profits" - on the contrary, they haven't even begun to figure out how to monetize almost any of the value their inference services provide, because they have zero visibility into how much any stream of tokens ends up being worth for their customers. They cannot tell whether the plumbing advice the LLM generated unclogged my toilet, or prevent a restaurant from having to close for the day; they cannot tell if the code their agent wrote won me a beer from a friend, or unblocked $2 000 000 dollar opportunity for my business. They capture none of the value from this either way.
No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.
I can't really tell if you are really evil and messed up inside or just stupid and can't reason.
If IP is real, then the AI companies have performed flagrant theft.
If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.
The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.
Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.
That's not to say that IP is useless, but treating IP infringement as "stealing" was always a problematic shorthand. Previously, calling IP infingement "stealing" was the domain of corporate interest groups like the RIAA or MPAA, but with AI the winds have turned and supposed anti-corporate lefties are keen to treat IP infringement as theft to attack AI companies. There's very little intellectual consistency on either side of aisle here.
When an AI company takes that type of value, and doesnt provide value back to the person who made it, thats theft.
Information / facts are real and free. But those who helped us get here should choose how to license their work, and an AI company picking it up and re-selling it is clearly unethical.
I can reasonably wrap my head around the idea that an author should be compensated for their work, but where in the social contract does this power to control how that work is used come from?
It's so destructive and tangles up courts and makes contracts complicated and we lose the original versions of the work because e.g., they have to change a background song due to complicated licensing.
I can respect protecting copy rights, but it should never be conditional; if you choose to make a work available to the world, you have a legal right to defend the unauthorized copy of it but you should not have a right to say how it gets used.
It reminds me of manufacturers like John Deer.
If you are freely posting sentences like this was on the internet you are giving away your IP for free. Anyone can read the sentence. If someone can make money off of it then that’s just markets at work. I can’t make money off of what I write here, for example.
But I also don’t think pirating a movie is theft either. You haven’t proven to me you’ve lost money. Maybe wouldn’t have watched it anyway.
Fun topic
Hell if a car was downloadable why would I not ? Your capability to make endless profit should be capped somewhere, if indeed it is Profit
It’s called Derivative Work and it’s a good feature for IP law: https://en.wikipedia.org/wiki/Derivative_work
You wouldn’t like a world where companies could copyright knowledge and then prevent anyone else from making a derivative of that knowledge.
That would be on what the AI model generates, not what it is trained on.
The latter is where the contention is, and it's a valid argument. So much so that some companies are not using stolen information to build their models.
IBM for example indemnifies its models for its customers and has detailed information on where the sources came from to train them.
It has been tried in court several ways already. Remember the lawsuit that forced Anthropic to use physical books? They tried to argue that the books couldn’t be trained on at all. It failed.
Arguably LLM companies could have made large-scale deals with libraries and got the exact same knowledge (much, much more slowly). I wonder if people would have the same issues then? My guess is probably. Goes back to the meme that if libraries were proposed today there's no way they would ever be allowed.
Your analogy is flawed because you are saying the information was free to begin with.
That analogy works for IBM granite models because the information they trained on is free to use.
The major LLMs did not do that.
People tend to forget that IP was not originally about digital distribution at all (copyright), it was about giving inventors exclusivity periods to profit without competition (patents).
It was a misguided attempt to stop sometimes literal theft of designs from rival inventors, by tying the design to the person instead of whoever possessed the schematic.
It's also a regime that in its modern incarnation protects businesses, not artists.
AI scraping for the goal of making a commercial LLM service is different from, say, a commercial file sharing platform.
The first difference is that LLM training is highly transformative. Let's say your LLM ingests the Harry Potter novels during its training. What you get at the other end is not the Harry Potter novels, it is a LLM that can talk to you about Harry Potter, it is not the same thing, and going from one to the other requires a significant amount of work, very expensive work in this case.
Not only that but there is no direct competition. People won't stop buying the Harry Potter novels because a LLM trained on it exists. If you want to read the books, you buy the books, you don't ask a LLM about it. A file sharing service on the other hand competes directly against the official channels, if you want to read the books, you can download it from this service instead of buying it on the official channels.
So, about how free you should be to get these weights from the AI companies. If you just share a 1:1 copy of the weights, that's the "file sharing" situation, not transformative, you took their work, didn't do any of your own. Usually considered unacceptable by IP laws.
Distillation is a more interesting case, you are using a LLM to train your own, it is transformative work, but you may also be competing directly against the LLM you are distilling. So, in a sense it is worse than scraping, but still, despite how much the likes of OpenAI and Anthropic are complaining, it seems to be legal.
So it is somewhat consistent: 1:1 copy and distribution is not allowed, be it source material or LLM weights, and training is, be it source material or another LLM (as in distillation).
This is how IP has always worked. It has never protected the little guy. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.
Nothing more and nothing less, and all of the normal debates about what should remain Capital and what should be The Commons apply.
Nothing short of a global revolution can fix this. Proprietary LLMs should be illegal. Either we achieve post scarcity within this generation or it's literally over.
Considering its become exceptionally difficult to search for things that used to be easy to find, I have to disagree with you there.
The web has been polluted with trash, covering up all the original media with messy imitations.
I doubt it, those are the same group of people anyway.
Orders of magnitude more people are using LLMs to brainstorm or generate content for them, which they are then using to solve their own problems and carry on with their life, increasingly solving many more problems for themselves this way, and not publishing any of it.
The world isn't made of "content creators" flinging crap around in hopes of monetizing it. Most people have actual jobs.
I guarantee it's scraped our work.
So where's our paychecks.
Not the least because I'm already getting many orders of magnitude more value each day from them providing me inference as a service.
Many many people didn't give permission for the data to be ingested and used this way.
We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.
All of these lag SOTA by ~6 months on average, so yes, indeed, the thing you could only pay for half a year ago, is now free for entire world forever, and it only keeps getting better with time.
https://en.wikipedia.org/wiki/Cultural_evolution
https://en.wikipedia.org/wiki/Memetics
The past is still there, but our culture, being a live thing, is pretty much being affected by AI. We now have an artificial intelligence in this loop of exchange of ideas. Slopifying everything in its way.
It's a robbery of our future culture, at least.
That is a different argument entirely, though. Not the one GP presented.
On this, I have no concrete views. Sloppifying is real, and AI is short-circuiting "this loop of exchange of ideas", for sure. On the other hand, we had several precedents on this within last 700 years - the printing press, the radio, and the Internet to name the largest three - and culture turned out fine. Different, but fine. AI feels more than this, though, so I'm not putting that much stock in argument from history here.
But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of denying it.
[1] Fuck you, Claude, for making me wince at that.
It isn't you that AI communism is stealing from.
This is good, but will not fly in court. Your honor... I only moved the Picasso...it is still there...
It's there in much diminished market, one that has been saturated overnight on a scale previously unfathomable. Their art is still there but it has been used against them to obliterate the playing field.
> nor are the people involved selling it back in any form.
Aren't all of these AI companies selling subscription models for people to create derivative art?
> And let's not forget what we got back for this
Get back for what, exactly? Didn't you just argue that no theft or sellback has occurred?
It used to be that, no matter what you want to do, there's a video on youtube of some guy who knows what he's doing showing you exactly how to do it. It could be, like, compressing a rear disc brake cylinder on a car or matching fiberglass gel coat colors, or carving a statue out of marble.
Theoretically those videos are still in there somewhere, all 15 years old at this point, but you'll not find them with youtube search instead you'll find a million AI slop videos no matter what you query.
There's a way that many people fear this is true: what you're getting for "free" has some external cost that you aren't taking into account.
Cost 1. The environment. Tech companies were once significantly interested in efficiency, their carbon footprint, etc. With the advent of LLMs, this was thrown out the window. Concerned people are now thinking about water footprint, heat footprint, noise footprint, and probably more. These things are hard to put a dollar figure on in a short comment, but the sustanability of a liveable climate for the billions in this planet is literally priceless.
Cost 2. Employment. We face a significant cull of employability -- graphic artists, programmers, mathematicians, paralegals, and more are rightly fearful for the end of their career. If we're getting intelligence for "free" in 2026 dollars, but we can't get jobs in 2030, will all that content still be affordable in 2035?
Cost 3. Infinite investor dollars. The AI companies are burning cash at a historically unprecedented rate. This has attracted a lot of attention from traditional investors, like pensions, banks, etc. If all of this ends up in a free product, that doesn't sound like it'll return those investments. And that can crash the economy -- I ask again about real affordability in 2035.
This has another side effect in that it breaks the power that consumers have in the market to have a say in how resources get allocated, by voting with their wallet.
The entire AI buildout is non-consensual. Individuals have zero say. We could all refuse to buy AI subscriptions and it would not matter because businesses will still buy, and, investors have decided that we are moving forward with this no matter what, seemingly whether anyone is actually buying or not.
AI isn't necessarily unique here, but it is one of the biggest new examples of wealth inequality and how society at large are no longer the ones who get to decide how resources are allocated under a capitalist system, especially when the investor dollars behind it amounts to the GDP of a small nation.
And this is fuel for climate change denial. See? Are the wealthy and powerful people acting like they are concerned about the climate? No. They are doubling down on energy use and consumption. It's all a fraud!
1. is mostly bullshit that's used by anti-AI crowd to gather further support. Data center companies did not stop caring about footprint, they're still pursuing that for the same reason they did it before: it aligns with overall lowering of costs. AI added an extra incentive to improve efficiency because the demand far outstrips supply of compute.
For 1., also the general argument stands: yes, these data centers use energy, just like everything else humans do. What matters is the value it provides to people, and in case of AI, it's one of the most objectively useful expenditure of watts on compute (your point 2. notwithstanding).
I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
> Your point about efficiency is particularly obtuse: efficiency doesn't mean burning less fuel, and getting more compute per Joule. Not reducing the overall carbon footprint, but accelerating the burn rate as fast as we can build.
Efficiency does mean getting more compute per Joule. They are doing that. They are also building out as fast as they can, because both of those address the same problem: demand for compute outstripping supply for it. They can't just pursue the build-out strategy, because hardware production is now supply-constrained too - that's where the "RAMpocalypse" came from.
Yes, this is a huge build-out, and little to none of that is 100% carbon-neutral, so ecological footprint adds up. But that's normal and expected and a kind of tradeoff humanity has been making forever: building new stuff, be it hospitals or airports or data centers, has an environmental footprint, and we only hope what we get from it is more important to us on the margin, and that we can offset the environmental costs some other way.
> I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
I'm not talking about financial "demand". I'm talking about real demand, for compute. There's nothing to doubt here, it's pretty clear that all major AI players are constantly running at capacity - it's obvious in the very structure and limits they place on even the highest tiers. Unless you believe half the data centers are just spinning busy-loops and burning energy to inflate stock prices and generate a fake reason to build out more compute capacity - if yes, then I don't know what to tell you.
--
[0] - Like the water story, where the reported datacenter usage numbers look big in isolation, but once you relate them to "how much water is there" and "how much is used by other industries", it turns out to be a nothingburger. The only numbers that survive are those showing the story is really about those other industries having a competitor for cheap water now, and having to pay a bit more than they're used to.
Water scarcity is a problem around the world; even my rainy hometown of Vancouver had steep water restrictions this year. The question "how much water is there" is incredibly disingenuous, to the point of feeling like a motte and bailey: yes, the universe, or the earth contains vast oceans of the stuff but datacenters are using drinking water which is quite scarce. Just because "other industries" use a lot of water doesn't make it okay -- I level exactly the same criticism at other industries which waste water at such an egregious degree.
To be clear: if data centers only charged up a pool of water to use as coolant, that's fine. But when the weather is warm, they run millions of gallons of tapwater down the drain just to cool off. And heat spikes are happening with increasing frequency, severity and duration.
We as a species need to conserve water and emit less CO₂, full stop. The tech industry doesn't get a pass just because other industries are dirty.
Again with the weasel words. Real people, including individuals? Pretty clear this is written from a pro-business perspective and can be safely dismissed.
>"robbery of all of our culture to sell it back to us at a mark-up".
Well, it is that. Why is it criminal for me to steal a big-business-movie for personal use, but they can just take all my blog posts and sell their deritives to others?
I respectfully request that you get all the way outta here with this take.
Artists have caught AI generating literally their own work, for free, at scale. There have been tons of articles on this website about the AI hyperscalers slurping up books, copyrighted works, etc through legal and questionable ways. Even TFA says that the NYT suffered absolutely devastating CTR drops.
If you create a blog post about something super esoteric, it will guaranteed end up in the training sets for every big-lab frontier model within 24 hours (probably less) and probably show up in Google's AI Summaries at around the same time. No audience for you!
You'll never agree but I think you underestimate how many people consider all that the AI companies have done to be the wealthy stealing and selling things back.
> nor are the people involved selling it back in any form
I am convinced that the longer one works in AI the less one has any grasp on reality
> too cheap to meter
Its the most expensive buildout in human history - you can't just split half the cost. The spending is the only significant growth in the US economy. Unless you think that all the datacenters and all that capex are for training?
However ordering a car with Uber through your phone just gets people to the destination faster, cheaper, more comfortably and reliably using those same roads.
This is theft I made up the route first >:(
It's actually not there. See how many websites have now closed doors or ceased to operate because of constantly being hammered and bombarded by robotic scrapers. For others, it made websites that turned a small profit into unsustainable money pits.
Sure, their content might now be ingested into an LLM training set sitting somewhere on proprietary servers. But the site itself (the origin of truth) now does not exist.
So no. It actually isn't there. Resting on this falsehood, the rest of your retort makes not much sense.
“On a chip” is also stretching the truth. The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them. Individuals who have free usage for a project commonly report that, had they paid, it would cost five or six figures. Companies have seen enormous bills, some approaching the size of their payroll. And all of this is true for tokens that are dramatically subsidized, by one of the most intense and largest concentrations of capital in history. It is the polar opposite - “almost too expensive to even do, and absolutely must be carefully metered”.
In summary, let us indeed not forget what we got back for this: extraordinary distortions of reality evenly intermingled with bald-faced lies.
It's not complete or that well-rounded. But it's something that was the domain of speculative science fiction only 5 years ago, and it's rounded enough to be applicable to ~everything to some degree.
> The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
By "on a chip" I meant more "in silica" than literally on a single chip" - though this actually is* true, but those chips aren't cheap.
> And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them.
You are looking at power users that use LLMs in agentic coding sessions. Most people just run off free tier of ChatGPT, which is free for them. There are equivalent open-weight models at this level, and while hardware to run one for yourself is expensive even for most westerners, the marginal inference cost is literally dirt cheap, which is why you can get that for near-free from smaller inference providers - or pony up some money, rent a bunch of compute with friends, and become an inference provider yourself.
(It's only a tough market because the major vendors are giving out better models than you can run for ~same or lower price than you can offer. Which either way is too cheap to meter in terms of solving useful problem for real people. Again, developers are a special case of power users, as usual.)
Not problems, laziness. The way I see students and colleagues use it is to get their work done with less effort. That's its selling point. There are not many real problems LLMs address.
> no one has actually been robbed
In our society, people get paid for work. If you think society is wrong, fine, but you must state so first. Under common assumptions, all that data has been produced through work, and that work represents value. Taking it for free is therefore theft.
I think this is quite the pollyannaish perspective and very much inline with those that think if you can take, then take and only apologize when caught.
We won't get anywhere in these discussion if good chunk of participants cannot admit to the trivially observable facts about the actual objective reality in which they live in.
The kind of "piracy" you're talking about wasn't depriving the authors of anything because you could always make the argument you weren't going to pay for it anyway. If, on the other hand, you were making copies and charging people for them, you could definitely say you were depriving the legitimate authors of that revenue. AI companies are very much doing the latter, not the former.
The other part of it is it's not just copying. Previously, if I decided to make a copy of a work without paying, I'm only copying the work, not the author's whole writing style. Now the AI companies are depriving authors of revenue from works they haven't even made yet.
If they were able to create a model de novo then they could truly claim it hasn't just been lifted from existing culture.
sell it back to us as markdown
The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
> The markup is the millions in training they committed and the connecting the knowledge
Sure, but gated behind a hallucinating idiot.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown, 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow
[2] https://sydneyuniversitypress.com/products/9781920898168
[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons
Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.
^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.
No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" can extend beyond collecting facts. Originality and creativity are also work.
However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.
[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...
And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.
There are ONLY TWO Stories and their details, that we collectively will never see.
1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...
And the other...
2) The Epstein Files.
AI is better at repackaging it back to the end user but ultimately I'm arguing it's the same thing.
(caveat: yes I know there were plenty of people that objected to Google et al indexing everything; famously, Gmail was scary to a lot of people because they read your email to give you ads)
At a mark-up would mean it’s more expensive. The outrage is that they’re taking knowledge that was expensive to access because you had to hire experts or otherwise pay a lot of money for it and making it accessible to anyone who signs up for the ChatGPT free tier.
Calling it “robbery” is also specious as no knowledge was taken away from anyone. The content in the training sets was out there in the world one way or another. It still is!
I’m really perplexed by this sudden swing toward the idea that knowledge is something that we should encourage or incentivize to keep locked away or that other people should be forced to pay for use of knowledge. Roll back the clock a few years and tech sites would be almost unanimous about knowledge being free and unrestricted for the benefit of humanity. I’m keeping knowledge separate from actual direct rote duplication of content.
Now we have this amazing era where I can download models to my computer, run them locally, and have enormous amounts of derived knowledge at my fingertips for the cost of some compute cycles. Except now it’s a “crime against humanity”?
In any case, Microsoft has stolen 25 billion from its employees in 2025, and OpenAI has got 13 billion in revenue from "stolen" content in the same period, so that'd make them about equally bad villains, except OpenAI has mostly stolen from other companies.
The web getting flooded with slop and drowning all original work is one way to go about that.
Another is paywalling and gatekeeping en masse.
A third is shutting down shadow libraries.
"Property" and "IP" discussions are distractions; no amount of it can rationally get us around the utter unfairness of what occurred and the way it will warp our economy at a basic level if not addressed.
It's not even enough to make the weights and models free; access should be free, and everyone who hitched their horse to this wagon should be on the hook for keeping the systems running, on their dollar. They took ownership of a venture that is short one (1) "Humanity's entire cultural corpus", and the only question is if we're going to issue a margin call.
first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft.
so, really, a lot will be coming out of it, we like it or not.
Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.
Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.
The copy part was a recognized right, then taken away.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
https://arxiv.org/html/2408.02487v3
I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.
https://en.wikipedia.org/wiki/Shared_Source_Initiative
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
https://www.theguardian.com/technology/2026/may/23/trump-ai-...
https://www.bbc.com/news/articles/c98r8r7dz5no
https://en.wikipedia.org/wiki/Commons
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
This sounds rather obvious, but I feel people forget this far too often.
"Never attribute to stupidity that which can be adequately explained by systemic incentives promoting malice."
Previously: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Honest question. There is a line in the sand somewhere apparently.
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
https://www.bbc.co.uk/news/world-australia-56163550
Suddenly, they don't like other people pirating.
For example, for software I usually use MIT license, which is a permissive license with attribution.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Look at who owns that site, no surprise here.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
Just spending money doesn’t mean it’s legal, for example. Criminals expect RoI too.
"Dial-a-Victim" / "Victims R Us" - startup founders of the future. Invincible in the face of legal challenges.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.
The profits and income is earned in America. The idea that America would pay manga artists whose work was copied is … beyond idealistic.
Most AI companies are not sharing it, though. They appropriated it and resell it.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.
The thing I produce does not replace demand for the original though?
BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.
Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.
Extrapolated further, this is a collapse of creative works as a profession.
What about the rest of that quote?
If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet
Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.
> Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
Uploading content online and getting a cut of ad revenue fits this criteria, no?
The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.
And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.
I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.
No. But it's not about current precedence or legality because the legal framework for accurately (according to general moral and societal acceptance) is decades behind where it needs to be. The courts will decide over the next few years.
If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?
In reality, the short-sighted greed is allowing widespread theft of intellectual property; do you think the number of writers would increase or decrease if there were no protections against content theft?
If you have such "immense gratitude", pay for the work.
Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.
If we take away that right, how would the Internet even work? It wouldn't.
Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.
If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!
People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.
My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.
But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
And profiting on it on a scale that’s hard to fathom. Someone who spent effort creating a great resource or doing some research and maybe got some income via donations, ads, whatever. Now that information from their resource is distilled into a big model. The original author is screwed, the model provider makes money through the effort of everyone else. It worked well for everyone before, because there was recognition, prestige, a sense of doing good for people, even a chance for some income. That’s completely eliminated with AI.
At least it’s consistent, is what I’m saying.
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.