It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free.
I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves
No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
Well, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post)
> companies spent a long time telling us downloading single songs via Napster was the worst thing ever
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owns anything they invent or create. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of 15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
In a communist society there is no profit (or incentive for), thus no need for copyright laws.
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
Example:
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.
That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were either not doing much to begin with or you were probably part of that bullshit.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft.
The copy part was a recognized right, then taken away.
In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).
It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
It is the largest democratization of knowledge that ever happened.
The 'sell it back to us' argument falls short in my view.
Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.
The comment here seems incredibly pessimistic and quite dramatical.
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves
really not the same entities here
No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
Yeah that's more fair
Well, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post)
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
Not at leaf level, but if you trace the trunk, pretty sure you end up on the same one.
> companies spent a long time telling us downloading single songs via Napster was the worst thing ever
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?
However, there is irony in a subscription to pirated material.
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
Came here to post this, got beaten by someone putting it far more succinctly than I would have.
I mean if you cite a copyrighted book verbatim. You are held liable. So should a company producing copyrighted work.
For instance a image/video generating model.
Regulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would.
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
Copy one book, and you're a thief. Copy thousands, and you're a VC.
And yet a third bucket is the license-laundering of GPL code when the entire github corpus was vacuumed up.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
The what thread, now? Got a link instead?
https://news.ycombinator.com/item?id=49737787
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
I am unclear how this would help anyone?
Any argument that writers and artists lose from these existing, would remain unchanged.
Regulation can mean all sorts of things, including declaring the models themselves illegal.
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
Scale matters.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owns anything they invent or create. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of 15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
In a communist society there is no profit (or incentive for), thus no need for copyright laws.
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
At least the Chinese AI companies are doing good service open-sourcing their models back to the public.
There are no open-source LLMs.
OLMo
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
What is DD supposed to mean?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
Due diligence — pretty standard acronym in this community.
No it's not, and I've been here more than a decade.
Copying data isn't a crime.
Agreed, but now they're trying to stop other people from copying data so that they can be the sole gatekeepers of humanities collective knowledge.
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
> sell it back to us at a mark-up.
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were either not doing much to begin with or you were probably part of that bullshit.
Defending the same entities getting large DoD contracts to use AI for killing?
They've been using computers for killing for decades, who's taking up pitchforks against computers?
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
“It is said that at the heart of every great fortune there is a great crime”
lol at this edgy 5th grade statement. So ridiculous.
“Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
Unpopular opinion on here
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
In case of programming.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft.
The copy part was a recognized right, then taken away.
Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
But it's copying...how is it theft? Your labour WASNT stolen was it?
Well, given these companies are trying to sell it back to you - seems like even worse than stealing. ;-)
In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?
Never ended, just changed in form.
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.
At least it’s consistent, is what I’m saying.
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
Yeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did!
If this was all open, I’d maybe half agree.
AI overall is the ultimate piracy crime.
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
Then let's make humans pay royalties to every author from whom they ever learned something, even if it was offered freely to them, only fair?
I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
they did say theft of labor, not theft of the things being trained on
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).
It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.
Lots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.
Copying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone?
Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
Probably indicative of America's wider downfall that they've all become so self interested
Hackers are irrationally anti-corporation. This is where the nonsensical AGPL came from, too.
The AGPL didn't go far enough because it didn't limit freedom 0 to natural born humans. In the 00s/10s it was corporations. Today it's AI.
If it has no soul to save and no body to torture it deserves no rights.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
How is this different from Microsoft scraping to build Bing?
Honest question. There is a line in the sand somewhere apparently.
Bing sent traffic to the original source. AI answers don't. That's the line. It's not complicated.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
Huge difference between building AI and a search index.
Why? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn't make.
Bing isn’t re-selling you back the content it took.
All the "LOL you wouldn't steal a car???" posts in this thread miss the point entirely.
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
Spiderman pointing
And here we are, just watching and doing nothing..
Copying isn't stealing you babies