926 points by pluc 2 days ago | 817 comments | View on ycombinator
haritha-j 1 day ago |
47282847 2 days ago |
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
sajithdilshan 2 days ago |
juvvel 2 days ago |
TutleCpt 2 days ago |
Kuyawa 1 day ago |
So no, your cries for regulating others because you are losing the race won't work this time.
tom2026hn 2 days ago |
Waterluvian 1 day ago |
leonidasrup 2 days ago |
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
ohrus 1 day ago |
The hypocrisy of this new world is already catching up to us.
heaney-555 1 day ago |
Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.
It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.
thunkshift1 1 day ago |
iamflimflam1 2 days ago |
But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.
cmiles8 1 day ago |
American87 2 days ago |
Weryj 2 days ago |
nullbio 2 days ago |
sebastiangrill 2 days ago |
IX-103 1 day ago |
I find it a little hard to be upset about AI. Supposedly stealing copyrighted works when the vast majority of those works. Probably should have been in the public domain to begin with. I have a faint hope that this scuffle between the AI companies and the publishing industry will result in more reasonable copyright laws, but I think it's more likely that exceptions will be made and AI will be treated as a special case.
pianoben 1 day ago |
hereme888 1 day ago |
jaybeavers 2 days ago |
markhahn 1 day ago |
Obviously, reproducing works in whole is infringement. That's not what AI is doing, so the question becomes: how is scraping different from ordinary reading? Is it just that site owners want to play back history and retroactively create high-cost licenses for scraping?
__bjoernd about 19 hours ago |
bentt 1 day ago |
juiceland 2 days ago |
levischoen 1 day ago |
GardenLetter27 2 days ago |
totetsu 2 days ago |
montjoy 1 day ago |
jacquesm 2 days ago |
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
keeda 1 day ago |
The laws at play here are related to Intellectual Property, specifically Copyright. Yes, it is terribly flawed, but it is the product of centuries of case law dealing with very hairy issues, and I believe it is fundamentally sound, and here's why.
As the name implies, it deals with only verbatim copies of works or subsantial portions thereof. It very expressly does not cover abstract things like concepts, ideas, themes, facts, or patterns, and rightfully so, because we really do not want anyone owning something that broad.
But these abstract things are precisely what have been extracted, at unimaginable scale, to build these models! Each pattern in the tokens derived from these works contributed imperceptibly tiny perturbations to randomly initialized weights, interacting in incomprehensible ways into vectors representing concepts and ideas and facts, the cumulative aggregate of which has somehow created a form of intelligence.
There is no copying, only gleaning, and so Copyright Law falls short. But what is the alternative, and do we want it?
To prevent something like this would require some sort of legal protection on the more abstract things. We do have a legal framework for those: Patents! But as is very clear on HN and in many Tech circles, those are an extremely contentious topic (even though they actually protect much narrower ideas than most presume.) I don't think anybody anywhere really wants any protection on broader abstractions, and rightfully so.
So: we as a society expressly decided these abstract things belong to the commons, and those are the exact things these labs harvested. This is probably the only logical culmination of our technological journey, and is within the very reasonable legal frameworks we have evolved over centuries.
As such, it is not productive to dwell on fighting this or bemoaning this. Instead we should focus on ensuring that this technology -- with its immense potential and opportunities and dangers -- benefits everybody as much as possible. That is a better way to compensate everybody's labor, and that is a much richer and fruitful discussion to be had.
gyosko 2 days ago |
alansaber 2 days ago |
rdsubhas 2 days ago |
These are two different things.
Note: am not an AI fanatic.
rietta 1 day ago |
undefined 1 day ago |
1vuio0pswjnm7 1 day ago |
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
p. 1
"This case is about, as Microsoft's Director of Applied Science put it, an astonishing theft of unprecedented proportions; SF1437, perhaps the largest theft of labor in human history. SF1652"
p.11
"As Microsoft recognized: millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions and admitted that almost no one intended for content they created to be used in this fashion, nor are they compensated for its use. SF1437."
p. 74
"As Microsoft's Dr. Glen Weyl put it, compensating creators is in the best interests of my employer, of my country, and of many other groups I belong to. SF1657."
Hyperbolic quotes from Microsoft employees are, IMO, the least interesting elements of this brief
Here is Microsoft's brief. Note how MSFT responds to the "web grounding" claims
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
It seems OpenAI does not want the public to know about (a) OpenAI's data collection and retention practices and (b) the number ChatGPT users have requested deletion of conversations
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
"OpenAI seeks to redact specific information about [(a)] the number of users who requested deletion of ChatGPT conversations and [(b)] OpenAI's related data collection and retention practices."
"Disclosure would give OpenAI's competitors insight into OpenAI's confidential business practices and customers and cause competitive harm to OpenAI. Yeats-Rowe Decl. 4."
Perhaps it would causes competitive harm because, upon learning about OpenAI's privacy practices, ChatGPT users might reduce their usage of ChatGPT
Declaration is sealed so we can only guess
wj 2 days ago |
Honest question. There is a line in the sand somewhere apparently.
AdamN 2 days ago |
There should be a carveout for non-profit or government AI.
Neil44 2 days ago |
proc0 2 days ago |
Arcuru 1 day ago |
Also that if you're going to use these things to write software, you should make it as virally copyleft as possible https://jackson.dev/post/moral-ai-licensing/
rafaelmn 2 days ago |
LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.
AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.
meerita 2 days ago |
fhn 1 day ago |
undefined 2 days ago |
fwlr 2 days ago |
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
sharts 1 day ago |
Fnoord 1 day ago |
dzink 1 day ago |
If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.
If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
undefined 1 day ago |
JohnFen 2 days ago |
123176 1 day ago |
“We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.”
All these failing companies are trying to bullshit their way out of the decline. Atlassian could have, you know, come up with a usable GitHub competitor. Instead they dream about coding by yapping.
b3lvedere 1 day ago |
So training can make it legal as well. Interesting...
sedan_baklazhan 2 days ago |
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
sinan-faizal 1 day ago |
j3th9n 1 day ago |
anon48293 1 day ago |
nullpoint420 about 14 hours ago |
gaigalas 1 day ago |
I think we'll not have people writing good content for a long time (there's no reason or incentive to), and the effects of this will splash back heavily on AI companies themselves.
You can see AI as a battery for intelligence that took a long time to charge and it's being used right now. For years, it was charged with all sorts of novel content that went undiscovered and AI is making available. That charge is the production of novel content, new insights, cross-pollination between areas, slowly driven by humans.
My view also draws a conclusion about recursive self-improvement: it is impossible for a battery to re-charge itself. I don't particularly think it can be done with this technology (LLMs).
I could be wrong though, but I don't think I am, and we'll know within our lifetimes. If things stall, it's likely because it has ran out of seeds/charge/substrate and not a technical limitation. It is in the long-term interest of AI companies to make incentives for people to generate novel public insights, they just don't know that yet.
1234letshaveatw 1 day ago |
danesparza 1 day ago |
pbasista 2 days ago |
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
kunley 2 days ago |
undefined 2 days ago |
WarmWash 1 day ago |
This same group of people, now being on the other side table, are screaming an crying that it's not fair.
Grow up and reap what you sow.
nathias 1 day ago |
joduplessis 1 day ago |
1p09gj20g8h 1 day ago |
jgalt212 1 day ago |
Sounds about right for the people and orgs involved.
shevy-java 1 day ago |
He does not see the moral dilemma here?
OrvalWintermute 1 day ago |
It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.
alex1138 2 days ago |
Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
seydor 1 day ago |
mike_bob 1 day ago |
UltraSane 2 days ago |
wrs 1 day ago |
Whew, good thing it’s too late to be accountable for that now, huh? Water under the bridge. Mistakes were made. Eggs, omelets.
undefined 2 days ago |
aitoolcrux 2 days ago |
aitoolcrux about 19 hours ago |
cineticdaffodil 1 day ago |
redsocksfan45 2 days ago |
jMyles 2 days ago |
aaron695 2 days ago |
bcjdjsndon 2 days ago |
rich_sasha 2 days ago |
dev1ycan 2 days ago |
FLeXMurphy 1 day ago |
jappgar 2 days ago |
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
noosphr 2 days ago |
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
vegnus 1 day ago |
nunez 1 day ago |
Big tech scrapes ALL OF THE WEBSITES CONSTANTLY to resell to you as knowledge? Shut up and take all of my money.
The 2020s is the wildest timeline indeed.
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.