Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

Microsoft exec called AI scraping 'the largest theft of labor in human history' (https://techcrunch.com)

926 points by pluc 2 days ago | 817 comments | View on ycombinator

haritha-j 1 day ago |

I just don't understand people saying "but a human learning from a book isn't illegal".

How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."

And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.

47282847 2 days ago |

“Information wants to be free“.

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

sajithdilshan 2 days ago |

If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.

juvvel 2 days ago |

I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.

TutleCpt 2 days ago |

The most shocking point is that they have a Microsoft exec who knows what he's talking about.

Kuyawa 1 day ago |

Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours

So no, your cries for regulating others because you are losing the race won't work this time.

tom2026hn 2 days ago |

The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.

Waterluvian 1 day ago |

I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.

leonidasrup 2 days ago |

In case of programming.

How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?

This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.

Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?

What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.

https://consortiuminfo.org/metalibrary/estimating-the-total-...

ohrus 1 day ago |

Yet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants.

The hypocrisy of this new world is already catching up to us.

heaney-555 1 day ago |

LLMs are not compression algorithms. From an information theory perspective, that's impossible given their size.

Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.

It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.

thunkshift1 1 day ago |

This will lead to a massive settlement between the big boys and most people who put stuff out in good faith will be left out of it. And that will be the end of it. We will never hear anything about this ever again and the ‘theft’ will continue like normal.

iamflimflam1 2 days ago |

I don’t mind these companies scraping my content.

But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.

cmiles8 1 day ago |

The evidence here is quite damning for OpenAI and Microsoft is clearly trying to distance themselves from OpenAI’s behavior here.

American87 2 days ago |

I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)

Weryj 2 days ago |

I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.

nullbio 2 days ago |

It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.

sebastiangrill 2 days ago |

I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.

IX-103 1 day ago |

I'm not sure how to be upset over this. For decades copyright has been extended and extended. Meanwhile, the ease and speed of spreading published works across the globe have increased massively. Don't forget that when copyright was first created, it could take multiple years for first editions to make it across the globe. In that environment, multiple decades of copyright makes sense. Whereas today that would actually stifle innovation and creativity rather than incentivizing it. I mean, because of the length of copyright we've gotten all of these live action remakes of Disney films or superhero movies rehashing the same stories.

I find it a little hard to be upset about AI. Supposedly stealing copyrighted works when the vast majority of those works. Probably should have been in the public domain to begin with. I have a faint hope that this scuffle between the AI companies and the publishing industry will result in more reasonable copyright laws, but I think it's more likely that exceptions will be made and AI will be treated as a special case.

pianoben 1 day ago |

As if the transatlantic slave trade wasn't a thing. What a perfect illustration of our industry's self-absorption and self-regard.

hereme888 1 day ago |

Microsoft is one to talk.... Remember when MS trained copilot on all your github code?

jaybeavers 2 days ago |

You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.

markhahn 1 day ago |

Has anyone found a meaningful discussion of how scraping is theft?

Obviously, reproducing works in whole is infringement. That's not what AI is doing, so the question becomes: how is scraping different from ordinary reading? Is it just that site owners want to play back history and retroactively create high-cost licenses for scraping?

__bjoernd about 19 hours ago |

But US labs have been doing it right, while Chinese labs are only distilling from their superior products, right?

bentt 1 day ago |

This is a great basis for a dividend from AI revenue to be paid back to society in a more inclusive form than stock. As more money flows into AI companies, data centers, and other related infrastructure, an amount should be extracted and redistributed in the name of balancing this equation.

juiceland 2 days ago |

Copyright infringement, if this even were that, is not theft. Chattel slavery is the largest theft of labor in human history.

levischoen 1 day ago |

Transatlantic slave trade calling - we’ll hold. I know there’s a memory shortage but history books are cheap.

GardenLetter27 2 days ago |

I think it's okay to advance humanity, but they can GTFO when they then try to ban distilling and open models.

totetsu 2 days ago |

Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?

montjoy 1 day ago |

What’s the Microsoft angle here? A few years ago they were ready to hire anyone from OpenAI that was willing to leave. Is it just catering to the anti-AI sentiment going around?

jacquesm 2 days ago |

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

keeda 1 day ago |

We may not like this, but let us contemplate what laws would be in place to prevent a thing like this; I suspect we would like those laws even less.

The laws at play here are related to Intellectual Property, specifically Copyright. Yes, it is terribly flawed, but it is the product of centuries of case law dealing with very hairy issues, and I believe it is fundamentally sound, and here's why.

As the name implies, it deals with only verbatim copies of works or subsantial portions thereof. It very expressly does not cover abstract things like concepts, ideas, themes, facts, or patterns, and rightfully so, because we really do not want anyone owning something that broad.

But these abstract things are precisely what have been extracted, at unimaginable scale, to build these models! Each pattern in the tokens derived from these works contributed imperceptibly tiny perturbations to randomly initialized weights, interacting in incomprehensible ways into vectors representing concepts and ideas and facts, the cumulative aggregate of which has somehow created a form of intelligence.

There is no copying, only gleaning, and so Copyright Law falls short. But what is the alternative, and do we want it?

To prevent something like this would require some sort of legal protection on the more abstract things. We do have a legal framework for those: Patents! But as is very clear on HN and in many Tech circles, those are an extremely contentious topic (even though they actually protect much narrower ideas than most presume.) I don't think anybody anywhere really wants any protection on broader abstractions, and rightfully so.

So: we as a society expressly decided these abstract things belong to the commons, and those are the exact things these labs harvested. This is probably the only logical culmination of our technological journey, and is within the very reasonable legal frameworks we have evolved over centuries.

As such, it is not productive to dwell on fighting this or bemoaning this. Instead we should focus on ensuring that this technology -- with its immense potential and opportunities and dangers -- benefits everybody as much as possible. That is a better way to compensate everybody's labor, and that is a much richer and fruitful discussion to be had.

gyosko 2 days ago |

And here we are, just watching and doing nothing..

alansaber 2 days ago |

Spiderman pointing

rdsubhas 2 days ago |

They are not selling the information. They are selling a service for easy access to that information.

These are two different things.

Note: am not an AI fanatic.

rietta 1 day ago |

I remember when open source software, and Linux in particular, was the threat to the world according to Microsoft execs.

undefined 1 day ago |

undefined

1vuio0pswjnm7 1 day ago |

Source:

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

p. 1

"This case is about, as Microsoft's Director of Applied Science put it, an astonishing theft of unprecedented proportions; SF1437, perhaps the largest theft of labor in human history. SF1652"

p.11

"As Microsoft recognized: millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions and admitted that almost no one intended for content they created to be used in this fashion, nor are they compensated for its use. SF1437."

p. 74

"As Microsoft's Dr. Glen Weyl put it, compensating creators is in the best interests of my employer, of my country, and of many other groups I belong to. SF1657."

Hyperbolic quotes from Microsoft employees are, IMO, the least interesting elements of this brief

Here is Microsoft's brief. Note how MSFT responds to the "web grounding" claims

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

It seems OpenAI does not want the public to know about (a) OpenAI's data collection and retention practices and (b) the number ChatGPT users have requested deletion of conversations

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

"OpenAI seeks to redact specific information about [(a)] the number of users who requested deletion of ChatGPT conversations and [(b)] OpenAI's related data collection and retention practices."

"Disclosure would give OpenAI's competitors insight into OpenAI's confidential business practices and customers and cause competitive harm to OpenAI. Yeats-Rowe Decl. 4."

Perhaps it would causes competitive harm because, upon learning about OpenAI's privacy practices, ChatGPT users might reduce their usage of ChatGPT

Declaration is sealed so we can only guess

wj 2 days ago |

How is this different from Microsoft scraping to build Bing?

Honest question. There is a line in the sand somewhere apparently.

AdamN 2 days ago |

Correct solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use.

There should be a carveout for non-profit or government AI.

Neil44 2 days ago |

I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.

proc0 2 days ago |

If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.

Arcuru 1 day ago |

It's the scale that's the problem. I've been saying for a while, the value of any individual piece of work is not terribly valuable for an LLM, but the aggregate value of all human work is obviously very valuable. I think this line of thinking can be the argument for why we should heavily tax these AI companies above and beyond how we tax other industries.

Also that if you're going to use these things to write software, you should make it as virally copyleft as possible https://jackson.dev/post/moral-ai-licensing/

rafaelmn 2 days ago |

Copyright is artificial scarcity rationalized by arguing that producing novel intellectual work is valuable, but requires substantial effort that can't be recouped, so we have to incentivize it somehow.

LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.

AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.

meerita 2 days ago |

There will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.

fhn 1 day ago |

But Windows telemetry, github code training, scanning all OneDrive documents/outlook emails is not theft

undefined 2 days ago |

undefined

fwlr 2 days ago |

The largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices.

It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.

If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.

sharts 1 day ago |

So Microsoft exec stating what most people have been saying already is newsworthy.

Fnoord 1 day ago |

Copyright infringement, not theft.

dzink 1 day ago |

Are we considering what is the shelf life of information?

If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.

If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.

undefined 1 day ago |

undefined

JohnFen 2 days ago |

As someone who thinks that genAI is harmful, I deeply resent that any of my work has been used to help train it. I will never forgive these companies for forcing me to contribute.

123176 1 day ago |

Atlassian, powered by this spy tool, is truly insane:

“We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.”

All these failing companies are trying to bullshit their way out of the decline. Atlassian could have, you know, come up with a usable GitHub competitor. Instead they dream about coding by yapping.

b3lvedere 1 day ago |

"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. "

So training can make it legal as well. Interesting...

sedan_baklazhan 2 days ago |

AI overall is the ultimate piracy crime.

I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.

sinan-faizal 1 day ago |

well it can be used for good as well, how we stopping it?

j3th9n 1 day ago |

Everyone benefits.

anon48293 1 day ago |

Funny. The end of copyright and patents is by far the best thing about AI to me. All of it is nonsense. Great that you drew a picture of a mouse once, I really fail to see why I couldn’t draw it and sell it either. It was moronic from the get go.

nullpoint420 about 14 hours ago |

One could call it the largest distillation attack in history

gaigalas 1 day ago |

Framing it as property is the wrong angle though. It's an ecosystem, not a cache of good writings that was stolen. That ecosystem was hurt severely and its recovery is uncertain.

I think we'll not have people writing good content for a long time (there's no reason or incentive to), and the effects of this will splash back heavily on AI companies themselves.

You can see AI as a battery for intelligence that took a long time to charge and it's being used right now. For years, it was charged with all sorts of novel content that went undiscovered and AI is making available. That charge is the production of novel content, new insights, cross-pollination between areas, slowly driven by humans.

My view also draws a conclusion about recursive self-improvement: it is impossible for a battery to re-charge itself. I don't particularly think it can be done with this technology (LLMs).

I could be wrong though, but I don't think I am, and we'll know within our lifetimes. If things stall, it's likely because it has ran out of seeds/charge/substrate and not a technical limitation. It is in the long-term interest of AI companies to make incentives for people to generate novel public insights, they just don't know that yet.

1234letshaveatw 1 day ago |

What if all this slurping and training empowers humanity to cure cancer? feed the starving? travel the stars? We are not only making the accumulated knowledge of the world accessible, we are making it actionable. Sure I'm ignoring all the possible bad outcomes lol, but if an independence day (movie) type scenario was playing out nobody would be batting an eye. I guess cancer is not as sexy though

danesparza 1 day ago |

I mean ... I know a few African nations that might disagree.

pbasista 2 days ago |

I do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet.

Publicly. Accessible.

Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.

Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.

kunley 2 days ago |

But what about M$ owning Github and doing the same with its content? Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...

undefined 2 days ago |

undefined

WarmWash 1 day ago |

The righteousness of the internet, the same internet that desperately called to end IP laws, championed piracy, ad-block everything, and always use proxy services to backdoor paywalls/login walls

This same group of people, now being on the other side table, are screaming an crying that it's not fair.

Grow up and reap what you sow.

nathias 1 day ago |

distillation is a human right

joduplessis 1 day ago |

Microsoft executives levelling "tone-deaf" up in realtime.

1p09gj20g8h 1 day ago |

It literally is. Anyone saying otherwise is deluding themselves

jgalt212 1 day ago |

> Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work.

Sounds about right for the people and orgs involved.

shevy-java 1 day ago |

Yet he also helps destroy all those jobs. The thief is calling "Catch the thief!".

He does not see the moral dilemma here?

OrvalWintermute 1 day ago |

Throughout history we’ve been able to retell stories, to copy content, to create shallow clones or synthesis

It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.

alex1138 2 days ago |

Information wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it

Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it

seydor 1 day ago |

Imagine the parthenon marbles. When they were looted it was even a celebrated act, but they are still stolen in the british museum centuries later.

mike_bob 1 day ago |

Go cry me a river of lies Mr. Microsoft Exec. Corporate culture is a blight on humanity.

UltraSane 2 days ago |

"The Net interprets censorship as damage and routes around it."

wrs 1 day ago |

>[Nadella said] if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

Whew, good thing it’s too late to be accountable for that now, huh? Water under the bridge. Mistakes were made. Eggs, omelets.

undefined 2 days ago |

undefined

aitoolcrux 2 days ago |

[flagged]

aitoolcrux about 19 hours ago |

[flagged]

cineticdaffodil 1 day ago |

[dead]

redsocksfan45 2 days ago |

[dead]

jMyles 2 days ago |

[dead]

aaron695 2 days ago |

[dead]

bcjdjsndon 2 days ago |

[flagged]

rich_sasha 2 days ago |

[flagged]

dev1ycan 2 days ago |

Because it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.

FLeXMurphy 1 day ago |

Apart from jacquesm's wonderful milquetoast comment, if anyone has any practical solutions to this problem that do not involve suspension of disbelief that voting (with or without wallet) and calling "your congresscritter" or whatever other nonsense people spout, now would be a great time to voice it.

jappgar 2 days ago |

All the "LOL you wouldn't steal a car???" posts in this thread miss the point entirely.

AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.

At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.

noosphr 2 days ago |

I'd call the introduction of copyright the largest theft of human labor in history.

No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.

That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.

The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.

At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.

vegnus 1 day ago |

Theyre just jealous that theyre being surpassed on their market capture of computing

nunez 1 day ago |

You or I scrape a website for casual use? Straight to jail, right away.

Big tech scrapes ALL OF THE WEBSITES CONSTANTLY to resell to you as knowledge? Shut up and take all of my money.

The 2020s is the wildest timeline indeed.