OpenAI and Microsoft knew they were starting a âdoom loopâ for the web
Recently unsealed court documents in the New York Timesâ case against OpenAI and Microsoft are pretty damning. The companiesâ own documentation warned that it was starting a âdoom loopâ that would damage the web, characterized its scraping of data to train its models as the âlargest theft of labor in human history,â and that it made a âcomplete mockery of the idea of fair use.â
OpenAI and Microsoft knew they were starting a âdoom loopâ for the web
The companies knew they were driving us toward Google Zero, and did it anyway.
The companies knew they were driving us toward Google Zero, and did it anyway.
Many of the most eye-catching quotes from the document come from Microsoftâs Director of Applied Science, Brent Hecht. Though, the company has tried to distance itself from Hechtâs assertions. Microsoft spokesperson Alex Haurek told The Verge that âThese comments reflect one employeeâs individual perspective, are not a legal analysis, and do not represent the companyâs views.â
In a separate court filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterized Hechtâs role as adversarial. He said that Hecht âholds divergent, academic, and forward-looking views about how data ecosystems for AI should operate and is employed at Microsoft to bring asymmetrical, futuristic, and academic points of view ⌠nor is he someone who speaks for Microsoft specifically as to his theoretical views on AIâs potential effect on content creators.â
But whether or not Microsoft wants to own these comments, itâs clear that this came true. Google Zero is real! AI is eating the web!
There are plenty more wild statements in NYTâs filing from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. Here are some highlights from the 92 page document.
âAn astonishing theftâ
The introduction quotes Hecht and OpenAIâs Head of ChatGPT (presumably Nick Turley) in a way that seems to show the companies knew they posed an âexistential threatâ to publishers like the New York Times. Hecht calls ChatGPT and Copilotâs harvesting of data the âlargest theft of labor in human historyâ and says that Microsoftâs defense makes a âcomplete mockery of the idea of âfair use.ââ
Itâs a âdoom loopâ
Satya Nadella admits that chatbots have basically replaced search and removed the need to go straight to the source for info. But perhaps more damning is an internal Microsoft document that says, âOur AI content strategy has started a âdoom loopâ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its âcontent supply chain.ââ
Thatâs not even a real number
Donât be fooled by OpenAI or Microsoftâs claims of altruistic intent. OpenAI cofounder Greg Brockman is more interested in the âgazillionsâ of dollars it he could potentially make through commercial AI.
Paywall shmaywall
Despite Nadella later being quoted as saying, âanything that is paywalled should be licensed,â An OpenAI representative admitted that he was âunawareâ of any effort to detect or remove paywalled content from training data.
âInsanely good at regurgitationâ
Internally, it seems that OpenAI was well aware of ChatGPTâs tendency to simply reproduce copyrighted material âverbatim.â Even though it acknowledged that the âprevention of memorizationâ was important to âminimize copyright violations,â employees admitted that GPT-4 âmemorized a ton of data and therefore will be insanely good at regurgitation.â
The filing then goes on to cite several examples of ChatGPT outputting long strings of copy straight from articles in the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to queries.
ââHoovering upâ all their workâ
Microsoft knew how its wholesale scraping of the internet would be perceived and admitted that âalmost no one intended for they [sic] content they created to be used in this fashion, nor are they compensated for its use.â
A âsubstitute for the labor of peopleâ
OpenAI Policy Director Jack Clark saw the writing on the wall, saying that it was âcreating systems that substitute for the labor of the people that define the âcultureâ of society.â Internal documents described ChatGPT as âthe modern newsstand.â OpenAIâs Nick Turley is later quoted as saying that once you get an answer from its chatbot, there is âno good reason to clickâ on a link to the source.
Destroying their own supply chain
Microsoft is quoted as admitting that âLLMs are a product that destroys its own supply chainâ because itâs a substitute for its own training data in many cases.
OpenAI knows its killing referral traffic
OpenAIâs own media and economic experts attributed the drop in referral traffic for sites like the Times directly to AI summaries like Googleâs AI Overviews. Theyâve speculated that search referrals may be down as much as 60 percent.
Microsoft spokesperson Haurek cautioned that âSatyaâs testimony and Microsoftâs position in this case are perfectly consistent. He spoke to broad principles and changes underway in how people find and consume information. Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.â
But it seems pretty clear based on this newly unsealed document that both Microsoft and OpenAI knew they were going to irreparably harm the publishing industry, the âmillions of peopleâ it employs, and, by extension, damage their own product, but carried forward anyway in pursuit of âgazillionsâ of dollars â doom loop be damned.
Most Popular
- The Apple Watch Series 12 is the start of a new wearable era
- The 2.5-hour AI-generated Odyssey movie is 2.5 hours too long
- OpenAI and Microsoft knew they were starting a âdoom loopâ for the web
- The iPhone 18 Proâs big camera update is all about the small gains
- This cartridge-playing Game Boy clone is smaller and cheaper than Analogueâs Pocket
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.