Destroying Books to Build a Mind
In early 2026, the owner of James Payne, Books and Prints, based in Brooklyn, began receiving order requests that he later described as âbrazen and bizarre.â The books themselves were strangeâa so-called legal survival guide from 1998 titled âHow to Win in Small Claims Court in New Yorkâ; an academic text from 2006 called âCultures of Glass Architectureââand so was the order method. Booksellers like Payne typically post their holdings on multiple websites, such as AbeBooks, Alibris, Amazon, and Biblio, with Alibris widely considered to be the sleepiest revenue stream of the bunch. Suddenly, Payne went from selling fewer than ten books annually on Alibris to fulfilling orders of ten books, averaging as much as seventy-five to eighty dollars each, at a time on the site. Other booksellers around the country reported similar happenings.
Sylvia Petras, who owns Leaf and Stone Books, in Toronto, was envious as she watched book dealers in the U.S. post about their surging sales. Then the orders started trickling into her store, too, and that envy was quickly replaced by confusion. Petras owns a vast inventory of printed books from the fifteenth to the seventeenth centuries, along with scarce and scholarly books; odd titles, on subjects such as sewage plants, were flying off the shelves. Like other booksellers who were inundated with requests for niche books, Petras noticed that the orders were not being placed under an individualâs name but, rather, by enigmatic L.L.C.s: Green Parrot Project and Red Sparrow Project. One bookseller, who asked to remain anonymous, decided to attach a G.P.S. tracker to one of the purchased books before shipping it out to the mystery buyer. After researching her options, she selected a SmartCard, a device resembling a credit card, slipped it into a small envelope, and sealed it onto the inner rear board of a hardcover book, behind the dust jacket. She then followed along, online, as the book moved from Grandview Heights, Ohio, to an industrial park in Addison, Illinois, where a company called ARC Document Solutions operates an industrial-grade scanning facility. In a promotional video, an ARC spokesperson explains that many of the companyâs sites (there are more than a hundred nationwide) run 24/7. He refers to a client who has hired ARC to scan âone football field with boxes stacked four-high every month.â
Anthropic, the artificial-intelligence company behind the family of large language models known as Claude, is trying to acquire as many printed books as possible. We know this because of Anthropicâs own legal filings, unsealed in a copyright suit that was brought against the company back in 2024. Having identified books as the âhighest quality source of training dataâ for Anthropicâs L.L.M.s, it initiated a covert program called Project Panama, defined succinctly in court documents as âour effort to destructively scan all the books in the world.â âDestructively scanâ means precisely that: Anthropic or its affiliates âstripped the bindings from the print books, cut the pages to workable dimensions, and scanned those pagesâdiscarding each print copy while creating a digital one in its place,â according to the court filings.
The process involves the use of a hydraulic-powered cutting machine, also known as a âbook guillotine.â Itâs a sinister-sounding practice that is, actually, relatively commonplace, with university-preservation departments routinely disbinding books in order to make them easier to scan. Destroying a book isnât the only way to scan it, of course. The library at Princeton makes a point to retain physical copies of the texts they are digitizing, even if it makes for a less efficient process. âItâs totally possible, though more expensive, to take images that preserve the integrity of the book without flattening it,â Meredith Martin, a professor of English and the faculty director of the Center for Digital Humanities at Princeton, told me.
Still, Martin and other experts emphasized that old books are thrown awayâand subsequently destroyedâall the time, even for non-academic reasons, by libraries that are simply culling their collections. Donation is obviously more palatable, but it can be difficult to off-load, say, an outdated instruction manual or textbook, or several worn-out copies of the same mass-market paperback. (Some of these books wonât even be accepted by prison libraries, which generally reject hardcovers or texts that are in poor condition.) As a result, these books might end up in a landfill, or, depending on their bindings, get broken down and pulped for recycling, or perhaps shredded. In a Medium post where she makes the case for libraries âweedingâ their shelves, Claire Sewell, an academic librarian in Houston, recognizes that âseeing a dumpster full of books can seem completely antithetical to everything libraries are supposed to stand for as repositories of knowledge.â And yet, âold, outdated, damaged, or simply low circulating books have to be weeded on a regular basis in order for us to make space for new books that youâll actually want to check out.â
What made Anthropicâs project controversial, though, wasnât the fear that it would destroy a water-damaged copy of âThe Secret,â but, instead, genuinely rare and valuable books. To satisfy its voracity for billions of pages, Anthropic started out by placing mass orders with wholesalers, before turning to individual and secondhand booksellers to help fill in the gaps. The companyâs interest in obscure texts led to headlines about A.I. companies buying and destroying ârare old books,â or âantique books,â which generally calls to mind first editions.
There are many legitimate ethical concerns associated with the rise of A.I. companies like Anthropic, but the destruction of valuable books isnât necessarily one of them. âNone of our data acquisition programs buy and destroy rare or antiquarian books,â an Anthropic spokesperson wrote, in a statement. (âAntiquarianâ refers to books that are at least a hundred years old, regardless of rarity.) The Antiquarian Booksellersâ Association of America said it has not been notified by any of its members that such books have been dealt to A.I. companies. Of the individual booksellers I spoke with, none has sold any antiquarian books to buyers that seemed to be A.I. companies. Rather, the majority of the books theyâve sold have ISBNs, unique numerical identifiers that were not instituted until 1970.
Petras, the bookstore owner in Toronto, said that she was generally comfortable selling to undercover buyers, but that she would never sell them a book with interesting marginalia or an important provenance. Joyce Kosofsky, one of the owners of Brattle Book Shop, in Boston, said that the books she sold all had multiple copies. âWe probably had them at the cheapest price,â she guessed. In general, she argued, books are no different from any other saleable good: âJust like when you go into a clothing store, and you buy a pair of jeansâtheyâre your jeans. You can wear them. You can decorate them. You can give them away. No one follows you around saying, âWhat are you going to do with your jeans?â â
Given that Anthropic is purchasing books that are âlikely not very rare,â Martin, the Princeton professor, said, the companyâs use of book guillotines shouldnât be considered inflammatory. But, if the books arenât rare, then what are they? The booksellers shared the names of more than six hundred titles that they believed they had sold to A.I. companies, and I sent the list to Melanie Walsh, an Assistant Professor in the Information School at the University of Washington, to process digitally. The texts were obscure in their subject matter and lack of popularity: these were books, sometimes with low-print runsâusually a thousand copies or fewerâto meet realistic market demands. Of the top ten publishers, eight were academic pressesâroughly a third of the sample over all. Most of the books were published between the nineteen-seventies and the twenty-tens, and the genres spanned history, biography, fiction, poetry, literary criticism, law, and the social sciences. âBased on this sample, it appears that A.I. companies may be interested in training models on peer-reviewed academic research across a wide range of subjects,â Walsh concluded. This dovetails with what Mycal Tucker, a research scientist at Anthropic, discovered while organizing training data for an A.I. model: âNonfiction works tend to be more valuable than fiction,â he said in his written testimony, adding that nonfiction-book data helped the model perform well in disciplines as diverse as philosophy and astronomy.
Walsh said that when it comes to nonfiction works that are specialized, as is the case with âUtilization of Municipal Wastewater Sludge,â a 1972 booklet that Petras recently soldââthereâs an argument to be made that these obscure academic books may make a bigger impact as part of a Claude model than they would otherwise.â After all, they were already headed toward obsolescence.
Anthropicâs attempts to get its hands on âall the books in the worldâ has attracted legal challenges. In 2025, the company agreed to pay $1.5 billion to settle a class-action lawsuit brought by a group of authors who accused the company of violating their copyrights, namely by using their books for A.I. training without their permission. William Alsup, the judge presiding over the case, reprimanded Anthropic for some of its actions, such as downloading over seven million pirated copies of books and keeping the files âas a permanent, general-purpose resource,â even if they werenât being used to train Claude. (âAnthropic seems to believe that because some of the works it copied were sometimes used in training L.L.M.s, Anthropic was entitled to take for free all the works in the world and keep them forever with no further accounting,â Alsup wrote in his decision.)
But the larger practice of using books to train A.I. models was fair use, according to Alsup. âAuthors cannot rightly exclude anyone from using their works for training or learning,â he wrote. âEveryone reads texts, too, then writes new texts. They may need to pay for getting their hands on a text in the first instance. But to make anyone pay specifically for the use of a book each time they read it, each time they recall it from memory, each time they later draw upon it when writing new things in new ways would be unthinkable.â
Brandon Butler, a copyright lawyer and the executive director of the Re:Create Coalition, an advocacy group that supports both the creators of proprietary works and their users, described A.I. training as âthe fairest use in the history of copyright law.â This is because of how it âtakes from existing culture only the stuff that actually belongs to all of usâfacts, ideas, elements of language and grammarâand makes it easier for all of us to use,â Butler argued. Copyright could be infringed upon if L.L.M.s regurgitated memorized material verbatim or conveyed protected expression, but they generally convey unprotected information in a way that varies from the original sourceâs wording. In this way, there is no marketplace competition between A.I. companies and authors, Butler said, because someone who encounters a paraphrased excerpt of a book is still incentivized to buy the book. (This is complicated by how some L.L.M.s, including Metaâs Llama, have occasionally âmemorizedâ entire novels nearly verbatim.)
According to Judge Alsup, Anthropic was also within its rights to destroy the print books that it had obtained legally. âAnthropic purchased millions of print copies âto build a research library,â â he wrote. âIt destroyed each print copy while replacing it with a digital copy for use in its library (not for sharing nor sale outside the company).â Since the format change was for convenienceâeasier storage and search functionalityâand did not result in more copies of the original books, it was fair use.
Legality aside, though, we are met with difficulties, even just on a linguistic level, if Anthropicâs understanding of the term âresearch libraryâ is so different from ours that it feels like a misnomer. Theyâre ânot building a library because somebody might, one day, want to read all these books,â Butler explained. âTheyâre not interested in helping anyone else read the books. What they want, really, is data.â Their digital collection is being composed with the sole intent to enlarge Anthropicâs L.L.M.sâ âmemory â and to hone their writing abilities.
The consequence is that, even if these models successfully preserve literary materials that might otherwise be lost to time or foundering interest, the source will always remain in the shadows. Researchers will not have access to the book data, and users cannot control or observe how information is being sourced to them. Mike Furlough, the executive director of HathiTrust, told me that it is our right to ask A.I. companies for more transparency about how their data sets are created, notwithstanding that their mission is not to distribute books and that publicizing this data might sacrifice their competitive edge. Libraries and academic researchers are held to a much more rigorous standard: âan academic researcher who wants to produce a small-scale language model,â he said, âwould be expected to disclose the entire list of books that they would have used.â He continued, âThatâs just common academic practice because you want to be able to reproduce that, or understand what went into that.â The same is true of basically any reputable tool that a layperson would use for research; even Wikipedia includes citations. Anthropicâs goal to corral âall the books in the worldâ could almost be exciting, were it not for the fact that this hypothetical repository of written knowledge would exist only in the digital depths of Claude.
âEvery generation rewrote the bookâs epitaph,â the historian Leah Price wrote, in a nonfiction history of the medium, in 2019. âAll that changes is whodunnit.â In the nineteenth century, the suspected murderer was libraries: in a pamphlet printed in London, titled âThe Truth About Giving Readers Free Access to the Books in a Public Lending Library,â an anonymous writer sought to expose the ghastly reality of free access, which could result in the theft and âmisplacement â of books. More recently, fears coalesced around the advent of the e-book and the Kindle.
In âWhat We Talk About When We Talk About Books,â Price points out that books have always had a chameleonic qualityâthat they adapt with technological innovation, and that the two things arenât necessarily always in tension. And yet, what it means to be a reader, a consumer of literature, is a concept thatâs beginning to shift, too. In an e-mail to me, Price expressed concerns about the looming threat that âA.I. will replace human readers,â since âdestructive scanning sacrifices multiple potential future human readings to a single instance of machine reading.â Books come with baggage, and A.I. shears them of that. Books demand our time, involvement, and deliberationâeach one has a unique set of readerly requirements. In a 2006 essay titled âThe Rise of Fictionality,â exploring the proliferation of fiction writing in the eighteenth century, the author Catherine Gallagher lists the ways in which early novels commanded the readerâs attention and engagement, asking them âto anticipate problems, make suppositional predictions, and see possible outcomes and alternative interpretations.â If one still believes in reading as an âimaginative, immersive experience,â Price told me, then there is a critical contrastâparticularly when it comes to hallmark works of literatureâbetween unmediated human reading versus bots doing the reading for users and extracting worthwhile content. âA novel is not just a bucket of sentences,â Price said. âThe order in which those sentences occur matters.â And yet Anthropicâs goal, it seems, is to accumulate and spit out ever more buckets.
Matthew Kirschenbaum, a Commonwealth Professor of Artificial Intelligence and English at the University of Virginia, warned of what will happen if the âgreat virtue of the bookââits existence as a self-contained object, with a human writer, and a human historyâis dissolved into data that siphons all context from the original. The boundaries of the book (e.g., its covers and materiality, publication information, and author) disappear when the text is âhomogenizedâ and âloses all of its relationships that were present in the original book.â Thus, the very thing protecting A.I. from copyright infringementâits ability to rephrase thingsâis also the problem: L.L.M.s transform literature into some hybridized, half-dead thing that is far removed, if not completely untethered, from its original source. It creates a product so drastically new that the text as an individual work is effectively destroyed. This less obvious form of erasure is much more hazardous and worrisome than destructive scanning.
Even still, I wonder if obtaining information by way of L.L.M.s will only ever intrigue a percentage of the population. There is a parallel here with the rise of the novel in the eighteenth century. After Daniel Defoeâs âRobinson Crusoeâ and Samuel Richardsonâs âPamelaâ were published, in 1719 and 1740, respectively, multiple print runs were exhausted throughout the remainder of the century. Competitors, noticing the copies being printed to keep up with demand, devised a way to make a quick buck: they published a variety of pirated editions. These second-rate versions included child-friendly copies, abridgements preserving the best-written or juiciest bits, shortened texts that were simply cheaper than the full-length ones, and unauthorized sequels. These pirated books were incredibly popular, too, even as people continued to seek out the original texts. We can divide these eighteenth-century consumers into two camps: those who bought and read the novels written by Defoe and/or Richardson, and those who were partial to altered or imitative copies. There have always beenâand there will perhaps always beâpeople who prefer slop over the primary source. âŚ
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.