Booksellers suspect AI firms are buying and then destroying rare books
Ars Technica · LC · trust 54/100

“Every collection is important” Booksellers suspect AI firms are buying and then destroying rare books AI firms quietly bulk buying rare books face resistance from booksellers.
346 Credit: hexvivo | iStock / Getty Images Plus Credit: hexvivo | iStock / Getty Images Plus Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul.
It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.
What makes this destruction extra painful, though, is that it doesn’t have to be this way.
Google patented a non-destructive book-scanning technology in 2009 that AI firms could use to efficiently scan books—if they were willing to slow down and invest in the process. It’s not perfect, however; studies have found that the curve of the page can distort text, and pages can be missed. As Wired reported , glitches can happen when workers move too quickly, including disembodied hands obscuring pages.
Overall, the trade-offs in cost and speed may not appeal to AI firms looking for the cheapest way to scan millions of titles, and Google’s method may not be the best way to handle rare books anyway. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts takes time and attention to limit handling, and that’s why it’s considered such a human job.
For AI companies bent on finding shortcuts, the Internet Archive’s method may feel almost alien.
But if you’re a book lover looking for a balm while parsing rumors of AI-driven book destruction, an old post from 2021 describes how the Internet Archive scans books to avoid damaging even the most precious rare books. In it, a book scanner who has been with the Archive since 2010, Eliza Zhang, offered a moment of zen by explaining what she likes so much about scanning books “the hard way.”
The Internet Archive did try going the automated route, the post said, even testing out “commercial book scanners that feature a vacuum-powered page-turning arm.” But “it turns out those automated scanners didn’t really work well for brittle books, rare volumes, and other special collections—the kinds of material our library partners ask us to digitize,” the Internet Archive said.
“The job requires keen concentration,” Zhang said, since the pages of “very old, fragile books” are “paper thin.” In the post, Andrea Mills, who helps lead the Archive’s book-scanning operations, explained that “clean, dry human hands are the best way to turn pages.”
To ensure each rare book only has to go through the scanning process once, Zhang takes her time. She carefully raises the scanner glass with a foot pedal each time she turns a page, then adjusts the cameras and ensures the page is readable. She also takes note of any fold-outs, setting a reminder to go back and scan the inserts so that bonus materials aren’t lost while documenting the main pages of the work.
The whole time, she understands that if a page is skipped or an image is too blurry, Internet Archive’s proprietary software will stop the process and prompt her to scan it again. Practice makes perfect, though, and she reports a low error rate after more than a decade of finding a rhythm in the job.
At the time, Zhang had scanned “more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books),” the post said, with the goal of guaranteeing “zero errors.”
Ars asked the Internet Archive for comment, but Chris Freeland, the director of library services, said the post detailing Zhang’s work is “still the best description of our scanning process today.”
The post came after a video of Zhang’s book scanning got 1.5 million views on what was then Twitter, accompanied by a caption that would resonate with book lovers appalled by AI-driven book destruction today.
“At the Internet Archive, this is how we digitize a book,” the tweet said. “We never destroy a book by cutting off its binding. Instead, we digitize it the hard way—one page at a time.”
Ever since a lawsuit in summer 2025 outed Anthropic for destroying millions of print books to train its AI models, book lovers have moved to defend some of the most precious collections from what feels like AI firms’ endless quest to feed all the books in the world into their large language models.
The biggest fear for people who want to see books preserved through the training process is that AI firms will callously pulp rare books that can never be replaced.
As The Atlantic reported last week, social media “raged” after two recent reports indicated that AI was already endangering rare books. First, a Telegraph report accused Silicon Valley of destroying millions of rare books and “shredding the originals,” then 404 Media reported that a book-database company called ISBNdb was advertising that it could help AI firms source books in bulk.
This backlash was expected, with ISBNdb reportedly warning its clients that the optics were bad. Anthropic started using a codename—“Project Panama”—for its destructive book-scanning in an effort to keep it hidden from the public.
There is no indication that Anthropic ever destroyed rare books, and the company has denied…
Read the original at Ars Technica →
Open in TruthVane →