AI companies are destroying rare books
60sShocking revelation about AI companies bulk-buying and shredding books sparks outrage and curiosity.
▶ Play Clip"The title is somewhat sensational but accurately reflects the video's core message about AI destroying books, though it could be more specific."
This video discusses a Washington Post investigation revealing that AI companies, particularly Anthropic, are bulk-purchasing rare books, scanning them, and destroying the physical copies to train their AI models. The video highlights the ethical and legal controversies surrounding this practice, including a judge's ruling that it constitutes fair use, and expresses strong criticism of the destruction of cultural artifacts.
The video begins by framing AI's impact on art, specifically literature, and introduces the topic of AI companies damaging books.
A Twitter post by Hedge Markets reveals that AI companies are bulk buying rare books, scanning them, and shredding the originals. The service ISBNdb orders up to a million books, keeping buyers anonymous, and focuses on pre-2022 books to avoid AI-generated text.
A federal judge ruled that this practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google's book scanning project.
The Washington Post article details Project Panama, hosted by Anthropic's Claude, which was kept hush-hush due to expected backlash. Anthropic agreed to pay $1.5 billion to settle a copyright lawsuit in August.
Anthropic hired Tom Turvey, former head of partnerships for Google's book scanning project, who contacted publishers about bulk purchasing books for AI training, letting licensing conversations die.
Anthropic spent millions to purchase print books, strip bindings, cut pages, and scan them into digital form, discarding the paper originals. The books were then sorted into data mixes for LLM training.
The video highlights that Anthropic could have used legal methods but instead resorted to downloading pirated books. Meta employees also questioned the practice, but it was escalated to Mark Zuckerberg who allegedly approved it.
Anthropic's co-founder Ben Mann personally downloaded books from Libgen and other copyright-infringing content over 11 days. He later shared a link to a new copyrighted book website with other employees.
Anthropic claimed in legal filings that it never trained a commercial AI model using Libgen data and never used the pirate library mirror to train any complete AI model.
A judge found that Anthropic's use of books for AI training was transformative, likening it to teachers training students to write well. Another judge in a Meta case found that book authors could be harmed by AI competition.
The video argues that training LLMs on books is not transformative, as it involves copying data. It also criticizes the competition from AI-generated books flooding platforms like Kindle Unlimited.
Anthropic was looking for a vendor to convert 500,000 to 2 million books over a 6-month period, using hydraulic-powered cutting machines to cut books for scanning and recycling.
The video presents an alternative: a post by Brian Romel shows a book scanner that can scan thousands of pages a day without destroying the book, which any large AI company can afford.
The video shares public reactions, including comparisons to the Burning of the Library of Alexandria and quotes from George Orwell's 1984 about rewriting history.
The video concludes that AI legislation is needed urgently, as the destruction of books and the fair use ruling are seen as absurd and harmful.
The creator expresses sadness and disappointment, emphasizing the importance of preserving books as historical artifacts and criticizing the destruction for AI training.
The video concludes that the destruction of books for AI training is a harmful practice that should be stopped, calling for immediate AI legislation and a more ethical approach to data collection.
What is the name of the project by Anthropic that involves scanning and disposing of books?
Project Panama
01:07
How much did Anthropic agree to pay to settle a copyright lawsuit in August?
$1.5 billion
01:32
Who was the former head of partnerships for Google's book scanning project hired by Anthropic?
Tom Turvey
01:32
What did Anthropic's co-founder Ben Mann allegedly do?
He personally downloaded books from a shadow library and other copyright infringing content from Libgen over 11 days.
03:10
What did a judge rule about Anthropic's use of books for AI training?
It was transformative under fair use, likened to teachers training school children to write well.
04:04
What is the scale of book conversion Anthropic was looking for?
500,000 to 2 million books over a 6-month period.
05:21
What alternative method for scanning books was mentioned?
A book scanner that can scan thousands of pages a day without destroying the book.
06:15
AI Companies Bulk Buying and Shredding Books
Reveals the shocking practice of AI companies destroying rare books for training data.
00:13Project Panama Details
Provides concrete details about Anthropic's secretive book scanning project.
01:07Transformative Use Ruling
Highlights the controversial legal ruling that allows AI companies to use copyrighted books without permission.
04:04Public Outrage and Historical Parallels
Connects the destruction to historical events like the Library of Alexandria, emphasizing cultural loss.
06:41Call for AI Legislation
Stresses the urgent need for regulations to prevent such practices.
08:05[00:01] to tell, because AI is really ruining everything. And I talk a lot about art on this channel, but one really big thing, a piece of art, is literature. Books, writing, it's also a form of art as well. And recently there was news
[00:13] that came out about how AI is damaging kicks off with this post on Twitter by Hedge Markets writing, "AI companies are bulk buying rare books, scanning them
[00:25] spines off, and shredding the originals." A service called ISBNdb orders of up to a million books and keeps buyers anonymous. And they're looking for pre-2022 books because they're free of AI-generated text. And a
[00:40] federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. And Anthropic hired the former head of quotes, "all the books in the world,"
[00:54] end quote. This sounds absolutely horrific and it is true. There is a titled, "Inside an AI Startup's Plan to Scan and Dispose of a Million Books," written January 27, 2026 by Aaron
[01:07] Schaer, Will Oremus, and Natasha Tiku. This details a project called Project Panama hosted by Anthropic's Claude. And it was something they really, really wanted to keep hush-hush because very likely people were going to be super mad
[01:20] quotes, "We don't want it to be known that we're working on this," end quote. And there was a copyright lawsuit that happened from some book authors whose books were used in this LLM training. And it highlights the company agreed to
[01:32] pay 1.5 billion to settle the case in August. So, as the lawsuit highlights, Anthropic in 2024 had hired former head of partnerships for Google's book scanning project, Tom Turvey. And Tom had, you know, sent emails to publishers
[01:47] books for AI training?" And he let the convos die. So, instead of licensing, he instead contacted them about bulk purchasing books for his AI firm's research library. They spent millions of dollars to purchase print books and then
[02:03] strip the books from their binding, cut the pages to size, and scan the books into digital form, discarding the paper originals. And then after that, the books were sorted into a bunch of different data mixes for different types
[02:15] of books for LLM training. And they were sought after for their data mixes because apparently authors like to use Claude to write books in the likeness of The people whose books were trained on. And going back to the Washington Post
[02:29] article, it's honestly kind of bonkers because it highlights that Anthropic met in other companies, you know, they could have gone the practical legal way to get the books, but instead they went through other ways by downloading pirated
[02:41] were many occasions where Meta employees were kind of like scratching their heads are we sure that we should be doing this?" But it was apparently escalated this?" But it was apparently escalated to Mark Zuckerberg and he apparently
[02:55] allegedly approved it. That was just from Meta. And talking about Anthropic, the co-founder Ben Mann apparently personally downloaded a whole bunch of books from a shadow library and other copyright infringing content from Libgen
[03:10] over 11 days. And then, you know, a year later in 2022, a new copyright book website was released and he sent that link to other Anthropic employees and was like, "Yay, just in time!" Presumably for more pirated material for
[03:23] Washington Post adds that apparently Anthropic said in legal filings that the company never trained a commercial AI model that generated revenue using its Libgen data. And apparently it never used the pirate library mirror, which is
[03:36] the, you know, that new copyrighted website, to train any complete AI model. allegedly if you didn't use the copyrighted material to train your LLMs, then why, you know, download all this stuff in the first place? Like, what are
[03:51] you downloading it for? Your own personal reading perusals? Even in that that you don't have to be doing this. But, what other excuse can you come up with? What's actually really sad is that according to this article, apparently,
[04:04] in June, a judge found that Anthropic was allowed to use those books for AI training models because it was processed in a transformative way. Where he likened the AI training process to teachers training school children to
[04:16] write well. And then another judge in the case with Meta found that the book the case with Meta found that the book authors could harm sales of their books. about your thoughts and I would love to hear it down below in the comment
[04:28] section. But, me personally, I'm not a judge, but I don't think, you know, training a whole bunch of books on LLMs is transformative by any means. You're of data. And same thing for the book authors that are concerned about harming
[04:42] the sales of their books, competition is healthy and it is necessary in spaces to promote growth, but this isn't a good type of competition. Competition because are graduating from college and writing really, really good peak literature that
[04:55] your books are competing with is good competition. Competing with a whole bunch of people who suddenly go into Claude and are like, "Claude, write me a hot fairy romance book." And they're flooding Kindle Unlimited with this
[05:07] That's not good competition. It's creating a lot of slop that people, customers that could be potentially your customers, are now like sifting through. Although, I guess a teeny bit of hope, the judge found that Anthropic may have
[05:21] they downloaded all that pirated material. And according to this article, of how many books Anthropic purchased, but Anthropic was apparently looking for a vendor who can convert 500,000 to 2
[05:34] million books over a 6-month period. And that they were using hydraulic-powered cutting machines to cut books to later scan it and then to schedule a recycling remnants. Again, the worst part about all of this is the fact that it doesn't
[05:49] don't need to be cut like this. They don't need to be destroyed. Like there's that don't involve just chopping the spines off the books, but
[06:01] money, whatever. So they're just going with the cheapest and most destructive option. Give me 2 million books in 6 months. Screw it, right? Another option for scanning books without destroying them. This is a post by Brian Romel
[06:15] writing, "In I can scan thousands of pages a day and never destroy a book. No guillotine spines, no ripping the bindings. Just a normal book ready to be treasured for centuries. Any large AI company can afford this both financially
[06:28] and ethically." And you can see here, it's scanning the book, flipping the pages, and as you can see it's scanning in here, and we don't have to actually are arguing, "Oh, well, it has to like be destroyed for the copyright law."
[06:41] why? Or do you think they're trying to be cheap and cut corners? And like you or me who is probably equally as horrified as this happening, a lot of people are really upset about this. In the original post posted on Twitter by
[06:54] easier to rewrite history and knowledge electronically the way they want us to history." And this is a quote from George Orwell's 1984. Quote, "Every record has been destroyed or falsified. Every book rewritten. Every picture has
[07:09] been repainted. Every statue and street building has been renamed. Every date has been altered. And the process is continuing day by day and minute by minute. History has stopped. Nothing exists except an endless present in
[07:22] responded, "Yep, and the courts are on their side. A judge already called it this either." Another person writes, "The physical copies are artifacts that need to be preserved, not destroyed. We
[07:35] careless, and thoughtless everywhere you look." Another person eyes, "Burning of the Library of Alexandria 2.0." And another user writes, "There's no reason effing evil. Share the books with the world digitally, sure, but then preserve
[07:49] read them after all our technology fails in a solar flare." And a quote retweet "Honestly, this is a crime against humanity. The worst book ever written permanently destroyed for whatever reason, let alone feed AI models. We do
[08:05] not hate AI enough. Action must be taken now." And I agree. This is extremely heartbreaking. And honestly, we needed AI legislations yesterday. Because the fact that this is allowed and that fact that this was ruled as fair use just
[08:21] seems absolutely bonkers. And again, you might be wondering, "Well, how is it absolutely off the rocker, and I agree. Well, according to the lawsuit, it says library copies to digital library copies was transformative under fair use. Yep,
[08:36] everybody's writing, but as long as you make it digital, that's fine. So, what? Somebody has a painting, and then I scan mine. That's fair use. That sounds absolutely absurd. I don't even know how
[08:48] this ruling went through. But this final comment I really resonate with on purchase a DVD and digitalize it, that's copyright infringement. But if an AI company buys my book and scans it to train their system, that's perfectly
[09:01] fine? Once again, corporations are allowed to play by their own rules. I am honestly on the floor, and I want to have hope for the future in humanity, this, it makes me so bummed out. Truly, truly, truly. So, at the end of the day,
[09:15] Because this for me personally was very sad to read as somebody who loves books, and I truly, truly believe in that books are preserving history, and that they should not be destroyed and chopped up to what? Have AI train on it for selfish
[09:29] capitalistic greed? We do not need more AI models, truly. I don't think so. But thoughts with respect in the comment section down below. And to answer your "Kat, why do you look like a lizard?" Well, on a complete side note, I went to
[09:44] warped tour and made myself the cutest crochet sweater ever. I thought this was super cute. Well, apply your sunscreen, kids, because I didn't and now I have a knows how long and I really hope this fades soon cuz it looks so stupid. But
[09:59] appreciate it if you like and subscribe. Apply your sunscreen. I have a vlog channel where I'm posting vlogs on there and I have two new vlogs, one about a day in the life and then one about an acting gig that I did. So yeah, it would
[10:11] there or I stream on Twitch three days a week, so catch me there as well. So I hope to see y'all there or in another one of my videos. Peace.
⚡ Saved you 0h 10m reading this? Transcribe any YouTube video for free — no signup needed.