A Unique Multilingual Media Platform

AI Articles Uncategorized

Shredding Knowledge, Feeding AI: The Dark Side of Destructive Scanning

  • August 23, 2026
  • 6 min read
Shredding Knowledge, Feeding AI: The Dark Side of Destructive Scanning

The 21st century is a living testament to Miltonic sin. John Milton famously argued in his 1644 text, Areopagitica, that destroying a good book is almost as bad as killing a man. The mass annihilation of books by AI language models to fully satiate the intellectual appetite of Artificial Intelligence is a very peculiar but, most importantly, an alarming situation.

A portrait of John Milton and the title page of John Milton’s Areopagitica, 1644 speech defending freedom of the press and opposing censorship in England.

Anthropic, one of the leading AI companies behind the language model Claude, has been found to have spent millions of dollars buying books, then stripping them of their bindings, cutting and scanning them, and ultimately discarding the physical copies. These books were converted into digital files and stored in an internal library for AI research and training. This practice, commonly described as ‘destructive scanning,’ is now all over the news and social media for all the wrong reasons.

Who are these digital goons, and what do they want from us?

Anthropic PBC, a software firm founded by former OpenAI employees in January 2021, is the parent company of the AI software service called Claude. Claude is currently under significant public and media scrutiny for its controversial acquisition and use of millions of copyrighted books, including copies obtained from online piracy.

A picture of Claude’s logo, developed by Anthropic.

The firm is accused of purchasing copyrighted books from major book distributors and retailers without consulting the authors. The writers Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson came to the forefront by filing a lawsuit against Anthropic and alerting the media to the injustice they believed AI companies had inflicted on their creative work through copyright infringement and the denial of their intellectual property rights.

Following this incident, the three authors sued Anthropic in Bartz et al. v. Anthropic PBC, alleging that the company had used pirated copies of their books and other copyrighted works in connection with its AI systems. The case subsequently expanded into a class action covering hundreds of thousands of books. In July 2026, a federal judge granted final approval to a $1.5 billion settlement between Anthropic and the authors whose works were at issue in the case.

To make things more complicated, Anthropic assimilated these digital copies into a curated central library of its own and used copies obtained from sources including Books3, LibGen, and the Pirate Library Mirror (PiLiMi). These materials were used in the development and training of large language models (LLMs). In his 2025 ruling, U.S. District Judge William Alsup distinguished between books Anthropic obtained lawfully and books it obtained through piracy: he found that using lawfully acquired books to train its LLMs constituted fair use, while acquiring and retaining pirated copies in its central library did not.

Anthropic’s co-founder, Benjamin Mann, was found to have downloaded large numbers of books from shadow libraries. Court records indicated that he downloaded 196,640 books from Books3 and later millions more from LibGen, while Anthropic also obtained material from PiLiMi. The company later pursued lawful acquisition of books, including purchasing print copies that could be digitised.

Benjamin Mann, Co-Founder of Anthropic.

According to a 2025 federal court ruling, Anthropic purchased millions of print books, used them for its purposes, and discarded the physical copies. The books were destructively scanned—stripped of their bindings, cut, and converted into digital files. But nobody can really determine the final goal. Is it merely for training Claude’s AI language models, or is there an entirely hidden agenda here? The question still lingers: who owns, censors, and edits these materials when they are converted into machine-readable data?

Destructive scanning has been adopted in certain large-scale digitisation initiatives because it allows books to be processed more quickly and yields uniformly flat, high-resolution images well suited to optical character recognition (OCR). Preservation-oriented digitisation generally prioritises methods that leave the physical book intact; however, this is not the case here due to the ethical dilemma involved.

The Internet Archive, a digital library widely recognised for its efforts to preserve web content and digitise books, has traditionally emphasised preservation and access in its digitisation work, making the preservation of the physical book an important consideration in contrast to destructive scanning. Unlike a process in which the physical copy is destroyed after digitisation, preservation-oriented scanning seeks to retain the original artefact.

Artificial intelligence has propelled the rise of computational assessment models, in which readers are perilously dependent on AI platforms, thereby altering a user’s innate relationship with words, vernacular, and language that they once inculcated through physical reading.

An image of Humans engrossed in their phones while AI-powered robots engage in creative activities.

In Walter Benjamin’s famous essay, The Work of Art in the Age of Mechanical Reproduction, he poses an important question to readers: “What happens when technological reproduction separates an artwork from its unique material existence?” The answer to the question he asked in his 1936 essay is now met with AI-generated responses in the year 2026.

The cover page of Walter Benjamin’s 1936 essay, The Work of Art in the Age of Mechanical Reproduction.

Humans of this day and age might argue, from a Marxian point of view, that what is happening now is the commodification of knowledge and literature, and that books are in transition towards becoming an input in technological production. But the important question to ask ourselves first is whether we really want to let this pass off as a technological leap in human history. We should be worried about how ruthlessly Anthropic decided that it was easier to scan and demolish books rather than to deal with the “legal/practice/business slog” of obtaining them through conventional licensing and acquisition channels.

A crowd experiencing the world through Virtual reality Headsets.

The rise of generative AI suggests that, as a community, we are ready to surrender our creative impulses to high-performing (sometimes sporadic) synthetic media based on prompts. The cultural damage caused by book destruction gives AI enormous power to control, amend, and propagate narratives, thereby stunting human inventiveness.

About Author

Anena S

Anena S is a English literature graduate and writer with an MA in English Literature From The University Of Delhi. She is interested in researching and understanding how Literature, Cinema, Media, Culture, and Gender, influence and shape social narratives. She is currently an intern at the AIDEM.

Subscribe
Notify of
guest
1 Comment
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Raj Veer Singh

Anena S raises a deeply disturbing question: what happens when the pursuit of AI knowledge begins with the destruction of the very books that carry human memory? This is not merely about scanning technology—it is about who owns knowledge, who controls its digital transformation, and whether human culture becomes nothing more than raw material for machines. A powerful warning against turning knowledge into a disposable commodity.

Support Us

The AIDEM is committed to people-oriented journalism, marked by transparency, integrity, pluralistic ethos, and, above all, a commitment to uphold the people’s right to know. Editorial independence is closely linked to financial independence. That is why we come to readers for help.

1
0
Would love your thoughts, please comment.x
()
x