AI Wants Your Old Books
Artificial intelligence companies are buying millions of older printed books to train their latest models because they are increasingly viewed as one of the last large sources of trustworthy, human-written text, highlighting a growing problem for the AI industry as the internet becomes saturated with machine-generated content.
Why Are AI Companies Buying Old Books?
For years, AI developers have relied heavily on vast quantities of online content to train increasingly capable language models. However, that approach is becoming more difficult because the web itself is changing.
Much of today’s online text is now created, edited or influenced by AI. As a result, companies developing the next generation of models are looking for sources of information that pre-date the arrival of generative AI and therefore contain only human-authored material.
One company hoping to meet that demand is ISBNdb, which has traditionally helped booksellers, libraries and distributors manage book inventories. It now offers AI companies the ability to acquire physical books in bulk specifically for AI training.
Its straightforward message is: “The world’s best AI training data is sitting on a shelf.”
The company argues that books are “dense, edited, authoritative” and contain structured knowledge that cannot easily be replicated by scraping the modern web.
Why Older Books Have Become So Valuable
The key factor here is when the books were published. For example, ISBNdb says books printed before 2022 are especially valuable because they were produced before large language models became widespread and therefore cannot contain AI-generated text.
That matters because researchers are becoming increasingly concerned about “model collapse”, a phenomenon where AI systems trained repeatedly on synthetic content gradually become less accurate, less diverse and more prone to errors.
Instead of learning from original human knowledge, successive generations of models risk learning from earlier AI outputs, reinforcing mistakes and reducing the overall quality of future models.
Printed books provide a rare alternative. Unlike websites, they cannot be quietly rewritten, updated or flooded with AI-generated material after publication. They represent fixed records of human knowledge created before today’s generative AI era.
ISBNdb also argues that these books avoid another growing problem, namely deliberate attempts by some authors to poison AI training data.
A New Front In The Copyright Battle
The growing demand for printed books also reflects the continuing legal and ethical debate surrounding AI training.
Rather than scraping copyrighted material from the internet, many AI companies are now purchasing physical books before scanning them into digital form.
ISBNdb argues this provides a clearer chain of custody and reduces legal uncertainty.
The company says: “Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received.”
The approach also follows recent US court decisions that distinguished between books acquired legitimately and pirated copies used for AI training, although wider copyright disputes between publishers, authors and AI developers remain far from settled.
Many publishers are also becoming more cautious about where their books ultimately end up, particularly as demand from AI companies continues to grow.
The Cost Of Creating Clean Data
Accessing millions of printed books is not as simple as placing them on a scanner. High-volume book digitisation typically involves removing the spine so that individual pages can pass rapidly through automated scanning equipment. The process is considerably faster and cheaper than scanning books intact, but it permanently destroys the physical copy.
ISBNdb openly acknowledges the reputational challenge this creates. As the company says: “The optics problem is real.”
It also notes that headlines about AI companies destroying millions of books are unlikely to generate public sympathy, even suggesting that clients present the process as digitally preserving knowledge rather than destroying physical books.
The practice has also raised concerns among booksellers, particularly where specialist, foreign-language or low-circulation books may become increasingly difficult to obtain after being converted into training data.
Human Knowledge Becomes A Scarce Resource
Perhaps the most interesting aspect of the story is what it says about the future economics of AI. For years, computing power, semiconductor chips and data centres have been viewed as the industry’s most valuable resources.
Increasingly, however, genuinely human-created information may be becoming just as important.
The rapid growth of AI-generated articles, websites, social media posts and marketing content means reliable human-written material is becoming harder to identify. As a result, books, journals and other carefully edited publications created before the AI boom are taking on new commercial value as trusted sources of knowledge.
Ironically, the AI industry now finds itself paying to recover information from a period before AI itself transformed the internet.
What Does This Mean For Your Business?
For businesses, this story highlights that data quality is rapidly becoming just as important as data quantity.
Organisations developing or deploying AI systems should increasingly expect questions about where their training data comes from, how reliable it is and whether it contains synthetic or manipulated content. Provenance, authenticity and transparency are likely to become competitive advantages as businesses seek AI systems that can be trusted.
The story also reinforces the continuing value of high-quality human expertise. As AI-generated content becomes increasingly common, carefully researched reports, specialist publications, books and proprietary business knowledge may become some of the most valuable assets organisations possess. Rather than reducing the importance of original human insight, the AI era may ultimately make it considerably more valuable.
The same principle applies within organisations themselves. Businesses that have spent years building well-written documentation, technical manuals, research and training materials, together with other original knowledge, may now possess a valuable resource for developing trustworthy internal AI systems. As organisations increasingly deploy AI using their own information, the quality of that underlying knowledge base is likely to become just as important as the sophistication of the AI model itself.



