Why are AI companies buying up old books just to shred them?
To train on books, AI companies need text in digital form — raw characters a model can process. Physical books must first be scanned, then converted to text using OCR (Optical Character Recognition) software. The problem is that scanners need books to lie flat, which often means cracking or damaging the spine.
Some labs have gone further, physically destroying books by cutting off their spines so pages can be fed through high-speed document scanners quickly. This is controversial because many of the books targeted are rare or out-of-print editions that exist in very few copies.
The deeper issue is one of data scarcity. As models have grown larger and hungrier for training data, the easily available digital text on the internet has largely been used up. Physical libraries represent one of the last large reserves of human knowledge not yet in digital form — making them a valuable, if ethically fraught, resource.