Recent Test Confirms What Happens to the Thousands and Thousands of Used Books Being Purchased Anonymously
- by Michael Stillman
VGT Logo shows alligator consuming, literally, a book.
404 Media recently performed a test to tell us something we already knew, or at least suspected. Still, it's valuable to have more evidence to confirm our suspicions. What was known is that someone (or more) was buying huge quantities of old books. The numbers were unheard of before, thousands or tens of thousands. The range of titles didn't make any sense. They didn't relate to a specific collection. They were random.
What 404 Media did was to get a bookseller who filled one of these thousand-book orders to put a tracking device on one of them. Now, 404 Media could track that book and see exactly where it went. What they found is that it went to an Amazon warehouse in Las Vegas, which goes by the name of VGT3. Most likely, you have never heard of it, and that undoubtedly is fine with Amazon.
They asked a few workers off the record what was going on. They confirmed what we all suspected. They found that a large numbers of books are being copied, but then destroyed, to train AI services. A lawsuit earlier brought out that Anthropic had copied a lot of books for their Claude AI service. This latest example relates to similar activities being carried out by Amazon, but also focused on the ugly underbelly of the process. Those tens or hundreds of thousands of books that are copied are being destroyed in the end.
What got Anthropic into some trouble, though not that much considering the amount of data they obtained, was illegally copying digital books. They used large datasets of books that had been stolen by others. A California court determined the fact that someone else stole the digital copies did not make it all right for Anthropic to use them. The court also made it obvious that as long as the company copying the books purchased or obtained them legally, copying was all right. Such copying, it decided, did not constitute a copyright violation. The only violation was using stolen books.
This is only one decision in one court, which could be appealed, but some companies may see this as possibly being a safe harbor. Instead of stealing books, they are being purchased. Based on that court decision, this would make copying their data for use in AI searches legal. However, copying many thousands of books can be an expensive, time-consuming business. One thing they quickly learned is that destructive copying is cheaper. The cheapest, fastest way to copy a book is to slice off the spine and feed the pages into a copier. Naturally, the book is now trash and is tossed away.
When Google was copying books for Google Books, many were valuable books borrowed from libraries. They needed to return them, and return them in as good a condition as when they were borrowed. So, they carefully opened the books to each page and then put them on a copier, one page after the other. It was a slow process, but preserved the books. Since the AI searchers own these books, they are under no obligation to preserve them. Most are more recent books or ones that never were valuable, or later editions. The cost of each book is not that much and not worth the time and effort to worry about preserving them. However, the concept of destroying books is very offensive to many book lovers. It is the equivalent of book burning (which may be the fate of these ripped apart books after hauled off for trash).
While it was fairly obvious that this is what was happening to these books, no one was talking about it. The AI companies certainly didn't want to talk about the ugly side of their business. It's not exactly the kind of publicity these companies want. When 404 Media got a bookseller who filled one of these thousand-book orders to put a tracking device on one of those books, they found out exactly where it went. That turned out to be the VGT3 Amazon warehouse. 404 Media then asked a few VGT3 employees off the record what was going on inside, and they confirmed that the books were being copied, and destroyed in the process.
404 Media asked Amazon about the process, but only received a terse, nonspecific answer, "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."
It would be easy to beat up on Amazon. No one likes to see books destroyed. However, it is also necessary to point out what is being created. If you have a computer and an internet connection, which is most of us, you have undoubtedly used one of these AI search engines, ChaptGPT, Claude, Google Gemini, or another. They may not be perfect, but neither are books. The amount of information they provide at your fingertips is incredible. You have access to all of the information in these books and many other sources to answer your questions. The answers come instantly and you don't even have to pay for them (although someday they may ask you to help monetize their investment). They have become your own personal library, except your virtual library contains far more information than any physical one and you don't have to search through the many books yourself to find your answers. They do that for you. These search engines have become an integral part of our lives. We are not going back, any more than pollution issues are going to get us to ditch our cars and hook up a horse and buggy. It's too late for that. We need to make sure authors are fairly compensated, and rare and significant books have their physical copies preserved.