AI companies destroy physical books – let's scan rare books before it's too late

110 points - today at 10:05 AM

Source

Comments

cladopa today at 12:32 PM
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.

By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.

If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.

ziyadb today at 12:11 PM
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

CamelCaseName today at 12:25 PM
You ask "Why destroy physical books?"

I ask "Why save physical books?"

If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.

I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.

pmoriarty today at 12:23 PM
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.

After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.

ZoomZoomZoom today at 12:09 PM
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
ryandvm today at 12:27 PM
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
carlosjobim today at 12:31 PM
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
ForHackernews today at 12:23 PM
Project Unica is an initiative by the University of Illinois libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
RcouF1uZ4gsC today at 12:22 PM
What often gets missed is that they are buy one physical copy and turning it into a digital copy.

They have done zero to destroy the durability. In fact, it’s probably more durable.

If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.

m00dy today at 12:16 PM
Since when books have become a supply limited asset ?
eulgro today at 12:03 PM
We've been seeing that headline for a few weeks now and I really don't understand the problem.

Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.

Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.

So what's the problem here exactly?

Also from the article:

> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.

I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.

warkdarrior today at 11:55 AM
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.

Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.

grammarisking today at 12:25 PM
[dead]
majke today at 11:53 AM
[dead]