I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway.
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.
The internet made this even worse by celebrating people who only have to say the right thing.
Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices.
No, that’s the hyperbolic reaction clickbait wants from you.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
A small change in the copyright law would fix this problem. Something like:
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?
How does this help anything, except create more work to throw in the trash?
If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.
1 point by jonhohle 0 minutes ago | edit | delete [–]
You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.
That’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.
They're not destroying rare manuscripts or incunables.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
> a digital copy with the ability of doing millions of copies is stored somewhere
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
>Plus having the info part of a LLM makes it immediately available to literally billions.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction?
I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
> permanently locking human knowledge inside private corporate servers
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
It’s just another lie of the type these threads tend to be filled with nowadays.
Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.
I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
Doubtful. A robust protocol would be to scan several of each edition (to ensure no scanning errors), and scan each edition. Then too, these books are being purchased in lots with accidental duplicates, and all the major labs are doing it. So we're likely talking about tens of each book. For rare books - anything over a couple of hundred years old or small print runs either, that could well be most or even all copies. This wouldn't be immediately noticed either, especially if the books aren't currently considered noteworthy or well known.
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
>If they are truly rare, then they are likely not valuable,
Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).
I ask "more clean air", you answer "why, did you deserve it, having clean air is not free, there's a real cost"
Somebody else asks "We need more accessible energy", you answer "why we bother with your needs, you're not efficient, energy belongs to more efficient purposes, you can live without that much energy".
I can continue with more, but hope you got another viewpoint.
Btw, I'm disgusted there are "humans" like you in existence. We definitely don't share the same cultural ancestry, and I hope ours will prevail at the end, rather than cold blooded, mechanical "brains" like those of your kind.
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.
So old books are valuable because old paper is valuable? I understand why the Gutenberg bible might be valuable but do we really need thousands of mass paperback novels?
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
archive.org is an incredible blessing that I never really respected enough until the last few years. I have bookmarks going back to the 90s, and for some reason, starting in 2015 or so, sites just started disappearing. I estimate at least 20% of my bookmarks are 404 now.
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
> The main question is why aren't they leaking it to AA themselves?
How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.
They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.
> The main question is why aren't they leaking it to AA themselves?
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
Exactly. They set up this operation specifically to comply with the letter of copyright law and defend against publisher law suits. “Leaking” to Anna’s Archive is the last thing they’re going to do.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
The other thing they could be doing is using some of that lobbying money to try to reform copyright law to allow them to release the scans that are still covered.
It would likewise earn goodwill from a lot of people.
They could be, however that would weaken their position as the sole source of knowledge which seems to me is half the purpose of their book destroying initiative.
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.
Which is really why the outrage cycle over Anthropic's actions is largely misplaced.
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.
The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.
As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.
But scanning the books also helps those AI companies because ultimately
they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
I guess they just cut the spines to speed scan the pages with a machine and then the resulting pile of papers is no longer worth anything, so they just dispose of it.
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
Is there any indication that Anthropic is destroying books from the 18th century? Even the BBC quote is a conditional. Emphasis added:
> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."
I agree with the sentiment of course but it is really a huge IF they are doing that.
IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.
I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?
The British Library has a large quantity of books that have literary, academic, cultural or bibliopolical significance but are too niche to ever see print new print runs.
This idea that there can only be merit in a work if it's commercially viable is incredibly ignorant, and if that becomes the standard for whether a work is preserved or not, we stand to lose a great deal of our cultural heritage.
The problem is that you probably do little research or read very few old books.
There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.
You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.
Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.
Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.
I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.
> We've been seeing that headline for a few weeks now and I really don't understand the problem.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
> because copyright law forces them to do stupid things.
Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.
> because copyright law forces them to do stupid things
This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?
Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?
Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.
It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯
I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.
But why do they have to do this at all? If it's a shit idea, and you cannot do something without negative side-effects of it, can't you just not do it? Why these companies absolutely have to do this?
Yeah, which is completely fine. There's a major difference between:
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway.
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.
The internet made this even worse by celebrating people who only have to say the right thing.
You’re equating something being ubiquitous and affordable with it not being valuable.
Yes, there are plenty of books, many were printed, many have lasted a very long time (plenty over 100 years!).
That says more about the success and utility of the technology than it does about whether individual books should be shredded.
Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
They are destroying them because it’s easier to scan them if you slice the binding.
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices.
What happens when the books don't sell?
The average lifespan of a library book is 26 loans. For a popular book that could be less than a year.
No, that’s the hyperbolic reaction clickbait wants from you.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
A small change in the copyright law would fix this problem. Something like:
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?
How does this help anything, except create more work to throw in the trash?
I love this idea. At least make them turn it into some form of a public good.
So now the Library of Congress has to manage all these submissions whenever someone scans something? How do you even go about enforcing such a thing?
If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.
So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)
Selling it is not allowed. The inability to sell it does not require its destruction.
But copyright law does, because otherwise you have two copies.
This isn't theoretical, AI companies have finished lawsuits about this and this was the ruling.
1 point by jonhohle 0 minutes ago | edit | delete [–]
You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.
Citation please? Bartz v Anthropic seems pretty clear, see also Authors Guild v Google and RIAA v Diamond Multimedia.
More importantly: Once Anthropic is gone, all knowlege is lost.
It will probably be actioned off in the bankruptcy proceedings.
That’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.
[delayed]
Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.
Obviously. Copyright infringement is settled law.
> working towards the benefit of humanity is not an exclusive right / domain of theirs
That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.
If they were just chopping them up without scanning them first, then sure, but I think that scanning books and burning books are polar opposites
Burning books and destroying them in a way that nobody else can access the content anymore is a distinction without a difference.
They're not destroying rare manuscripts or incunables.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
Yep. I was personally misinformed the same way at some point. Never expected evaporative cooling to be so popular, but it is.
> This is just part of a CCP-aligned moral panic, along with the water use nonsense
What was nonsense about water use?
It was never that dramatic, and it’s declining day by day. It’s a panic over a real but small problem.
Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.
Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
The outcry over data centers using a fraction of the water used for things like golf courses or growing alfalfa in a desert.
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
we also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up
Nobody will ask you, they'll be taken from you, in case you missed what happens around.
The age-old cry of the aging population, faced with tech that didn’t exist when they were young. Turn back time!
I’m reminded of the screeds about the dangers of novels.
They have already gotten to Archive.org. Books that were available to borrow are no longer 'available'. SMH
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
> a digital copy with the ability of doing millions of copies is stored somewhere
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
> somewhere were we can't access it
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
>Plus having the info part of a LLM makes it immediately available to literally billions.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
> Is it "locked"? Yes, by copyright laws,
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
> permanently locking human knowledge inside private corporate servers
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
> History tells us that very few "permanent" situations are truly permanent.
If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.
Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
Where did you see they tend to buy all the copies? This comment is the first time I've heard of this.
I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?
It’s just another lie of the type these threads tend to be filled with nowadays.
Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.
I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them.
But where did you hear that they’re buying “all copies”? And to what end?
This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
what is your source?
Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
Somehow I don't think they're looking for the books that have been copied over and over.
First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
Doubtful. A robust protocol would be to scan several of each edition (to ensure no scanning errors), and scan each edition. Then too, these books are being purchased in lots with accidental duplicates, and all the major labs are doing it. So we're likely talking about tens of each book. For rare books - anything over a couple of hundred years old or small print runs either, that could well be most or even all copies. This wouldn't be immediately noticed either, especially if the books aren't currently considered noteworthy or well known.
The dissonance when a cause you support is loudly represented by disingenuous types.
At some point they’ll hit on data center water consumption as yet another reason to support the cause.
You ask "Why destroy physical books?"
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
>If they are truly rare, then they are likely not valuable,
Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
Libraries tend to regularly destroy books as well. And they aren't scanning them first either.
Doesn't that make them even worse?
"State libraries" such as the Library of Congress. (Which is not regularly destroying books, AFAIK.)
Good point, why waste time deciphering old badly burnt scrolls when anything worthwhile should have been preserved.
If the Romans had the printing press we'd have a lot more of those scrolls.
You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).
> If they are truly rare, then they are likely not valuable
"likely" being the keyword here, what about heavily censored books?
I ask "more clean air", you answer "why, did you deserve it, having clean air is not free, there's a real cost" Somebody else asks "We need more accessible energy", you answer "why we bother with your needs, you're not efficient, energy belongs to more efficient purposes, you can live without that much energy". I can continue with more, but hope you got another viewpoint. Btw, I'm disgusted there are "humans" like you in existence. We definitely don't share the same cultural ancestry, and I hope ours will prevail at the end, rather than cold blooded, mechanical "brains" like those of your kind.
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.
Same can be said about many books.
Ask: valuable to whom and why?
- Reader: narrative
- Collector: scarcity of the physical artifact
- AI Company: language samples (quantity, variety), facts
So old books are valuable because old paper is valuable? I understand why the Gutenberg bible might be valuable but do we really need thousands of mass paperback novels?
I agree with you for the same reason I think McDonald's is the best restaurant in the world!
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
archive.org is an incredible blessing that I never really respected enough until the last few years. I have bookmarks going back to the 90s, and for some reason, starting in 2015 or so, sites just started disappearing. I estimate at least 20% of my bookmarks are 404 now.
archive.org itself reminds me of the Library of Alexandria. It is irreplaceable, and this is a disaster waiting to happen.
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
> The main question is why aren't they leaking it to AA themselves?
How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.
They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.
> The main question is why aren't they leaking it to AA themselves?
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
Because that’s illegal.
Exactly. They set up this operation specifically to comply with the letter of copyright law and defend against publisher law suits. “Leaking” to Anna’s Archive is the last thing they’re going to do.
The illegal part is (or ideally should be) using the books in their training. The morally right action is to make it public afterwards.
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
PR only cares that the company is percieved as unstoppable & inevitable.
If their PR teams did have a response to these actions, it would be something to the effect of:
"would you rather china destroy all the books and gatekeep the knowledge?"
The other thing they could be doing is using some of that lobbying money to try to reform copyright law to allow them to release the scans that are still covered.
It would likewise earn goodwill from a lot of people.
They could be, however that would weaken their position as the sole source of knowledge which seems to me is half the purpose of their book destroying initiative.
That's why the "but it's illegal" propaganda is so unconvincing for me. It's really "but it's illegal (and we want to keep it that way)".
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.
Which is really why the outrage cycle over Anthropic's actions is largely misplaced.
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
Value is entirely in the eye of the beholder - money itself only has value because enough of us agree that it does.
How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
Anna's Archive merely indexes copyrighted content. It doesn't host.
Project Unica is an initiative by the University of Illinois libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
What often gets missed is that they are buy one physical copy and turning it into a digital copy.
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
> turning it into a digital copy.
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
> Or are they just using it for training
You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.
They are. They can’t release it; that would be copyright infringement. But they’re absolutely planning to make further use of the book later.
The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.
As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.
But scanning the books also helps those AI companies because ultimately they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
I guess they just cut the spines to speed scan the pages with a machine and then the resulting pile of papers is no longer worth anything, so they just dispose of it.
During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions
They did rather bring that on themselves though
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
Since when books have become a supply limited asset ?
Try and read a book that’s been burnt and find out
Ever since someone dredged up this old nothingburger of news from 2024 and made it the latest outrage bait.
We've been seeing that headline for a few weeks now and I really don't understand the problem.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
https://www.bbc.com/news/articles/cp3rprx2wl4o
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
Is there any indication that Anthropic is destroying books from the 18th century? Even the BBC quote is a conditional. Emphasis added:
> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."
I agree with the sentiment of course but it is really a huge IF they are doing that.
IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.
I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?
Rare and out of print does not mean important or valuable.
It doesn't automatically mean pointless or worthless, either. And it's too late to judge either way once it's been destroyed.
Usually it means the opposite. Books that are old and valuable tend to be out of copyright, so they do see new printing runs.
The British Library has a large quantity of books that have literary, academic, cultural or bibliopolical significance but are too niche to ever see print new print runs.
This idea that there can only be merit in a work if it's commercially viable is incredibly ignorant, and if that becomes the standard for whether a work is preserved or not, we stand to lose a great deal of our cultural heritage.
Presumably, being from the 18th Century, copyright law wouldn't apply?
Detestable if they're doing it anyway to prevent competitors getting hold of it.
The problem is that you probably do little research or read very few old books.
There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.
You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.
Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.
Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.
I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
I’m confused - are you intending to say Anthropic and Google are buying ALL the books?
I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.
The problem is we don't know what we're losing, due to lack of transparency.
> We've been seeing that headline for a few weeks now and I really don't understand the problem.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs.
At least there will be a copy left for us. The AI companies won't share these books in their original form.
Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things.
Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.
The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place
> because copyright law forces them to do stupid things.
Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.
Copyright law didn't force them to torrent terabytes of books what they did and got caught doing so.
I'm not convinced they destroy the books to obey the law, lol.
Google probably wanted to sell the whole of Google Books.
> because copyright law forces them to do stupid things
This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?
Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?
Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.
It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯
I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.
But why do they have to do this at all? If it's a shit idea, and you cannot do something without negative side-effects of it, can't you just not do it? Why these companies absolutely have to do this?
Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.
Yeah, which is completely fine. There's a major difference between:
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
I much prefer B.
B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to.
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.
[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...
[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court