- 42comments
- 19comments
- 46comments
- 7comments
- 85comments
- 41comments
- 232comments
- 351comments
- 20comments
- 33comments
- 740comments
- 67comments
- 2comments
- 31comments
- 166comments
- 56comments
- 14comments
- 1comments
- 200comments
- 8comments
- 182comments
- 38comments
- 308comments
- —discuss
- 184comments
- 42comments
- 11comments
- 111comments
- 26comments
- 56comments
Could this be for AI training data?
It’s on the headline if you visit the link.
A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.
Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.
The value in these AI companies will be more in their proprietary training data than the models.
The sad part is, imagine how positive this could be if the scans were made available to the public.
Might add insult to injury for the publishers/authors though?
A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?
OCR Scanned for training, then tossed away or burnt. Great for nature.
Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.
It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.
How does one "consume knowledge"?
Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.
That's not consuming knowledge, that's consuming cellulose and ink.
Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.
Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.
I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.
Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.
Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.
What do you think happened to all these used books before the AI companies showed up?
Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?
Writing style maybe?