Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. F-Droid 2.0: A New Chapter for Android Freedom(f-droid.org)
    66comments
  2. GitHub has not removed malicious imitation software after 3 weeks(successfulsoftware.net)
    22comments
  3. I Have a Confession: I Built This Site with AI – Please Forgive Me(dynamicallytyped.org)
    25comments
  4. Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering(blog.madhukaraphatak.in)
    30comments
  5. The science of Monkey Island: can grog dissolve a metal mug that fast?(jgeekstudies.org)
    15comments
  6. Two-tier encryption in the UK(macanorak.com)
    226comments
  7. Enjoy Every Sandwich(bradmontague.substack.com)
    57comments
  8. Nokia Design Archive (2025)(aalto.fi)
    86comments
  9. Linux support is coming to Snapdragon X2 Series(qualcomm.com)
    238comments
  10. LinkedIn wins court order blocking mass scraping of user data(therecord.media)
    11comments
  11. Ideas on modernizing the open-source desktop(lwn.net)
    385comments
  12. QR Codes That Route to the Appropriate App Store(matthuggins.com)
    6comments
  13. B5-BJ2 – Ice Cream Barges – Concrete Ship Constructors(thecretefleet.com)
    discuss
  14. RAM: the forgotten history (2024)(coredump.cx)
    2comments
  15. ArXiv receives multiyear commitments to support it as an independent nonprofit(arxiv.org)
    37comments
  16. What Is RLCD? The Secret Behind Jev(di-zhang-llm.github.io)
    3comments
  17. WaveDigger: Dig into wireless signals to discover their physical locations(github.com/christianrowlands)
    2comments
  18. When the Debugger Lies(danielmangum.com)
    16comments
  19. The newest ESP32 can run Linux and it's getting close to a Raspberry Pi(xda-developers.com)
    69comments
  20. VSCode's SSH Agent Is Bananas (2025)(fly.io)
    185comments
  21. Coulomb's law remains tricky to test at home(chillphysicsenjoyer.substack.com)
    9comments
  22. Contrastive Language Models(contrastive-lm.notion.site)
    40comments
  23. The "Windows XP Box" (2003)(mini-itx.com)
    43comments
  24. The Year of Internal Tools(geocod.io)
    12comments
  25. Fixing the Portobello Police Station Clock(pointinthecloud.com)
    112comments
  26. Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest(github.com/nestrilabs)
    59comments
  27. Hackers influence ChatGPT and Gemini to direct users to scam centers(medium.com/arielsimon)
    33comments
  28. Humans Are Reading Your ChatGPT Chats, Lawsuit Claims(openclassactions.com)
    7comments
  29. Owners mourn spoiled food after firmware update bricks Samsung smart fridges(arstechnica.com)
    207comments
  30. Why 'What's Opera, Doc?' looks like that(animationobsessive.substack.com)
    22comments

Japanese used bookstores see 5x sales surge as books are being bought by the ton

58 pointsby 2h agotomshardware.com
70 comments
1h agoHN ↗

It’s on the headline if you visit the link.

1h agoHN ↗

A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.

1h agoHN ↗

Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.

43m agoHN ↗

I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?

36m agoHN ↗

We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.

22m agoHN ↗

This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.

1h agoHN ↗

The value in these AI companies will be more in their proprietary training data than the models.

1h agoHN ↗

The sad part is, imagine how positive this could be if the scans were made available to the public.

1h agoHN ↗

Might add insult to injury for the publishers/authors though?

1h agoHN ↗

A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?

OCR Scanned for training, then tossed away or burnt. Great for nature.

1h agoHN ↗

Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.

46m agoHN ↗

Unfortunately copyright law does not have a squatters-rights exception

32m agoHN ↗

I've believed for years that it really should have one, or at least a "you aren't selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it" regime.

1h agoHN ↗

It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.

57m agoHN ↗

Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.

52m agoHN ↗

That's not consuming knowledge, that's consuming cellulose and ink. If it were, then printing another copy of a book would be "producing knowledge".

BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.

1h agoHN ↗

Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.

1h agoHN ↗

Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

51m agoHN ↗

Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.

45m agoHN ↗

My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.

Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.

We were taught respect for books, but not everything is worth preserving.

25m agoHN ↗

I didn't go this year but my local town library has a book sale where books go for something like $10/bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I'm sure a huge number get pulped.

41m agoHN ↗

Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.

35m agoHN ↗

That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?

32m agoHN ↗

Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

Training an LLM on copyrighted works is not illegal.

This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

19m agoHN ↗

That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

14m agoHN ↗

That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.

13m agoHN ↗

They legally cannot scan them in entirety either but they are.

12m agoHN ↗

Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.

39m agoHN ↗

Why do they shred them at all instead of donating or selling them again?

38m agoHN ↗

A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.

25m agoHN ↗

The judge ruled that format shifting was fair use. The fair use argument was enhanced because the original was destroyed. Hypothetically, one could do only the format shifting and keep the original - but the original could not be sold or donated or otherwise given away because then the format shifted copy would not be fair use.

I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.

36m agoHN ↗

My understanding is they cut off the binding for scanning. They'd have to resell it donate by the page

13m agoHN ↗

They have to destroy the copy either way for it to be fair use (according to Bartz v Anthropic). Cutting the bindings off is already the faster method to scan, and once you're required to pulp the original anyway, it becomes a no-brainer.

31m agoHN ↗

I have a similar thought too each time I walk past piles of discount books.

I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

22m agoHN ↗

Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

Yeah, now instead old books being turned into toilet paper, we'll get turned into toilet paper.

And Sam Altman will become richer than God, and isn't that what really matters?

But don't worry! You'll still have access to ChatGPT until your savings run out.

16m agoHN ↗

Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?

11m agoHN ↗

A lot of these books are reference titles that are outdated to the point of uselessness and/or weren't that great/interesting when they were new. Not many people have interest in or use for textbooks from the 50s.

1h agoHN ↗

Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.

1h agoHN ↗

What do you think happened to all these used books before the AI companies showed up?

45m agoHN ↗

Not quite:

But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.

https://www.offthebeatenshelf.com/blog/pulp-fiction-is-real

35m agoHN ↗

TIL that's where the term "pulp fiction" comes from!

23m agoHN ↗

That's an American take. In Latin America they sit in boxes and you can see it, because the pages will have been discolored after so many years of sitting somewhere waiting to be sold. They're only recycled if they're given away at book exchange events and no one wants them.

41m agoHN ↗

Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.

24m agoHN ↗

The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.

1h agoHN ↗

Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?

46m agoHN ↗

Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.

46m agoHN ↗

Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.

45m agoHN ↗

The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

Books are more likely to be about a specific topic or story or time or setting and be more information dense

45m agoHN ↗

They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.

45m agoHN ↗

It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

34m agoHN ↗

the information density of published books is a lot higher than most internet text

I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

31m agoHN ↗

They're targeting non-fiction books, so romance novels would be out.

17m agoHN ↗

Is that a policy change after o4 got a little out of hand?

33m agoHN ↗

The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.

29m agoHN ↗

Also the ability to set the training input limit in the past could be useful.

40m agoHN ↗

I'm guessing the more integration tables they consume, the better they become at integration.

I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

35m agoHN ↗

but now a few thousand books are what is needed to run a successful AI company?

They're scanning millions of books.

It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

33m agoHN ↗

I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.

30m agoHN ↗

why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

21m agoHN ↗

Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

12m agoHN ↗

They aren't but these companies have more money than sense.

38m agoHN ↗

Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.

27m agoHN ↗

You are a moron: vinyl records came and went in about 25 years, books have been the engine of human progress for at least the last 3 millenia, the books being burned by these misanthropic lunatics are not available in any other medium, this is not about fascination with some particular mediumn of transmission, it's about the contents

23m agoHN ↗

"the engine of human progress"? I think you mean capitalism.

the books being burned by these misanthropic lunatics are not available in any other medium

prove it, name one title

this is not about fascination with some particular mediumn of transmission

it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.