Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. F-Droid 2.0(f-droid.org)
    105comments
  2. Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design(github.com/devdotfast)
    6comments
  3. Two-tier encryption in the UK(macanorak.com)
    269comments
  4. Fearless SIMD v1.0(linebender.org)
    2comments
  5. WaveDigger: Dig into wireless signals to discover their physical locations(github.com/christianrowlands)
    5comments
  6. Toyota is taking the Corolla electric(electrek.co)
    33comments
  7. Nokia Design Archive (2025)(aalto.fi)
    96comments
  8. The forgotten battle of East Lansing(eastlansinginfo.news)
    —discuss
  9. Google’s Project Suncatcher to put ML infrastructure in space(blog.google)
    18comments
  10. GitHub has not removed malicious imitation software after 3 weeks(successfulsoftware.net)
    61comments
  11. B5-BJ2 – Ice Cream Barges – Concrete Ship Constructors (2023)(thecretefleet.com)
    3comments
  12. Book review: Is parallel programming hard, and, if so, what can you do about it?(ahelwer.ca)
    —discuss
  13. Experiencing writing at our recent Chinese calligraphy workshop(viewsproject.wordpress.com)
    1comments
  14. Linux support is coming to Snapdragon X2 series(qualcomm.com)
    244comments
  15. Apple iPhone 4 “Antennagate” Q&A (2010) [video](youtube.com)
    11comments
  16. Search – A small, fast WebKit browser for macOS(github.com/driceroland)
    —discuss
  17. Show HN: AgentRun: DSL to turn agents into workflows(github.com/parcha-ai)
    —discuss
  18. Motor Characterization for Small Running Robots (2016)(robot-daycare.com)
    —discuss
  19. Show HN: Treepeat – Code similarity detection using Tree-sitter(github.com/dsummersl)
    —discuss
  20. My Weird New Hobby: Wandering Around Tokyo on Google Maps(ahmedhossamdev.com)
    3comments
  21. Ideas on modernizing the open-source desktop(lwn.net)
    423comments
  22. Geothermal heat map of US hot springs(soakingsprings.com)
    3comments
  23. The science of Monkey Island: can grog dissolve a metal mug that fast?(jgeekstudies.org)
    22comments
  24. RAM: the forgotten history (2024)(coredump.cx)
    3comments
  25. ArXiv receives multiyear commitments to support it as an independent nonprofit(arxiv.org)
    38comments
  26. When the Debugger Lies(danielmangum.com)
    17comments
  27. Enjoy Every Sandwich(bradmontague.substack.com)
    80comments
  28. Coulomb's law remains tricky to test at home(chillphysicsenjoyer.substack.com)
    16comments
  29. The newest ESP32 can run Linux and it's getting close to a Raspberry Pi(xda-developers.com)
    87comments
  30. VSCode's SSH Agent Is Bananas (2025)(fly.io)
    190comments

Japanese used bookstores see 5x sales surge as books are being bought by the ton

63 pointsby 3h agotomshardware.com
83 comments
2h agoHN ↗

It’s on the headline if you visit the link.

2h agoHN ↗

A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.

2h agoHN ↗

Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.

1h agoHN ↗

I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?

1h agoHN ↗

We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.

1h agoHN ↗

This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.

2h agoHN ↗

The value in these AI companies will be more in their proprietary training data than the models.

2h agoHN ↗

The sad part is, imagine how positive this could be if the scans were made available to the public.

2h agoHN ↗

Might add insult to injury for the publishers/authors though?

2h agoHN ↗

A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?

OCR Scanned for training, then tossed away or burnt. Great for nature.

1h agoHN ↗

Why would they be tossed or burnt? Tons of discarded books are simply recycled like any other paper object (that's why people keep saying they were "pulped")

2h agoHN ↗

Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.

1h agoHN ↗

Unfortunately copyright law does not have a squatters-rights exception

1h agoHN ↗

I've believed for years that it really should have one, or at least a "you aren't selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it" regime.

2h agoHN ↗

It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.

2h agoHN ↗

Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.

2h agoHN ↗

That's not consuming knowledge, that's consuming cellulose and ink. If it were, then printing another copy of a book would be "producing knowledge".

BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.

2h agoHN ↗

Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.

2h agoHN ↗

Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

2h agoHN ↗

Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.

1h agoHN ↗

My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.

Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.

We were taught respect for books, but not everything is worth preserving.

1h agoHN ↗

I didn't go this year but my local town library has a book sale where books go for something like $10/bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I'm sure a huge number get pulped.

1h agoHN ↗

Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.

1h agoHN ↗

That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?

1h agoHN ↗

Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

Training an LLM on copyrighted works is not illegal.

This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

1h agoHN ↗

Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you'd think people would be happier about that...

38m agoHN ↗

Wait, if you’re talking about Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.

1h agoHN ↗

That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

1h agoHN ↗

That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.

1h agoHN ↗

They legally cannot scan them in entirety either but they are.

1h agoHN ↗

Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.

1h agoHN ↗

Why do they shred them at all instead of donating or selling them again?

1h agoHN ↗

A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.

1h agoHN ↗

The judge ruled that format shifting was fair use. The fair use argument was enhanced because the original was destroyed. Hypothetically, one could do only the format shifting and keep the original - but the original could not be sold or donated or otherwise given away because then the format shifted copy would not be fair use.

I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.

1h agoHN ↗

My understanding is they cut off the binding for scanning. They'd have to resell it donate by the page

1h agoHN ↗

They have to destroy the copy either way for it to be fair use (according to Bartz v Anthropic). Cutting the bindings off is already the faster method to scan, and once you're required to pulp the original anyway, it becomes a no-brainer.

57m agoHN ↗

The wording doesn't exactly say that. https://www.akingump.com/a/web/h6WFidTYyTnoNEPehPXMYu/auvY7D...

The relevant part of the ruling starts on page 27.

    For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
    The third fair use factor favors fair use for the purchased library copies converted from print to digital.

Fair use favors the destruction to avoid accidental surplus copying of the source material. It doesn't require it. If Anthropic put everything in a warehouse, they could keep it in the warehouse... but they couldn't do anything with them afterwards. They couldn't sell them as books or donate them as that would mean that it wasn't fair use and copyright infringement would have taken place once that additional copy was distributed again. At that point, fair use and economics both favor destroying the original.

If I make a DVD copy of an old VHS tape that I own, that's fair use format shifting. I can keep the VHS tape without issue. I cannot donate it to the library or put it out in a garage sale. When I got rid of my VHS player, I threw out VHS tapes too since they were of no use and only cluttered my shelf.

1h agoHN ↗

I have a similar thought too each time I walk past piles of discount books.

I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

1h agoHN ↗

Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

Yeah, now instead of old books being turned into toilet paper, we'll get turned into toilet paper.

And Sam Altman will become richer than God, and isn't that what really matters?

But don't worry! You'll still have access to ChatGPT until your savings run out.

1h agoHN ↗

Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?

1h agoHN ↗

Yes, you missed it. Perhaps you should read it again until you get it?

54m agoHN ↗

Again, do you have an actual objection to the subject of the article, or are you just mad it enriches people you don't like?

1h agoHN ↗

A lot of these books are reference titles that are outdated to the point of uselessness and/or weren't that great/interesting when they were new. Not many people have interest in or use for textbooks from the 50s.

2h agoHN ↗

Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.

2h agoHN ↗

What do you think happened to all these used books before the AI companies showed up?

1h agoHN ↗

Not quite:

But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.

https://www.offthebeatenshelf.com/blog/pulp-fiction-is-real

1h agoHN ↗

TIL that's where the term "pulp fiction" comes from!

1h agoHN ↗

That's an American take. In Latin America they sit in boxes and you can see it, because the pages will have been discolored after so many years of sitting somewhere waiting to be sold. They're only recycled if they're given away at book exchange events and no one wants them.

19m agoHN ↗

In central America for example I went to Panama's book fair and you could see old books for sale. There's a lot of fear among booksellers that a book will not sell, because they purely resell them so any losses fall entirely on booksellers. They do not have any partnerships with publishers

14m agoHN ↗

That only lasts 10-30yrs. They're cheap paper, they were never made to last.

1h agoHN ↗

Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.

1h agoHN ↗

The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.

58m agoHN ↗

Please provide proof for the billions of dollars claim, thank you.

2h agoHN ↗

Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?

1h agoHN ↗

Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.

1h agoHN ↗

Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.

1h agoHN ↗

The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

Books are more likely to be about a specific topic or story or time or setting and be more information dense

1h agoHN ↗

They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.

1h agoHN ↗

It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

1h agoHN ↗

the information density of published books is a lot higher than most internet text

I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

1h agoHN ↗

They're targeting non-fiction books, so romance novels would be out.

1h agoHN ↗

Is that a policy change after o4 got a little out of hand?

1h agoHN ↗

Booksellers noticed a huge uptick in non-fiction purchases, that is not the same thing as them not targeting fiction at all.

Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.

1h agoHN ↗

The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.

1h agoHN ↗

Also the ability to set the training input limit in the past could be useful.

1h agoHN ↗

I'm guessing the more integration tables they consume, the better they become at integration.

I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

1h agoHN ↗

but now a few thousand books are what is needed to run a successful AI company?

They're scanning millions of books.

It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

1h agoHN ↗

I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.

1h agoHN ↗

why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

1h agoHN ↗

Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

1h agoHN ↗

They aren't but these companies have more money than sense.

44m agoHN ↗

I think a book is like one completion of the mind behind it, so in a way, this is just distillation.

1h agoHN ↗

Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.

1h agoHN ↗

Before you buy the apologia that these books are not valuable or interesting: if they weren't valuable or interesting the AI labs would not be spending billions of dollars to purchase and scan them. Yes, everyone has a personal anecdote about pallets of garbage books but I have three insights for you that you may not have because you don't read or sift through pallets of garbage books:

1) The labs don't want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.

2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that's why the labs are starting to purchase by the pallet, it's because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.

3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don't get destroyed or thrown away.

Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don't expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.

59m agoHN ↗

I have sifted through everything from piled up junky independent book shops, to cast offs from research libraries, to dumpsters of end of life books. Most books really are not worth the paper they are printed on.

Most books of value don't get destroyed or thrown away.

Of value to who? Most used bookstores are boutiques that over-curate and will happily refuse or recycle books that they deem are inferior/irrelevant. This is a big reason why I prefer Half Price Books over most any other used bookstore, they sell most everything.

21m agoHN ↗

AI firms buying and destroying the sources of knowledge is a weird way to get us to Fahrenheit 451 but maybe the USA will get there before it switches to metric after all.