Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Show HN: Majority – Find new music and vote on your favorites(majority-eight.vercel.app ↗)
    discuss
  2. Warren Buffet's last Berkshire letter: "Father Time always wins." [pdf](berkshirehathaway.com ↗)
    discuss
  3. What Using AI Therapy Gets Wrong(emilylee293105.substack.com ↗)
    discuss
  4. Show HN: Jev beats Astra, fable, Opus at RF engineering task(twitter.com/cohavygal ↗)
    discuss
  5. An Empirical Study of Harness Design for Coding Agents(arxiv.org ↗)
    discuss
  6. The Trap of "Phantom TAM"(marketinvestigation.beehiiv.com ↗)
    discuss
  7. I Tried to Leave the Terminal Eight Times. The Ninth Stuck(thoughts.jock.pl ↗)
    discuss
  8. Germany turns on Brussels as Chinese car sales on track to exceed 1M in 2026(euronews.com ↗)
    1comments
  9. Vulkan Documentation Project(vulkan.org ↗)
    discuss
  10. Gabor Fields: Orientation-Selective Level-of-Detail for Volume Rendering(arcanous98.github.io ↗)
    discuss
  11. 12-Factor Agents – Principles for building reliable LLM applications(github.com/humanlayer ↗)
    discuss
  12. Handling 100M time series with Arrow, Parquet, and object storage(parseable.com ↗)
    discuss
  13. TLDR; what the heck is Jev?(jrzs.dev ↗)
    discuss
  14. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    discuss
  15. DJ Shadow looks back at "Entroducing" and other early work(msn.com ↗)
    discuss
  16. When the Debugger Lies(danielmangum.com ↗)
    discuss
  17. Putin envoy and far-right AfD prepare talks to get Russian gas back for Germany(reuters.com ↗)
    discuss
  18. Maker's Schedule, Manager's Schedule(paulgraham.com ↗)
    discuss
  19. Clocky: Alarm clock on wheels that runs away from you(clocky.com ↗)
    discuss
  20. Show HN: Bastionskill – scan an AI agent skill for malicious code
    discuss
  21. The Return of the Utah Teapot – Siggraph 2026 [video](youtube.com ↗)
    discuss
  22. AI eyes in the sky: New satellites and AI transforming wildfire detection(theguardian.com ↗)
    discuss
  23. Sales calculation tool for Excel price lists
    discuss
  24. Qbix Server – PHP 100× faster than Nginx and PHP-fpm(github.com/qbix ↗)
    discuss
  25. The Input Layer(mg-crea.com ↗)
    discuss
  26. Show HN: Explore 2D semantic space with the Jev model(semanticspace.dev ↗)
    discuss
  27. King Charles has hesitations about AI(techcrunch.com ↗)
    discuss
  28. The Harms of Modern Lighting and the Fight to Bring Back Incandescent Bulbs(midwesterndoctor.com ↗)
    discuss
  29. Show HN: Using a diffusion model to write docs quickly(cortee.ai ↗)
    discuss
  30. Making a game for the GBA and PC from the same codebase(mattgreer.dev ↗)
    discuss

Microsoft exec called AI scraping 'the largest theft of labor in human history'

379 pointsby 3h agotechcrunch.com
306 comments
3h agoHN ↗

I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.

3h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

2h agoHN ↗

At least the Chinese AI companies are doing good service open-sourcing their models back to the public.

2h agoHN ↗

There are no open-source LLMs. There are only downloadable LLMs, with no source for the model being provided. The source for the training and inference programs is not the source for the model itself, which would be the training set, training program, and random seeds.

2h agoHN ↗

Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?

2h agoHN ↗

What is DD supposed to mean?

Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things

2h agoHN ↗

Due diligence — pretty standard acronym in this community.

2h agoHN ↗

No it's not, and I've been here more than a decade.

1h agoHN ↗

We aren’t on WSB. DD can mean datadog, domain driven, design driven, data driven, etc

2h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up

Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.

With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".

2h agoHN ↗

With regulation and compensation, only rich companies would be able to do that

Well, with some imagination, you can have regulation that forces companies to open up, not just close down.

Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.

Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.

2h agoHN ↗

I am unclear how this would help anyone?

Any argument that writers and artists lose from these existing, would remain unchanged.

14m agoHN ↗

The argument was "It's the robbery of all of our culture to sell it back to us at a mark-up".

Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.

Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.

I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.

2h agoHN ↗

We don’t have to treat people reading books and companies stealing all human knowledge the same.

Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.

2h agoHN ↗

companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves

really not the same entities here

2h agoHN ↗

No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.

2h agoHN ↗

Not at leaf level, but if you trace the trunk, pretty sure you end up on the same one.

1h agoHN ↗

Morally speaking, one of the issues of modern society is the idea that knowledge should be free which was partially started by the file sharing movement, which didn't really move society in the right direction imo.

Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.

When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.

Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.

Same applies for Steam & videogames.

Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.

Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.

This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.

And once nobody cares, people even forget what a quality product is like.

1h agoHN ↗

Interesting argument. At first sight it looks like an argument for scarcity, but I think it's more an argument for relationship - the value of art isn't in the capitalist concept of 'a for-profit content object in a corporate inventory' but in the social relationships and shared experiences it creates.

Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.

This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.

Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.

The value is relationally both fleeting and persistent in ways that content consumption experiences - including live and recorded media of all kinds - aren't.

1h agoHN ↗

Nowadays you just fish out a song ... Spotify didn't make me enjoy music more.

Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!

If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.

2h agoHN ↗

We don’t have to treat people reading books and companies stealing all human knowledge the same.

We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.

(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)

2h agoHN ↗

I mean if you cite a copyrighted book verbatim. You are held liable. So should a company producing copyrighted work.

For instance a image/video generating model.

2h agoHN ↗

Came here to post this, got beaten by someone putting it far more succinctly than I would have.

2h agoHN ↗

companies spent a long time telling us downloading single songs via Napster was the worst thing ever

One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.

2h agoHN ↗

However, there is irony in a subscription to pirated material.

1h agoHN ↗

Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?

1h agoHN ↗

No, for obvious reasons. "acknowledging" implies it's a self-evident truth that people are just refusing to see as such, when it just isn't. Completely different companies - in fact one is the Recording Industry Association of America, an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.

While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.

41m agoHN ↗

Moreover, it's the people/companies from RIAA and adjacent circles (news publishing) that are stoking the "AI training is theft" arguments that people here are so breathlessly repeating - in a twist of irony, it's the people that suddenly decided to align themselves with media/news conglomerates, the same ones that were considered scum of the Earth before ChatGPT debuted.

1h agoHN ↗

Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves.

So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."

1h agoHN ↗

These arguments always boil down to "it's fine when I do it, but wrong when a company does."

The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?

1h agoHN ↗

I don’t believe any of these companies have paid for all the books they have trained on.

Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.

It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?

1h agoHN ↗

Napster was a company. Pretty sure that company didn’t tell you that.

2h agoHN ↗

Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.

2h agoHN ↗

The what thread, now? Got a link instead?

2h agoHN ↗

Well theres at least two different buckets of this.

First is the scraping of the open internet.

The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.

Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.

The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.

2h agoHN ↗

And yet a third bucket is the license-laundering of GPL code when the entire github corpus was vacuumed up.

2h agoHN ↗

Copy one book, and you're a thief. Copy thousands, and you're a VC.

2h agoHN ↗

Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data

We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.

2h agoHN ↗

and they would definitely not give it back for free.

...not like they are doing it for free now either.

open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.

I put "stolen" in quotation marks because it's still unclear if we can call that stealing

It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.

This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.

They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.

So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.

All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.

Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.

Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.

2h agoHN ↗

Regulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would.

If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.

2h agoHN ↗

Regulation can mean all sorts of things, including declaring the models themselves illegal.

Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.

1h agoHN ↗

Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".

It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.

We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.

EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.

So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.

1h agoHN ↗

The multiplication comparison makes no sense because humans don’t learn by multiplying numbers to change weights. It’s like comparing the lubrication oil consumption of a car to the cooking oil consumption of a human to compare the carrying capacity. That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn. Otherwise, a human would learn much less in their whole life than a GPU does in one second.

51m agoHN ↗

That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn.

The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.

41m agoHN ↗

"It is stealing. A human paid for the book, compensated the author and learnt from it."

I went to the library. Didn't pay a cent.

Now what?

37m agoHN ↗

EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.

News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.

1h agoHN ↗

Break up the multi trillion, multifaceted companies into the their separate facets. These companies all thrived on far significant smaller portfolios in the past. Cap the size a company can grow. Regulate the amount of compute they are allowed to use.

53m agoHN ↗

None of this actually stops the profiteering, or reverses the harm already done.

49m agoHN ↗

Legally mandating that companies open the weights of, say, 18-month-old models might change their minds a little.

I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.

44m agoHN ↗

When a human learns, they carry that forward into their future ventures. They might use their learning to recreate the original work and profit from it without attribution. In the West we view that badly and have legislation to protect from some abuses. But the same person might later collaborate with the original author, or make a derrivaive work that improves on the original (a la most science).

AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)

31m agoHN ↗

Would regulation help with that?

I mean it's already happened, right?

I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...

17m agoHN ↗

A simple law stating that laundering data through an LLM does not constitute "fair use", that the outputs can be subject to copyright claims of the original authors, and that the outputs themselves do not qualify for copyright protections would go a long way.

It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.

2h agoHN ↗

Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.

1h agoHN ↗

Scale matters.

The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.

1h agoHN ↗

Yes, the whole IT sector is built on it. Robbing jobs, money, power, opportunities from billions of people and delegating them in to shitty jobs.

Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.

1h agoHN ↗

It is called PROGRESS.

There is no such thing as PROGRESS for progress' sake.

And FYI, agriculture is a wonderful invention.

Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.

Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.

I won't even go into climate change.

And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.

"ripping off", "uprooting", "leeching (read: creating)" wealth

Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.

* * *

And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.

May the odds be ever in your favor.

1h agoHN ↗

Yet our world can feed billions of people. More people are alive today than more people ever lived in this floating rock. Most people are living better than kings from 200 years ago.

We invent. We make progress. We make course correction. We end up in a better place.

Climate change? Lmao. We were told by 2000 20% of our country will be under water... It's 2026. Not under water yet.

What amount of pesky human intervention (positive or negative) will affect earth waking up from it's cold climate? Not luddites and their cow farts are global warming.

The real fight with "climate change" will be done with planetary scale technology derived from massive amounts of technical progress and energy production.

Your argument is as good as I should throw trash in the road or plastic in the sea. I'm just one individual. It doesn't matter. I'm not making the whole society do it!!

I fully agree with you. It is survival of the fittest in the face of progress. Competition is essential. That's how society grows. Humans prosper. May the odds be in your favor too.

1h agoHN ↗

Theres a big difference.

We have social conventions regulating this.

Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.

You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…

1h agoHN ↗

Social conventions are regressive and are followed blindly by luddites. Real progress comes from doing what is right and breaking idiotic conventions.

Pirating is fine.

Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.

Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.

2h agoHN ↗

sell it back to us at a mark-up.

What if it was for free, like Wikipedia?

Crimes this large are crimes against humanity.

jfc no, sit down.

Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.

At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)

It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too)

  Example:

Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.

That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?

Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?

God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.

(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)

If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.

2h agoHN ↗

Defending the same entities getting large DoD contracts to use AI for killing?

2h agoHN ↗

They've been using computers for killing for decades, who's taking up pitchforks against computers?

Who's taking up pitchforks against THEM?

How do we keep getting deflected into hating the TECH instead of the people who abuse it??

2h agoHN ↗

Agreed, but now they're trying to stop other people from copying data so that they can be the sole gatekeepers of humanities collective knowledge.

1h agoHN ↗

And they should get effed. Distilling should be a right.

1h agoHN ↗

"copyright". its right there in the name of the rights.

2h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity.

Yeah the introduction of copyright was truly criminal.

So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

Oh wait ...

2h agoHN ↗

“It is said that at the heart of every great fortune there is a great crime”

lol at this edgy 5th grade statement. So ridiculous.

2h agoHN ↗

Owning ideas with copyright and patents is what separates the United States from communism.

The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.

It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:

To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;

I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.

[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...

1h agoHN ↗

In a communist society there is no profit (or incentive for), thus no need for copyright laws.

It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.

1h agoHN ↗

Well it's worth reading the linked wiki article section, which includes a link to another article "Copyright law of the Soviet Union", flatly contradicting you unless you maintain that the USSR wasn't truly communist.

1h agoHN ↗

unless you maintain that the USSR wasn't truly communist.

Didn’t they themselves say they were a socialist society on the path to communism?

Does anyone think the USSR was communist?

53m agoHN ↗

The USSR was a socialist state, which is different.

1h agoHN ↗

In the United State, the individual or corporation owns the invention. In the Soviet Union the state automatically owned the invention. That clause is what ensures private ownership.

The clause is what ensures profits from market sales or licensing of ideas go to the creator.

Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.

50m agoHN ↗

I have private ownership over many things that I've bought which are not inventions and are not infringing copyrights.

1h agoHN ↗

In the Soviet Union the state owned all media (de facto) and you had to get permission from a party official before you published or copied anything. It's not the same thing as a free for all.

It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.

In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.

2h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.

Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.

1h agoHN ↗

You're ignoring that there are options other than theft. Heard of buying things?

1h agoHN ↗

Do you think OpenAI has a New York Times subscription? Do you think that's actually relevant to what the New York Times is arguing here?

The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.

1h agoHN ↗

This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn.

Incorrect. You overlooked consideration.

1h agoHN ↗

Learning isn’t stealing, but if you’d set up a forum or hotline where paid staff would answer questions and write essay directly based on NYT content without a license for that content, that would probably be deemed illegal.

1h agoHN ↗

It is the largest democratization of knowledge that ever happened.

The 'sell it back to us' argument falls short in my view.

Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.

The comment here seems incredibly pessimistic and quite dramatical.

1h agoHN ↗

Free versions are abundant,

Awesome, can I make my own competitive LLM, just like I can make my own open source software?

and in some time useful models will ship preinstalled on all mobile phones.

Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.

1h agoHN ↗

No. I also cannot make a competitive computer, yet can use one for my work to remain competitive.

Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.

But what are 10 years in the grand scheme of things?

Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

1h agoHN ↗

Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

Nah, we should have:

1. invested less, in a more targeted way

2. ideally like ARPANET, with the benefits given to all humanity

3. with fair royalties paid to all (where relevant)

4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM

ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.

Exactly like the internet.

1h agoHN ↗

and never developed AI, because it will take a decade to disseminate the benefits to everyone?

Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.

You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"

We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.

The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.

1h agoHN ↗

I couldn't disagree more.

What I expect to see in 10 years is breakthroughs in science and medicine, with major diseases becoming treatable.

You can disagree, but on what basis do you think your predictions are more likely to be correct?

Humanity has turned out fine despite the invention of the book, the TV, and then social media. I reckon kids will develop just fine with AI too.

1h agoHN ↗

It's absurd to equate books and social media to being akin. And you're being facetious in your comments.

What if you are wrong? What are the risks vs rewards?

53m agoHN ↗

The main risk you have presented is a supposed negative effect on children's development. I am not convinced that there will be a negative effect at all.

On the reward side we have potentially curing most major disease, automating labor and freeing humanity from having to work for a living, as well as perhaps generally advancing science at an unprecedented pace. And of course improving education.

1h agoHN ↗

Can I make my own competitive LLM?

Yes, yes you can.

See the number of startups that have finetuned or trained an OSS model to build their business on.

“But you have to have compute!”

Ok, and you’re writing OSS on a rock with no internet connection?

The world has all kinds of barriers, but if anyone has a chance of competing on the LLM from it’s not going to be by getting rid of fair use.

1h agoHN ↗

It is woefully pessimistic indeed.

There will be new jobs coming from this.

The bountiful abundance of intelligence is truly the best thing that has happened this decade.

1h agoHN ↗

Wiki already democratized it just fine and was legitimately free for people who know how to read.

It's asinine that you think the sell it back to us argument falls short.

Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.

From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.

1h agoHN ↗

Wikis made knowledge available.

Making it accessible, understandable, and usable is another matter.

How LLMs sound is not a fundamental limitation of the technology.

The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.

You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.

1h agoHN ↗

Making it accessible, understandable, and usable is another matter.

I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.

1h agoHN ↗

There are no viable free versions with sufficient computing power. The do-it-yourself AI is dangled as a carrot in front of users to camouflage the lock-in and rent seeking by Big AI. Go prove something new like Navier Stokes (replicating N-S itself no longer counts due to scraping and plagiarism) on your Mac Pro!

Paid influencers who perpetuate the open narrative are a whole new industry.

Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".

1h agoHN ↗

The comment seems in line with real life, while your catchphrase about democratization is more like wishful thinking.

9m agoHN ↗

Ah, this remembers me when I was also young and naive.

Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..

1h agoHN ↗

Works produced by AI are not copyrightable. Are they in the public domain? Can an entity sell unique works which are in the public domain?

1h agoHN ↗

Yes? I own a lot of books that were public domain when published (as reprints). I could read them on Gutenberg, but I'm paying for the nice paper formatting.

1h agoHN ↗

Our "culture" has long been the province of corporations. In prior epochs it was still the product of patronage and power.

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?

1h agoHN ↗

Flourishing life? Where is your data leading to the conclusion that AI models are leading to the world populace leading a ‘flourishing life’? Are taxes on revenue of companies like OpenAI somehow being collected and turned into a UBI and I just didn’t hear about it?

1h agoHN ↗

I said a glimpse. I didn't say it's here now or it would be easy.

It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.

1h agoHN ↗

Like trickle down economics?

We all experience the substance which is trickling down.

1h agoHN ↗

Or, perhaps, the companies forming the for profit LLMs could be made to pay a license fee for the copyrighted data they lifted from behind the paywall and then used to form a for profit entity with.

Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?

1h agoHN ↗

This is solid but how many people you reckon would benefit if they paid every single penny for any copyrighted data? I am in my 50's with 30 years in the industry and I know maybe 2 people that might benefit from this... certainly not something general public would benefit from

1h agoHN ↗

We’re likely looking at a new world where physical and mental work from humans holds much less value.

So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.

So our future world probably looks like a system where effort/knowledge/skill has little to no reward and capital (e.g. inheritance, passive income) has all the reward.

So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.

1h agoHN ↗

Why more and more stories from "Pump Six and Other Stories" are becoming real step by step?

59m agoHN ↗

You are missing the obvious answer.

The state runs the system. Or really at a certain point, the state runs itself. In 100 years or so, the CEO becomes a barbaric relic of the past age, like our ancestors who beat each other over the head with clubs. Nevertheless, such a future will necessarily owe a debt to both.

Wealth and states are often connected. In this case if the state is properly aligned (this is what the coming battles are to be about - as they always have been) to be egalitarian, then redistribution can occur unblocked by human enclaves of greed.

48m agoHN ↗

We’re likely looking at a new world where physical and mental work from humans holds much less value.

Or, perhaps, we're regressing to the mean after a century where this work was over-valued, because companies like Disney succeeded in regulatory capture and created artificial protections to maximize their own revenue.

How much money did Bach earn from royalties (ok, there's a pun there, but I mean payment for reproduction of his work)? How much power did Melville have over who published Moby Dick, and where, and how it was used (hint: very little in the US, none at all overseas).

We are seeing a weakening of control and revenue extraction from copyrighted works. But on the chart of history, the 1900's were a very anamolous spike in that area. And it saddens me that so much of HN is unhappy about more of a return to the commons.

So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.

On the bright side, I think you're wrong here. Or (as Claude likes to tell me), you're half-wrong. If this were true, it would already be the norm and only big companies could bring new products to market. Yet startups are a thing, and many succeed (more fail, but still.

The trick in entrepreneurship and creative work has always been knowing what to ask the system to do. You might as well say that music is dead because the synthesizer and music software companies can produce as much as they want at zero cost. It turns out that owning the means of production only loosely correlates to producing things people want.

So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.

Already done. There are kids in Africa studying with AI tutors today, getting insight that they would likely never have had access to before. One-person shops are releasing board games and business services that they would never have had the capital to do before.

Where you see centralization of capital and control, and the masses reduced to abject poverty... I see the complete collapse of barriers to entry and switching costs in many fields.

But who am I? Some random guy. But I've heard very senior execs, at Microsoft and other companies, in absolute panic that the IP and systems they've spent billions of dollars to build over decades of work are suddenly subject to disruption by teenagers who are great at using AI. Seriously, panic.

It's a complex subject, and sorry for writing a book, but your prompt apparently got me. None of us know for sure what will happen but I think yours is a needlessly pessimistic view, ironically informed by the exact abuses of th past century that you're worried might not continue.

1h agoHN ↗

It is achievable in theory yes. But it was also achievable without LLMs and yet, here we are.

1h agoHN ↗

I mean I'm producing more and better art and software each month than I used to do in a year. I feel like I'm flourishing because of AI. Tell me I'm wrong?

1h agoHN ↗

Considering that we don't seem able to appropriately tax the biggest corporations, why do you think we'll suddenly be able to redistribute LLM gains?

Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.

1h agoHN ↗

These doomsday theories sound solid in practice but in order for someone to accumulate a lot of wealth there have to be consumers to chip into this. If we are all slaves, where is that wealth going to come from?!

1h agoHN ↗

The Amazon warehouse worker is a modern day slave.

Don’t believe me? Try it. I have.

26m agoHN ↗

Not necessarily as there are methods to redistribute wealth without the consumers having any choice.

Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.

Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.

There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.

1h agoHN ↗

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Let me fix that for you: A world in the the value of the labor of each human approaches zero.

1h agoHN ↗

Right, and the consequence is going to be, what?

Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.

It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.

1h agoHN ↗

We are not deluding ourselves that we will maintain our standard of living. If the boss’ dream of replacing all workers by AI that does their job poorly but cheaply will be realized. We will go to the gutter with the rest of the economically worthless people.

There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.

1h agoHN ↗

Correct.

Humans are animals. We have rules to keep humans in check but that’s it.

The wealthy do not care nor need to care. They will justify it as ‘you lot were too weak to organise and stop us’ and there is an element of truth to that.

54m agoHN ↗

I still see a lot of denialism.

It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.

As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.

Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.

39m agoHN ↗

“I find hope in your acknowledgment of your own imminent devaluation.”

Gave me a chuckle. I think this is my personal ‘best sentence of the year’.

Great way to summarize how we will likely look back on 2026.

52m agoHN ↗

You assume there will still be free elections and democracy. There is also the possibility that the form of government will change into something else.

1h agoHN ↗

The value will tend towards infinity [1] because we can keep building more data centers. This will allow us to apply tokens to solving every problem we can conceive.

The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.

If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.

Industrialization led to iPhone, PlayStation, Spotify, and Waymo.

AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...

AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.

[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.

[2] Until the robots leave earth, anyway

1h agoHN ↗

In the 50's there were predictions that in 10 years no-one will have to work again because of the advances made in automation, like the washing machine for example. Any predictions of less work this time around are a complete joke.

1h agoHN ↗

For thousands of years humans have dreamed of reaching the stars.

Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.

On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.

You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?

1h agoHN ↗

When was the last time you saw a Chinese laundry?

42m agoHN ↗

That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.

1h agoHN ↗

Now the mother and father both have to get jobs, in the future mother, father, and child will get jobs to support the trillionaires.

1h agoHN ↗

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Fantasy lets us glimpse anything.

1h agoHN ↗

US has been procrastinating reparations for slavery. LLM redistribution can't happen until the capitalism machines recursively solve "sins of our fathers".

1h agoHN ↗

a world in which the labor required of each human to lead a flourishing life approaches zero.

The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.

1h agoHN ↗

a world in which the labor required of each human to lead a flourishing life approaches zero.

Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?

Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.

1h agoHN ↗

Is this "escape route" in the room with us now?

55m agoHN ↗

Nobody believes that nonsense any longer. Will we also get 72 virgins each?

1h agoHN ↗

If piracy isn't stealing training definitely isn't stealing.

1h agoHN ↗

At the very least we should foribly confiscate the models & make them available free-for-all as open-weights downloads. Failing that, bring back the guillotine.

1h agoHN ↗

So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.

Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.

The more fundamental objection seems to just to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.

1h agoHN ↗

Musicians and writers build new works on top off millennia of literary and musical history, then copyright their works and sell it back to us. Is this so bad? If Taylor Swift, consciously or subconsciously, gets an harmonic idea from a 1970's song and a fragment of a melody from some 1990's song... is that theft?

I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.

But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \

It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.

I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?

58m agoHN ↗

The problem is not Taylor Swift is being "inspired" from other artists. We have tons of examples this throughout history.

The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.

This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...

We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.

Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.

43m agoHN ↗

Ok, thought experiment: at what point in the past 1000 years do you think the balance of collective versus individual benefit from creative works was at its most fair?

getting things were free was part of sending a message, not monetary gain.

As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.

57m agoHN ↗

when you do your job you take ideas from others before from university from calculus etc

therefore you have to work for me for free

42m agoHN ↗

Oh I'm sorry, who's forcing you to work without pay?

49m agoHN ↗

I agree with you, but there is one aspect that is different here - the scale that is way beyond what any human can do.

I don't know what to think about that though.

47m agoHN ↗

Fair point. Think of it this way: the scale is more than any human, but it is probably akin to what humanity in general can do.

37m agoHN ↗

Yes. And even if (when) AI surpasses humanity in general I don't know if that changes the argument.

33m agoHN ↗

You might argue that since those LLMs build on all of humanity's knowledge, they should belong to everyone. But where do you draw the lines? Or make a practical case.

20m agoHN ↗

They kind of do belong to everyone. The technology and majority of training corpus is available to essentially everyone.

The training and inference infra are owned and operated, but I can't see an argument that any machine that processes public domain (or stolen, if you prefer) info should be available to everyone for free.

But those training and inference costs will go to zero. Think about your cell phone today versus $1m+ supercomputers in the 1980's.

We're living in a transitory blip where capitalists and gold rushers are getting rich arbitraging the cost of processing against the non-cost of corpus. We can argue about morals (it doesn't bother me much) but it is a narrow window and it will be remembered the way Compuserve is: a precursor to the actual revolution, worth a footnote.

5m agoHN ↗

No, that's what I mean by where to draw the line. An artist taking inspiration from others is not obliged to give up the work to the public.

Sidenote: It may be tricky/impossible in the future to uphold intellectual property laws. If anyone is able (for instance) to prompt-create all their software, a software patent is worthless.

1h agoHN ↗

Robbery implies taking it away from you so you can't have it anymore. How is that the case?

52m agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up

Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.

49m agoHN ↗

Yes, you are getting a laundered version for a monthly payment to the rent seekers, who are 1000x worse than the RIAA.

And the internet is destroyed by slop spam in the process, so our culture is diluted and crowded out.

47m agoHN ↗

That's true. However it does sometimes feel like our culture is being buried in slop.

42m agoHN ↗

Imagine someone asks you how to do something at work, you tell them. Then they create a huge packet filled with bullshit about how it got done and now they are your boss.

40m agoHN ↗

Isn't that the same point piracy advocates have been making for a while? If I watch a pirated film, I didn't really consume a physical resource. Nothing physical is lost. Therefore it's not theft?

Learning isn’t stealing.

This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid. It's a public good to have information available, but it's not really acceptable to circumvent established ways of compensating the creator of the work you're benefiting from.

16m agoHN ↗

Isn't that the same point piracy advocates have been making for a while?

No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.

Learning isn’t copyright infringement.

Learning isn’t theft.

Copyright infringement isn’t theft.

These are all different statements, and all are true.

There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.

“Gaining exposure to” isn’t theft either.

Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.

There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.

39m agoHN ↗

Partially true.

If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.

If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.

If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.

If people stop making art because they can't outcompete generated art, that is also lost to LLMs.

So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.

27m agoHN ↗

If $X dies because ..., the entirety of its knowledge is now only availabe through LLMs that trained on it.

And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.

The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.

13m agoHN ↗

You are mixing up culture, data, and the creation of new works as if they are the same thing when they are very different.

51m agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up.

Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.

46m agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up.

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.

Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.

And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.

There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".

37m agoHN ↗

Creators are unwillingly and contra to economic systems that have evolved over a few millennia entered in the Borg or the Matrix. So their achievements are reused for private and public benefits without their permission. Piracy is theft. And this piracy is a bigger theft than piracy on an individual download basis. Was there any doubt? (Linguistically I agree that you can’t “rob culture”. You can rape or reap culture though, and that is the point at hand.)

17m agoHN ↗

HN crowd: Check out my Plex server and 132TB media collection! Also HN crowd: Reading publicly posted information on the open internet is morally outrageous theft of the highest order

7m agoHN ↗

Also,

The crowd: "Knowledge should be free! Culture belongs to all of us! Information must be democratized!"

AI companies *proceed to copy literally everything and put it through virtual blender, until an universal general-purpose problem-solver tool comes out, then give it out for free or serve for peanuts to literally everyone on the planet with Internet connection *

Crowd: bbb..buuut not like that!

34m agoHN ↗

What do you mean, it's already robbing me of the ability to discuss techniques with a number of my peers since they decided the mediocre output of LLMs is good enough instead of understanding what they are doing.

In addition, a great number of people have decided to stop publishing their code publicly, and discuss techniques except in spaces they can be sure it's a one-on-one with a human being.

Some things already got lost.

31m agoHN ↗

I still do open-source. You just get the code on a CD like the good old times!

23m agoHN ↗

You're discussing your peers' behaviors, not AI.

It's the same as saying "People don't kill people; guns do." Maybe guns(AI) should be regulated, but the primary problem is murderers(lazy peers), not weapons(AI).

17m agoHN ↗

Just like with guns, AI is a crime of scope.

I can be a murderer with a knife, sure, but how many people can I kill? How many people can I kill if I have a class three license an a full auto gun?

I can pirate a news article here or there, sure, but how many news articles can I pirate? How many news articles can I pirate with AI?

10m agoHN ↗

How many news articles can I pirate with AI?

Not much more than with a bash/curl loop. Which, on one hand, AI can write for you, but on the other hand, AI will make you no longer need to pirate those articles in the first place.

Also it's not end-user piracy being discussed - it's the act of training AI itself that's accused here to be "robbery of our culture".

17m agoHN ↗

That's people problems though, namely:

1) Lazy peers, and

2) Spiteful peers with "dog in the manger" mentality.

The latter case personally irks me, mostly because this often involves personal benefit (economical or moral) due to them claiming to give away knowledge for altruistic reasons, which later behavior reveals was a lie, mislabeling proprietary as open and free, to reap unfair gains.

But that's all off-topic for this thread anyway.

28m agoHN ↗

it's still there

There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:

Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.

To use another Slashdot cliche: "But you knew that already."

16m agoHN ↗

AI didn't steal any culture. It just made mediocre culture more accessible. Turns out, most people do like mediocre culture. Previously, public TV channels could at least pretend that people enjoy educational content or classical music. AI exposed that lie. And now we're shocked that the emperor is naked.

36m agoHN ↗

sell it back to us at a mark-up

sell it back to us as markdown

29m agoHN ↗

“It is said that at the heart of every great fortune there is a great crime“. That quote is from a fiction author - you’re essentially quoting Spiderman “with great power comes great responsibility”.

24m agoHN ↗

Are you seriously putting Balzac and Spiderman on the same level to try to ridicule thoughts?

Ad hominem.

17m agoHN ↗

They are both quotes that people say things like "they say" to imply someone of import said it to give it credence and truth to it but were in fact from fictional sources.

Exact same pattern - take it as an ad hominem all you want.

20m agoHN ↗

Other than the rare books than have been ruined, all the same knowledge is still out there though. So it’s not robbed in the sense of a bank heist. Maybe in the sense of pirating a movie.

The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.

2h agoHN ↗

And here we are, just watching and doing nothing..

2h agoHN ↗

I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.

2h agoHN ↗

Lots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.

2h agoHN ↗

Copying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone?

Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"

Probably indicative of America's wider downfall that they've all become so self interested

2h agoHN ↗

Hackers are irrationally anti-corporation. This is where the nonsensical AGPL came from, too.

2h agoHN ↗

The AGPL didn't go far enough because it didn't limit freedom 0 to natural born humans. In the 00s/10s it was corporations. Today it's AI.

If it has no soul to save and no body to torture it deserves no rights.

1h agoHN ↗

Show me where this "soul" is. Never seen one yet.

1h agoHN ↗

I don't believe in souls. But I believe in metaphorical souls, and an LLM ain't got one.

1h agoHN ↗

I want to share the open source code for the good of everyone - which is why I put it under the GPL, for this right of everyone to share it to be protected. Therefore, any AI model that ingested GPL code should logically also have all its output GPL licensed. Not problem with that.

2h agoHN ↗

I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.

1h agoHN ↗

It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.

1h agoHN ↗

We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).

1h agoHN ↗

I suppose it's theft in so far as you're deriving value from something that others labored for. I think that sticks in peoples throats. You can go to the library and do that already, gain knowledge, start a business or whatever. But the scale feels impersonal and monstrous in comparison.

1h agoHN ↗

The same point was made by many in the music industry about huge scale Internet music piracy vs people dubbing CDs onto mixtapes.

Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.

Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.

1h agoHN ↗

The original might no longer be there as the site might have already shut down due to AI scrapper bot overload. Or the original was a book Anthropic scanned and then shredded. Or the artist stopped publishing their works or doing art all together after all their creations were ingested into the model blob without their permission.

It is really insane to compare individuals copying data to big corporations parasiting on the Internet.

2h agoHN ↗

If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.

2h agoHN ↗

If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.

2h agoHN ↗

Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours

2h agoHN ↗

But it's copying...how is it theft? Your labour WASNT stolen was it?

1h agoHN ↗

In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?

1h agoHN ↗

Even better question: would you be willing to train the AI-driven robot which will end your profession altogether?

1h agoHN ↗

I work in infosec and I would give up everything I own and die happy if infosec became a solved problem and the profession died forever.

It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.

1h agoHN ↗

Do we normally consider recorded music to be theft from musicians who would have been paid to play music live if we never allowed (or invented) recorded music?

Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.

But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?

Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.

1h agoHN ↗

Well, given these companies are trying to sell it back to you - seems like even worse than stealing. ;-)

1h agoHN ↗

For intellectual property, copying without permission is theft.

1h agoHN ↗

Depending on the context, it’s fair use.

As it is, in this case.

14m agoHN ↗

It's as fair use as this scenario:

I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.

2h agoHN ↗

No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters

2h agoHN ↗

If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.

Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.

46m agoHN ↗

Nobody normal talks like this.

There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.

1h agoHN ↗

Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft. Then you get the serfdom in the middle ages. Up to recent times humans have been brutally exploited.

The copy part was a recognized right, then taken away.

2h agoHN ↗

I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)

2h agoHN ↗

they did say theft of labor, not theft of the things being trained on

2h agoHN ↗

In case of programming.

How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?

This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.

Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?

What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.

https://consortiuminfo.org/metalibrary/estimating-the-total-...

2h agoHN ↗

They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.

IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.

1h agoHN ↗

I guess you're not writing code much or haven't done much so as your professional career?

2h agoHN ↗

The most shocking point is that they have a Microsoft exec who knows what he's talking about.

27m agoHN ↗

In my experience many execs know what they're talking about.

Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.

Knowledge alone is less often a factor.

Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.

Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.

2h agoHN ↗

How is this different from Microsoft scraping to build Bing?

Honest question. There is a line in the sand somewhere apparently.

2h agoHN ↗

Huge difference between building AI and a search index.

2h agoHN ↗

Why? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn't make.

2h agoHN ↗

Bing isn’t re-selling you back the content it took.

2h agoHN ↗

Bing sent traffic to the original source. AI answers don't. That's the line. It's not complicated.

1h agoHN ↗

Under “conduct requirements” imposed by the CMA in June, UK websites are able to activate an opt-out to stop Google from scraping their content to power search features such as AI overviews - very similar conceptually to the news law passed in Australia.

https://www.bbc.co.uk/news/world-australia-56163550

1h agoHN ↗

I remember when this was the prevailing thinking online until about 2024. But that's when everyone was trying to justify their own piracy of GTA or whatever.

Suddenly, they don't like other people pirating.

2h agoHN ↗

I'd call the introduction of copyright the largest theft of human labor in history.

No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.

That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.

The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.

At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.

1h agoHN ↗

That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.

Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.

At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.

Look at who owns that site, no surprise here.

1h agoHN ↗

Vectoralist News doesn’t have the same ring to it

2h agoHN ↗

“Information wants to be free“.

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

1h agoHN ↗

Probably because it claims there’s no problem, and then makes a tiny little mention of the BIG problem at the very end.

1h agoHN ↗

The unpopular part is the hypocrisy where their work has already or must be rewarded while other people’s work is not.

1h agoHN ↗

Information wants to be free“. It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it. My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

1h agoHN ↗

"I just wished the collected data was public. "

That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.

11m agoHN ↗

On the flip side, output from an LLM is not copyrighted.

1h agoHN ↗

I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.

1h agoHN ↗

Sharing information is an act of love

Most AI companies are not sharing it, though. They appropriated it and resell it.

1h agoHN ↗

Well the future we seem to be getting is “information wants to be free for the first ten thousand tokens, then $1 per million tokens after”.

1h agoHN ↗

There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.

49m agoHN ↗

The key problem is that IP is either proprietary to the creator or it is a commons type of situation.

Even if you agree with the former exploiting the commons for personal profit is... not good.

One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.

24m agoHN ↗

"Information wants to be free".

I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.

2h agoHN ↗

It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.

2h agoHN ↗

Yeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did!

If this was all open, I’d maybe half agree.

2h agoHN ↗

It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.

At least it’s consistent, is what I’m saying.

1h agoHN ↗

Genocide vs non-destructive copying of digital information - I'd say that's pretty different.

1h agoHN ↗

Ask Anthropic how non-destructive their copying is, when the clandestinely buy rare books for cheap & shred them after scanning and not sharing the result.

35m agoHN ↗

First, most of these books are old and plentiful, not rare at all. Think "Master Windows Exxcel '95!", not first editions of "Master and Margarita". Libraries routinely destroy plentiful old books that nobody is reading any more. Anything actually rare and valuable is too pricey to hand over.

Second, they HAVE to destroy the books because US copyright law REQUIRES it. They're only allowed to digitize the books if it's considered "transformation" and not "copying". That's only allowed if they destroy the book afterwards.

2h agoHN ↗

AI overall is the ultimate piracy crime.

I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.

1h agoHN ↗

Then let's make humans pay royalties to every author from whom they ever learned something, even if it was offered freely to them, only fair?

1h agoHN ↗

Indeed. If I take somebody's work, transform it somewhat and sell it, I should pay royalties (unless the author explicitly allowed me to do so). That is exactly the case.

40m agoHN ↗

Sadly, they're just following in the footsteps of every large media/publishing/music conglomerate that already screwed over the vast majority of artists/musicians/writers. For every Taylor Swift striking it rich, there's 999,999 who can't even pay their bills with what their copyright gets them.

I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here

37m agoHN ↗

It is not just about artists. It is about every kind of intellectual work: scientific research, essay, fiction, painting, software, etc - the list goes on..

2h agoHN ↗

All the "LOL you wouldn't steal a car???" posts in this thread miss the point entirely.

AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.

At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.

1h agoHN ↗

I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.

1h agoHN ↗

Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?

1h agoHN ↗

I do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet.

Publicly. Accessible.

Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.

Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.

1h agoHN ↗

A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.

If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.

If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.

1h agoHN ↗

A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.

Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them.

An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible.

If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.

Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it.

Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.

1h agoHN ↗

Just because something is publicly accessible doesn't mean you can use it for free, or that it gives you rights to do whatever you want with it.

I can access a public park, but that doesn't necessarily give me the right to also bike on its sidewalks, or walk on the grass, or take some of the plants home with me.

... or has been obtained illegally

Similarly, content that is _accessible_ publicly may be illegal to _obtain_, these aren't mutually exclusive.

On the internet, you'll find there are terms of services and licenses. These restrict how you can use even publicly accessible material. Public availability doesn't give you a license to use it however you want.

53m agoHN ↗

Public availability doesn't give you a license to use it however you want.

Of course. You listed complex examples from the real world. A park where walking is allowed but damaging the plants is not, for instance. There it makes sense to distinguish various activities that can be done in there and treat them separately.

But a website offers not much activities that you can do with it. You can read it. And that is about it.

58m agoHN ↗

If I make the path behind my home publicly accessible so people in our town can get where they’re going easier, then Amazon builds a warehouse in our town and starts driving their trucks through my path 24/7, are you going to tell me “well you made it public access, they have a right to it same as everyone”?

With the advent of trillion dollar corporations selling extremely powerful general purpose imitation as a service, the meaning of “public access” has substantially changed, potentially invalidating the original agreement.

45m agoHN ↗

are you going to tell me “well you made it public access, they have a right to it same as everyone”?

Yes. You can of course always close it to the public if the public usage bothers you. Or require those using it to agree to your terms and conditions where you restrict the speed, the weight, the time of the day, etc. Anything you want.

But if you choose to make it public with no restrictions, you have to be prepared to face the consequences.

23m agoHN ↗

Well, I posted a sign saying that but they’re still doing it, so I guess I’ll just close it to public access. Sucks for all of the people who now have to take a longer route. I hope they blame Amazon and not me.

1h agoHN ↗

The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.

1h agoHN ↗

One leads to the other so its simpler to point the root issue.

58m agoHN ↗

And the cruel irony is that they stole our work to train the AI that devalues our work going forward, and will cause many of us to lose our jobs.

I regret every line of open source code I ever wrote.

12m agoHN ↗

This process can even feel a bit like parasitism, it empties out the host, like in <Alien>.

1h agoHN ↗

The largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices.

It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.

If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.

1h agoHN ↗

I think it's okay to advance humanity, but they can GTFO when they then try to ban distilling and open models.

1h agoHN ↗

But what about M$ owning Github and doing the same with its content? Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...

1h agoHN ↗

You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.

1h agoHN ↗

I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.

1h agoHN ↗

There will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.

1h agoHN ↗

As someone who thinks that LLMs are harmful, I deeply resent that any of my work has been used to help develop them. I will never forgive these companies for forcing me to contribute.

1h agoHN ↗

Because it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.

1h agoHN ↗

I don’t mind these companies scraping my content.

But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.

1h agoHN ↗

"The Net interprets censorship as damage and routes around it."

1h agoHN ↗

Copyright is artificial scarcity rationalized by arguing that producing novel intellectual work is valuable, but requires substantial effort that can't be recouped, so we have to incentivize it somehow.

LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.

AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.

1h agoHN ↗

If there is no legal guarantee that human creativity can pay off we are starving art and humanity from its inception. If ordinary people cannot participate in the act of creation, you get exactly what hollywood has become.

Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.

1h agoHN ↗

They are not selling the information. They are selling a service for easy access to that information.

These are two different things.

Note: am not an AI fanatic.

1h agoHN ↗

Correct solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use.

There should be a carveout for non-profit or government AI.

58m agoHN ↗

First of all it can. Second I'm talking about attribution to copywritten material (and royalties being paid accordingly). This would only be for for-profit AI implementations though.

1h agoHN ↗

Copyright infringement, if this even were that, is not theft. Chattel slavery is the largest theft of labor in human history.

54m agoHN ↗

Information wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it

Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it

42m agoHN ↗

I just don't understand people saying "but a human learning from a book isn't illegal".

How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."

And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.

39m agoHN ↗

"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. "

So training can make it legal as well. Interesting...

32m agoHN ↗

Theyre just jealous that theyre being surpassed on their market capture of computing

30m agoHN ↗

Yet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants.

The hypocrisy of this new world is already catching up to us.

26m agoHN ↗

I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.

25m agoHN ↗

Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours

So no, your cries for regulating others because you are losing the race won't work this time.

16m agoHN ↗

Microsoft is one to talk.... Remember when MS trained copilot on all your github code?

15m agoHN ↗

Imagine the parthenon marbles. When they were looted it was even a celebrated act, but they are still stolen in the british museum centuries later.