Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Code now reads AGENTS.md if there is no Claude.md(claude.com ↗)
    135comments
  2. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    187comments
  3. Saving another 100TB of RAM(cloudflare.com ↗)
    33comments
  4. Cloudflare Quick Tunnels(cloudflare.com ↗)
    221comments
  5. Xcode 27.1 Beta Release Notes(developer.apple.com ↗)
    58comments
  6. US troop deaths during Iran war exceed Pentagon count by at least four(reuters.com ↗)
    23comments
  7. How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip(ieee.org ↗)
    discuss
  8. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)(arxiv.org ↗)
    11comments
  9. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    45comments
  10. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    72comments
  11. How to Write with an LLM(sockpuppet.org ↗)
    243comments
  12. OpenJev(openjev.com ↗)
    238comments
  13. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    35comments
  14. Cyclomatic Complexity in C#(ndepend.com ↗)
    9comments
  15. The Implications of Linguistic Illegibility for LLM Security(arxiv.org ↗)
    17comments
  16. Two parallel neural ectoderm progenitors contribute to the developing brain(newscientist.com ↗)
    51comments
  17. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    168comments
  18. From Geometry to Algebra and Back Again: 4000 Years of Papers (2023) [video](youtube.com ↗)
    discuss
  19. Minimal Phone 2(minimalcompany.com ↗)
    148comments
  20. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    24comments
  21. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    13comments
  22. A search-and-inference database from scratch in pure Zig(antfly.io ↗)
    16comments
  23. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    89comments
  24. Size-Specialized Memory Allocation(go.dev ↗)
    3comments
  25. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    175comments
  26. Korea raises data breach fines to 10% of revenue(koreajoongangdaily.com ↗)
    73comments
  27. The US 'Kill Chain' That Destroyed an Iranian School(bloomberg.com ↗)
    discuss
  28. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  29. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    281comments
  30. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    148comments

Microsoft exec called AI scraping 'the largest theft of labor in human history'

851 pointsby 13h agotechcrunch.com
701 comments
13h agoHN ↗

I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.

13h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.

12h agoHN ↗

At least the Chinese AI companies are doing good service open-sourcing their models back to the public.

12h agoHN ↗

There are no open-source LLMs. There are only downloadable LLMs, with no source for the model being provided. The source for the training and inference programs is not the source for the model itself, which would be the training set, training program, and random seeds.

12h agoHN ↗

Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?

12h agoHN ↗

What is DD supposed to mean?

Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things

12h agoHN ↗

Due diligence — pretty standard acronym in this community.

12h agoHN ↗

No it's not, and I've been here more than a decade.

11h agoHN ↗

We aren’t on WSB. DD can mean datadog, domain driven, design driven, data driven, etc

12h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up

Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.

With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".

12h agoHN ↗

With regulation and compensation, only rich companies would be able to do that

Well, with some imagination, you can have regulation that forces companies to open up, not just close down.

Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.

Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.

12h agoHN ↗

I am unclear how this would help anyone?

Any argument that writers and artists lose from these existing, would remain unchanged.

10h agoHN ↗

The argument was "It's the robbery of all of our culture to sell it back to us at a mark-up".

Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.

Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.

I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.

12h agoHN ↗

We don’t have to treat people reading books and companies stealing all human knowledge the same.

Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.

12h agoHN ↗

companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves

really not the same entities here

12h agoHN ↗

No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.

12h agoHN ↗

Not at leaf level, but if you trace the trunk, pretty sure you end up on the same one.

11h agoHN ↗

Morally speaking, one of the issues of modern society is the idea that knowledge should be free which was partially started by the file sharing movement, which didn't really move society in the right direction imo.

Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.

When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.

Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.

Same applies for Steam & videogames.

Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.

Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.

This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.

And once nobody cares, people even forget what a quality product is like.

11h agoHN ↗

Interesting argument. At first sight it looks like an argument for scarcity, but I think it's more an argument for relationship - the value of art isn't in the capitalist concept of 'a for-profit content object in a corporate inventory' but in the social relationships and shared experiences it creates.

Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.

This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.

Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.

The value is relationally both fleeting and persistent in ways that content consumption experiences - including live and recorded media of all kinds - aren't.

11h agoHN ↗

Nowadays you just fish out a song ... Spotify didn't make me enjoy music more.

Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!

If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.

5h agoHN ↗

Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it.

The same can be said for time. You said yourself it was an investment. There's no need for money to be involved since how much time you spend on a thing will establish the value of that information to you. If you must have a ritualistic act to form a connection to music, let it be the act of listening instead of paying.

Music discovery also takes in investment in time. Having near instant access to so much music also makes discovering a new song or artist you love very rewarding. I can easily listen to hundreds of songs before even one makes it into rotation in my current playlist.

12h agoHN ↗

We don’t have to treat people reading books and companies stealing all human knowledge the same.

We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.

(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)

12h agoHN ↗

I mean if you cite a copyrighted book verbatim. You are held liable. So should a company producing copyrighted work.

For instance a image/video generating model.

12h agoHN ↗

Came here to post this, got beaten by someone putting it far more succinctly than I would have.

12h agoHN ↗

companies spent a long time telling us downloading single songs via Napster was the worst thing ever

One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.

12h agoHN ↗

However, there is irony in a subscription to pirated material.

11h agoHN ↗

Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?

10h agoHN ↗

No, for obvious reasons. "acknowledging" implies it's a self-evident truth that people are just refusing to see as such, when it just isn't. Completely different companies - in fact one is the Recording Industry Association of America, an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.

While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.

10h agoHN ↗

Moreover, it's the people/companies from RIAA and adjacent circles (news publishing) that are stoking the "AI training is theft" arguments that people here are so breathlessly repeating - in a twist of irony, it's the people that suddenly decided to align themselves with media/news conglomerates, the same ones that were considered scum of the Earth before ChatGPT debuted.

6h agoHN ↗

an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.

This is the most unintentionally hilarious misunderstanding of what the RIAA does, and the power relationship between artists and publishers I've read in years. In practice the RIAA exists to maintain the copyright monopoly of a few major labels. Rent seeking from the non-artist owned catalogues of the enormous majority of musicians who never 'recoup' their initial record deal.

While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.

It's far more seductive (since it's the default) to assume class relations don't exist, and wealth distribution is meritocratic. There may not be 'goodies and baddies', but there absolutely are rentiers and workers, billionaires and plebs.

All the major labels are public companies. Which means it's the very same people - the investment class, who claim ownership and extract wealth from say Warner and Open AI (should it make any money - obviously the whole house of cards could come down first).

11h agoHN ↗

Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves.

So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."

11h agoHN ↗

These arguments always boil down to "it's fine when I do it, but wrong when a company does."

The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?

8h agoHN ↗

I'm 100% happy with either viewpoint.

1) AI companies all get sued out of existence.

2) AI companies can train networks with piracy, those networks don't fall under copyright, so give me a copy to do what I want with.

11h agoHN ↗

I don’t believe any of these companies have paid for all the books they have trained on.

Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.

It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?

8h agoHN ↗

I mean, I don't seem to get to say, "Oh I want this but I can't buy it so I'll just steal it LOL", so why should AI companies?

11h agoHN ↗

Napster was a company. Pretty sure that company didn’t tell you that.

12h agoHN ↗

Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.

12h agoHN ↗

The what thread, now? Got a link instead?

12h agoHN ↗

Well theres at least two different buckets of this.

First is the scraping of the open internet.

The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.

Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.

The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.

12h agoHN ↗

And yet a third bucket is the license-laundering of GPL code when the entire github corpus was vacuumed up.

12h agoHN ↗

Copy one book, and you're a thief. Copy thousands, and you're a VC.

12h agoHN ↗

Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data

We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.

12h agoHN ↗

and they would definitely not give it back for free.

...not like they are doing it for free now either.

open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.

I put "stolen" in quotation marks because it's still unclear if we can call that stealing

It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.

This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.

They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.

So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.

All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.

Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.

Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.

12h agoHN ↗

Regulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would.

If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.

12h agoHN ↗

Regulation can mean all sorts of things, including declaring the models themselves illegal.

Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.

11h agoHN ↗

Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".

It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.

We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.

EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.

So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.

11h agoHN ↗

The multiplication comparison makes no sense because humans don’t learn by multiplying numbers to change weights. It’s like comparing the lubrication oil consumption of a car to the cooking oil consumption of a human to compare the carrying capacity. That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn. Otherwise, a human would learn much less in their whole life than a GPU does in one second.

10h agoHN ↗

That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn.

The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.

9h agoHN ↗

The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard.

Yeah, which is why it's only used to compare cars etc. among each other. Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.

In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?

Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?

8h agoHN ↗

It was a marketing thing at the time to make steam engines look good.

Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.

In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?

Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).

Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?

Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.

EDIT: Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?

To answer this in more detail: HP/CC and all other measurements were created to equalize with human specific metrics. Bridges, for example, have safety measured based on how much weight it can sustain at any given time (called load limit). Weight, in this specific case, is an equalizing measurement (it can be in tonnes, kN, PSF, Pa etc). A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.

8h agoHN ↗

Except in your own examples it can easily be shown that it is not handled differently.

I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.

Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?

A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.

And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?

EDIT: Maybe let's just go back to the original point

We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.

So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs, if the way they learn is fundamentally different?

Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?

10h agoHN ↗

"It is stealing. A human paid for the book, compensated the author and learnt from it."

I went to the library. Didn't pay a cent.

Now what?

9h agoHN ↗

Library paid the authors.

Didn't pay a cent.

Taxpayers did pay on your behalf by funding the Library via the Government (if Library is public).

Nothing is free. Except ofcourse stealing, which is free.

8h agoHN ↗

Do you not think that stealing requires the original owner to lose access?

If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?

Can we not just stick to calling it copyright infringement?

8h agoHN ↗

Do you not think that stealing requires the original owner to lose access?

No.

If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?

Not mine. But the company that made the coffee machine. It is stealing IP.

Can we not just stick to calling it copyright infringement?

It is just a fancy way of saying you stole someone's IP. You can call it infringement if it makes you feel good. But the act is the same end of the day.

8h agoHN ↗

Sunshine, air, wind, gravity, radioactivity, public domain, knowledge passed on

8h agoHN ↗

I would not classify sunshine and air as "free". Sunshine requires the Sun to undergo continuous fusion. It is invaluable as opposed to "free". Air is invaluable asset too. Without both we would be dead. The price of both is literally price of my being alive every second of my existence.

Public domain on the other hand is legally only possible if/when copyright has expired. That means the owner has enjoyed proceeds from copyright protection for more than his own lifetime. That is fair. It is still not comparable.

EDIT: since you tacked on more like "wind, gravity, radioactivity" etc, I would still not classify them as "free". They are invaluable to very existence of life.

"Knowledge passed on" is also after someone (in ancestry) has paid for it through blood, sweat and tears. It isn't "free". "Public domain" is legally recognized form of "knowledge passed on".

7h agoHN ↗

They all have “some price”, because of universal laws, but i doubt you would be friends with a corporate entity claiming the ip of them.

10h agoHN ↗

EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.

News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.

9h agoHN ↗

News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.

What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.

EDIT: Let me clarify this further. A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well. Else it would be, using your own words, "xenophobic" against humans to have their copyrightable works be consumed by any species (other than humans) while the reverse is not recognized by Law.

https://www.reinhartlaw.com/news-insights/only-humans-can-be...

9h agoHN ↗

A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well.

That does not follow in any reasonable way.

9h agoHN ↗

That does not follow in any reasonable way.

"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.

11h agoHN ↗

Break up the multi trillion, multifaceted companies into the their separate facets. These companies all thrived on far significant smaller portfolios in the past. Cap the size a company can grow. Regulate the amount of compute they are allowed to use.

10h agoHN ↗

None of this actually stops the profiteering, or reverses the harm already done.

10h agoHN ↗

Legally mandating that companies open the weights of, say, 18-month-old models might change their minds a little.

I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.

10h agoHN ↗

When a human learns, they carry that forward into their future ventures. They might use their learning to recreate the original work and profit from it without attribution. In the West we view that badly and have legislation to protect from some abuses. But the same person might later collaborate with the original author, or make a derrivaive work that improves on the original (a la most science).

AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)

10h agoHN ↗

Would regulation help with that?

I mean it's already happened, right?

I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...

10h agoHN ↗

A simple law stating that laundering data through an LLM does not constitute "fair use", that the outputs can be subject to copyright claims of the original authors, and that the outputs themselves do not qualify for copyright protections would go a long way.

It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.

12h agoHN ↗

sell it back to us at a mark-up.

What if it was for free, like Wikipedia?

Crimes this large are crimes against humanity.

jfc no, sit down.

Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.

At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)

It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too)

  Example:

Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.

That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?

Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?

God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.

(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)

If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.

12h agoHN ↗

Defending the same entities getting large DoD contracts to use AI for killing?

12h agoHN ↗

They've been using computers for killing for decades, who's taking up pitchforks against computers?

Who's taking up pitchforks against THEM?

How do we keep getting deflected into hating the TECH instead of the people who abuse it??

12h agoHN ↗

Agreed, but now they're trying to stop other people from copying data so that they can be the sole gatekeepers of humanities collective knowledge.

11h agoHN ↗

And they should get effed. Distilling should be a right.

11h agoHN ↗

"copyright". its right there in the name of the rights.

12h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity.

Yeah the introduction of copyright was truly criminal.

So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

Oh wait ...

12h agoHN ↗

“It is said that at the heart of every great fortune there is a great crime”

lol at this edgy 5th grade statement. So ridiculous.

12h agoHN ↗

Owning ideas with copyright and patents is what separates the United States from communism.

The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.

It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:

To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;

I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.

[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...

11h agoHN ↗

In a communist society there is no profit (or incentive for), thus no need for copyright laws.

It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.

11h agoHN ↗

Well it's worth reading the linked wiki article section, which includes a link to another article "Copyright law of the Soviet Union", flatly contradicting you unless you maintain that the USSR wasn't truly communist.

11h agoHN ↗

unless you maintain that the USSR wasn't truly communist.

Didn’t they themselves say they were a socialist society on the path to communism?

Does anyone think the USSR was communist?

10h agoHN ↗

The USSR was a socialist state, which is different.

11h agoHN ↗

In the United State, the individual or corporation owns the invention. In the Soviet Union the state automatically owned the invention. That clause is what ensures private ownership.

The clause is what ensures profits from market sales or licensing of ideas go to the creator.

Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.

10h agoHN ↗

I have private ownership over many things that I've bought which are not inventions and are not infringing copyrights.

11h agoHN ↗

In the Soviet Union the state owned all media (de facto) and you had to get permission from a party official before you published or copied anything. It's not the same thing as a free for all.

It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.

In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.

12h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.

Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.

11h agoHN ↗

You're ignoring that there are options other than theft. Heard of buying things?

11h agoHN ↗

Do you think OpenAI has a New York Times subscription? Do you think that's actually relevant to what the New York Times is arguing here?

The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.

11h agoHN ↗

This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn.

Incorrect. You overlooked consideration.

11h agoHN ↗

Learning isn’t stealing, but if you’d set up a forum or hotline where paid staff would answer questions and write essay directly based on NYT content without a license for that content, that would probably be deemed illegal.

7h agoHN ↗

I would highly doubt it, because news outlets and blogs rehash each others' reporting all the time. How many times have you seen an article start with "Today <XYZ publication> broke the news that..."?

Since this is a copyright fight, rights extend only to verbatim copies of the full work or significant portions thereof. Abstract things like facts, ideas, concepts, themes, and patterns are explicitly not protected, and rightfully so. Yet those abstract things are what get repeated and distributed, and are what get encoded into model weights.

This is probably plaintiffs' biggest challenge because it has been very hard to get models to regurgitate entire works except for a very small handful of extremely popular works (and now there are guardrails against even that.)

11h agoHN ↗

It is the largest democratization of knowledge that ever happened.

The 'sell it back to us' argument falls short in my view.

Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.

The comment here seems incredibly pessimistic and quite dramatical.

11h agoHN ↗

Free versions are abundant,

Awesome, can I make my own competitive LLM, just like I can make my own open source software?

and in some time useful models will ship preinstalled on all mobile phones.

Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.

11h agoHN ↗

No. I also cannot make a competitive computer, yet can use one for my work to remain competitive.

Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.

But what are 10 years in the grand scheme of things?

Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

11h agoHN ↗

Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?

Nah, we should have:

1. invested less, in a more targeted way

2. ideally like ARPANET, with the benefits given to all humanity

3. with fair royalties paid to all (where relevant)

4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM

ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.

Exactly like the internet.

11h agoHN ↗

and never developed AI, because it will take a decade to disseminate the benefits to everyone?

Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.

You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"

We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.

The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.

11h agoHN ↗

I couldn't disagree more.

What I expect to see in 10 years is breakthroughs in science and medicine, with major diseases becoming treatable.

You can disagree, but on what basis do you think your predictions are more likely to be correct?

Humanity has turned out fine despite the invention of the book, the TV, and then social media. I reckon kids will develop just fine with AI too.

11h agoHN ↗

It's absurd to equate books and social media to being akin. And you're being facetious in your comments.

What if you are wrong? What are the risks vs rewards?

11h agoHN ↗

Can I make my own competitive LLM?

Yes, yes you can.

See the number of startups that have finetuned or trained an OSS model to build their business on.

“But you have to have compute!”

Ok, and you’re writing OSS on a rock with no internet connection?

The world has all kinds of barriers, but if anyone has a chance of competing on the LLM from it’s not going to be by getting rid of fair use.

8h agoHN ↗

Edit: There.is.no.hope.reality.has.left.the.building.

Oh, well, I guess I'll just wait for the bubble to pop for the over-excited SF fans to cool down.

I'm having a TON of flashbacks about cryptocurrency discussions.

11h agoHN ↗

It is woefully pessimistic indeed.

There will be new jobs coming from this.

The bountiful abundance of intelligence is truly the best thing that has happened this decade.

11h agoHN ↗

Wiki already democratized it just fine and was legitimately free for people who know how to read.

It's asinine that you think the sell it back to us argument falls short.

Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.

From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.

11h agoHN ↗

Wikis made knowledge available.

Making it accessible, understandable, and usable is another matter.

How LLMs sound is not a fundamental limitation of the technology.

The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.

You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.

11h agoHN ↗

Making it accessible, understandable, and usable is another matter.

I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.

9h agoHN ↗

Cool. Then don't use LLMs and keep waiting for that prophesied "model collapse" to happen. See how that works out for you.

8h agoHN ↗

I'm more worried about the collapse of humanity a la idocracy. The model collapse might just be secondary. Even just conversing on HN, that's the vibe I feel we're headed towards. Or, perhaps, I'm just running into more and more posters that have financial ties to the success of LLMs.

11h agoHN ↗

There are no viable free versions with sufficient computing power. The do-it-yourself AI is dangled as a carrot in front of users to camouflage the lock-in and rent seeking by Big AI. Go prove something new like Navier Stokes (replicating N-S itself no longer counts due to scraping and plagiarism) on your Mac Pro!

Paid influencers who perpetuate the open narrative are a whole new industry.

Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".

5h agoHN ↗

There is no lock in. Anyone with sufficient budget can download the weights for GLM 5.3. On OpenRouter, I count 29 different providers for this model [1].

In terms of purely local LLMs, one can run GLM 5.3 flash on a beefy workstation.

[1] https://openrouter.ai/z-ai/glm-5.3

11h agoHN ↗

The comment seems in line with real life, while your catchphrase about democratization is more like wishful thinking.

10h agoHN ↗

Ah, this remembers me when I was also young and naive.

Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..

11h agoHN ↗

Works produced by AI are not copyrightable. Are they in the public domain? Can an entity sell unique works which are in the public domain?

11h agoHN ↗

Yes? I own a lot of books that were public domain when published (as reprints). I could read them on Gutenberg, but I'm paying for the nice paper formatting.

11h agoHN ↗

Our "culture" has long been the province of corporations. In prior epochs it was still the product of patronage and power.

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?

11h agoHN ↗

Flourishing life? Where is your data leading to the conclusion that AI models are leading to the world populace leading a ‘flourishing life’? Are taxes on revenue of companies like OpenAI somehow being collected and turned into a UBI and I just didn’t hear about it?

11h agoHN ↗

I said a glimpse. I didn't say it's here now or it would be easy.

It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.

11h agoHN ↗

Like trickle down economics?

We all experience the substance which is trickling down.

11h agoHN ↗

Or, perhaps, the companies forming the for profit LLMs could be made to pay a license fee for the copyrighted data they lifted from behind the paywall and then used to form a for profit entity with.

Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?

11h agoHN ↗

This is solid but how many people you reckon would benefit if they paid every single penny for any copyrighted data? I am in my 50's with 30 years in the industry and I know maybe 2 people that might benefit from this... certainly not something general public would benefit from

11h agoHN ↗

We’re likely looking at a new world where physical and mental work from humans holds much less value.

So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.

So our future world probably looks like a system where effort/knowledge/skill has little to no reward and capital (e.g. inheritance, passive income) has all the reward.

So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.

11h agoHN ↗

Why more and more stories from "Pump Six and Other Stories" are becoming real step by step?

10h agoHN ↗

You are missing the obvious answer.

The state runs the system. Or really at a certain point, the state runs itself. In 100 years or so, the CEO becomes a barbaric relic of the past age, like our ancestors who beat each other over the head with clubs. Nevertheless, such a future will necessarily owe a debt to both.

Wealth and states are often connected. In this case if the state is properly aligned (this is what the coming battles are to be about - as they always have been) to be egalitarian, then redistribution can occur unblocked by human enclaves of greed.

11h agoHN ↗

It is achievable in theory yes. But it was also achievable without LLMs and yet, here we are.

11h agoHN ↗

I mean I'm producing more and better art and software each month than I used to do in a year. I feel like I'm flourishing because of AI. Tell me I'm wrong?

9h agoHN ↗

You’re not wrong. Me too. I spend most of my waking hours using AI for new product development.

But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?

I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.

Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.

I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?

9h agoHN ↗

is QWEN a 4d chess game to replace capitalism with communism?

No. The Chinese labs pay much less to train their models because they distill Western models. Western labs can't do it that way because they'd immediately be sued.

To distill a model requires high-quality prompts (and responses). Here's how the Chinese labs get these prompts: when a user sends a prompt to www.kimi.ai, Kimi immediately forwards it to Claude or another leading US model (using a vast network of laundered subscriptions), waits for the response from Claude, then forwards that response to the user via www.kimi.ai. All the Chinese labs do this. Of course, they never inform the user that their conversation is being routed to a US model.

Beijing is angry at them because they did this indiscriminately without filtering out the conversations of high Chinese government officials and military officers, so now the US labs have those conversations, which contain information useful to US intelligence because the officials and officers assumed the conversations would remain inside China's borders and that the Chinese labs would care about confidentiality.

The Chinese labs release model weights under permissive licenses to get any attention and usage share at all for the model.

8h agoHN ↗

The Chinese labs pay much less to train their models because they distill Western models.

That's nonsense.

6h agoHN ↗

I'm going to take your comment as a request for supporting details.

I believe what I wrote is an accurate summary of information in a report published this month by Anthropic. I haven't seen the report, but have seen this next summary written by someone who I trust to summarize accurately: https://thezvi.substack.com/p/the-bad-guy-with-an-ai-named-c...

OpenAI reported similar behavior last year. The assertion that the other (two) leading US models are being distilled in a similar manner is an assumption on my part. Where I wrote, "Claude or another leading US model" I initially had just "Claude", but it felt silly to imply that the same thing is not happening to the other leading US models.

1h agoHN ↗

I'm happy. You may think my art sucks, but five years ago I didn't have the resources to create it, whether or not it sucks.

11h agoHN ↗

Considering that we don't seem able to appropriately tax the biggest corporations, why do you think we'll suddenly be able to redistribute LLM gains?

Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.

11h agoHN ↗

These doomsday theories sound solid in practice but in order for someone to accumulate a lot of wealth there have to be consumers to chip into this. If we are all slaves, where is that wealth going to come from?!

11h agoHN ↗

The Amazon warehouse worker is a modern day slave.

Don’t believe me? Try it. I have.

10h agoHN ↗

Not necessarily as there are methods to redistribute wealth without the consumers having any choice.

Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.

Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.

There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.

11h agoHN ↗

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Let me fix that for you: A world in the the value of the labor of each human approaches zero.

11h agoHN ↗

Right, and the consequence is going to be, what?

Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.

It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.

11h agoHN ↗

We are not deluding ourselves that we will maintain our standard of living. If the boss’ dream of replacing all workers by AI that does their job poorly but cheaply will be realized. We will go to the gutter with the rest of the economically worthless people.

There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.

11h agoHN ↗

Correct.

Humans are animals. We have rules to keep humans in check but that’s it.

The wealthy do not care nor need to care. They will justify it as ‘you lot were too weak to organise and stop us’ and there is an element of truth to that.

10h agoHN ↗

I still see a lot of denialism.

It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.

As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.

Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.

10h agoHN ↗

“I find hope in your acknowledgment of your own imminent devaluation.”

Gave me a chuckle. I think this is my personal ‘best sentence of the year’.

Great way to summarize how we will likely look back on 2026.

9h agoHN ↗

Well... it's true, right?

People can't help themselves when they don't believe they're being exploited.

10h agoHN ↗

You assume there will still be free elections and democracy. There is also the possibility that the form of government will change into something else.

11h agoHN ↗

The value will tend towards infinity [1] because we can keep building more data centers. This will allow us to apply tokens to solving every problem we can conceive.

The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.

If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.

Industrialization led to iPhone, PlayStation, Spotify, and Waymo.

AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...

AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.

[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.

[2] Until the robots leave earth, anyway

9h agoHN ↗

but the S-curve will look like it for awhile.

Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.

Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.

9h agoHN ↗

[2] this is why we need georgism, i.e. shifting taxes to finite resources like land, minerals, etc.

9h agoHN ↗

AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...

You need to show how having X more data centers is somehow going to translate into affordable, highly advanced robots in the immediate future.

There are hard problems about robotics we don't yet know how to solve. Not to mention, who the fuck is going to buy them if AI takes their jobs?

7h agoHN ↗

You think people are going to stop watching celebrities and influencers just because there are AIs and robots?

Do you think people will stop dating, trying to impress mates, buying luxury, etc.? Neither candlelight dinner to a robot no a Dior manufactured by robots have the same allure. There's a huge industry around this.

Sports aren't going to go away. Huge industry.

People aren't going to stop making art. I know a ton of artists who have embraced AI that are doing even bolder work using the tools. (I was a filmmaker pre-AI, and I know a lot of people in this field.)

People aren't going to stop traveling. And consuming. And eating human food and consuming human experiences.

There are going to be all new kinds of businesses and opportunities that spring up. OpenAI and Anthropic are not going to be the only two employees. They won't be staffed by only agents.

7h agoHN ↗

The economy cannot just transition over night, or even within a few years, to quickly absorb a labor shock the size you're describing.

Also, you have not addressed how we're going suddenly to develop the robots required for your future.

and consuming human experiences.

Consumption requires money. If AI is going to massively impact employment, that reduces the amount of people in the economy capable of consuming these "human experiences".

30m agoHN ↗

the value will tend towards infinity

It may be the case that the potential output of a worker engaging in the productive process goes arbitrarily high. But the that won't matter, if they're not permitted to. Given a choice between involving a human worker who will demand compensation, and a fully general robot, which will the owner class choose?

The vast majority of human intelligence is already squandered: millions of potential geniuses in impoverished places, suppressed by lack of opportunity to flourish. It's not about the quality or quantity of intelligence. It's about who controls it.

The fruits of the industrial revolution didn't end up in the hands of the worker by divine grace, or by some natural law. They were won by the hard struggles of the labour movements, by leveraging their indispensibility to the process of production.

If we want the utopia you imagine to be accessible to ordinary people, workers must cease being so eager to build their replacements, and be prepared to collectively struggle for their share. But if we are lulled to complacency by the notion that this will be a passive process, that struggle will not be necessary, then prospects are grim.

11h agoHN ↗

In the 50's there were predictions that in 10 years no-one will have to work again because of the advances made in automation, like the washing machine for example. Any predictions of less work this time around are a complete joke.

11h agoHN ↗

For thousands of years humans have dreamed of reaching the stars.

Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.

On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.

You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?

11h agoHN ↗

When was the last time you saw a Chinese laundry?

10h agoHN ↗

That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.

11h agoHN ↗

Now the mother and father both have to get jobs, in the future mother, father, and child will get jobs to support the trillionaires.

11h agoHN ↗

At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.

Fantasy lets us glimpse anything.

11h agoHN ↗

US has been procrastinating reparations for slavery. LLM redistribution can't happen until the capitalism machines recursively solve "sins of our fathers".

11h agoHN ↗

a world in which the labor required of each human to lead a flourishing life approaches zero.

The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.

9h agoHN ↗

Money will be created to justify the valuation.

11h agoHN ↗

a world in which the labor required of each human to lead a flourishing life approaches zero.

Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?

Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.

11h agoHN ↗

Is this "escape route" in the room with us now?

11h agoHN ↗

If piracy isn't stealing training definitely isn't stealing.

11h agoHN ↗

So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.

Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.

Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.

The more fundamental objection seems to just be to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.

8h agoHN ↗

Would that resolve the injury caused by the intellectual osmosis?

Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.

What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.

Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.

11h agoHN ↗

Musicians and writers build new works on top off millennia of literary and musical history, then copyright their works and sell it back to us. Is this so bad? If Taylor Swift, consciously or subconsciously, gets an harmonic idea from a 1970's song and a fragment of a melody from some 1990's song... is that theft?

I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.

But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \

It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.

I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?

10h agoHN ↗

The problem is not Taylor Swift is being "inspired" from other artists. We have tons of examples this throughout history.

The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.

This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...

We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.

Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.

10h agoHN ↗

Ok, thought experiment: at what point in the past 1000 years do you think the balance of collective versus individual benefit from creative works was at its most fair?

getting things were free was part of sending a message, not monetary gain.

As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.

5h agoHN ↗

We’re about the same age, probably. On this side of the pond, all of this stuff was rooted in inability to get something legally first, and if it was possible to get, it was prohibitively expensive to buy for us. So we got it for “free”. The thing is we were not using the software to earn money, so companies don’t care (this what I heard from a couple of companies directly). Music was in a similar position.

It was a righteous moral stand because we wanted to get things relatively affordable for us.

I for one prefer to buy my software and music nowadays, because it’s affordable and I can get it at the quality I want.

For your other question, the answer is probably 1800s, because the current model was not entrenched everywhere and creators had sane rights for what they created. So creators had to get what they made out to masses to show it, but lost their rights in relatively shorter times, so things were free to use for everyone. So you can’t excessively milk something till proverbial eternity.

1h agoHN ↗

On this side of the pond, wow were the college kids excited when they discovered they could reframe piracy as a political and countercultural statement, despite things being available paid.

At the time I pirated a lot of stuff. If we're of the same age, you may well have played video games I cracked in your teens. And, TBH, for me it was mostly collector mentality (I have have all the games!) and a bit of poverty (I can have the games I'd like to buy but can't afford), and about zero politics.

That said, I'm generally with you on 1800's. It's astounding how fast things changed. IIRC it wasn't until 1890 or so that international copyright was even a thing; you published in your home country and publishers in other countries just copied and published without permission or payment.

And then just 100 years later, copyright was essentially perpetual and global.

10h agoHN ↗

when you do your job you take ideas from others before from university from calculus etc

therefore you have to work for me for free

10h agoHN ↗

Oh I'm sorry, who's forcing you to work without pay?

9h agoHN ↗

I write books for a living, when you buy them off Amazon I get 7 USD. when someone else gets them off libgen I get 0 USD.

However if I put a gun to their head and demand to live in their house or give me food for free they get mad.

Its lovely how everyone deserves to get paid except writers / musicians cause we are only fucking hippies when it comes to that.

9h agoHN ↗

The point of the argument is that "forcing you to work without pay" is silly.

10h agoHN ↗

I agree with you, but there is one aspect that is different here - the scale that is way beyond what any human can do.

I don't know what to think about that though.

10h agoHN ↗

Fair point. Think of it this way: the scale is more than any human, but it is probably akin to what humanity in general can do.

10h agoHN ↗

Yes. And even if (when) AI surpasses humanity in general I don't know if that changes the argument.

10h agoHN ↗

You might argue that since those LLMs build on all of humanity's knowledge, they should belong to everyone. But where do you draw the lines? Or make a practical case.

10h agoHN ↗

They kind of do belong to everyone. The technology and majority of training corpus is available to essentially everyone.

The training and inference infra are owned and operated, but I can't see an argument that any machine that processes public domain (or stolen, if you prefer) info should be available to everyone for free.

But those training and inference costs will go to zero. Think about your cell phone today versus $1m+ supercomputers in the 1980's.

We're living in a transitory blip where capitalists and gold rushers are getting rich arbitraging the cost of processing against the non-cost of corpus. We can argue about morals (it doesn't bother me much) but it is a narrow window and it will be remembered the way Compuserve is: a precursor to the actual revolution, worth a footnote.

10h agoHN ↗

Sure, that's what I mean by where to draw the line. An artist taking inspiration from others is not obliged to give up the work to the public.

Sidenote: It may be tricky/impossible in the future to uphold intellectual property laws. If anyone is able (for instance) to prompt-create all their software, a software patent is worthless.

10h agoHN ↗

Robbery implies taking it away from you so you can't have it anymore. How is that the case?

10h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up

Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.

10h agoHN ↗

That's true. However it does sometimes feel like our culture is being buried in slop.

10h agoHN ↗

Imagine someone asks you how to do something at work, you tell them. Then they create a huge packet filled with bullshit about how it got done and now they are your boss.

9h agoHN ↗

It's a shitty move, but ultimately, in between the bullshit narrative, they also did the thing - not you - so the promotion rightfully belongs to them.

Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?

9h agoHN ↗

Sure thats obviously how it works today.

Now what do I do next time someone comes and asks how to do something?

3h agoHN ↗

Of course you tell them, because you're nice person, and not jealous of someone else succeeding in a thing you aren't even pursuing? You wouldn't want to act like "the dog in the manger" from childhood stories.

3h agoHN ↗

LOL I have no idea about the dog in the manger. You have my full attention, please tell me the story.

10h agoHN ↗

Isn't that the same point piracy advocates have been making for a while? If I watch a pirated film, I didn't really consume a physical resource. Nothing physical is lost. Therefore it's not theft?

Learning isn’t stealing.

This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid. It's a public good to have information available, but it's not really acceptable to circumvent established ways of compensating the creator of the work you're benefiting from.

10h agoHN ↗

Isn't that the same point piracy advocates have been making for a while?

No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.

Learning isn’t copyright infringement.

Learning isn’t theft.

Copyright infringement isn’t theft.

These are all different statements, and all are true.

There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.

“Gaining exposure to” isn’t theft either.

Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.

There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.

10h agoHN ↗

Partially true.

If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.

If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.

If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.

If people stop making art because they can't outcompete generated art, that is also lost to LLMs.

So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.

10h agoHN ↗

If $X dies because ..., the entirety of its knowledge is now only availabe through LLMs that trained on it.

And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.

The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.

10h agoHN ↗

You are mixing up culture, data, and the creation of new works as if they are the same thing when they are very different.

10h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up.

Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.

10h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up.

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.

Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.

And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.

There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".

10h agoHN ↗

Creators are unwillingly and contra to economic systems that have evolved over a few millennia entered in the Borg or the Matrix. So their achievements are reused for private and public benefits without their permission. Piracy is theft. And this piracy is a bigger theft than piracy on an individual download basis. Was there any doubt? (Linguistically I agree that you can’t “rob culture”. You can rape or reap culture though, and that is the point at hand.)

10h agoHN ↗

What do you mean, it's already robbing me of the ability to discuss techniques with a number of my peers since they decided the mediocre output of LLMs is good enough instead of understanding what they are doing.

In addition, a great number of people have decided to stop publishing their code publicly, and discuss techniques except in spaces they can be sure it's a one-on-one with a human being.

Some things already got lost.

10h agoHN ↗

I still do open-source. You just get the code on a CD like the good old times!

10h agoHN ↗

You're discussing your peers' behaviors, not AI.

It's the same as saying "People don't kill people; guns do." Maybe guns(AI) should be regulated, but the primary problem is murderers(lazy peers), not weapons(AI).

10h agoHN ↗

Just like with guns, AI is a crime of scope.

I can be a murderer with a knife, sure, but how many people can I kill? How many people can I kill if I have a class three license an a full auto gun?

I can pirate a news article here or there, sure, but how many news articles can I pirate? How many news articles can I pirate with AI?

10h agoHN ↗

How many news articles can I pirate with AI?

Not much more than with a bash/curl loop. Which, on one hand, AI can write for you, but on the other hand, AI will make you no longer need to pirate those articles in the first place.

Also it's not end-user piracy being discussed - it's the act of training AI itself that's accused here to be "robbery of our culture".

9h agoHN ↗

Yet countries that have strong regulation of firearms have significantly less murders (and accidental shootings) than countries that don't regulate them.

That does suggest that along with murderers, the easy accessibility of firearms is a problem (assuming that the ratio of murderers in different populations is approximately the same).

9h agoHN ↗

If this is your stance on guns, I’m not surprised by your calloused take on AI

10h agoHN ↗

That's people problems though, namely:

1) Lazy peers, and

2) Spiteful peers with "dog in the manger" mentality.

The latter case personally irks me, mostly because this often involves personal benefit (economical or moral) due to them claiming to give away knowledge for altruistic reasons, which later behavior reveals was a lie, mislabeling proprietary as open and free, to reap unfair gains.

But that's all off-topic for this thread anyway.

9h agoHN ↗

Spiteful peers with "dog in the manger" mentality.

Last night my daughter told me her friend (a boy) signed up to take the last spot of a baking class at camp just to spite his sister, who wanted it.

Turns out he liked baking!

---

I think this whole discussion shows that there's a lot of Calvinball going on with IP, where the creator keeps trying to change the rules as a way to defend their power and social order.

9h agoHN ↗

People creating free software for everyone to use, only to pull it away when suddenly there's a technology that can instantly let everyone use it at the drop of a hat, is not something I quite understand. If you're doing it so you show people, hey I made this, then you can still do that. If you're doing it to help the world, then that's very much happening. Of if you're doing it just for fun, that's possible too, though I totally understand how much less fun it is when AI can rip it out in a few minutes.

9h agoHN ↗

People have multiple motivations at the same time. Many OSS contributors have moved on after LLM’s and that’s completely consistent with their motivations.

Some of the major OSS licenses require attribution when someone copies the work in question, AI companies fail to do so. People getting annoyed when others break the social contract has nothing to do with any other motives they can always benefit the world by doing something else.

You may prefer someone continue maintaining infrastructure you use but they could be just as happy delivering food to the elderly. The desire for people to continue helping you is little more than entitlement here.

8h agoHN ↗

Some of it might be loss of external validation, no more stars in github. Some of it might be an aversion to a non-consensual contribution to the end of their chosen profession as a means to make a good living.

Much of it might be because their license (MIT, BSD etc) was ignored completely.

7h agoHN ↗

is not something I quite understand

I can only speak for myself, but right now working on open-source is even more thankless and shitty. Now any douche-knuckle comes to your repo and starts spamming you with bad PRs or issues as if their favourite LLM du jour has some fantastical insight they have to share with you, when it has already been explained ad nauseam that what they want is out of scope, or isn’t technically possible, or has tradeoffs XYZ.

You had lazy people before, sure, but those interactions were less frequent, shorter, and for each person I used to craft a careful response that would explain everything in detail and teach those I interacted with. But why would I waste my time crafting a reply that the person on the other side is just going to shove into their LLM and learn nothing?

Regarding open-source, LLMs make people more disrespectful of the people doing the work. And if you’re sharing your expertise for free to be disrespected, why bother?

7h agoHN ↗

People were always able to "instantly use it at the drop of a hat" (paraphrased). What people are objecting to, is instant use without any acknowledgement or attribution (the minimum courtesy). They also object to being mistreated by people who have never actually had to type characters of code into a text-editor, or figure out how to use a build-system and version-control-system. And finally, they object to having their _free_ and _generous_ labor turn millionaires into billionaires (and billionaires into trillionaires), while they themselves still have to pay that one mortgage they have, and medical bills, and tuition, and so on.

They object to the systematized devaluing of their work. They object to becoming a class of invisible laborers (who only become visible when there is need to blame someone).

8h agoHN ↗

That's people problems though

And the people make the culture. So, in effect, the culture suffered. I think you’re taking the original comment too literally with your “piracy is theft” comparison. Obviously you can’t literally rob culture¹, because culture is a concept. But you can lose something from it, and that loss can have a cause you can point to, so in a way elements of it can be “robbed”.

That’s what I get from the original comment when I steelman it, anyway.

¹ Sounds like the name of a Rob Zombie associate.

9h agoHN ↗

The majority of code has always been proprietary and never published. But there's also more open source code being published than ever before.

8h agoHN ↗

a great number of people have decided to stop publishing their code publicly

Stopped publishing anything and stopped creating anything.

But in all fairness this started long ago. I remember a conversation where I asked someone who read a niche blog every day why he never posted a comment. (the blog had zero comments) If you like what they wrote and have thoughts about it, why not "honor" them with a few words? How is that to much to ask? They clearly just never bothered to but in stead made up all kinds of excuses.

Imagine having to explain that you can respond if someone talks to you? If there was a culture surely it ended there?

7h agoHN ↗

> In addition, a great number of people have decided to stop publishing their code publicly,

A lot of my peers who have been in tech for more than a decade are all doing this. They're effectively backing out of the industry. Their attitude now is they're ready for retirement in their 40's. They're only building and working on side projects and only sharing with people they know, nobody else.

One thing I'm seeing right now is the amount of institutional and domain knowledge being lost is a on a massive scale. The more startling fact is nobody seems to care - as if vibe coding and LLM's are going to fill this gap.

7h agoHN ↗

It's interesting to me that the actual motivations for releasing code open-source were very different than I had assumed for the past ~30 years. I always liked to give stuff away so folks could play with and try stuff. As far as I can tell, the only reason not to still give away code is the desire for credit, which was never all that interesting to me.

But then again, I haven't devoted my life to developing and releasing a huge open source project used by millions, so perhaps my opinion should be discounted accordingly.

2h agoHN ↗

No, I don't think your opinion should be discounted. It just turns out that a bulk of open-source movement had ulterior motives, and were actually in this for credit or downstream opportunities, not to actually give things away in a pay-it-forward fashion.

The reaction to LLMs pretty clearly revealed who was giving things away, and who was indirectly selling them and using "giving away" for unfair advantage.

10h agoHN ↗

Agreed. Jacques is usually a good poster with insightful comments. This one fell short of that mark. It is suprising.

10h agoHN ↗

it's still there

There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:

Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.

To use another Slashdot cliche: "But you knew that already."

9h agoHN ↗

which most people use them for

Citation needed

10h agoHN ↗

AI didn't steal any culture. It just made mediocre culture more accessible. Turns out, most people do like mediocre culture. Previously, public TV channels could at least pretend that people enjoy educational content or classical music. AI exposed that lie. And now we're shocked that the emperor is naked.

9h agoHN ↗

If we ignore the massive amount of books being destroyed, the outages and increased usage bills for once free websites (some resulting in closure), the increased difficulty to access public information such as Reddit and Twitter, and lastly ignore the amount of conversations being taken away from the public in favor of LLMs, I would agree.

9h agoHN ↗

Except books are already destroyed in physical form every single day. My local library has a section near the front for free and/or extremely cheap books ($1) they are desperately trying to get rid of. My impression is that if they are untaken/unsold they will end up in a dumpster.

[Edit] Here's an article from 2003 talking about how Random House destroyed up to 25,000 books a day: https://www.theguardian.com/books/2002/mar/19/fiction.stephe...

8h agoHN ↗

They are destroying the foundations of civilization

No, they are not. For all the hype around that 404 media article, no one came forward with titles or authors. Turns out the booksellers that sold them those books said they were "dead inventory".

If you're going to assert that companies are destroying the foundations of civilization, you're going to have to produce quite a bit more evidence than the (poorly cited) 404 Media article.

8h agoHN ↗

Evidence ? Good lord, how about the US district court of justice ? Anthropic physically destroyed millions of books after scanning them - are you really, truly questioning that ?

Here is what was signed by Judge Alsup at https://docs.justia.com/cases/federal/district-courts/califo...

Anthropic spent many millions of dollars to purchase MILLIONS of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.

Dismissing the above as "dead inventory" is atrocious reasoning. The more obscure/low-circulation books are arguably MORE valuable as training material because they contain information that is less likely to be duplicated across the ordinary web.

And 404 Media's reporting wasn't merely based on a bookseller's terminology of your "dead inventory". There are now unsealed Anthropic internal documents describing Project Panama as an effort to "destructively scan all the books in the world"

https://www.washingtonpost.com/podcasts/post-reports/the-que...

7h agoHN ↗

There are now unsealed Anthropic internal documents describing Project Panama as an effort to "destructively scan all the books in the world"

If I do something destructive to my copy of Accelerate -- feed it through a scanner, leave it where the munchkin can reach, whatever -- that does nothing to deprive you of your copy. Books are funny like that; there tend to be lots of identical ones.

9h agoHN ↗

the massive amount of books being destroyed

Surely that's not about AI? Or you mean the small fraction of that that's result of "destructive format shift" process[0], which itself is a consequence of copyright regulation that would otherwise prevent anyone from accessing these works?

the outages and increased usage bills for once free websites (some resulting in closure)

That's assumed to be AI companies for some reason, even if they have no real incentive to do that, while the usual business underbelly of people scraping web for whatever reasons (which now may include some AI upstart wannabies too, to be fair) is forgotten about. Not helping is people confusing AI agents acting as user-agents and doing one-off fetches with "AI scrappers".

the increased difficulty to access public information such as Reddit and Twitter

It was never public, and they started locking down before AI, when both platforms (as well as all other social media platforms) run out of VC subsidsies and realized they need to start monetizing; first step they did was to wage a war on third-party clients and API users. AI came later.

ignore the amount of conversations being taken away from the public in favor of LLMs

You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online? I can understand that, it's usually much better experience (particularly on sites like StackOverflow - they dug their own grave here, and they know it; there have been memes about this way before LLMs were a thing).

You're also not considering the amount of questions answered that would not have been asked otherwise. I for one don't ask many questions on-line, so anything I ask LLMs that they solve for me, is a question that would've remained unanswered for me otherwise. Not everyone is gregarious online, many people are more self-reliant and only answer questions, but solve their own problems without asking for help (it's probably not optimal thing to do, but that's another topic).

--

[0] - Digitizing and retaining physical copy is clear infringement, digitizing but destroying the physical original can be argued to be fair use.

8h agoHN ↗

You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online?

Oh, please, AI is single most hated technology to the level we did not seen before. And that is after staggering propaganda going out of these CEO that hiding behind human agency is beyond lie. It was pushed on us, whether we like it or not, because people with a lot of money are the only ones who matter.

And people with a lot of money have that weird cult belief in emerging AI god and singularity, so they dont care what they destroy in the process.

9h agoHN ↗

Your talking about a bunch of books that are out of print and sat unsold for years in a warehouse somewhere. The rights holders could print another run tomorrow if they wanted, or better yet digitize them but they don’t. Where’s your ire for the rights holders who sit on these books and don’t do anything with the ?

7h agoHN ↗

It's the collective action of losing access to information that's the issue, not the potential of the book itself.

Regardless, these books were available. Researchers, hobbyists - many people have and will find value.

Justifying the destruction of information for a for-profit company is bizarre.

7h agoHN ↗

Destruction of books is normal and totally OK. Publishing houses do it daily in the thousands. When you own something, you can choose how to dispose of it. Come back when you have the title of actually valuable rare books that were bought and destroyed. The 404 article was not well researched in this regard.

The emotional reaction we're conditioned to have hearkens back to 1940s Germany, where the goal was to have people burn all the copies of existing works. But this is not that. They are going for "one of each book" so they have a comprehensive corpus of works to train from. They are willing to buy books to obtain this corpus. It sounds like the system working exactly as intended, with even the book sellers selling these books saying they are "dead inventory".

There's a time to be breathless about "destroying all of humanity's knowledge", but let's wait until it's actually happening.

Again, I'm am very anxious for anyone to come forward with credible information about any valuable/rare/irreplaceable book that was destroyed in this process. So far zero people have been able to do this, probably because they are buying the books in bulk and want to obtain them as cheaply as possible. Book sellers will know the value of any actually rare/valuable books, and would charge accordingly, which would be pointless for e.g. Amazon to pay. They don't need some crazy first edition, they just need some copy to digitize. The incentives just don't add up to "Amazon is buying priceless art and burning it!" the way the narrative suggests.

35m agoHN ↗

The publishers and rights holders still have the text, and the books themselves weren’t being accessed just sitting in a warehouse. This is a trivial problem to solve, digitize them.

8h agoHN ↗

Most of these things are bad. None of them is robbery, though. Some are of questionable legality, but most are not.

People need to face reality. The old Internet was fragile and couldn't have lasted long. It was dying slowly due to closed social media and LLMs are accelerating that death.

I'd rather it die quickly and be replaced by something more robust than decay slowly. I was sick of how it was before LLMs even came among.

9h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.

This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.

Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.

They work very hard to protect their own IP so they obviously believe that IP is worth protecting.

9h agoHN ↗

Yep: the problem with IP is exactly the "Property". The texts are economically (even if only for leisure) valuable, so the producer deserves a fair comoensation unless he has explicitly dismissed it.

So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.

9h agoHN ↗

And if attribution is their coin then they should be properly attributed even if they aren't paid.

9h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.

It hasn't been literally stolen but indirectly it has been.

I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.

Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.

AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.

9h agoHN ↗

It hasn't been literally stolen but indirectly it has been.

There actually was theft. See how some big corporations slurped data from Libgen and Anna's Archive. Where did they pay for this?

See also this article:

https://www.theguardian.com/books/2025/apr/03/meta-has-stole...

And similar articles.

It is theft - there is no denying in that.

All these companies owe us a ton of money. But Trump protects the AI mafia so we won't get compensation for the damage they cause here.

9h agoHN ↗

It hasn't been literally stolen but indirectly it has been.

It's literally stolen IP.

One minus epsilon of the corpus did not give informed consent, or get compensated. The fact that it's laundered doesn't make it any less literally stolen.

2h agoHN ↗

Someone once told me, any illegal thing can become legal if you add extra steps. This is not _legally_ true, but it is _practically_ true. Evading a tariff is illegal, until you start using a proxy-country. Firing an employee, and not giving them severance is illegal, until you find a way to make their job so miserable, that they quit on their own. Really, if you can turn signal into _noise_, it becomes difficult and expensive for the enforcement organs to actually enforce their own rules.

It is unclear to me just what percentage of tech-companies, are in the business of adding extra steps to an illegal process, of turning signal into noise. For example, the vapes made by Juul are a kind of hack around public health laws. The Amazon marketplace shields merchants who sell counterfeit goods. Uber and Lyft bypassed the local laws that applied to taxi services and the medallion system. Airbnb did something similar with the laws governing hotels, and delineating who owns and who rents.

It is a special kind of disappointment, to be someone who loves technology, and to have to work in the technology industry -- because it appears to be run by people who hate everything that is not money.

1h agoHN ↗

I'm sad to hear your business model didn't survive the march of technology, but at the same time, this is what's been happening to everyone ultimately. And rightfully, no one is entitled to a business model working in perpetuity.

I know this is a real harm to individuals, and bringing this up is not some kind of anti-technological thinking (the luddites were OG with that, and I always argue in favor of them - they had a very valid point).

I also realize that in a few years, I very well may be in the same position here. So will most of us.

But that's a different argument than GP was making, different one from what I replied to. That was about losing culture, and losing a business model is not losing culture.

Also I was with you all the way until the last paragraph:

AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet.

AI is doing exactly what users want it to - what I too use it for - it bypasses the spam and scams that stand between the user and the solution to their problem. Good riddance, in this.

Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.

And that is just bullshit. AI is not "taking everyone's content and selling it back to users", and the vendors are definitely not "keeping all of the profits" - on the contrary, they haven't even begun to figure out how to monetize almost any of the value their inference services provide, because they have zero visibility into how much any stream of tokens ends up being worth for their customers. They cannot tell whether the plumbing advice the LLM generated unclogged my toilet, or prevented a restaurant from having to close for the day; they cannot tell if the code their agent wrote won me a beer from a friend, or unblocked $2 000 000 dollar opportunity for my business. They capture none of the value from this either way.

9h agoHN ↗

Except, of course, no one has actually been robbed

No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.

9h agoHN ↗

I only see 2 consistent world views: either intellectual property is real, or it is a false concept and all information should be free.

If IP is real, then the AI companies have performed flagrant theft.

If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.

The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.

9h agoHN ↗

I only see 2 consistent world views

Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.

Instead we should try to find the most practical solution that has the most benefit - which will probably be more complex than yes/no.

8h agoHN ↗

Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.

Intellectual property is as real as private property (which is a lot more elaborate and weird than possession and territoriality, which is the most that has any natural basis). It's really foolish to claim one doesn't exist and should be abolished and the other this fundamental sacred thing that should be respected absolutely (as many do).

8h agoHN ↗

Yes, that's why I said property (not just IP) is a social construct.

8h agoHN ↗

We made it up. We made all this up. We do it for an outcome.

Intellectual property exists as a concept to foster the creation of more intellectual property. That is not the case for all private property, because I cannot copy your land or your car infinitely. That’s why IP rights expire at some point, or have fair use that doesn’t harm the IP rights holder, something other forms of property don’t have. Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.

I think AI training falls most likely in the fair use category of intellectual property: there is some societal benefit* that requires no actual harm** to the IP holder, therefore it’s a good trade-off for society if we poke a hole in the social construct of property to get that benefit.

*Let’s put aside the question of whether AI is good for society or private ownership of AI models is good for society. Important questions but separate from the theory. IF IT IS GOOD, it follows the above. If it is not good, then of course it does not.

**Also not a fully settled question. Again, an important debate to have and the tradeoffs here matter. If the harms are small enough, the societal benefit could be worth it. Both notes have to be true for this to be worthy of “fair use”.

8h agoHN ↗

This is going to blow your mind, but people created music, literature, art long before intellectual property laws existed. IP laws are nothing but rent seeking.

7h agoHN ↗

Yes, and we lived in a world of guilds, Orders and secrets which kept their knowledge tightly within themselves.

Music and art was created by patronage. Entire periods o where the art that mattered was made by the powers that moved the world.

For my whole life the goal was to move away from that era, not to see it recreated.

Also, this is a debate that English speakers can enjoy on a website for an US based accelerator. Most of humanity doesn’t even have the standing to be heard in this conversation.

8h agoHN ↗

Intellectual property exists as a concept to foster the creation of more intellectual property.

The idea of IP as an economic tool to foster creation comes out of the UK and subsequently the US. In mainland Europe, IP comes out of the French Revolution and the idea that copyright is like a "moral right" that you intrinsically deserve for putting in the effort to create something. This viewpoint has de facto won out because as global commerce and global culture has become more and more widespread, everybody has standardized on the longest durations (the standard has long been "life + 50 years" in Europe) so one country doesn't have to worry about freeing up its works for "exploitation" by another country.

7h agoHN ↗

That is not the case for all private property, because I cannot copy your land or your car infinitely.

All property is fundamentally about exclusion and control to allow for private exploitation. That common bit of rhetoric about copying misses that point. The reason private property exists is not because land (for instance) can't be copied.

Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.

You should note that you misread me: I didn't say they were identical in every respect, I said their "reality" is the same. They're both made up social constructs. One isn't more fundamental than the other.

7h agoHN ↗

It's really foolish to claim one doesn't exist and should be abolished

Hear me out. Property rights exist to assign stewardship and use rights over rivalrous goods. goods where the use of the good precludes the use of that good for its purpose by another party. If you take my bike, I cannot ride it to work. This concept exists to prevent violent conflict over non-shareables and to prevent the tragedy of the commons (see the highly successful fisheries rights, NOx and SOx emissions markets as propertization schemes)

"Intellectual property" (except for trademark if you want to get pedantic), is not rivalrous, and its primary purpose is to be shared, not hoarded. Therefore intellectual property isn't a thing, creating it as a legal construct was a mistake that has hamstrung society for a long time.

5h agoHN ↗

Copies are not rivalrous, but the underlying creativity absolutely is. If I pay someone to draw a picture, that's labor, and they've got a finite number of hours to sell. Except it's also very inconvenient to pay for creative works this way: drawings and artists are not fungible with one another. More importantly, quality and desirability of the work is incredibly variable. The buyer of the art is bearing the risk of the art being bad.

What copyright lets you do[1] is offload that risk onto a publisher[0]. Instead of having to pay to commission every piece of art, a publisher can do that, and then sell the now-monopolized copies of whatever art turns out to actually be valuable.

A lot of hay was made during the Piracy Wars over filesharing tools breaking this bargain. A bunch of data hoarders with an interest in sharing media made it possible to just get the shit for free. This created a social dynamic where artists were annoyed about it, but publishers were Fucking Pissed. You could even measure how publisher-brained an artist got by how angry they were over Napster[2].

AI generated art also breaks this bargain, by making creative labor nearly non-rivalrous. The only cost is electricity and GPUs. This has created nearly the opposite reaction: artists are pissed while publishers don't care, because AI is to publishers like tort reform is to insurance companies. A publisher that gets art for free doesn't care if everyone else has it, because they have the payola dividend: they can push whatever slop they want onto the market and the market will eat it because they're big and powerful.

This concept exists to prevent violent conflict over non-shareables and to prevent the tragedy of the commons

The copyright maximalists would argue that free reuse of creative works is a tragedy of the commons. I certainly remember hearing that phrase bandied about a lot during the Piracy Wars.

It's also important to note that "tragedy of the commons" is not a natural law, but a specific framing that is used to justify antisocial ends. The communal ownership so decried worked perfectly well in England for hundreds of years until the ruling classes found it inconvenient and abolished it. The kind of ecological collapse the tragedy attempts to invoke did not happen because there were already communal means of preventing overuse of the land. In fact, an emissions market is probably closer to communal management than enclosure.

Also none of this changes the underlying logic that AI companies are trying to enclose the intellectual commons, and that their business model relies on being able to replace human brains with machine intelligence they can rent out by the megatoken.

[0] Individuals who self-publish included.

[1] To be clear, copyright was created as a censorship regime, it just happens to be useful for other things.

[2] In Lars Ulrich's defense, they weren't just angry that Metallica songs were on Napster, they were specifically angry that Napster had their latest album before it was in stores.

8h agoHN ↗

Property is at best a social construct that is possibly supported by instinctual behavior.

That very much falls within the "IP is not real" category. Taking a utilitarian approach here is exactly what the sam altman / effective altruist crowd is doing (or claiming, at least).

Believe it or not, but there are different philosophies. The US Constitution, for example, is written under the framework that all rights are innate, and the government merely endorses, not grants, those which are described. This framework creates a moral basis to rights, such that things like property are not merely social constructs but moral goods. To violate them is itself an immoral act.

8h agoHN ↗

That very much falls within the "IP is not real" category.

AH. Given that framing, then I suppose anything that varies from "these axioms are perfect there can be no others" must fall into the "not real" category.

But isn't there some debate about the moral axioms themselves? A space that is much more complex than "yes/no"?

EDIT: I always wonder, when I get two downvotes in the same moment, if someone is cheating. It happens so often.

EDIT2: And now! 3 upvotes in the same moment. I daresay someone is confessing.

8h agoHN ↗

Social constructs are bought and paid for by the wealthy.

Most people want a clean environment while the current US administration is pulling back environmental regulations to allow for more pollutants. The current US Supreme Court has taken gifts from those that have invested interest in the outcome of their judgements.

Social axioms are mutable while mathematical axioms are immutable.

Human trafficking is bad. How many of the wealthy that partook of Epstein's trafficking have been prosecuted? This shows that human trafficking being bad is mutable based on wealth and power.

7h agoHN ↗

no, I think people are just annoyed that you're trying to have an irrelevant and kinda cheesy 'is epistemology really real - think about it bro' conversation

I'm also a fan of nuance and non-black&white thinking but that is quite literally not how the rule of law generally works

8h agoHN ↗

Your wording is pretty revealing:

This framework creates a moral basis to rights

The Constitution (and the thinkers that it was based off of) creates a moral basis to rights. At the time the Constitution was written, this was a pretty radical idea. The most prevalent moral framework at the time was the divine right of kings, the idea that the king was ordained by God to rule and anyone who questioned that right was speaking heresy.

7h agoHN ↗

My wording was actually wrong, as the constitution recognizes a divine source for the rights, it doesn't create the moral basis. Perhaps 'establishes as the formal basis' would have been more appropriate.

You're right that it was a relatively radical idea at the time, but not without precedent. The Dutch, Venice and Genoa were republics, the Swiss Confederation was a union of cantons and territories, for example, not to mention the influence of the Iroquois Confederacy, John Locke, Montesquieu, and the assorted Greek and early Roman city states.

7h agoHN ↗

Funny how, when Aaron Swartz did the exact same thing (in a much narrower and more focused way), he was driven to suicide by the US government, but when Sam and Dario do it, people argue about whether (or not) "IP is real".

Note that society is more about justice than it is about philosophy or logic. Whether you believe that IP makes logical sense or not, is not terribly relevant through the lens of justice. To allow the Altmans and Amodeis of the world to escape the kinds of consequences that Swartz could not (despite having purer intentions than either -- despite being a genuine utilitarian, instead of an aspiring oligarch _pretending_ to be one), would be the height of injustice.

It would be a grotesque and cruel insult to all the people who remember Swartz, and share his values and optimism about the internet and computers.

7h agoHN ↗

The funny thing about having a divine source for rights is also recognizing that humans are fallible and thus inevitably going to be unable to live up to that standard. This argument is the basis for people who want government to have limited powers- the less power it has, the less it can abuse it.

There's a flip side to this as well- President Lincoln famously suspended habeus corpus unilaterally in violation of the Constitution, outright ignoring a Supreme Court ruling in the process. It became moot when Congress retroactively approved the suspension and allowed it to continue for the remainder of the civil war. Lincoln remains celebrated, despite what could rightfully be argued as blatantly unconstitutional power grabs.

7h agoHN ↗

What happened to Swartz was injustice.

Why would I want it to happen to more people? This is, at least to me, an insane point of view.

"Oh my god, the police just shot a black man for no reason. To make this fair they should also shoot more white people for no reason". Do you see how absolutely unhinged that sounds?

5h agoHN ↗

It sounds unhinged because it's a straw man.

3h agoHN ↗

I didn't want to bring up Aaron Swartz myself first, but now that you did and proved the point I did not want to post first, I nevertheless will:

It's a cruel and sad irony that people invoke Aaron Swartz to argue directly opposite worldviews depending on prevailing fashion, making him into some kind of perpetuum mobile in his grave. Literally the same arguments, using his memory to argue position that's opposite to what was argued two years ago, back before AI fully took off.

I personally doubt Aaron would be against AI companies doing what they did on the grounds of intellectual property. There's plenty of things to hate about how AI is transforming the world, but this is not it.

7h agoHN ↗

You are assuming that social constructs are not real.

6h agoHN ↗

"A social construct" and "not real" are not the same thing. Money, laws, countries, gender, even civilization itself are all social constructs. More importantly, all forms of ownership are social constructs, and that doesn't provide any information as to whether we should respect them or not.

1h agoHN ↗

Perhaps you need to look up what "real" actually means.

Inches aren't real. Laws aren't real. Language isn't even real.

Have you ever tripped over a line of longitude?

People live such abstract lives now that they're losing track of what is concrete.

5h agoHN ↗

Who decides what's a right? By what criteria?

7h agoHN ↗

Property is a social construct in the sense that anything is a social construct. Like, sure, the fact that I "own" this spoon is societally constructed, but it's so foundational that almost every other aspect of human society is more abstract. Even animals have some sense of "property" - whether that's territory they defend or their stash of food. The social construct part is that we respect each other's property rights (mostly) without resorting violence, where animals resort to individual violence or threat displays.

7h agoHN ↗

anything is a social construct

Well, no, the physical laws aren't social. An apple falls whether there is a social group to see it or not.

And yes! Animals are a perfect example. Mother Nature came up with a nice strategy for maximal benefit in a social group, and she encoded it in our instincts and intuitions.

But that doesn't mean it's perfect. I have an instinctual fear of fire - but it's still better to cook my own food.

7h agoHN ↗

Sovereignty is the ultimate form of property ownership in a funny way.

4h agoHN ↗

while you mentioned how tribal property management looks like in some places (that things are dynamically moving between who needs them at the time), it does not help to solve issue at hand, that is how to call what ai companies did:

- amassed then destroyed/denied largest knowledge corpora seen so far

- poisoned knowledge for decades to come, deliberately or not, which would stop new developments by independent teams

- now they are crying that they are being robbed of their precious models, which - at this point in time - are neither truly good not economical

(not to mention actual environmental and social harm they cause)

this is truly first of a kind situation, that was never seen nor heard before

8h agoHN ↗

Yeah, this whole "we're being attacked by China, they're distilling our models!" thing is completely and utterly absurd. Insulting, even.

Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.

8h agoHN ↗

Let me introduce you instead to the one golden rule. He who has the gold, makes the rules.

7h agoHN ↗

Let me introduce you to the power of cynicism to normalize corruption...

8h agoHN ↗

IP is obviously not "real" like physical property is real. Moreso, IP is a restriction on free speech as it walls off some expression as "Copyrighted" so you can't legally express it yourself without paying someone for the privilege.

That's not to say that IP is useless, but treating IP infringement as "stealing" was always a problematic shorthand. Previously, calling IP infingement "stealing" was the domain of corporate interest groups like the RIAA or MPAA, but with AI the winds have turned and supposed anti-corporate lefties are keen to treat IP infringement as theft to attack AI companies. There's very little intellectual consistency on either side of aisle here.

8h agoHN ↗

It's actually worse than that. IP laws trump property laws because they prevent you from using your property in ways you would otherwise be able to. It's nothing more than base rent seeking.

8h agoHN ↗

There are a wide variety of middle-ground views between them. Copyright law itself is one of them: it provides exceptions for transformative "fair uses", and there are some courts that have ruled that AI is one of them. Another might be Harbinger taxes on IP.

7h agoHN ↗

I think the blend of the two is that value added to any information, such as the first time its distilled or positioned in a certain way, that's IP and should be real.

When an AI company takes that type of value, and doesnt provide value back to the person who made it, thats theft.

Information / facts are real and free. But those who helped us get here should choose how to license their work, and an AI company picking it up and re-selling it is clearly unethical.

16m agoHN ↗

But those who helped us get here should choose how to license their work

I can reasonably wrap my head around the idea that an author should be compensated for their work, but where in the social contract does this power to control how that work is used come from?

It's so destructive and tangles up courts and makes contracts complicated and we lose the original versions of the work because e.g., they have to change a background song due to complicated licensing.

I can respect protecting copy rights, but it should never be conditional; if you choose to make a work available to the world, you have a legal right to defend the unauthorized copy of it but you should not have a right to say how it gets used.

It reminds me of manufacturers like John Deer.

7h agoHN ↗

Both can be true. IP is real and they created new IP off of stolen IP.

7h agoHN ↗

Intellectual property is a legal fiction, not an empirical fact about the universe. It's as real as societies want it to be, and that view can change over time.

7h agoHN ↗

It’s real but whether you can make a profit or not is dependent on whether you can protect the information. It has been this way since civilization as we know it started, I believe.

If you are freely posting sentences like this was on the internet you are giving away your IP for free. Anyone can read the sentence. If someone can make money off of it then that’s just markets at work. I can’t make money off of what I write here, for example.

But I also don’t think pirating a movie is theft either. You haven’t proven to me you’ve lost money. Maybe wouldn’t have watched it anyway.

Fun topic

7h agoHN ↗

It’s real but whether you can make a profit or not is dependent on whether you can protect the information.

That's what contracts and licenses [0] are for. Or perhaps you're arguing for a world where the only possible protection is that of trade secrets... that once protected information is made public by any means anyone can do anything they wish with it? If you are, then this quote [1] seems relevant:

  If IP is real, then the AI companies have performed flagrant theft.
  
  If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.

[0] ...which are contracts in disguise...

[1] <https://news.ycombinator.com/item?id=49754448>

1h agoHN ↗

To read this post you must pay me $5. Please leave your contact information as a reply or you'll be hearing from my lawyer. [1]

Those contracts and licenses are just a form of protection enforced by the State. In general once knowledge or information is widely available it's also freely available. Whether someone can further protect the distribution or the usage of that knowledge to some ends is up to them.

I disagree with both of those premises.

IP is real, and scraping freely given comments or published material on the Internet and doing something economically useful with it doesn't entitle you to retrospectively go back and decide that you are owed some money. If that were the case you have to prove material damages. How much is my post here worth? one quadrillionth of a cent?

I think it can become different if, for example, a book was scraped or cataloged but that's primarily because it's likely that you can enforce or make a case within a given legal system to enforce your copyright or IP claims. But what if I read a book and then thought the main idea was my own, or just spoke to someone else about it and they spoke to someone else about it and it winds up in an LLM? There's nothing wrong with that or anything you can do.

To expose information is to put that information at risk of being used by others. It's up to you to protect that information or enforce claims on it. If Reddit's public site gets scraped and OpenAI does something economically useful with it, well, that's just life. You can't put information out in the public for free and then demand payment later. Reddit comments are freely accessible by the public, yes? Companies are part of the public.

If you want to argue that it's IP theft then distilling weights or otherwise reverse engineering the models is a violation of IP protections too.

[1] Rhetorical

7h agoHN ↗

You would not download a car, would you ?

Hell if a car was downloadable why would I not ? Your capability to make endless profit should be capped somewhere, if indeed it is Profit

7h agoHN ↗

If IP is real, then the AI companies have performed flagrant theft.

It’s called Derivative Work and it’s a good feature for IP law: https://en.wikipedia.org/wiki/Derivative_work

You wouldn’t like a world where companies could copyright knowledge and then prevent anyone else from making a derivative of that knowledge.

Imagine Wikipedia being taken down for sharing knowledge that another company wrote about first. It’s a bad idea.

7h agoHN ↗

It’s called Derivative Work

That would be on what the AI model generates, not what it is trained on.

The latter is where the contention is, and it's a valid argument. So much so that some companies are not using stolen information to build their models.

IBM for example indemnifies its models for its customers and has detailed information on where the sources came from to train them.

7h agoHN ↗

The latter is where the contention is, and it's a valid argument

It has been tried in court several ways already. Remember the lawsuit that forced Anthropic to use physical books? They tried to argue that the books couldn’t be trained on at all. It failed.

6h agoHN ↗

Yes. That judge misunderstood badly, and made a bad ruling.

5h agoHN ↗

I'm not sure I understand this argument. If I go to the library every day for 10 years and learn everything there is to know about subject x I shouldn't be able to sell my skills to the world about it later because I didn't give the creators of the books I read any money?

Arguably LLM companies could have made large-scale deals with libraries and got the exact same knowledge (much, much more slowly). I wonder if people would have the same issues then? My guess is probably. Goes back to the meme that if libraries were proposed today there's no way they would ever be allowed.

4h agoHN ↗

it's of wrong scale, you cannot claim derivative work when you literally ingressed sum of knowledge while destroying it in the process so your "competitors" cannot do the same

7h agoHN ↗

Objectively, IP exists and is acted upon, so IP is real, obviously, but IMO it is bad, and should not exist.

People tend to forget that IP was not originally about digital distribution at all (copyright), it was about giving inventors exclusivity periods to profit without competition (patents).

It was a misguided attempt to stop sometimes literal theft of designs from rival inventors, by tying the design to the person instead of whoever possessed the schematic.

It's also a regime that in its modern incarnation protects businesses, not artists.

6h agoHN ↗

There are nuances.

AI scraping for the goal of making a commercial LLM service is different from, say, a commercial file sharing platform.

The first difference is that LLM training is highly transformative. Let's say your LLM ingests the Harry Potter novels during its training. What you get at the other end is not the Harry Potter novels, it is a LLM that can talk to you about Harry Potter, it is not the same thing, and going from one to the other requires a significant amount of work, very expensive work in this case.

Not only that but there is no direct competition. People won't stop buying the Harry Potter novels because a LLM trained on it exists. If you want to read the books, you buy the books, you don't ask a LLM about it. A file sharing service on the other hand competes directly against the official channels, if you want to read the books, you can download it from this service instead of buying it on the official channels.

So, about how free you should be to get these weights from the AI companies. If you just share a 1:1 copy of the weights, that's the "file sharing" situation, not transformative, you took their work, didn't do any of your own. Usually considered unacceptable by IP laws.

Distillation is a more interesting case, you are using a LLM to train your own, it is transformative work, but you may also be competing directly against the LLM you are distilling. So, in a sense it is worse than scraping, but still, despite how much the likes of OpenAI and Anthropic are complaining, it seems to be legal.

So it is somewhat consistent: 1:1 copy and distribution is not allowed, be it source material or LLM weights, and training is, be it source material or another LLM (as in distillation).

6h agoHN ↗

The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.

This is how IP has always worked. It has never protected the little guy. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.

6h agoHN ↗

And by constructing that false dichotomy, you're failing to see the actual truth, which is that Intellectual Property is just a legal construct to turn ideas into Capital, with the goal of incentivizing its creation.

Nothing more and nothing less, and all of the normal debates about what should remain Capital and what should be The Commons apply.

6h agoHN ↗

IP could be real and LLM training could be considered fair use. IP protects published information, so even abolishing all IP would not help us since the LLM weights are secret.

Nothing short of a global revolution can fix this. Proprietary LLMs should be illegal. Either we achieve post scarcity within this generation or it's literally over.

4h agoHN ↗

Intellectual property rights only protect against certain things. IP can be real while there exists fair use like training AI models.

9h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen

Considering its become exceptionally difficult to search for things that used to be easy to find, I have to disagree with you there.

The web has been polluted with trash, covering up all the original media with messy imitations.

9h agoHN ↗

LLMs have made it easier than ever so search for things. I remember fumbling around with Google, trying different combinations of search terms and never hitting useful results. Now I can give a vague or even partially wrong prompt to an LLM and it sort of magically figures out exactly what I actually want and takes me right there, along with a helpful summary and analysis.

9h agoHN ↗

Content creators are taking steps to prevent bot from walking away without compensation for their visit as if they’d been stolen from. At least Google, in the beginning sent traffic back that could potentially be monetized by the creator.

8h agoHN ↗

Perhaps some. But I'd wager orders of magnitude more are using LLMs to brainstorm if not outright generate content for them, which they are then reframing as their own creation.

1h agoHN ↗

orders of magnitude more are using LLMs to brainstorm if not outright generate content for them, which they are then reframing as their own creation

I doubt it, those are the same group of people anyway.

Orders of magnitude more people are using LLMs to brainstorm or generate content for them, which they are then using to solve their own problems and carry on with their life, increasingly solving many more problems for themselves this way, and not publishing any of it.

The world isn't made of "content creators" flinging crap around in hopes of monetizing it. Most people have actual jobs.

9h agoHN ↗

Here's the thing. If any of my work is in their model, then I co-own their model as well as everything it makes.

I guarantee it's scraped our work.

So where's our paychecks.

45m agoHN ↗

I don't know, and I don't care. Scaled proportionally to actual contribution, it'll be like $0.000001 / year, and I'm not going to begrudge the AI companies for not paying me a fractional cent, and I'm definitely not going to cry foul and demand stopping progress over that fractional cent they plausibly owe me.

Not the least because I'm already getting many orders of magnitude more value each day from them providing me inference as a service.

9h agoHN ↗

It is robbery because you're not paying the person that produced the knowledge, you're paying the person that took the knowledge and put a chatbot interface on it.

It's different from search engines because you're not even getting the chance to monetize it yourself, you lose the ability to even show some ad for a pittance. Hell you don't even get credited most of the time. Many many people didn't give permission for the data to be ingested and used this way.

We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.

8h agoHN ↗

We got? We didn’t get anything last time I checked, we are offered to pay rent to access this intelligence, at a discount

8h agoHN ↗

Why does everyone ignore open-weight models in these discussions?

1h agoHN ↗

We have free plans and open-weight models you can host yourself, or pay someone for hosting, or host yourself and sell others to recoup the costs.

All of these lag SOTA by ~6 months on average, so yes, indeed, the thing you could only pay for half a year ago, is now free for entire world forever, and it only keeps getting better with time.

1h agoHN ↗

It's a robbery of our future culture, at least.

That is a different argument entirely, though. Not the one GP presented.

On this, I have no concrete views. Sloppifying is real, and AI is short-circuiting "this loop of exchange of ideas", for sure. On the other hand, we had several precedents on this within last 700 years - the printing press, the radio, and the Internet to name the largest three - and culture turned out fine. Different, but fine. AI feels more than this, though, so I'm not putting that much stock in argument from history here.

8h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form.

But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of trying to deny it.

And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.

We don't have a society where "we got back" anything, and if you weren't aware, this "intelligence" is definitely getting metered. If anything "we" got undercut, and "our" assets lost a lot of their value, while others will get very rich from that loss. Come back with communism, and then maybe you'll have a point.

[1] Fuck you, Claude, for making me wince at that.

8h agoHN ↗

Yet IP is processed (without permission, nor payment) and the software benefiting from it has a monthly fee

8h agoHN ↗

Usually that stance means that you're not one of the people making culture.

It isn't you that AI communism is stealing from.

8h agoHN ↗

The problem is the hypocrisy. They scrape, but don't want others to scrape them. If they act like they've been stolen from when they get scraped, that is a confession of guilt to when they did it themselves. They should be punished according to the same rules they wish to impose on others.

7h agoHN ↗

I really agree with this critique! It's very hard (without hiding behind ToS) to claim that creating the models was A-OK, but distilling them is some sort of ethical breach or attack.

1h agoHN ↗

That I can't argue with. "Distillation is theft" is one of the silliest and most hypocritical things uttered in any of the debates surrounding AI.

8h agoHN ↗

> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there

This is good, but will not fly in court. Your honor... I only moved the Picasso...it is still there...

8h agoHN ↗

the culture has not been stolen - it's still there

It's there in much diminished market, one that has been saturated overnight on a scale previously unfathomable. Their art is still there but it has been used against them to obliterate the playing field.

nor are the people involved selling it back in any form.

Aren't all of these AI companies selling subscription models for people to create derivative art?

And let's not forget what we got back for this

Get back for what, exactly? Didn't you just argue that no theft or sellback has occurred?

8h agoHN ↗

This is true but it gets people who tied their value to outcomes, wealth accumulation, status accumulation upset. The road for this has closed.

7h agoHN ↗

"It's still there" only in a technical sense because it's being buried by the automatic slop machines.

It used to be that, no matter what you want to do, there's a video on youtube of some guy who knows what he's doing showing you exactly how to do it. It could be, like, compressing a rear disc brake cylinder on a car or matching fiberglass gel coat colors, or carving a statue out of marble.

Theoretically those videos are still in there somewhere, all 15 years old at this point, but you'll not find them with youtube search instead you'll find a million AI slop videos no matter what you query.

7h agoHN ↗

There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".

There's a way that many people fear this is true: what you're getting for "free" has some external cost that you aren't taking into account.

Cost 1. The environment. Tech companies were once significantly interested in efficiency, their carbon footprint, etc. With the advent of LLMs, this was thrown out the window. Concerned people are now thinking about water footprint, heat footprint, noise footprint, and probably more. These things are hard to put a dollar figure on in a short comment, but the sustanability of a liveable climate for the billions living on this planet is literally priceless.

Cost 2. Employment. We face a significant cull of employability -- graphic artists, programmers, mathematicians, paralegals, and more are rightly fearful for the end of their career. If we're getting intelligence for "free" in 2026 dollars, but we can't get jobs in 2030, will all that content still be affordable in 2035?

Cost 3. Infinite investor dollars. The AI companies are burning cash at a historically unprecedented rate. This has attracted a lot of attention from traditional investors, like pensions, banks, etc. If all of this ends up in a free product, that doesn't sound like it'll return those investments. And that can crash the economy -- I ask again about real affordability in 2035.

7h agoHN ↗

Infinite investor dollars

This has another side effect in that it breaks the power that consumers have in the market to have a say in how resources get allocated, by voting with their wallet.

The entire AI buildout is non-consensual. Individuals have zero say. We could all refuse to buy AI subscriptions and it would not matter because businesses will still buy, and, investors have decided that we are moving forward with this no matter what, seemingly whether anyone is actually buying or not.

AI isn't necessarily unique here, but it is one of the biggest new examples of wealth inequality and how society at large are no longer the ones who get to decide how resources are allocated under a capitalist system, especially when the investor dollars behind it amounts to the GDP of a small nation.

7h agoHN ↗

And Businesses will buy, because the AI promise is a promise to automate workers away.

7h agoHN ↗

the sustanability of a liveable climate

And this is fuel for climate change denial. See? Are the wealthy and powerful people acting like they are concerned about the climate? No. They are doubling down on energy use and consumption. It's all a fraud!

1h agoHN ↗

The idiocy of those arguments around AI and climate are indeed a fuel for climate change denial, because the only consistent views are that either the AI is not an environmental disaster, or the whole climate change thing must be overblown. The numbers don't add up, but if you're so blind with hatred towards "AI companies" while still desiring to have a remotely consistent worldview, it's seriousness of climate situation that has to go.

6h agoHN ↗

Item 2 is the big one. Because not only is AI about to destroy jobs wholesale; it's doing so in an environment that abhors wealth and income redistribution that could blunt the harms.

2h agoHN ↗

2. and 3. I agree with.

1. is mostly bullshit that's used by anti-AI crowd to gather further support. Data center companies did not stop caring about footprint, they're still pursuing that for the same reason they did it before: it aligns with overall lowering of costs. AI added an extra incentive to improve efficiency because the demand far outstrips supply of compute.

For 1., also the general argument stands: yes, these data centers use energy, just like everything else humans do. What matters is the value it provides to people, and in case of AI, it's one of the most objectively useful expenditure of watts on compute (your point 2. notwithstanding).

7h agoHN ↗

real problems of real people, including individuals and non-profits

Again with the weasel words. Real people, including individuals? Pretty clear this is written from a pro-business perspective and can be safely dismissed.

"robbery of all of our culture to sell it back to us at a mark-up".

Well, it is that. Why is it criminal for me to steal a big-business-movie for personal use, but they can just take all my blog posts and sell their deritives to others?

7h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen

I respectfully request that you get all the way outta here with this take.

Artists have caught AI generating literally their own work, for free, at scale. There have been tons of articles on this website about the AI hyperscalers slurping up books, copyrighted works, etc through legal and questionable ways. Even TFA says that the NYT suffered absolutely devastating CTR drops.

If you create a blog post about something super esoteric, it will guaranteed end up in the training sets for every big-lab frontier model within 24 hours (probably less) and probably show up in Google's AI Summaries at around the same time. It's also well-documented that their spiders don't honor robots.txt either and blocking them is basically impossible unless you give Cloudflare protection money (since spiders can super easily solve Anubis hashing challenges).

If that doesn't sound like theft-as-a-service to you, I don't know what to tell you. All I know is that hosting content ANYWHERE AI labs can get a hold of it is rapidly becoming a fool's errand, and we will all be at a loss for it.

7h agoHN ↗

Stealing in this case is not about physically removing something from one's possessions, but more about disguising attribution and in many cases also reducing a future income stream.

7h agoHN ↗

I have artist friends - asking AI to draw something similar to subjects they've drawn ends up making, line for line, the exact image they created, hallucinated alongside a couple others. They spent decades getting good at their craft. They spent so many years getting paying patrons and customers to commission them.

You'll never agree but I think you underestimate how many people consider all that the AI companies have done to be the wealthy stealing and selling things back.

nor are the people involved selling it back in any form

I am convinced that the longer one works in AI the less one has any grasp on reality

too cheap to meter

Its the most expensive buildout in human history - you can't just split half the cost. The spending is the only significant growth in the US economy. Unless you think that all the datacenters and all that capex are for training?

6h agoHN ↗

I have horse & wagon friends - they spent decades understanding the ins and outs of roads, some of them also creatively made up their own routes and roads.

However ordering a car with Uber through your phone just gets people to the destination faster, cheaper, more comfortably and reliably using those same roads.

This is theft I made up the route first >:(

1h agoHN ↗

Yknow, I guess you’re right, human creativity really has no place in your world.

6h agoHN ↗

Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there

It's actually not there. See how many websites have now closed doors or ceased to operate because of constantly being hammered and bombarded by robotic scrapers. For others, it made websites that turned a small profit into unsustainable money pits.

Sure, their content might now be ingested into an LLM training set sitting somewhere on proprietary servers. But the site itself (the origin of truth) now does not exist.

So no. It actually isn't there. Resting on this falsehood, the rest of your retort makes not much sense.

6h agoHN ↗

    reified intelligence on a chip, almost too cheap to meter, 

“Reified intelligence” is stretching the truth. It is certainly closer to intelligence than anything we’ve come up with before, it can certainly match or beat actual intelligence in some domains, but just as certainly it clearly lacks many features and capabilities of intelligence.

“On a chip” is also stretching the truth. The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.

And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them. Individuals who have free usage for a project commonly report that, had they paid, it would cost five or six figures. Companies have seen enormous bills, some approaching the size of their payroll. And all of this is true for tokens that are dramatically subsidized, by one of the most intense and largest concentrations of capital in history. It is the polar opposite - “almost too expensive to even do, and absolutely must be carefully metered”.

In summary, let us indeed not forget what we got back for this: extraordinary distortions of reality evenly intermingled with bald-faced lies.

49m agoHN ↗

It is certainly closer to intelligence than anything we’ve come up with before, it can certainly match or beat actual intelligence in some domains, but just as certainly it clearly lacks many features and capabilities of intelligence.

It's not complete or that well-rounded. But it's something that was the domain of speculative science fiction only 5 years ago, and it's rounded enough to be applicable to ~everything to some degree.

The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.

By "on a chip" I meant more "in silica" than literally on a single chip" - though this actually is* true, but those chips aren't cheap.

And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them.

You are looking at power users that use LLMs in agentic coding sessions. Most people just run off free tier of ChatGPT, which is free for them. There are equivalent open-weight models at this level, and while hardware to run one for yourself is expensive even for most westerners, the marginal inference cost is literally dirt cheap, which is why you can get that for near-free from smaller inference providers - or pony up some money, rent a bunch of compute with friends, and become an inference provider yourself.

(It's only a tough market because the major vendors are giving out better models than you can run for ~same or lower price than you can offer. Which either way is too cheap to meter in terms of solving useful problem for real people. Again, developers are a special case of power users, as usual.)

5h agoHN ↗

because of how many real problems of real people, including individuals and non-profits, it addresses.

Not problems, laziness. The way I see students and colleagues use it is to get their work done with less effort. That's its selling point. There are not many real problems LLMs address.

no one has actually been robbed

In our society, people get paid for work. If you think society is wrong, fine, but you must state so first. Under common assumptions, all that data has been produced through work, and that work represents value. Taking it for free is therefore theft.

5h agoHN ↗

reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West.

I think this is quite the pollyannaish perspective and very much inline with those that think if you can take, then take and only apologize when caught.

2h agoHN ↗

It's not a perspective, it's a literal fact.

We won't get anywhere in these discussion if good chunk of participants cannot admit to the trivially observable facts about the actual objective reality in which they live in.

5h agoHN ↗

It's become a bit clichéd to say it but I think it's different this time.

The kind of "piracy" you're talking about wasn't depriving the authors of anything because you could always make the argument you weren't going to pay for it anyway. If, on the other hand, you were making copies and charging people for them, you could definitely say you were depriving the legitimate authors of that revenue. AI companies are very much doing the latter, not the former.

The other part of it is it's not just copying. Previously, if I decided to make a copy of a work without paying, I'm only copying the work, not the author's whole writing style. Now the AI companies are depriving authors of revenue from works they haven't even made yet.

If they were able to create a model de novo then they could truly claim it hasn't just been lifted from existing culture.

5h agoHN ↗

Every time AI is used it is a theft from artists, programmers, writers, and all creative workers whose work was used to train it without permission or compensation and are now unable to find work or buyers. It is absolutely theft and you're being intentionally disingenuous to suggest otherwise.

10h agoHN ↗

sell it back to us at a mark-up

sell it back to us as markdown

10h agoHN ↗

“It is said that at the heart of every great fortune there is a great crime“. That quote is from a fiction author - you’re essentially quoting Spiderman “with great power comes great responsibility”.

10h agoHN ↗

Are you seriously putting Balzac and Spiderman on the same level to try to ridicule thoughts?

Ad hominem.

10h agoHN ↗

They are both quotes that people say things like "they say" to imply someone of import said it to give it credence and truth to it but were in fact from fictional sources.

Exact same pattern - take it as an ad hominem all you want.

10h agoHN ↗

Other than the rare books than have been ruined, all the same knowledge is still out there though. So it’s not robbed in the sense of a bank heist. Maybe in the sense of pirating a movie.

The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.

7h agoHN ↗

all the same knowledge is still out there though

The markup is the millions in training they committed and the connecting the knowledge

Sure, but gated behind a hallucinating idiot.

9h agoHN ↗

Nobody has made much money from AI yet unless you count the shovel sellers (Nvidia). Not sure closed models would make any money ever and eventually the benefits should flow to everyone.

9h agoHN ↗

If LLMs scraping the web is robbery then piracy is stealing.

9h agoHN ↗

"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.

An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]

AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).

If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:

1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?

2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?

3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.

"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.

[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow

[2] https://sydneyuniversitypress.com/products/9781920898168

[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons

8h agoHN ↗

I wasn't aware of the "sweat of the brow" doctrine until you mentioned it. But you seem to be using it incorrectly. The doctrine only states that creativity or originality isn't required to make a work copyright-able.

Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.

7h agoHN ↗

The parent comment I replied to is concerned with "life's work got appropriated without consideration, compensation or consent". To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions. Today in most jurisdictions copyright laws do not care the slightest about an LLM ingesting databases -- phone directories, sport fixtures and results, someone's life work measuring the dimensions of frogs, etc. 100% of the original factual data could be learned by the LLM, and 100% could be output all at once.

^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.

6h agoHN ↗

To alleviate this concern^, "sweat of the brow" doctrine would be required, but it doesn't exist in most jurisdictions

No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" is a superset of that. Originality and creativity, which are required to copyright something, are also work.

5h agoHN ↗

The US government's official position on LLMs is (very simply paraphrased) that LLMs are sufficiently transformative and do not hamper the potential market of authors of training material, therefore, copyright claims arising from training material should not be successful.[1] For original and creative training material, for example, a Harry Potter novel, seemingly the US government is asking the courts to set aside some previous questionable findings such as copyright existing very loosely in the likeness of fictional characters (impacting the likes of fan fiction). Can a human -- or LLM -- create a story about children travelling on a train from New York to a school of magic in the "wild west", with many loose similarities to Harry Potter for those familiar with those books? The US government appears to be saying this is OK, especially with the view that the market for Harry Potter is not diminished by a "wild west magic school" book in its likeness.

However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.

[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...

4h agoHN ↗

LLMs are sufficiently transformative and do not hamper the potential market of authors of training material

And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.

Furthermore inference indirectly affects the market for the artist's future work. Don't need the writers and artists the LLM trained on anymore, when it can do similar work for free.

8h agoHN ↗

We know what is going inside black holes in other galaxies, we know the details of Israel nuclear program..., the Windows source code got leaked, we got the NSA tools and full details and locations of the Echelon architecture... Phds in Maths warns us daily about the terrible secrets of the evil AI inside their labs. How, their numeric matrices and gradient descent Python scripts, are about to kill 10% of us all, I guess the sick or genetically less interesting ones...and use the rest, as some meat/metal hive drones part of some Borg collective...

There are ONLY TWO Stories and their details, that we collectively will never see.

1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...

And the other...

2) The Epstein Files.

8h agoHN ↗

I haven't paid a dime for any of the LLMs I've used (electricity and internet aside)

7h agoHN ↗

So why is there so much outrage now when Google and other search engines did it 30 years ago? They literally said their objective was to collect and index all the world's knowledge, and when they scraped the internet clean a thousand times over they invested in scanning and digitizing everything that wasn't on the internet yet.

AI is better at repackaging it back to the end user but ultimately I'm arguing it's the same thing.

(caveat: yes I know there were plenty of people that objected to Google et al indexing everything; famously, Gmail was scary to a lot of people because they read your email to give you ads)

6h agoHN ↗

It's not the same at all though. Google indexes the original content, hosted at the original location. It just gives you a way to find it. LLMs ingest the original data, throw it away, and spit it back out in whatever form they like back to the user. And since it's gotten so popular, many peoples' interactions with the content are no longer in the original form at all; their site traffic / book sales diminished.

6h agoHN ↗

Google didn't immediately create $20 and $200 per month paid tiers. The free thing lasted a long time (and you didn't need to sign in), and the initial monetization with ads, when it eventually came, was not immediately shoved in the user's face. the enshittification and bean counting came way later compared to the monetization blitz that occured with the AI companies.

5h agoHN ↗

Because they aggregated links to that information rather than copy it outright. It’s like the difference between an encyclopedia and its table of contents.

5h agoHN ↗

Google (and other search engines) gave you an easy and reliable way to opt-out of having your content indexed and linked to. People did in fact object pretty loudly every time Google attempted to surface the information on google.com rather than sending traffic to the source.

7h agoHN ↗

It's the robbery of all of our culture to sell it back to us at a mark-up.

At a mark-up would mean it’s more expensive. The outrage is that they’re taking knowledge that was expensive to access because you had to hire experts or otherwise pay a lot of money for it and making it accessible to anyone who signs up for the ChatGPT free tier.

Calling it “robbery” is also specious as no knowledge was taken away from anyone. The content in the training sets was out there in the world one way or another. It still is!

I’m really perplexed by this sudden swing toward the idea that knowledge is something that we should encourage or incentivize to keep locked away or that other people should be forced to pay for use of knowledge. Roll back the clock a few years and tech sites would be almost unanimous about knowledge being free and unrestricted for the benefit of humanity. I’m keeping knowledge separate from actual direct rote duplication of content.

Now we have this amazing era where I can download models to my computer, run them locally, and have enormous amounts of derived knowledge at my fingertips for the cost of some compute cycles. Except now it’s a “crime against humanity”?

7h agoHN ↗

I don't know about "crimes against humanity", seems to diminish a bunch of actual crimes against humanity. And Microsoft is one to complain! Talk about a glass house. Just checking the annual report for 2025, the median employee compensation was 200k per year, but the total dividends divided by number of employees was 100k, meaning that each employee gets only about 66% of the value they've generated. Isn't that theft?

In any case, Microsoft has stolen 25 billion from its employees in 2025, and OpenAI has got 13 billion in revenue from "stolen" content in the same period, so that'd make them about equally bad villains, except OpenAI has mostly stolen from other companies.

6h agoHN ↗

Since they didn't improve the culture, it makes no sense to buy anything back at a mark-up. We can keep using the one we still have.

5h agoHN ↗

It’s a crime against humanity if the only way to retrieve the digital version of the works is through an LLM.

The web getting flooded with slop and drowning all original work is one way to go about that.

Another is paywalling and gatekeeping en masse.

A third is shutting down shadow libraries.

5h agoHN ↗

People are getting lost in the weeds and ignoring the simple, fundamental fact that these companies are making fortunes off of labor that they didn't attempt to compensate the laborers for.

"Property" and "IP" discussions are distractions; no amount of it can rationally get us around the utter unfairness of what occurred and the way it will warp our economy at a basic level if not addressed.

It's not even enough to make the weights and models free; access should be free, and everyone who hitched their horse to this wagon should be on the hook for keeping the systems running, on their dollar. They took ownership of a venture that is short one (1) "Humanity's entire cultural corpus", and the only question is if we're going to issue a margin call.

5h agoHN ↗

you can forget about anything coming of this.

first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft.

so, really, a lot will be coming out of it, we like it or not.

12h agoHN ↗

And here we are, just watching and doing nothing..

12h agoHN ↗

I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.

12h agoHN ↗

Lots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.

12h agoHN ↗

Copying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone?

Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"

Probably indicative of America's wider downfall that they've all become so self interested

12h agoHN ↗

Hackers are irrationally anti-corporation. This is where the nonsensical AGPL came from, too.

12h agoHN ↗

The AGPL didn't go far enough because it didn't limit freedom 0 to natural born humans. In the 00s/10s it was corporations. Today it's AI.

If it has no soul to save and no body to torture it deserves no rights.

11h agoHN ↗

Show me where this "soul" is. Never seen one yet.

11h agoHN ↗

I don't believe in souls. But I believe in metaphorical souls, and an LLM ain't got one.

9h agoHN ↗

Corporations cannot act. All corporation actions are performed by natural born humans. All corporations are ultimately owned by natural born humans. All corporation actions serve to benefit natural born humans. All corporations are created and destroyed by natural born humans.

This idea that corporate personhood is some perverse idea is silly. Corporations are just people acting in groups for economic benefit. Everything they do is done by members of those groups (or people they hire).

Similarly, AI is an inanimate tool, like a keyboard. No comment is posted by “bots” - comments are posted by humans running software.

11h agoHN ↗

I want to share the open source code for the good of everyone - which is why I put it under the GPL, for this right of everyone to share it to be protected. Therefore, any AI model that ingested GPL code should logically also have all its output GPL licensed. Not problem with that.

9h agoHN ↗

Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"

If you can't tell the difference between "I want to share all information freely with my fellow mankind" and "I want to share all information freely, even to billion dollar corporations that are making the human-replacer machines that threaten my fellow mankind" I really just don't know what to tell you

12h agoHN ↗

I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.

11h agoHN ↗

It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.

11h agoHN ↗

We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).

11h agoHN ↗

I suppose it's theft in so far as you're deriving value from something that others labored for. I think that sticks in peoples throats. You can go to the library and do that already, gain knowledge, start a business or whatever. But the scale feels impersonal and monstrous in comparison.

11h agoHN ↗

The same point was made by many in the music industry about huge scale Internet music piracy vs people dubbing CDs onto mixtapes.

Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.

Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.

11h agoHN ↗

The original might no longer be there as the site might have already shut down due to AI scrapper bot overload. Or the original was a book Anthropic scanned and then shredded. Or the artist stopped publishing their works or doing art all together after all their creations were ingested into the model blob without their permission.

It is really insane to compare individuals copying data to big corporations parasiting on the Internet.

12h agoHN ↗

If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.

12h agoHN ↗

If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.

6h agoHN ↗

You're right that it never ended.

But it didn't change form much. Still around 50 million people enslaved nowadays.

5h agoHN ↗

Are you comparing modern workplace aches and gripes to literal 1800s slavery?

12h agoHN ↗

Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours

12h agoHN ↗

But it's copying...how is it theft? Your labour WASNT stolen was it?

11h agoHN ↗

In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?

11h agoHN ↗

Even better question: would you be willing to train the AI-driven robot which will end your profession altogether?

11h agoHN ↗

I work in infosec and I would give up everything I own and die happy if infosec became a solved problem and the profession died forever.

It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.

9h agoHN ↗

I'm not sure how to square

It's hard for me to imagine a profession that should exist in a utopian society

with

People should have instant access to food when they're hungry and housing when the weather gets bad

You don't think that farming, baking, and building are professions?

4h agoHN ↗

Farming and baking are already very thoroughly automated and their labor force completely proletarianized.

11h agoHN ↗

Do we normally consider recorded music to be theft from musicians who would have been paid to play music live if we never allowed (or invented) recorded music?

Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.

But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?

Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.

9h agoHN ↗

Musicians are paid to record music. If they were not, as is the case for ai data, we would indeed consider it theft.

8h agoHN ↗

The few who get paid to record are replacing far more who would have otherwise been paid to play if that was the only way to listen to any music. Is it okay if one musician replaces another but not if a non-musician does so? The person getting replaced didn't get paid either way.

Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?

8h agoHN ↗

Whether it is okay or not has no bearing on it being theft. And no need to create a hypothetical, this is exactly what happened to musicians.

There is broad evidence that labs have used substantial amounts of pirated data, no need to reach for a new definition of theft.

7h agoHN ↗

Musicians are paid to record music. If they were not, as is the case for ai data, we would indeed consider it theft.

I remember reading literature from the 1930s, and there were quite a few folks who thought that the musicians-doing-recordings were stealing from the old-timers who played for live audiences.

History does not exactly repeat itself, but it rhymes

5h agoHN ↗

Normally this isn't the case for any technology except for the time it first comes around.

Well, sure, there's no one left to fight once everyone already lost their job and moved to another career. But, that's sort of a "might makes right" resolution. The workers don't have the political sway needed to get the government to intervene.

Businesses succeed in getting that sort of market intervention all the time. Most of the modern changes to copyright law are driven by business lobbying.

11h agoHN ↗

Well, given these companies are trying to sell it back to you - seems like even worse than stealing. ;-)

11h agoHN ↗

For intellectual property, copying without permission is theft.

11h agoHN ↗

Depending on the context, it’s fair use.

As it is, in this case.

10h agoHN ↗

It's as fair use as this scenario:

I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.

7h agoHN ↗

No, it’s been decided by the courts to be fair use.

I’m not sure why people think they understand IP law better than the courts just because they don’t like the answer

37m agoHN ↗

Ah yes because “the courts” have an unblemished historical record of never getting anything wrong

12h agoHN ↗

No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters

11h agoHN ↗

Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft. Then you get the serfdom in the middle ages. Up to recent times humans have been brutally exploited.

The copy part was a recognized right, then taken away.

12h agoHN ↗

I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)

12h agoHN ↗

they did say theft of labor, not theft of the things being trained on

12h agoHN ↗

In case of programming.

How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?

This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.

Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?

What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.

https://consortiuminfo.org/metalibrary/estimating-the-total-...

12h agoHN ↗

They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.

IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.

11h agoHN ↗

I guess you're not writing code much or haven't done much so as your professional career?

8h agoHN ↗

My argument is if AI companies are ignoring copyright law and looking at all training data as commons, then we should look at LLM output as something that is not protected by copyright law.

Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).

https://www.theguardian.com/technology/2026/may/23/trump-ai-...

https://www.bbc.com/news/articles/c98r8r7dz5no

https://en.wikipedia.org/wiki/Commons

8h agoHN ↗

I wanted to understand your background first because what you initially said is a very oversimplified view of LLM mechanics, and generally not quite the way how software is in practice written. Since you didn't answer that question, I will assume that you're not a SWE by a call. To give you an example of what I am trying to convey is: imagine a data-intensive workload hitting your storage/database/kernel implementation, and it's painfully slow, your customers are not happy. Then you as engineer sit down, spend days profiling and understanding the code, researching about existing algorithmic solutions to the same or similar issues found in the wild, you read some open-source implementations of viable approaches, you ditch some, some you take, you also read books, articles, other peoples experiences etc. And finally you end up, let's put it bluntly, with some sharded data structure by which you solve the bottleneck. It's not novel, the technique is so common and is already implemented across many many different products in slightly different flavors so I am wondering why do you think this is not a copyright breach but the LLM, which does more or less the same thing, is?

12h agoHN ↗

The most shocking point is that they have a Microsoft exec who knows what he's talking about.

10h agoHN ↗

In my experience many execs know what they're talking about.

Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.

Knowledge alone is less often a factor.

Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.

Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.

8h agoHN ↗

I think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed.

This sounds rather obvious, but I feel people forget it far too often.

5h agoHN ↗

You can see the same dynamic playing out in this very thread too.

4h agoHN ↗

Oh, I think MS execs (and all other execs, top-level politicians, and pundits for that matter) know what they're talking about, but they'll tell whatever lies they need to enrich and empower themselves.

12h agoHN ↗

How is this different from Microsoft scraping to build Bing?

Honest question. There is a line in the sand somewhere apparently.

12h agoHN ↗

Huge difference between building AI and a search index.

12h agoHN ↗

Why? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn't make.

12h agoHN ↗

Bing isn’t re-selling you back the content it took.

9h agoHN ↗

They absolutely are; you pay by viewing the ads.

9h agoHN ↗

Yes but the product is different. One is a shovel for you to dig, another one is an artificial intelligence that competes with or replaces the digger.

12h agoHN ↗

Bing sent traffic to the original source. AI answers don't. That's the line. It's not complicated.

11h agoHN ↗

Under “conduct requirements” imposed by the CMA in June, UK websites are able to activate an opt-out to stop Google from scraping their content to power search features such as AI overviews - very similar conceptually to the news law passed in Australia.

https://www.bbc.co.uk/news/world-australia-56163550

12h agoHN ↗

I'd call the introduction of copyright the largest theft of human labor in history.

No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.

That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.

The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.

At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.

11h agoHN ↗

That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.

Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.

At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.

Look at who owns that site, no surprise here.

11h agoHN ↗

Vectoralist News doesn’t have the same ring to it

12h agoHN ↗

“Information wants to be free“.

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

11h agoHN ↗

"I just wished the collected data was public. "

That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.

10h agoHN ↗

On the flip side, output from an LLM is not copyrighted.

9h agoHN ↗

That has not been decided. The only thing that's been decided is that the LLM itself does not have copyright on its output.

9h agoHN ↗

I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.

49m agoHN ↗

I think the difference is that people are investing significant amounts of time, effort, and money into transforming their work into something useful, so naturally they would like some return on that investment. Giving away that work for free is not a particularly good business model if those people expect to be compensated for the value they create.

9h agoHN ↗

Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.

9h agoHN ↗

I don't think this trend of open sourcing LLM will continue for a simple reason: Money.

8h agoHN ↗

I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.

7h agoHN ↗

IP does not protect the little guy. This is nothing new. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.

5h agoHN ↗

Legislators have consistently picked a lane. Protect the rich and powerful.

5h agoHN ↗

AI training is the clearest example yet that companies are allowed to get away with what is treated as a serious crime only when an individual does it. There are many more examples of this, of course, but this one seems to be the most stark and obvious.

2h agoHN ↗

Huey Freeman: "Kim Dotcom was pissed."

4h agoHN ↗

Anthropic getting angry other AIs are trained on their AIs output is, to me, one of the stupidest things I’ve read in a while.

11h agoHN ↗

I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.

9h agoHN ↗

I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.

Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.

7h agoHN ↗

Nope. You are making one or both of these mistakes. (1) Overlooking that the same logic can lead to different outcomes depending on the premises to which the logic is applied. (2) Overlooking that real life is analog, not digital, and so thinking the premises are the same when they are not.

7h agoHN ↗

The targeted outcome is the redistribution of wealth away from the capital class to the working class. Capitalist exploitation of labor was analog to begin with, and the same premise does apply: the capital class absorbs the fruits of labor, training, and education that is performed by the masses in order to enrich themselves. The capital class owes a tremendous debt to society and if they don't plan on paying we should plan to seize it.

6h agoHN ↗

And how has exproproation been working out for you? What is that? Massive poverty and nobody wants to trade with you?

6h agoHN ↗

The exploitation of labor has has a tremendous negative impact on the environment and the health of humans. Recall that it took dragging the factory bosses from their and beating them to death to get an 8 hour work day, a weekend, and restrictions on child labor.

We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.

1h agoHN ↗

Pretty sure we had to sanction/assassinate/goad into self-destructive military campaigns the Red Terror to beat it. And even then, the major survivor still beat us to cyberpunk dystopia (the cool one with hologram skyscrapers, not the uncool one with decaying suburbs).

45m agoHN ↗

the problem at this point is none of those AI companies are profitable or are even flirting with the possibility of being profitable

11h agoHN ↗

Sharing information is an act of love

Most AI companies are not sharing it, though. They appropriated it and resell it.

11h agoHN ↗

Well the future we seem to be getting is “information wants to be free for the first ten thousand tokens, then $1 per million tokens after”.

35m agoHN ↗

Your point notwithstanding, that's still a bargain.

11h agoHN ↗

There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.

10h agoHN ↗

The key problem is that IP is either proprietary to the creator or it is a commons type of situation.

Even if you agree with the former exploiting the commons for personal profit is... not good.

One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.

10h agoHN ↗

"Information wants to be free".

I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.

9h agoHN ↗

Just because “information wants to be free” is a thing people say doesn’t mean it’s true.

8h agoHN ↗

anyone producing content, everyone’s creativity, is fed by something that others did before

The thing I produce does not replace demand for the original though?

I can’t recite the original for a million people

7h agoHN ↗

So by definition creative labor cannot be stolen? Seems like flawed logic to me.

7h agoHN ↗

Taking the original of a painting from your house is stealing. Copying is only potentially violating the government-granted limited-time exclusivity that allows you to decide who can copy your work.

BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.

6h agoHN ↗

It's telling that you need to fall back to non-creative information in your argument.

Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.

Extrapolated further, this is a collapse of creative works as a profession.

7h agoHN ↗

“Information wants to be free“

What about the rest of that quote?

7h agoHN ↗

Yikes. That’s some deep entitlement.

Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.

If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet

These AI companies actively chose not to get consent from creators and earn billions from their content with no compensation.

5h agoHN ↗

In the past, but today fewer people are getting paid less this way.

Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?

3h agoHN ↗

In the past, but today fewer people are getting paid less this way.

There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.

Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?

Uploading content online and getting a cut of ad revenue fits this criteria, no?

The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.

33m agoHN ↗

There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.

And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.

4h agoHN ↗

IIUC, the question at hand is: does training require a special, separate, license or can you legally acquire a work and then use it for training?

I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.

6h agoHN ↗

You wouldn't be standing on the shoulders of giants without IP laws, bub. Your "immense gratitude" is a farce.

4h agoHN ↗

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?

2h agoHN ↗

I think you’re confusing fan in and fan out.

12h agoHN ↗

It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.

12h agoHN ↗

Yeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did!

If this was all open, I’d maybe half agree.

9h agoHN ↗

That's what this boils down to: What are our rights?

Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.

If we take away that right, how would the Internet even work? It wouldn't.

Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.

If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!

People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.

The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.

My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.

But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.

12h agoHN ↗

AI overall is the ultimate piracy crime.

I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.

11h agoHN ↗

Then let's make humans pay royalties to every author from whom they ever learned something, even if it was offered freely to them, only fair?

11h agoHN ↗

Indeed. If I take somebody's work, transform it somewhat and sell it, I should pay royalties (unless the author explicitly allowed me to do so). That is exactly the case.

10h agoHN ↗

Sadly, they're just following in the footsteps of every large media/publishing/music conglomerate that already screwed over the vast majority of artists/musicians/writers. For every Taylor Swift striking it rich, there's 999,999 who can't even pay their bills with what their copyright gets them.

I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here

10h agoHN ↗

It is not just about artists. It is about every kind of intellectual work: scientific research, essay, fiction, painting, software, etc - the list goes on..

11h agoHN ↗

All the "LOL you wouldn't steal a car???" posts in this thread miss the point entirely.

AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.

At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.

11h agoHN ↗

I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.

11h agoHN ↗

Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?

11h agoHN ↗

I do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet.

Publicly. Accessible.

Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.

Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.

11h agoHN ↗

A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.

If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.

If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.

11h agoHN ↗

A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.

Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them.

An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible.

If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.

Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it.

Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.

8h agoHN ↗

If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.

Where is the loss? Apart from trust?

11h agoHN ↗

Just because something is publicly accessible doesn't mean you can use it for free, or that it gives you rights to do whatever you want with it.

I can access a public park, but that doesn't necessarily give me the right to also bike on its sidewalks, or walk on the grass, or take some of the plants home with me.

... or has been obtained illegally

Similarly, content that is _accessible_ publicly may be illegal to _obtain_, these aren't mutually exclusive.

On the internet, you'll find there are terms of services and licenses. These restrict how you can use even publicly accessible material. Public availability doesn't give you a license to use it however you want.

10h agoHN ↗

Public availability doesn't give you a license to use it however you want.

Of course. You listed complex examples from the real world. A park where walking is allowed but damaging the plants is not, for instance. There it makes sense to distinguish various activities that can be done in there and treat them separately.

But a website offers not much activities that you can do with it. You can read it. And that is about it.

10h agoHN ↗

If I make the path behind my home publicly accessible so people in our town can get where they’re going easier, then Amazon builds a warehouse in our town and starts driving their trucks through my path 24/7, are you going to tell me “well you made it public access, they have a right to it same as everyone”?

With the advent of trillion dollar corporations selling extremely powerful general purpose imitation as a service, the meaning of “public access” has substantially changed, potentially invalidating the original agreement.

10h agoHN ↗

are you going to tell me “well you made it public access, they have a right to it same as everyone”?

Yes. You can of course always close it to the public if the public usage bothers you. Or require those using it to agree to your terms and conditions where you restrict the speed, the weight, the time of the day, etc. Anything you want.

But if you choose to make it public with no restrictions, you have to be prepared to face the consequences.

10h agoHN ↗

Well, I posted a sign saying that but they’re still doing it, so I guess I’ll just close it to public access. Sucks for all of the people who now have to take a longer route. I hope they blame Amazon and not me.

11h agoHN ↗

The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.

11h agoHN ↗

One leads to the other so its simpler to point the root issue.

10h agoHN ↗

And the cruel irony is that they stole our work to train the AI that devalues our work going forward, and will cause many of us to lose our jobs.

I regret every line of open source code I ever wrote.

10h agoHN ↗

This process can even feel a bit like parasitism, it empties out the host, like in <Alien>.

9h agoHN ↗

I regret every line of open source code I ever wrote

And every stack overflow post, every reddit post, everything.

I regret participating in the open Internet.

Here I am anyways, I guess. It's just in my genes or something.

3h agoHN ↗

Same. I've scrubbed my presence as much as possible from the majority of the internet (except HN for whatever reason) to prevent future models from being trained on my output, but the damage is already done. Part of me exists in pretty much all AI models now, without my consent. And those models are actively stealing my career and passions.

7h agoHN ↗

I think you just described "technology" -- would you outlaw all technology then? ;-)

5h agoHN ↗

Right? I've been automating things and eliminating jobs using technology for over 20 years. No you do not need a team of five to manage your monthly reporting. We can build automated reports that are delivered automatically to everyone who needs them! No you shouldn't be managing all of this data in an excel file on a shared computer. I don't care if 90% of what Dave does during the week is maintain that file. Let's streamline that.

Suddenly it's very different when it's our jobs being automated away.

1h agoHN ↗

Yes, I don't like it much either -- though I continue to be absolutely fascinated by all this -- but I realize this is unstoppable, and I must trust that on the balance technology has been extremely good for humanity and the source of all our progress, and so I must adapt to a new future.

But I also realize the impact of technology, good or bad, entirely depends on how society uses it, and that is where our focus must lie.

11h agoHN ↗

The largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices.

It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.

If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.

11h agoHN ↗

I think it's okay to advance humanity, but they can GTFO when they then try to ban distilling and open models.

9h agoHN ↗

Is AI advancing humanity, though? Or is it only advancing technology while divorcing it from the human?

9h agoHN ↗

Define, "advancing humanity."

Are we talking about turning everyone into philosophers and somehow ascending to a higher plane of existence? Yeah, AI isn't going to help with that (probably).

Or are we talking about useful, positive benefits to every day people like better speech recognition, tools for the visually impaired, disease research, physics research, science in general, and loads of other areas where AI is improving things?

11h agoHN ↗

But what about M$ owning Github and doing the same with its content? Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...

11h agoHN ↗

You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.

7h agoHN ↗

There is nothing funny about $9Tn in FOSS getting misappropriated, and sold as isomorphic plagiarism tokens... or the estimated $4.6Tn in loan debts these 7 companies incinerated when the bubble pops.

Anyone that lived through the dot-com or housing market bubble know what a collapsing Ponzi scheme does to real businesses, and peoples retirement funds.

Popcorn ready =3

11h agoHN ↗

I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.

8h agoHN ↗

It's the age old privatize the profits and socialize the losses.

People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...

I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?

2h agoHN ↗

People are like sheep. Plus those who work are preoccupied and are too tired to react to these changes.

11h agoHN ↗

There will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.

9h agoHN ↗

Absolutely. The parrots will get so clever that they'll extract knowledge from pure vaccuum.

11h agoHN ↗

As someone who thinks that genAI is harmful, I deeply resent that any of my work has been used to help train it. I will never forgive these companies for forcing me to contribute.

11h agoHN ↗

Because it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.

11h agoHN ↗

I don’t mind these companies scraping my content.

But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.

11h agoHN ↗

"The Net interprets censorship as damage and routes around it."

11h agoHN ↗

Copyright is artificial scarcity rationalized by arguing that producing novel intellectual work is valuable, but requires substantial effort that can't be recouped, so we have to incentivize it somehow.

LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.

AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.

11h agoHN ↗

If there is no legal guarantee that human creativity can pay off we are starving art and humanity from its inception. If ordinary people cannot participate in the act of creation, you get exactly what hollywood has become.

Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.

9h agoHN ↗

Lack of way to capture value from intellectual labor is considered a market failure that leads to suboptimal market results for consumers, but with AI that argument becomes very weak.

Sorry but the point isn't to create artificial scarcity just so intellectual labor is well off, that's a negative side for the consumer that was considered necessary tradeoff. Market economy should be about providing the most value to the consumer.

Disclaimer - I was never a fan of IP laws despite them working in my favor, with AI I can see them finally being abolished.

8h agoHN ↗

    We could imagine, as an extreme case, a technologically highly advanced society, containing many complex structures, some of them far more intricate and intelligent than anything that exists on the planet today – a society which nevertheless lacks any type of being that is conscious or whose welfare has moral significance. In a sense, this would be an uninhabited society. It would be a society of economic miracles and technological awesomeness, with nobody there to benefit. A Disneyland with no children.
11h agoHN ↗

They are not selling the information. They are selling a service for easy access to that information.

These are two different things.

Note: am not an AI fanatic.

10h agoHN ↗

Correct solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use.

There should be a carveout for non-profit or government AI.

10h agoHN ↗

First of all it can. Second I'm talking about attribution to copywritten material (and royalties being paid accordingly). This would only be for for-profit AI implementations though.

10h agoHN ↗

Copyright infringement, if this even were that, is not theft. Chattel slavery is the largest theft of labor in human history.

7h agoHN ↗

Trademarks are intellectual property, and every LLM knows what Mickey Mouse looks like. =3

10h agoHN ↗

Information wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it

Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it

6h agoHN ↗

"And all that" is a linguistic trick to imply "yes, this is being discussed -"

10h agoHN ↗

I just don't understand people saying "but a human learning from a book isn't illegal".

How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."

And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.

9h agoHN ↗

I agree that there is not an orange to orange comparison, but I have still not seen a law that could scale as well. The example that comes to mind with a proposed law like this is how would you license work that build on another work? What if I decide to publish a blog post after taking some course, that distills whatever I learned in the course for free?

8h agoHN ↗

We already have laws that cover this.

If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place.

If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original.

And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property.

Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans.

I do think copyrights should be shortened to 20 years but that’s another discussion

9h agoHN ↗

Because it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book.

Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.

5h agoHN ↗

It's the same bad faith argument as:

"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."

4h agoHN ↗

I don't know if it's even OK for a cop to just randomly watch a random street corner without a specific reason to do so. If a cop was following a given person around all day every day without probable cause would this not be considered harassment? If we view the Flock Camera network as a single system is this not harassment?

3h agoHN ↗

Just to add context and not to comment on rightness, that was one of the early core purposes of policing--to ensure that a people in a particular public space (where either the people or the space are particularly vulnerable) are free from unwanted disturbance. Here is that exact concept shown in a painting from 172 years ago: https://en.wikipedia.org/wiki/The_Gleaners_(Jules_Breton).

1h agoHN ↗

I haven't seen this articulation before, and I'm going to remember this one!

In return, I'll share my own reasoning:

Human Life is finite, and there is a real opportunity cost to learning (say) to copy Picasso's style -- you could've been doing something else in that same Time.

However, if you're training a model on the entire career output of dozens of artists, it only needs more electricity + GPUs to do this.

And, transposing this back to the Human realm, there's no way one person can put in that kind of wide effort to learn how to copy dozens of artists' styles in one lifetime.

I really like the way you put it too -- it's a bit more succinct.

8h agoHN ↗

This is a similar problem with data brokers. They take public data laws to an extreme and resell easy access to the aggregated data. This easy access has created a tremendous number of problems unforeseen by the original intent of the access.

Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.

Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.

8h agoHN ↗

I would argue artists imitating, for example, the artists who revolutionize a genre of music are also in fact massively reducing the demand for the originals by 99%. Imagine if no one ever made a song after the Beetles that sounded any newer. The demand for Beetles music would have remained somehow even more enormous than it was for the last 50 years. Or if no one did abstract paintings after Kandinsky, or impressionism after Monet or cubism after Picasso.

7h agoHN ↗

Because copyright is a tradeoff. Society wants information to be free. But authors won’t publish if anyone can republish their work for free. Hence the distinction: The ideas you write are not protected - only your expression is.

In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.

6h agoHN ↗

Society profits from having AI

Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.

4h agoHN ↗

authors won’t publish if anyone can republish their work for free

There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).

[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.

4h agoHN ↗

Allowing AI companies to get away with this because we're "better off"¹ is like eating our seeds. Sure, we built something cool with the massive corpus available, but how will it affect future decisions to develop or share creative work?

If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).

1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.

2h agoHN ↗

Agreed. The most damning behavior was Meta.

Specifically, reaching out and discussing licensing material, then pirating because it was too expensive/slow to legally acquire it.

Heaven forbid Meta have to pay for something.

6h agoHN ↗

Being anti-permission culture is not the same thing as bad faith. It is just rejecting your arguement as stating scale suddely changes the rules. Plus the whole argument based on supposed impact is missing the point like saying that free speech should be treated differently for being more impactful than anticipated with mass publication. Something throughly rejected.

6h agoHN ↗

I think that last thing is key. The authors of these things published them with the express intent that other people would read them. Not that they would be used as training data.

6h agoHN ↗

"How do people not understand that some laws only make sense at a certain scale? "

Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.

4h agoHN ↗

Murder, terrorism and genocide.

Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.

I'm sure there's other examples too.

4h agoHN ↗

Theft was historically very scale-sensitive. In England you'd face execution for "grand larceny" if the goods in question were worth more than twelve pence. This was on the books from the 13th century until 1832 (!).

36m agoHN ↗

Jesus Christ on a pogo stick, execution for stealing 12 pence worth of stuff? That's crazy! Adjusting for inflation[1][2], that would be the same as nowadays being executed for stealing £5!

[1] According to https://www.bankofengland.co.uk/education/education-resource..., 12 pence (1 shilling, or 1/20 of a pound)) in the old £sd system is equal to 5 pence (or £0.05) in the later decimal system.

[2] According to https://www.bankofengland.co.uk/monetary-policy/inflation/in..., £0.05 (decimal) in 1832 would be equivalent to £5.02 in 2026 Aug.

4h agoHN ↗

Laws do consider scale and we can see that in everyday law.

4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.

Moving around with a couple of bills is treated differently from carrying huge bundles of cash.

Those laws or exceptions were probably added later as a reaction to abuse of existing laws.

The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.

It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.

2h agoHN ↗

Laws are typically implicitly considerate of scale. Very rarely are they explicitly considerate of scale.

To wit, sharing music digitally, Google scanning books, or Uber/Lyft providing unlicensed taxi services.

It's pretty rare that laws consider what should happen if it were suddenly possible to 10x or 100x preexisting throughput.

And specifically where laws balance multiple, often-opposed, stakeholders' interests, that change can drastically upset the previously negotiated compromise.

Which is why piracy at scale, before it's banned, tends to be a successful foundation for many businesses.

5h agoHN ↗

That's eerily similar to the argument used by mass surveillance systems like Flock around expectations of privacy in public.

Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.

4h agoHN ↗

How do people not understand that some laws only make sense at a certain scale?

Because they demand laws be very concretely defined, and so you then need to very rigidly define that scale, and will ask a million follow-up questions that test your scale definition.

But of course, it's all bullshit. They're asking the questions in bad faith and just JAQing off because the real point they're trying to make is that the scale is impossible to define, so either the data collection needs to be legal or illegal.

1h agoHN ↗

I would hate to live in a world where questions of "legal or illegal" are "impossible to define."

4h agoHN ↗

My problem with your solution to require author permission to train is that it doesn’t solve anything long term.

What happens when all the licensed information still leads to the creation of demand hoarding AI? Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.

37m agoHN ↗

Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.

Well then they cannot possibly stealing this, and short of creating laws that directly discriminate between algorithmic processing and human consumption - regulating the process, not the subject - this argument is quite literally nonsense.

3h agoHN ↗

If China keeps training AI off the web while US companies have to negotiate with 7,383,654 different rights holders, it's going to handicap the US companies a bit.

3h agoHN ↗

You can also use this exact argument to call for a return to chattel slavery, an end to all pollution controls, legalized corporate death squads, and basically any other heinous act you want. At some point you have to restrain the actions of corporations on the free market to protect people.

2h agoHN ↗

Yep. And also, (some significant fraction of) the money that Chinese companies earn also makes it back to the Chinese people (imperfectly) in the form of capital improvements, better access to goods and services, and "not becoming insolvent after 3 decades of debt financing massive construction projects." Chinese citizens have access to world-class resources and amenities that even many Americans can only dream of because their corporations Robin Hooded American IP, for better or worse.

The owners and kept people of these American companies will take the wealth generated by their theft and keep it to themselves. They talk about a "permanent underclass" with disguised glee. Break their operation until they learn noblesse oblige.

10h agoHN ↗

"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. "

So training can make it legal as well. Interesting...

10h agoHN ↗

Theyre just jealous that theyre being surpassed on their market capture of computing

10h agoHN ↗

Yet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants.

The hypocrisy of this new world is already catching up to us.

10h agoHN ↗

I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.

10h agoHN ↗

Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours

So no, your cries for regulating others because you are losing the race won't work this time.

9h agoHN ↗

How much money have people paid to use these services over the years?

None?

Ok, now you understand the business model.

1h agoHN ↗

I understand the business model, we are data providers, they are aggregators. That's fine.

What's not right is that they want to limit the use of such data when it's not theirs in first place. They just store it but that doesn't give them a license to prohibit the use by a third party since we all are owners of that data

10h agoHN ↗

Microsoft is one to talk.... Remember when MS trained copilot on all your github code?

10h agoHN ↗

Imagine the parthenon marbles. When they were looted it was even a celebrated act, but they are still stolen in the british museum centuries later.

10h agoHN ↗

Throughout history we’ve been able to retell stories, to copy content, to create shallow clones or synthesis

It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.

10h agoHN ↗

I remember when open source software, and Linux in particular, was the threat to the world according to Microsoft execs.

9h agoHN ↗

Microsoft executives levelling "tone-deaf" up in realtime.

9h agoHN ↗

Are we considering what is the shelf life of information?

If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.

If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.

9h agoHN ↗

Yet he also helps destroy all those jobs. The thief is calling "Catch the thief!".

He does not see the moral dilemma here?

9h agoHN ↗

Transatlantic slave trade calling - we’ll hold. I know there’s a memory shortage but history books are cheap.

8h agoHN ↗

Slavery was widespread and universally practiced across the world throughout history. Everyone posting on this board is a descendent of someone who was a slave at some point. Transatlantic slave trade is a drop in the bucket.

9h agoHN ↗

Funny. The end of copyright and patents is by far the best thing about AI to me. All of it is nonsense. Great that you drew a picture of a mouse once, I really fail to see why I couldn’t draw it and sell it either. It was moronic from the get go.

8h agoHN ↗

Draw a mouse then. But, I do agree in principle that disney's mouse is theirs. Forever. Whenever I see the little rodent, I think 'disney', and if you drew a similar mouse and you aren't disney, then you've cheated me.

9h agoHN ↗

I mean ... I know a few African nations that might disagree.

8h agoHN ↗

The black folks still in Africa were the ones selling the slaves so maybe not.

9h agoHN ↗

The righteousness of the internet, the same internet that desperately called to end IP laws, championed piracy, ad-block everything, and always use proxy services to backdoor paywalls/login walls

This same group of people, now being on the other side table, are screaming an crying that it's not fair.

Grow up and reap what you sow.

8h agoHN ↗

It literally is. Anyone saying otherwise is deluding themselves

8h agoHN ↗

What if all this slurping and training empowers humanity to cure cancer? feed the starving? travel the stars? We are not only making the accumulated knowledge of the world accessible, we are making it actionable. Sure I'm ignoring all the possible bad outcomes lol, but if an independence day (movie) type scenario was playing out nobody would be batting an eye. I guess cancer is not as sexy though

8h agoHN ↗

Apart from jacquesm's wonderful milquetoast comment, if anyone has any practical solutions to this problem that do not involve suspension of disbelief that voting (with or without wallet) and calling "your congresscritter" or whatever other nonsense people spout, now would be a great time to voice it.

8h agoHN ↗

The evidence here is quite damning for OpenAI and Microsoft is clearly trying to distance themselves from OpenAI’s behavior here.

8h agoHN ↗

Framing it as property is the wrong angle though. It's an ecosystem, not a cache of good writings that was stolen. That ecosystem was hurt severely and its recovery is uncertain.

I think we'll not have people writing good content for a long time (there's no reason or incentive to), and the effects of this will splash back heavily on AI companies themselves.

You can see AI as a battery for intelligence that took a long time to charge and it's being used right now. For years, it was charged with all sorts of novel content that went undiscovered and AI is making available. That charge is the production of novel content, new insights, cross-pollination between areas, slowly driven by humans.

My view also draws a conclusion about recursive self-improvement: it is impossible for a battery to re-charge itself. I don't particularly think it can be done with this technology (LLMs).

I could be wrong though, but I don't think I am, and we'll know within our lifetimes. If things stall, it's likely because it has ran out of seeds/charge/substrate and not a technical limitation. It is in the long-term interest of AI companies to make incentives for people to generate novel public insights, they just don't know that yet.

7h agoHN ↗

Yes. The resulting LLM is much easier to find information than ever before. To me, that's an improvement on the source data, much like the hitchhiker's guide to the universe. If it results in breakthroughs especially in health, steal away baby.

4h agoHN ↗

information was easy to find since wikipedia and google (when both were unpoisoned), LLM-s were already good as "books you can talk with", now when LLM-s start cosplay experts they are not (and try make decisions they are not fit to take responsibility for) it gets literally insane, like looking at car wreck in slow motion

7h agoHN ↗

LLMs are not compression algorithms. From an information theory perspective, that's impossible given their size.

Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.

It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.

7h agoHN ↗

It’s irrelevant if it’s legal today or not. This is new technology and may be new precedent.

5h agoHN ↗

... but you have to pay to read the NYT. You paid for the information. Guess who didn't.

7h agoHN ↗

You or I scrape a website for casual use? Straight to jail, right away.

Big tech scrapes ALL OF THE WEBSITES CONSTANTLY to resell to you as knowledge? Shut up and take all of my money.

The 2020s is the wildest timeline indeed.

6h agoHN ↗

Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work.

Sounds about right for the people and orgs involved.

6h agoHN ↗

As if the transatlantic slave trade wasn't a thing. What a perfect illustration of our industry's self-absorption and self-regard.

6h agoHN ↗

Has anyone found a meaningful discussion of how scraping is theft?

Obviously, reproducing works in whole is infringement. That's not what AI is doing, so the question becomes: how is scraping different from ordinary reading? Is it just that site owners want to play back history and retroactively create high-cost licenses for scraping?

24m agoHN ↗

A key consideration in most legal definitions of theft is “intent to permanently deprive the owner”. While this has historically meant scraping is not theft (because copying doesn’t erase the original, nobody is deprived), in this specific case the AI companies’ business plan (copy a person’s content and train on it to make their model more capable of replacing that person) could very well meet the bar of intent to deprive.

6h agoHN ↗

I'm not sure how to be upset over this. For decades copyright has been extended and extended. Meanwhile, the ease and speed of spreading published works across the globe have increased massively. Don't forget that when copyright was first created, it could take multiple years for first editions to make it across the globe. In that environment, multiple decades of copyright makes sense. Whereas today that would actually stifle innovation and creativity rather than incentivizing it. I mean, because of the length of copyright we've gotten all of these live action remakes of Disney films or superhero movies rehashing the same stories.

I find it a little hard to be upset about AI. Supposedly stealing copyrighted works when the vast majority of those works. Probably should have been in the public domain to begin with. I have a faint hope that this scuffle between the AI companies and the publishing industry will result in more reasonable copyright laws, but I think it's more likely that exceptions will be made and AI will be treated as a special case.

1h agoHN ↗

I'm upset because I have to pay to use it or else be hit with onerous rate limiting for stuff that "should have been in the public domain to begin with."

6h agoHN ↗

It's the scale that's the problem. I've been saying for a while, the value of any individual piece of work is not terribly valuable for an LLM, but the aggregate value of all human work is obviously very valuable. I think this line of thinking can be the argument for why we should heavily tax these AI companies above and beyond how we tax other industries.

Also that if you're going to use these things to write software, you should make it as virally copyleft as possible https://jackson.dev/post/moral-ai-licensing/

6h agoHN ↗

But Windows telemetry, github code training, scanning all OneDrive documents/outlook emails is not theft

5h agoHN ↗

This is a great basis for a dividend from AI revenue to be paid back to society in a more inclusive form than stock. As more money flows into AI companies, data centers, and other related infrastructure, an amount should be extracted and redistributed in the name of balancing this equation.

5h agoHN ↗

What’s the Microsoft angle here? A few years ago they were ready to hire anyone from OpenAI that was willing to leave. Is it just catering to the anti-AI sentiment going around?

4h agoHN ↗

Atlassian, powered by this spy tool, is truly insane:

“We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.”

All these failing companies are trying to bullshit their way out of the decline. Atlassian could have, you know, come up with a usable GitHub competitor. Instead they dream about coding by yapping.

4h agoHN ↗

[Nadella said] if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

Whew, good thing it’s too late to be accountable for that now, huh? Water under the bridge. Mistakes were made. Eggs, omelets.

4h agoHN ↗

Go cry me a river of lies Mr. Microsoft Exec. Corporate culture is a blight on humanity.

4h agoHN ↗

This will lead to a massive settlement between the big boys and most people who put stuff out in good faith will be left out of it. And that will be the end of it. We will never hear anything about this ever again and the ‘theft’ will continue like normal.

3h agoHN ↗

Source:

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

p. 1

"This case is about, as Microsoft's Director of Applied Science put it, an astonishing theft of unprecedented proportions; SF1437, perhaps the largest theft of labor in human history. SF1652"

p.11

"As Microsoft recognized: millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions and admitted that almost no one intended for content they created to be used in this fashion, nor are they compensated for its use. SF1437."

p. 74

"As Microsoft's Dr. Glen Weyl put it, compensating creators is in the best interests of my employer, of my country, and of many other groups I belong to. SF1657."

Hyperbolic quotes from Microsoft employees are, IMO, the least interesting elements of this brief

Here is Microsoft's brief. Note how MSFT responds to the "web grounding" claims

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

It seems OpenAI does not want the public to know about (a) OpenAI's data collection and retention practices and (b) the number ChatGPT users have requested deletion of conversations

https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...

"OpenAI seeks to redact specific information about [(a)] the number of users who requested deletion of ChatGPT conversations and [(b)] OpenAI's related data collection and retention practices."

"Disclosure would give OpenAI's competitors insight into OpenAI's confidential business practices and customers and cause competitive harm to OpenAI. Yeats-Rowe Decl. 4."

Perhaps it would causes competitive harm because, upon learning about OpenAI's privacy practices, ChatGPT users might reduce their usage of ChatGPT

Declaration is sealed so we can only guess

2h agoHN ↗

So Microsoft exec stating what most people have been saying already is newsworthy.