- 107comments
- 20comments
- 244comments
- 10comments
- 79comments
- 226comments
- 81comments
- 22comments
- 82comments
- 26comments
- 26comments
- 23comments
- 3comments
- 46comments
- 631comments
- 102comments
- 41comments
- 100comments
- —discuss
- 11comments
- 180comments
- 30comments
- 38comments
- —discuss
- 269comments
- 57comments
- 75comments
- 12comments
- 2comments
- 20comments
Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics - i.e. 'openai uses copyrighted data from sketchy russian website’ showing up on HN would be unfortunate."
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
It's good to know we all live rent-free in Dario Amodei's head
You might be thinking of Demis Altman.
Think they're referring to the following, when Amodei was still working for OpenAI:
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
I was curious what he was responding to. Per https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff... it was "On July 19, 2019, McCandlish wrote in an OpenAI Slack channel: “We’re not sure if we’re going to release the Foresight LM Scaling paper publicly, but if we do we were thinking about removing all mentions of LibGen, since it's a bit of a sketchy data source." The paper may or may not be https://arxiv.org/abs/2001.08361 where they write "we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-available Internet Books."
Same as all the world’s copyright text lives rent-free in Claude and ChatGPT.
As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.
This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
"A lobby organisation?" Of course a single author would not be able to afford facing a multi billion dollar company on their own? And the "sketchy russian website" quote is from OpenAI employees themselves? What are you on about?
First, it's still a lobbying organization, so it's their job to make exaggerated claims, like a union in a company or any other organization with a political purpose. My first point was to highlight that it's not neutral or news related. It's fine that they have their opinion, but it's also my right to say that they're biased.
The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
No, you are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.
They are having it in the subtitle. In general their whole article is about two main points, first the use of stuff from Libgen, second the points of making people jobless. Three of their points are about the jobless thing two about the LibGen.
About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?
There is nothing to discuss about LibGen, really. I don't read this as OpenAI employees even believing LibGen is sketchy. They were worried about optics, because LibGen itself is Russian and does look a bit sketchy, and at the time - much like today - it was easy to make it a headline that makes people pattern-match to "troll farms".
(And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)
Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.
I agree with you, but I acknowledge that some people (not myself) might feel different about copyright.
I acknowledge that, but this part isn't even about copyright!
The whole "sketchy russian website" bit resolves entirely about being seen as associated or supporting troll farms and Putin.
EDIT: look at it this way: no one is calling Internet Archive "a sketchy US website".
Any source that this was a library? Even then that would still raise a question if OpenAI is a Russian organisation or not to access that library with good faith.
I think they just used a Russian torrent site.
They mention it in the text, its this: https://en.wikipedia.org/wiki/Library_Genesis
Just go look: https://libgen.im/ (may be blocked by your isp, just find an alternate URL somehwere)
(note: https://z-library.sk/ is prettier/nicer)
They're talking about LibGen, per the article.
apparently, the website in question is libgen.io (appears to have been taken down now), currently it has lots of mirrors like libgen.im , libgen.com.de, etc.
"This is a lobby organisation using only the pieces and bits they like to push their own agenda."
Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
Simply look at the Wikipedia article?
Have you looked at their name?
It's just sad that this is the top comment on Hacker News. Why are we giving free pass to these tech companies? Why are we trusting these CEOs when they have repeatedly broken laws? Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
Then you misunderstood my point. I think copyright law should be significantly changed and the current system hurts us all.
Sure. But the fact is they broke current existing laws, with known punishments with precedents. Same as a new law doesn't retroactively punish someone, then a new law shouldn't absolve someone before it's passed.
The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
They did. They were found guilty. The case I looked at* was a civil case so this was settled out of court before the court imposed a settlement.
The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.
IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.
* the Anthropic one
How? The current system enables the GPL. The GPL protects many open source projects.
All the GPL does is restrict rights, unlike permissive licenses.
So.. Linux is bad and we should all use BSDs?
Why has restrictive Linux succeeded far more than any BSD ever has?
No, they were responding to the post defending OpenAI that you wrote. If you meant to communicate something other than “criticism of OpenAI in this context is unwarranted” then it looks like you forgot to do that and wrote something else instead
Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws are treated by these companies as mere instructions.
Big tech will enable access to free information and will help people reach new heights. Do you see how wrong that sounds?
Suppose you have three propositions:
A: "Information is free"
B: "Information is not free"
C: "Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility"
You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you/Swartz to prefer C over B while A is not true?
Sure it’s not free for anyone and both companies and individuals are treated similarly. It’s not like you will be jailed for pirating movies. And neither should OpenAI. What part of this is hard to understand
We're literally talking in a thread about someone who committed suicide because the US government was hellbent on ruining his life with a felony conviction for piracy.
Sounds like a gross misapplication of the law enforcement system. The free exchange of information should be legalised.
As we can see in the good article, the copyright system might almost have cost us a lot of AI capabilities. The damage it has done in cases where the lawyers got ahead of the builders is incalculable.
His case was materially different
1. Unauthorised network access
2. Intent to distribute licensed material
The labs aren’t doing this. They are doing something similar to you and I downloading torrents. Look, I also think laws should apply somewhat equally to individuals and companies. But this is different.
We tend to accept lawbreaking if it is for a product we want.
There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
Aaron was sued for distributing and not for pirating. They are different offences
Calling out equivocation and out-of-context quotes is not “giving a free pass”. If there is a case to be made (and to be clear, I believe there is), then shouldn't be resorting to rhetorical slight-of-hand to “prove” it.
Because this website is full of bootlickers without real friends. People cut off from social life. Losers that hate humans.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
100% of the rulings agree with me.
The piracy is not in question. It is unarguably copyright violation.
But that's not what anyone means in this context. Training is what everyone means.
I didn't say otherwise. That's a straw man.
So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?
No.
So, which is it? First you said it was fair use, now you’re saying it’s unarguably copyright violation. You can’t have it both ways.
Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete which that kind of copying and pasting (in encoding from) from copyright-protected material. And the scale is what hurts authors most: millions of copies on demand floods the market without paying.
Exactly analogous to a human reading the material.
Please give me the location of any copyrighted work in the weights of any open source model or a prompt that will retrieve it.
They didn't believe it was sketchy. They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.
Judging by how AI threads look like for the past year, they were absolutely right to be worried.
In fact, you're doing exactly that right here.
You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.
Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
The comment you're replying to is citing almost verbatim [1] Microsoft’s director of Applied Science, Brent Hecht, who called OpenAI's data collection practices "the largest theft of labor in human history" in an internal memo.
[1] https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s...
Yes, one man at Microsoft agrees with you.
It's not citing, it's regurgitating, like a stochastic parrot.
You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.
Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve of libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.
“Sketchy Russian website” is part of a quote by Sam McCandlish (who worked at OpenAI), not the Author’s Guild characterisation.
Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
I do see a problem with a company loudly announcing that they are going to make people's lives miserable purely for profit. Leaving aside that it goes against OpenAI's stated mission ("to ensure that artificial general intelligence benefits all of humanity"), the disdain for the lives they are intentionally trying to ruin makes it a problem.
And even if you believe that the transition is inevitable, as it is the case with phasing out combustion engines in cars, anyone reasonable would see that the transition is gradual to give people time to adapt. Instead of doing that, OpenAI is burning cash at astonishing rates, polluting the environment, and killing personal computing with the only aim of being the only ones left atop the ruins. I do see a problem with that.
https://web.archive.org/web/20260927062011/https://authorsgu...
Thank you.
Not even openai want to get flamed on show hn
I could not care less
Can we fix the title? This is clickbait for HN, the actual article's title is "Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work"
TFA is by the plaintiff. This is PR spin, not click bait.
Pulled datasets off libgen for a corpus once and it's overwhelmingly in-copyright textbooks, the public domain framing doesn't really hold up.
Is our opinion that important ? Nice !
It's wild how much weight "they're destroying jobs" has.
I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.
Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
Why is book piracy a better way of doing things?
You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
What's twisted? They pirated books to train their models.
Clearly these books had value to AI companies but they were too weak and too dishonest to pay for that value. That's not impressive.
Yawn
Because you want as many people as possible to be exposed to your ideas, to the point that Christianity used to fund armies to go to other places so they could force the teachings of Jesus Christ upon them. If your thoughts aren't at least as good as that, why are you wasting eink on them?
That is a pretty irrelevant point considering they could have just bought all the books and we would still be in the same situation.
Obviously the authors should sue them to bankruptcy though.
I wouldn't put AI on the same level as the inventions of the industrial revolution. Automation replaced specific tasks. AI is replacing humans themselves. What are humans supposed to do in a world where everything a human can do, a machine can do too?
Machines still cannot make babies ... Just sayin'.
There's a book about that, here is its ACX review:
https://www.astralcodexten.com/p/book-review-deep-utopia
There are a lot of tasks machines can do today that humans do anyway. Think baristas and bartenders. People pay more for human-made goods too, despite mass-manufacturing being a couple centuries old at this point. Having a human face is a competitive advantage in lots of stuff.
And this is why we see ever better attempts to mimic human skin, faces, movement, emotions, etc, with robot manufacturers.
It's interesting that this argument wasn't popular decades ago, when the internet basically made publishers and other copyright rent seekers redundant. Instead the response was harsher enforcement and more laws to protect the rent seekers.
Amplification of commentary and opinion is much easier now. It’s also a lot easier to play into people’s fears.
There’s a wide range of political motivations for this sort of thing too. Not just AI taking jobs, but convincing people to vote against their best interests, or even supporting radical religious groups whose values would persecute you. This sort of rhetoric feels noisier than it ever did decades ago.
Everybody gets a platform now, even the bots.
The issue is the breadth of scale. At the time of the industrial revolution those who we would consider knowledge workers were a much smaller percentage of the population. Today, they are a significant part of the workforce, and it is them who are being displaced first, and in increasing numbers.
The second issue is one of migration between skill sets. By the numbers, today is likely harder as non Ai-affected skilled professions take longer to train into in a general sense.
The third is the synthesis of AI and robotics. Humanoid help is good. As the Venn diagram overlap of capabilities between humans and humanoid robots increases, many professions defined by complex vision and environmental manipulation that are currently safe will not be.
There are, of course, upsides. But the world is right to wonder what the future looks like, and where people fit within that future in terms of earning an income if whatever path they choose seems to be able to be replaced by a machine within their lifetime. How do they achieve personal stability and safety if, even when thinking about their multitude of career options, the machines are so good that they can do almost anything.
Yes, but we're the horse in that scenario.
Lol, the fear mongering is just part of the hype train at this point.
Yes it’s all hype to cash in the ipo before the bubble bursts. The labs are spreading spooky things about AI so that people think these models are very capable. It is a coordinated effort amongst all labs. The reality is that these models can’t do basic mathematics.
If this doesn’t put anyone in prison then we might as well declare copyright dead. They knew they broke the law and then they deliberately covered it up. How much more evidence is required here?
It’s 2030, I open HN and top result is from 7 hours ago, 2100 points, 700 comments:
The optics are bad. The submission marks a turning page in human history.
Can some explain why would a billionaire like Altman care about the opinions expressed on HN?
Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
You forget how many influential tech professionals are present here (especially those in Silicon Valley) and that Altman was president of Y Combinator for several years.
Elon Musk seems to care a great deal about a lot of stuff that frankly isn't any of his concern, not really. Trump care immensely about shallow stuff that's really below what the president of the United States of America should spend time on.
Altman has some reason to worry, because developers and the users of HN are then ones he needs to promote/hype his products. OpenAI still needs customers and they need to sell tokens and a large number of users on HN are not only not buying, but are becoming increasingly hostile towards the product Altman is selling. My guess would be that he'd assume that the users on HN are in some sense his peers, and they are currently extremely divided in the question about the value and dangers of LLMs, and are increasingly critical of the business practises of OpenAI (and other AI companies).
the AI isnt going to destroy jobs in numbers large enough to blip the unemployment stats more thna a qtr a percent.
the economics of those companies will cause a catastrophic wipe out of jobs across the board.
HN is their AB testing playground.