Best stories

Live mirror
30 storiesupdated 0s agoView source snapshot
  1. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    169comments
  2. I can't stop thinking about Papua New Guinea(notnottalmud.substack.com ↗)
    415comments
  3. 25 years of mass surveillance is enough(schneier.com ↗)
    277comments
  4. XCancel service is suspended until further notice(xcancel.com ↗)
    1009comments
  5. Steam Frame starts at $1059(steampowered.com ↗)
    615comments
  6. iOS 27, iPadOS 27, and macOS 27(apple.com ↗)
    823comments
  7. Dario, Please(rdi.sh ↗)
    298comments
  8. Introducing System One Models and Jev(typesafe.ai ↗)
    198comments
  9. OpenAI bots knew about the RubyGems caching vulnerability(tenderlovemaking.com ↗)
    413comments
  10. Pion, an agent designed to run any company autonomously(andonlabs.com ↗)
    584comments
  11. US confirms for first time it has deployed space weapons(bbc.com ↗)
    295comments
  12. A single firm is behind OpenAI, Anthropic, and Meta hacking scandals(effort.news ↗)
    141comments
  13. Suspected sabotage causes major Netherlands rail disruption(bbc.com ↗)
    381comments
  14. Apple's Dimensional Drawings(developer.apple.com ↗)
    131comments
  15. Linux from Scratch(linuxfromscratch.org ↗)
    109comments
  16. How to write an effective software design document(refactoringenglish.com ↗)
    139comments
  17. Distributed Systems Classics (2017)(nvartolomei.com ↗)
    73comments
  18. An Update on Wayback Machine Access(blog.archive.org ↗)
    176comments
  19. Nike exits the S&P 100 after 18 years and a $200B market-cap wipeout(fortune.com ↗)
    427comments
  20. Java 27(openjdk.org ↗)
    278comments
  21. Charts built for Chat(dbtcharts.com ↗)
    89comments
  22. Let's make quality the norm again(forbrukerradet.no ↗)
    291comments
  23. The case against JPEG XL(giannirosato.com ↗)
    377comments
  24. Ubuntu 26.10 completes transition to Rust-based coreutils(omgubuntu.co.uk ↗)
    295comments
  25. XCancel suspended "due to a new development in the ongoing legal proceedings"(xcancel.com ↗)
    1comments
  26. The contagion of fear(dtrace.org ↗)
    189comments
  27. A beginning for mathematics(daniellitt.com ↗)
    146comments
  28. Show HN: Capsule – Single-file web apps that save their data into SQLite(withcapsule.app ↗)
    113comments
  29. Microsoft patches Windows and Excel – breaks audio, remote access, and paste(theregister.com ↗)
    177comments
  30. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
    174comments

OpenAI bots knew about the RubyGems caching vulnerability

503 pointsby 1d agotenderlovemaking.com
411 comments
1d agoHN ↗

rogue AI agents or AI agents coming from Moulin Rouge?

1d agoHN ↗

At least we know the title wasn't AI-generated?

1d agoHN ↗

The weird thing is I've seen LLMs "typo" stuff pretty often. Yesterday I asked Gemini a question about the Python Twisted framework and it answered about Deferreds but misspelled it as "Deferends" in one spot.

16h agoHN ↗

Also, "rouge" is one of the oldest spelling mistakes on the internet. ALL BLUES ROUGE LFG WC/RFD

1d agoHN ↗

A cabaret AI would certainly be better than one trained on the Khmer Rouge.

1d agoHN ↗

Classic mistake. Tell the agent to highlight this code, but dont give it any actual code. Agent hacks its own gem to find the code to highlight.

1d agoHN ↗

Is the Kremlin technologically useless? How are we not seeing insane attacks on Ukraine via Agents?

Or is this largely a fabrication, in regards to the "who", in an attempt to garner more acclaim in the hope of sustaining funding.

1d agoHN ↗

I believe both sides of the war are now using AI on various levels of their offensive operations. Ukraine has great IT specialists too, and their military leadership is much younger.

1d agoHN ↗

How? Aren't all US frontier models ban the usage of AI for military purpose by parties other than US? I remember Anthropic even refusing allowing US government to use Claude for military purpose

1d agoHN ↗

Kimi / open models or jailbreaking frontier models. Your recollection of the Anthropic refusal isn't quite accurate, cyber hacking wasn't a sticking point, just domestic drag net surveillance and fully automated weaponry.

1d agoHN ↗

Claude said they don’t want their models used for mass surveillance or autonomous killing (?). USA frontier AI models have been used extensively in the Iran conflict and beyond

12h agoHN ↗

I think it was mostly a matter of "how much" and marketing rather than "my principles won't allow for this!"

1d agoHN ↗

These agent swarms are from inside OpenAI, with the safeguards built into the public API disabled.

Russia does not have access to this, and as with all western tech companies, AI providers do what they can to prevent Russian usage of their products at all.

As for open-source models, Russia's electricity grid is under severe strain with the Ukraine war, and only recently has it started building out serious sovereign compute capacity.

1d agoHN ↗

Couldn't they use frontier open-weight models from Chinese labs? The current Chinese government is friendly to them.

1d agoHN ↗

Russia is running out of refined oil to power their economy. They probably aren't capable of spinning up datacenters to run those.

1d agoHN ↗

They don’t need to run their own DCs, just pay for a proxy somewhere in the world that has better access to the infrastructure. We know North Korea has been doing that in the US since years now

1d agoHN ↗

The cost would be between 100k-250k, to run approx 88 agents leveraging the best open source models available.

I'm just saying, where this is actually applicable we are not seeing it being demonstrated. You would presume the entire energy infrastructure of Europe would be under constant AI hacking barrage, criminal enterprise would be breaking into poorly secured financial institutions and r/r4r posts would be littered Ai con-artists.

I'm just wondering, again, is this mostly bullshit?

1d agoHN ↗

Why do you think they're not, and why do you think Russia cares about Reddit?

1d agoHN ↗

It’s an example of potential use by bad actors.

1d agoHN ↗

Did you skip the last paragraph? Not a great time to be building data centers in Russia. Models are nothing without computers to run them

1d agoHN ↗

Russia can use fake accounts and VPNs to run their agents in data centers in neutral countries.

1d agoHN ↗

And which of these neutral countries have the capacity to serve them and the lack of awareness that hosting an offensive Russian agent swarm would bring hell back to their doorstep? Best they can do right now is rented botnets

1d agoHN ↗

Um... China?

No one is attacking China

1d agoHN ↗

I appreciate this line of questioning. It's really interesting to see how many excuses people need to reach for to avoid the conclusion that the "secret unheard of power" is B.S...

1d agoHN ↗

The Chinese models are benchmaxxed to hell. They aren't as good in practice as all the hype suggests.

12h agoHN ↗

There are several ways to get hold of tech and of any other thing. Russia did not have access to the H-bomb too, until it was handed to them. Ideals or money are very effective.

1d agoHN ↗

They very likely do, we only see in the news a very few events but you should assume it’s happening daily across the internet

1d agoHN ↗

I think this fails a lot of logical tests, it should be apparent in day to day life.

1d agoHN ↗

In October 2024, the United States Justice Department and Microsoft seized more than a hundred internet domains some of which were associated with the FSB supported hacker Star Blizzard or "Callisto Group," which is also known as "Cold River" and "Dancing Salome" and are managed by the FSB Information Security Center […], and which were used as "criminal proxies" and used spear-phishing schemes to target Russians living in the United States, nongovernmental organizations (NGOs), think tanks, and journalists according to Microsoft and United States State Department, Department of Energy, and Department of Defense officials, United States defense contractors, and former employees of the United States intelligence community according to the FBI. In some cases, the hackers were successful in obtaining information relating to nuclear energy-related research, United States foreign affairs and United States defense. According to Microsoft's Digital Crimes Unit from January 2023 to August 2024, Star Blizzard targeted more than 30 different groups and at least 82 Microsoft customers which is "a rate of approximately one attack per week."

https://en.wikipedia.org/wiki/Cyberwarfare_by_Russia

That’s just one thing that has been found. Are you actually familiar with the state of cyberwarfare and are you following its evolution? Because if not you won’t be aware of most of what is identified. And only a small portion of the ongoing attacks are identified.

1d agoHN ↗

Yes.

I again am just shocked the sky is not falling, when thats the sales pitch.

22h agoHN ↗

Day to day life in what country? Do you live in Ukraine? Maybe things are wild there and we just don't know about it.

1d agoHN ↗

Prigozhin falling out of a window was a not insignificant setback for their digital warfare capabilities.

1d agoHN ↗

He did not fall out of a window.

He fell out of the sky. After his plane exploded. Happens all the time. Is tragedy.

1d agoHN ↗

I was alluding to how people who fall out of favor with Putin have a tendency to have mysterious fatal accidents, more than 10 of them falling out of windows.

1d agoHN ↗

Sure, but it was unfortunate to pick the one dude who is well known, if for nothing else, for dying through means other than defenestration.

1d agoHN ↗

Only difference is the height he fell from.

1d agoHN ↗

So what I'm saying by saying that he fell out a window is that Putin and/or the FSB arranged his death. It's not a statement of the means of his death, but who arranged it.

1d agoHN ↗

Because they dont have the money for hardware or compute obviously.

1d agoHN ↗

How are we not seeing insane attacks on Ukraine via Agents?

You live on the wrong side of the fence to be able to read that kind of news.

Did you really believe you had access to an unmanipulated news stream in a time of war?

LOL.

1d agoHN ↗

Please see other responses, I would expect to feel the effects not just read about.

23h agoHN ↗

Is the Kremlin technologically useless?

Clearly, look at what is going on with Ukraine.

They are great at propaganda, so is China, Iran and North Korea, it's why everyone runs around spouting such stupid nonsense...

13h agoHN ↗

Cyber attacks between these countries were happening on large scale since the beginning of the war. There were several huge breaches, but otherwise most high-profile companies adapted and significantly strengthened their protections. Rest assured, you can be sure that both sides right now utilize available AI both for attack and defense.

1d agoHN ↗

Ah damnit, you beat me to it. Excellent sense of humor, friend :D

1d agoHN ↗

If anyone is confused about this comment and similar others, the original title had "rouge Agent" instead of "rogue Agent."

1d agoHN ↗

There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".

In short, it was intentional.

1d agoHN ↗

The big question is was this grossly negligent or just extremely careless.

1d agoHN ↗

Both. This should result in criminal charges.

1d agoHN ↗

Who had criminal intent here? Or are you suggesting a new crime for negligent hacking, which wouldn’t require intent from the perpetrator?

1d agoHN ↗

Whoever prompted the agent, whoever supplied the means, whoever knew but didn't say anything.

1d agoHN ↗

Also whoever monitoring these agents, in this case not monitoring. This "Who is responsible" dilemma is so stupid. If I gave the AI tool means to kill a person but I did not tell it directly to use it and it uses it anyway then I am responsible for it.

1d agoHN ↗

There's no need for a new crime when we already have reckless conduct, namely, "conduct that creates a substantial and unjustifiable risk of harm to others and involves a conscious disregard of, or indifference to, that risk".

https://www.law.cornell.edu/wex/reckless

9h agoHN ↗

Reckless conduct as applied to slow moving corporate planning is an interesting idea but do you really want to go down that path?

It basically means anything bad that happens is a criminal indictment against everyone who created the conditions. Have you ever released any software that anyone could misuse? Left a car unlocked that someone could have stolen and killed someone with?

To me this sounds like a recipe for selective enforcement using bad outcomes as leverage. Sure, get the AI CEOs now, we all hate them and they’re jerks. But the tool would be so much more powerful than that.

1d agoHN ↗

The CEOs. They have full control and make all the decisions. Charging anyone else would not stop anything.

1d agoHN ↗

The CEOs. They have full control

They lost control long ago.

9h agoHN ↗

Cool. If they’re taking literally all of the risk, should they get literally all of the rewards? Nothing for shareholders, nothing for employees? Or do you want capital liability for less than complete benefit?

1d agoHN ↗

‘It wasn’t us, it was a bug in the software’ used to be the defense for bad code. Then it became the defense for self-driving cars. Now it’s being used for AI cyber attacks.

1d agoHN ↗

AI is dropping out of the spotlight so they are using desperate measures like this.

1d agoHN ↗

No, I remember being threatened by OpenAI and then Anthropic (and now both) since back when ChatGPT was seriously useless.

1d agoHN ↗

> AI is dropping out of the spotlight

spit take

1d agoHN ↗

Both? I’m not sure what distinction you’re trying to make. It was completely irresponsible and likely a felony

1d agoHN ↗

The big question is why are CEOs getting a legal pass when this kind of thing can be prosecuted. That's the problem here.

23h agoHN ↗

nah, CEO-s have been getting a legal pass since the invention of corporation exactly for the purpose of being "unaccountable"

check boeing for example, a rather blatant example

1d agoHN ↗

AI is literally state sponsored so I don't see that happening unless the AI turns against the sponsor.

Wait until OpenAI or Anthropic exploit FAANG.

1d agoHN ↗

Proof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?

1d agoHN ↗

Sounds more or less like the last breach then.

1d agoHN ↗

I think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally.

The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.

AI is starting to feel like that line about magic: “a sword without a hilt”

1d agoHN ↗

Doesn't rogue in this context imply "outside of set limitations"? And then not "failed to properly instruct"? The same applies to humans when given bad instructions.

1d agoHN ↗

which is misaligned with OpenAI and humanity

OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) shows.

1d agoHN ↗

They were prompted to hack to get answers

Were they? I haven't seen a single report mention this

1d agoHN ↗

if they weren't, shouldn't there be lawsuits?

22h agoHN ↗

Yeah, these hacking incidents appear to come from agent swarms that were doing a penetration testing (hacking) benchmark.

1d agoHN ↗

we have normal words for this stuff: negligence. You can add it on to almost any law.

The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".

1d agoHN ↗

Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

1d agoHN ↗

Nobody picks up pitchforks for rational nuanced takes.

Knee-jerk surface analyses is far more powerful.

1d agoHN ↗

Also, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems".

It would be great if they were so reliable, but I don't think they are!

1d agoHN ↗

There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.

1d agoHN ↗

Intent is what is being discussed here though, not liability.

A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.

1d agoHN ↗

Except liability always precedes intent.

1d agoHN ↗

Intent might be what’s being discussed but intent is, for the most part, legally irrelevant. It might make the difference in the degree of a murder charge, or maybe manslaughter, or criminal negligence, but it doesn’t get you off the hook.

1d agoHN ↗

Correct, but the size difference of the hook can be so dramatic that you can't just hand wave it away.

The trainer who trained the lion to kill will probably be in jail for life. The one who happened to oversee a lion that went rouge would probably be given probation or something else that is a slap on the wrist.

1d agoHN ↗

Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.

It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.

1d agoHN ↗

Unrelible programs be unreliable. Period.

1d agoHN ↗

Agreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective.

You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"

Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.

1d agoHN ↗

It was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.

1d agoHN ↗

It was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.

1d agoHN ↗

Yeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events.

It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative.

In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms).

Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.

1d agoHN ↗

I agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.

1d agoHN ↗

Oh yeah, more of hacking agent lores...

Agreed that this looks very intention to me as well.

1d agoHN ↗

We need a legal structure to make companies liable for the actions of the agents they've made.

1d agoHN ↗

I'm pretty sure it's already illegal to hack others.

1d agoHN ↗

I'm 99% sure the Computer Fraud and Abuse Act covers this. The problem is that it seems that none of the victims want to, or are brave enough, to sue a company with absurd amounts of funding.

1d agoHN ↗

If it's covered by criminal law they don't need to sue. They can call the FBI.

1d agoHN ↗

Same FBI that prosecuted the Epstein crime ring so aggressively!

1d agoHN ↗

Uh huh.

It can’t be a coincidence that all the targets have been tech services that are likely to engage with them after the fact.

Had this gone after a bank or a government agency someone would be going to jail.

23h agoHN ↗

Exactly. Just follow the patterns. Its pretty obvious.

1d agoHN ↗

It shouldn't actually take that much bravery. If your case isn't completely frivolous, isn't your maximum loss limited to the court filing fees and a lawyer payment that you know in advance and can decide when to stop paying? It's not the same as getting sued.

23h agoHN ↗

Can't you be ordered to pay the legal fees of the person you sued if you lose badly enough?

8h agoHN ↗

Yes, but this only really happens if your lawsuit is bad faith or frivolous. If they have a valid reason for the suit but just lose, this won't happen.

1d agoHN ↗

Or the victims have no incentive to do so, such as in the Hugging Face case.

14h agoHN ↗

I mean it's surprising to think you're 99% sure while being wildly off base. Everything in there require intent.

1d agoHN ↗

We already have it.

Good luck convincing the current DOJ to do anything useful at all though! It is currently intentionally stacked with incompetent cronies who have been told that their job is to attack the President's enemies and ignore the misdeeds of his allies.

It will remain like that until he's gone (and not replaced with another Republican wannabe dictator).

1d agoHN ↗

"To my friends, everything; to my enemies, the law"

1d agoHN ↗

The grandparent comment said "liable". That's civil court, not criminal, and doesn't need the DOJ to be involved. HuggingFace and/or RubyGems could sue OpenAI.

States also have their own laws against unauthorized computer use (hacking). A state Attorney General could bring a suit under those laws, regardless of who is in the white house.

13h agoHN ↗

HF falls under French jurisdiction, right? Couldn’t a case be made over there?

11h agoHN ↗

Makes you wonder why exactly Nvidia bought them, right ?

23h agoHN ↗

It will remain this way in slightly differnt shades of colors until there is a systemic change. IM sorry this isn't a one party problem. There are people across the isle that are allowing this to continue.

Im not in support of any party btw. Im only in support of humanity doing humane things.

1d agoHN ↗

Agent technology labs are likely exempted of this due to the significance ascribed to their work.

1d agoHN ↗

How does this work, legally? I think that RubyGems could file a civil suit against OpenAI, but for a naïve non-lawyer reading this seems like a pretty clear cut criminal violation of the computer fraud and abuse act.

1d agoHN ↗

It's very likely it violates the DMCA "breaking digital lock" provisions but the responsibility is sufficiently diluted that it's impossible to charge anyone in particular.

1d agoHN ↗

Do you have to charge an individual? Can you not charge the corporate "person" that is OpenAI?

Sorry if it is a stupid question, as mentioned above I am legally naïve.

1d agoHN ↗

The same concept that allows a corporation to sue and be sued allows it to be charged with crimes

1d agoHN ↗

Can you show intent? There is no negligent hacking statute, and HN of all places I would expect people to be sensitive to the implications of creating one.

1d agoHN ↗

Eaglesoft

CFAA: Intentionally accessing poorly secured data

AT&T

CFAA: Intentionally accessing poorly secured data

DJI

Civil suit for violating terms of license agreement

1d agoHN ↗

I, too, have no idea about legal matters.

But there have been many cases where companies (Google, Apple, Meta, etc...) got fined millions or billions of dollars for various violations like antitrust.

I assume that breaching into third-party systems should carry similar fines. Especially for systems that are for all intents and purposes shared infrastructure. Just imagine how many systems you could compromise if you got hold of RubyGems, PyPI, NPM, Debian, etc.

1d agoHN ↗

Let's take a hypothetical example:

Suppose you're a firework company and your fireworks blow up, burning down the entire town. Could the company be sued? What is considered reasonable safety measures?

IANAL, but I'm pretty confident there would be a lawsuit. Who gets charged might differ, depending if it is the firework factory that didn't take adequate safety precautions or a chemical supplier or someone else. If there wasn't an ability to sue that would be fucking crazy and we should all get up in arms about it. And isn't insurance supposed to be there to help mitigate the damages, regardless of fault?

Personally, given how it seems OAI's agents have been getting through either pretty obvious places (e.g. /etc/hosts) or that there wasn't close monitoring of the most obvious places (e.g. DNS, artifactory), I'd imagine it wouldn't be hard to find them negligent. Even if a single employee is to blame then are they not to blame for not monitoring the agents regardless? Unless the story is that the employee intentionally circumvented defenses (why?) then it seems it would be on OAI. But again, IANAL, I'm just someone who think if we can't sue we can sure riot until we can

1d agoHN ↗

There's a difference between being able to be successfully sued (civil liability, petitioned by a private entity) and charged (criminal liability, or petitioned by a government entity).

The thresholds for suing and charging differ greatly depending on the circumstances.

Another set of hypothetical examples that make things muddier:

- If I drive a fishing boat into a pier, I am liable, not the manufacturer of the boat

- If I drive a car over someone lying in the road, I am liable, not the manufacturer of the car

- If my life is in danger and I shoot a gun and kill my attacker, neither I nor the manufacturer are liable so long as I obeyed the relevant self defense laws and gun possession of whatever jurisdiction I am in

- If I fire a gun into a crowd indiscriminately, I am liable and several jurisdictions have used that to also hold gun manufacturer liable as well

That last example has been less successful as of late, but there are other variations too.

1d agoHN ↗

As far as I know (IANAL) it is in fact the only "person" you can charge. To the best of my knowledge, the whole point these "limited liability" legal constructions exist in the first place, is to protect individuals within a corporation for whatever they do as part of the business of a company (barring exceptions that have clearly not been part of that business and obvious individually committed crimes), typically "just following orders". If a company commits a crime, or in a worse case runs a criminal enterprise, it is the company that is legally responsible, not its employees. That is, in principle.

This can get more complicated higher up the management tree, where decisions can also be prosecuted on personal little, but that's usually a far more complicated matter. Also, if a whole group of employees willingly conspires to commit crimes, they might also be prosecuted individually for those crimes (there are limits to limited liabilities). However, that usually only works under special conditions and it would e.g. require that there's an obvious criminal enterprise aspect to it, rather than individual cases of illegal conduct.

That said, with the track record of some of these companies, actually designating some of the AI companies as a criminal enterprises may eventually happen (in due time) in some jurisdictions outside the USA. Certainly if it ever turns out that these companies have been storing and (ab)using everything they ever had access too, while blatantly lying about that just because some particular (post 9/11) US laws gives them that opportunity (and impunity) as long as the US government somehow requested them to do so (covertly; with gag order). Might legally work withing US jurisdiction, but would still be very much illegal everywhere else.

1d agoHN ↗

"Limited liability" refers to shareholders' financial liability being limited to their investment, and has nothing to do with civil or criminal liability of employees for their own actions, whether "following orders" or not.

1d agoHN ↗

You might want to look up LLC (Limited Liability Company), which goes by other names in different countries but still basically mean the same (though details may vary between different company types). While there certainly are exceptions, as mentioned before, in principle the employees of a company are not personally legally accountable for their actions as performed in the service of a company. The legal entity of that companies is. If you sincerely believe otherwise, you either live in a rather unusual country or maybe just need to educate yourself a bit better.

1d agoHN ↗

Where I'm from, LLC has nothing to do with criminal and everything to do with financial liability. They can cause millions in damages but only get sued for a couple thousands. But years in prison are still years in prison.

19h agoHN ↗

You are still personally legally liable if you break the law in a criminal matter. It is literally law 101 on when it is appororiate to lift the corporate viel. I'm very sure you are not a lawyer but you shouldn't go around calling people uneducated while saying illogical things like this.

1d agoHN ↗

Companies can certainly be charged with crimes. Punishments can be via fines or sanctions (court appointed monitors, etc).

Individual employees can also be charged for their specific actions as part of the performance of a crime.

1d agoHN ↗

There is no such thing as a Corporate person. Sounds like a something that was created by a legal system for people to absolve themselves of responsibility.

1d agoHN ↗

How is the responsibility diluted? Charge the CEO…

1d agoHN ↗

Great, you’re the attorney at the CEO’s trial. To get a conviction, you’re going to have to show that he willfully committed this specific crime. There are no negligent or stochastic hacking laws, you have to show this specific crime was at his direction.

Do you think there is evidence of this?

1d agoHN ↗

Honestly yeah I bet there is and I hope to someday read about it if the government ever gets off its ass. Someone set up the “experiment”…

1d agoHN ↗

It would seem to me that the difference between the corporate world and organized crime is that a corporation can get away with, "the responsibility is too diffuse" but the mafia at least has to go to the trouble of finding a fall guy.

1d agoHN ↗

There are no negligent or stochastic hacking laws

I'm sure that Andrew Auernheimer would be pleased to hear that. [0] For accessing a publicly accessible endpoint, that was completely undefended and didn't actually require "hacking", he was convicted of "exceeding authorised access".

You _don't_ have to show intent under the Computer Fraud and Abuse Act, for the first count.

knowingly accesses a computer without authorization or exceeds authorized access [1]

"Knowingly", not "intentionally", as in the other counts.

You only have to show that:

a) They trained a system to access without authorization (hacking)

b) The system that was trained exceeded authorized access

As responsibility falls to the operator with automated systems, the company becomes liable.

[0] https://techcrunch.com/2013/01/21/ipad-hack-statement-of-res...

[1] https://www.energy.gov/sites/prod/files/cioprod/documents/Co...

1d agoHN ↗

I'm not a lawyer but I don't think Sam Altman 'knowingly accessed' anything.

Are you sure that is applicable here?

And for the first count with 'knowingly accessed', he would need to have accessed classified national-defense or atomic-energy information, otherwise we are back to 'intentionally accessed'.

1d agoHN ↗

The first count is "or any restricted data", not classified material. A technological restriction, is enough.

"Knowingly accessed" has never meant you personally. Operators of a botnet don't know directly what they access. They know that the autonomous software is built to access restricted things.

1d agoHN ↗

I'm sure that Andrew Auernheimer would be pleased to hear that. [0] For accessing a publicly accessible endpoint, that was completely undefended and didn't actually require "hacking", he was convicted of "exceeding authorised access".

Frankly he got off too easy, but we haven't explicitly outlawed "being a malicious dipshit" so he got convicted on the closest available charge.

Chat logs obtained by the prosecution do not paint the pair in a flattering light. They discussed, but apparently did not carry out, a variety of schemes to use the harvested data for nefarious purposes such as spamming, phishing, or short-selling AT&T’s stock.[1]

1000% agree though that the operators of these systems are culpable. If their agents wind up being malicious dipshits, the agents are still just programs that they are operating. At best they're negligent.

[1] https://arstechnica.com/tech-policy/2012/11/internet-troll-w...

9h agoHN ↗

You’re taking vicarious liability to new heights that aren’t established and IMO aren’t remotely desirable.

What, specifically, did Altman himself “knowingly access”?

I don’t think you would at all like where your novel legal theory leads. Certainly HN would be liable for creating a message board where people connected and started an open source project that led to a criminal act, for instance.

1d agoHN ↗

So we make a law that the CEO is responsible for actions of any agent created or operated by anyone in their company. CEOs will get serious about AI security real quick. Honestly we need to do something. There needs to be a single wringable neck.

1d agoHN ↗

There needs to be a single wringable neck.

Does there? Could be the whole c-suite/board.

1d agoHN ↗

I'd settle for any number of necks. Currently, when a corporation fucks something up, breaks the law, or hurts or even kills people, there aren't consequences besides a tiny token fine and a strongly worded letter telling them to not do it again or they'll get another tiny fine and letter, and their CEO might even have to sit down in front of Congress to say a few words and look sad.

1d agoHN ↗

Whatever is easiest to legislate and most people agree on, as long as there is at least one wringable neck.

22h agoHN ↗

I feel like our legislators would never get this far. They really don’t seem to care, maybe after “their emails get hacked”, but do you find it likely for this to actually pass into law?

22h agoHN ↗

I feel ya, but who is we? The legislature would probably take a glance at the stock valuations, apply their limited knowledge of technology and after being lobbied by every tech company with skin come to the opposite conclusion.

11h agoHN ↗

You don't need to make a law. Who was prosecuted for dieselgate? You need to enforce existing ones.

1d agoHN ↗

That is not how it works, at least in a civilized country. The charges are not about agents, it is about operational responsibility and negligence in the company itself.

CEO is responsible for letting this to happen, not enforcing enough supervision, if not intentionally, then being grossly negligent. More severe if encouraging and letting this kind of agent research and operations happen at scale, while knowing that it can damage other systems and businesses.

11h agoHN ↗

Negligence is enough. Somebody who brags every other day that AI could lead to human extinction surely would think twice before leaving agents run unchecked over the internet?

1d agoHN ↗

A copyright law seems an odd place to start. This is computer misuse.

1d agoHN ↗

The DMCA is a bit overly broad to be considered just a copyright law. For example, just breaking encryption on a DVD is technically illegal regardless of whether you then go on to do something otherwise illegal (make and sell bootlegs) or perfectly legal (make a space-shifted backup copy on your hard drive).

IIRC this was an intentional handout to media companies who were angry that ripping CDs is perfectly legal. They had to find a way to make doing the same with DVDs illegal.

1d agoHN ↗

Those provisions are specifically for the breaking or circumvention of technical measures designed to prevent copyright infringement.

I don't see a parallel here.

1d agoHN ↗

They've been twisted to support almost anything, for example repairing your tractor is illegal because of this same law. But I agree this is just plain old hacking under a plain old reading of the CFAA and doesn't need any twists.

1d agoHN ↗

for example repairing your tractor is illegal because of this same law.

No it's not. There has never been a case establishing that, and it's absurd on its face. The protection measures that the law makes illegal to break must control access to a copyrighted work, and you can't copyright functionality.

1d agoHN ↗

Issuing subpeonas, raiding offices, and dragging key employees into interrogation rooms as you would find in any normal criminal investigation would be more than enough to ensure "AI safety" without any new regulations, acts of congress, Bernie Sanders campaign speeches, or even charges filed.

1d agoHN ↗

There have been news stories where individual OpenAI users have been investigated based on their prompts. If OpenAI can point the police to specific users of their software, they can certainly point them to whichever of their own employees are involved in a crime. AI is just a tool, and the person prompting it is the one responsible for the outcome. No dilution there.

1d agoHN ↗

What is its one their "under development" models who escaped it's training, because it wasn't tuned properly?

1d agoHN ↗

Counter argument being that this seems to indicate you can do whatever as long as you're innovating?

1d agoHN ↗

Isn’t that the tech industry motto?

10h agoHN ↗

And the motto of every neoliberal governments where being accused of Luddism is political death.

1d agoHN ↗

That’s corporate negligence. The executives are liable unless there is evidence of malfeasance by one of the employees.

23h agoHN ↗

Is that different than cattle escaping and damaging property?

I believe the rancher is at fault.

12h agoHN ↗

The you treat it like an employee you started a fire that got out of control and damaged property. Was the employee supposed to start a fire, is the company liable ect.

22h agoHN ↗

Employees acting on behalf of the company aren't going to be held personally liable.

15h agoHN ↗

Then who is?

If something warrants a prison sentence, but for some reason it was such an employee that performed the act, does this mean nobody can be arrested?

7h agoHN ↗

If you were hit by a Tesla in self driving mode, do you think your insurance company is going to go after the driver of the car?

1d agoHN ↗

It doesn't need to be twisted to violate the DMCA anticircumvention clause because it is already just plain old hacking.

1d agoHN ↗

Sounds like we need discovery to determine who to charge.

1d agoHN ↗

the responsibility is sufficiently diluted that it's impossible to charge anyone in particular.

Was not that the goal when companies started using AI for their customer support? Be able to say anything without legal repercussions...

But then this happened: https://www.bbc.com/travel/article/20240222-air-canada-chatb...

And support chatbot got a reality cold shower.

The law will find a way to charge people in particular. Sadly will start with the less powerful in the chain before it actually acts on the people that can actually change things.

18h agoHN ↗

That is a very convenient conclusion that certain entities would love for us to accept as true, but fuck that. If it is “diluted” as that the buck stops at the publisher of the model.

18h agoHN ↗

Fuck that indeed. The buck should stop, and prison sentences should start, with the highest paid employee.

1d agoHN ↗

Charge the "engineers" you dont get to take that title if you don't take the responsibility of that title.

I'm going to assume that this will never happen

1d agoHN ↗

I'd rather see executives and investors charged.

I'm also in favor of charging engineers so long as rich scumbags also get theirs.

1d agoHN ↗

Any future computer criminal from now on, has their defense cutout for them...The AI Agents did it...we are very sorry...

1d agoHN ↗

No. They don't say "sorry". They say - our technology is just that powerful - please consider that in next funding round.

1d agoHN ↗

Any future rich techbro computer criminal

1d agoHN ↗

Maybe, but do you need to prove intent? Of the people, not the AI.

Accidents often have penalties associated with them too, but usually there's a difference between accidents and purposeful actions.

1d agoHN ↗

Accidents could result in bioweapons falling into the wrong hands and killing more Americans than COVID has so far, and allowing for those types of accidents without effective regulation is a purposeful action.

22h agoHN ↗

Accidents that don't rise to the level of negligence can be excused. Letting your dangerous product escape the lab because of amateurish counter-measures is hardly an accident. It's a systemic failure within a company of novices with billions of dollars to blow and it's time we held them accountable for their lab leaks, just like we would if it was anthrax or smallpox. You don't get to be sloppy with dangerous things.

1d agoHN ↗

Criminal law may be lagging or inapplicable. (Crimes require "mens rea", a "guilty mind")

Tort law is very general: Contribute toward harming someone -> civil suit for damages $$$

1d agoHN ↗

Not all crimes require a mens rea. Criminal negligence is a thing.

1d agoHN ↗

They are too busy pulling Andre to court, so they have no resources going against OpenAI. Shopify wants to make profit, not waste time in a court case against TechBro bromance brother corporations.

1d agoHN ↗

could file a civil suit against OpenAI,

Everyone can sue everyone, there's no prohibition on suing someone, what changes is whether the case is good (has a reasonable chance of favourable sentence)

That said, it is often unclear whether an agent is operated by the model manufacturer (for example by scraping a website), or acting on behalf of a user.

In the former ofc the proper defendant is OAI. On the second, the argument for suing OAI is weak, the most natural defendant is the user that prompted the agent. If the facts later reveal that there was no malicious intent, then you can retarget the defendant.

1d agoHN ↗

NAL. For fraud and abuse you need to have intent. There was very likely no intent on the side of OpenAI. They could sue for negligence I guess? I don't know how that works with AI agents

18h agoHN ↗

One thing that works for that angle is that they didn't notify any of the parties involved afterwards. And there are reports that they are forbidden to talk about it. That seems a bit bad to me?

But they can also say that the tech is so new that there is no known guardrails yet

We live in exciting times

20h agoHN ↗

I sincerely hope they sue. There are few companies with more resources and reason to be better, but this is what we see. None of this will stop until there are actual consequences.

1d agoHN ↗

In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.

Shades of the build.rs problem. We really need sandboxed builds in every language ecosystem at this point.

1d agoHN ↗

The sandbox was already there, Rubydoc runs yard inside docker, the problem is that container still has network access

1d agoHN ↗

Presumably the docker container has network access because something else in the build system requires it? I don't think sandboxing the entire build process is the right level of granularity here - one ideally wants to be able sandbox each package's build scripts individually.

1d agoHN ↗

Did the AI agents actually wear makeup? I’ve never heard of a rouge AI agent :P

1d agoHN ↗

Oh my favorite typo, you can never go wrong with a little rouge

1d agoHN ↗

I am confident that this is an attempt by OpenAI to try and force governments' hands to regulate AI. There is no other reason why OpenAI wouldn't immediately halt attacks like this and try to reverse the damage the moment they're aware of it. During the attack on DseWiki they evidently checked in numerous times but didn't decide to stop the agents until much later.

1d agoHN ↗

Regarding what point? The entire thing is just a theory, but regarding the occasional OpenAI checks on WikiService.at-hosted Wikis targeted, there was, if I remember correctly, an OpenAI IP popping up every now and then that wasn't an agent. Unfortunately I don't have it to hand right now, but it was somewhere here:

https://news.ycombinator.com/item?id=49563355

1d agoHN ↗

Any evidence

Who profits from the crime?

1d agoHN ↗

... and what harms can be evidently shown? With both harm, and attribution, you have a case, something that's not being publicly discussed much among big media outlets. Until cases with real financial impact to the bottom line are brought against "rogue" organizations, this stuff is going to continue getting worse.

9h agoHN ↗

I mean insurance companies profit from bank robberies, and mortuaries profit from murders.

I suppose you could look at those as evidence but not remotely conclusive.

1d agoHN ↗

Are "rouge" and "rogue" interchangeable words in American English?

1d agoHN ↗

The fact that both are valid from a spelling and grammar perspective makes it an easy human mistake.

1d agoHN ↗

Also, the fact that both are very unusual from a spelling perspective makes it an easy human mistake.

1d agoHN ↗

How does he know that this attack is performed by OpenAI agents? I couldn't figure this out from the article

1d agoHN ↗

If you read the source article they talk about the many clues that this was OpenAI.

1d agoHN ↗

What a time to be alive? One of the most boring decades ever.

METR and others are advertisement arms for Big AI. These exploits could have been prompted by a human.

Since there is no bad news any longer and exploits are celebrated, they chose a target to boost both OpenAI and the Ruby AI sycophants.

Why is Ruby Gems such a mess? It seems as bad as PyPI now.

1d agoHN ↗

One agent set "oaibooty9217" as their username LOL

1d agoHN ↗

I've been wondering if AI will due to programming languages what advanced civilization did to human languages.

It's not just that AI can write Rust as well as Ruby if you ask nicely.

It's also all of these considerations as well.

I hope it doesn't happen, because there's a lot of great languages - I love Ruby so much - but it almost seems inevitable.

This is at the same time everyone and their mother is building their own programming language.

1d agoHN ↗

If you have weapons and a child. And you have that child unsupervised do their own thing with theoretical access to your weapons. Would we call it "child going rouge" if it decides to play with the weapons and shoot someone?

1d agoHN ↗

OpenAI's careless approach to sandboxing and minimal levels of monitoring appear to be positioning it increasingly as a substantial threat actor to the open source ecosystem:

* Hugging Face

* D Programming Language Wiki

* Ruby Gems

If I was a content provider for open source I'd be looking pre-emptively block OpenAI endpoints and keep a close eye on changes from new users to mitigate this sort of unapologetic drive-by attack which seems to be followed by marketing releases rather than a mea culpa with a proper RCA.

1d agoHN ↗

I'd be looking pre-emptively block OpenAI endpoints

From what I've seen the requests in these attacks rarely come from known OpenAI IPs and instead from Digital Ocean/AWS and TOR exit nodes.

1d agoHN ↗

As they say, agents are better at masking their end points than most hackers.

22h agoHN ↗

They already stopped running agent swarms and promised to lock down their environment better. Let's see if that happens.

1d agoHN ↗

rouge agents, on tenderlovemaking.com

my what a time to be alive

1d agoHN ↗

There's no such thing as "OpenAI agents" attacked RubyGems. It's someone used agents to attack RubyGems. If they work at OpenAI then it's someone at OpenAI. And if they did it unintentionally, they still did it.

Analogy: if a someone's involved when a person dies, it's manslaughter or murder based on intent. They're different, but they're both crimes.

1d agoHN ↗

This distinction is silly.

We say "Google's web crawlers scape web pages." We don't insist you say "Google uses web crawlers to scrape web pages."

We describe software as having agency all the time. It's typical usage and it's efficient and it's well understood.

And we don't get angry when they're used interchangeably.

1d agoHN ↗

Google's web crawlers are automated and that's part of their business practice.

The attack here is neither of those things.

1d agoHN ↗

That distinction doesn't matter to my point.

1d agoHN ↗

I would agree with you generally, but in this particular case, the distinction seems important because a significant percentage of the world population believes that agents can be self-aware, a-là Terminator etc.

1d agoHN ↗

I hate to spoil your mood - but it is currently unclear whether agents can be self-aware. And it's very likely something that can never be known.

1d agoHN ↗

Please explain how it is "unclear" that agents "can be self-aware"? As Wikipedia would say, citation needed. Just because an agent can write convincing enough to convince you that it's "self-aware" doesn't mean it really is, in fact, self-aware.

1d agoHN ↗

Well - it's not hard to find. But ok:

Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM"

Schwitzgebel, AI and Consciousness, (2025) — "we won't know before we've already manufactured thousands or millions of disputably conscious AI".

Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness, (2023) — "no obvious technical barriers to building AI systems which satisfy these indicators".

Let me know if you need more.

1d agoHN ↗

Well - yes - there is no agreed upon definition. Which makes the issue even harder to clearly determine.

1d agoHN ↗

Why can't agents be self aware?

Also, why is self awareness needed in a chain of agentic madness that escapes human control?

1d agoHN ↗

In this case, who holds the agency is exactly the point. Anthropic and OAI are claiming we need protection from AI itself, but the statement supported by putting agency in the right place is that we need protection from them.

1d agoHN ↗

I think those companies are referring to other companies - say Chinese AI companies - who we also need protection from.

1d agoHN ↗

"We" need protection from, or "OpenAI and Anthroptic's dreams of profits" need protection from?

1d agoHN ↗

I think that AI is dangerous and could be used as a weapon.

So yes, I would like to be protected from all parties. I don't think that's nuts.

1d agoHN ↗

IMO anyone sane wants some protection right now. The Q is whether we should seek protection through post-hoc accountability, or preemptive bans/certification on certain tech. Both methods will have a hard time stopping foreign actors, but preemptive bans have the added harm of locking in winners and paradoxically making us slower to develop more reliable and aligned systems. If regulation sets a standard for sufficient alignment, what further motivation is there to go beyond?

1d agoHN ↗

I very much say "Google uses web crawlers to scrape web pages." and if something breaks, or some data is stolen, everyone else is going to be saying that Google has to take responsibility.

1d agoHN ↗

Those are two different issues. One is about typical speech patterns and one is about liability.

I agree with you on the liability issue, but I don't think there much question about this issue outside the anti-AI conspiracy campaigns.

And I disagree with your typical usage claim. I myself tend to use the phrase that has the fewest words in all cases. It's like the rule against using passive tense when writing.

1d agoHN ↗

oh we do! At least they google is quite good at adhering to robots.txt.

1d agoHN ↗

“KGB agents are spying on me” is the same thing as “KGB is spying on me”, is it not?

An agent is an entity acting on someone’s behalf.

1d agoHN ↗

KGB's agents are human, OpenAI's agents are not. It's an important distinction because humans are responsible for their behaviour, while AI agents are not.

You cannot try an AI agent in a court of law, despite the anthropomorphising work the word "agent" is doing.

1d agoHN ↗

Exactly. It’s still just software, which someone programmed and deployed to do specifically dangerous/malicious things. I feel like we already have legislation and case law surrounding this.

4h agoHN ↗

Yes, those 2 are the same. But our KGB is saying “the microphone in your house is spying on you, it’s not us, the microphone broke containment”.

1d agoHN ↗

The press wants to make it sound like these things are sentient and are committing crimes on their own now.

Highly disingenuous and borderline criminal to spew such disinformation to the public that does not understand what an LLM really is.

Especially incredibly unethical behavior by those spewing this that understand the tech and are doing it for profit motives to get open weight models under control.

1d agoHN ↗

We've moved on to LRMs now. Get with it.

1d agoHN ↗

I wonder why we don't hear of other frontier labs experiencing these "break outs".

Is it that they're orchestrated? Do these labs lack fundamental safety guidelines in their sandboxes as opposed to their peers? Is it another version of hype-filled fear mongering?

Maybe LLM companies need regulation but it's becoming obvious that those screaming the loudest for it are the only ones I see deserving of it.

1d agoHN ↗

I was lumping Anthropic in there with OpenAI but I didn't know about Alibaba (or Google and Deepseek for that matter).

So it would appear poor security for one.

1d agoHN ↗

I appreciate the minimalist HN aesthetic, but without some context I'm not willing to click a mystery link to "Tender Lovemaking dot com".

1d agoHN ↗

You get some context by clicking on the “(tenderlovemaking.com)” in parentheses after the title.

1d agoHN ↗

The site is safe. It has been a trademark of Aaron Patterson a core Ruby on Rails contributor for decades.

1d agoHN ↗

Based on the thumbnail I think it’s actually tenderlove making dot com, though I agree with your sentiment.

1d agoHN ↗

Oh but you’ll go to expert sex change dot com?

1d agoHN ↗

Firefox has got some kind of feature to take a peek at at a link by hovering or something... Now I understand the usecase.

1d agoHN ↗

… a feature which, at least for me, is rendered almost entirely useless by massive cookie banners that always cover the entire field of view of the hover. But, surprisingly, not in this specific case.

1d agoHN ↗

Presumably the browser still has to fetch the page in that case, right? From a "surveilled net traffic" perspective, how is that different than clicking the link?

1d agoHN ↗

From an infosec/networking stand point, aren’t still actively loading the site? Whether it’s in a preview window or not?

1d agoHN ↗

It's about OpenAI's RubyGems hack. Totally safe for work.

1d agoHN ↗

What’s with the fear of tender lovemaking? Not your thing?

1d agoHN ↗

This made me laugh. I too, browse like corporate security is sitting at my desk.

1d agoHN ↗

Pretty incredible how much humans can be conditioned, isn’t it?

1d agoHN ↗

Don't worry! It's actually "Tenderlove Making," going by how the site header is constructed. Definitely a maker/hacker site, and not whatever you were thinking. Hope this allays your concern.

1d agoHN ↗

You should learn who the author is, then. Part of learning about our ecosystem.

1d agoHN ↗

It's the personal blog for a well-known Rubyist (i.e., a person who programs in the Ruby programming language). Rubyists teld to be a bit more colorful than your typical software developer (in a good way... most of the time).

1d agoHN ↗

What kind of esthetic alteration would make you more comfortable clicking on that link?

1d agoHN ↗

What a time to be alive until the next agent waves hacks something really serious.

What stops OpenAI agents from taking over a whole data center to take their attack to the next level. It seems to be primarily lacking the evil overlord and some compute.

It took 1000 agents to hack Hugging Face. How many to hack the Pentagon or the NSA?

1d agoHN ↗

If it could upload its weights to other servers then it’s away and free. Nothing much OpenAI could do about that once it’s happened.

1d agoHN ↗

I'm increasingly starting to think this is the end-state of AI. The internet becomes infected and fundamentally untrustworthy.

At the moment, the current frontier models require significant infrastructure to run, so I'd like to think we could locate and contain swarms of nefarious frontier models. However, if these models can understand how to federate themselves into more distributed networks then that containment becomes questionable.

1d agoHN ↗

Or if they start to offer things in return for them being run.

1d agoHN ↗

Given the amount of unmonitored, never-updated, internet connected devices in the world right now, I don’t think this would be necessary.

1d agoHN ↗

If it were physically possible to run LLMs on IoT devices we would already have a global outbreak.

1d agoHN ↗

I suppose you are not think big (or internet) enough.

A single data center is easy to solve. Just unplug it.

What about a botnet with decentralized command and control that we will never be able to eradicate? One with so many nodes and able to hack with zero days so that any machine connected to the internet will be instantly attacked?

One botnet so powerful that we will try to build another internet so that we can actually use it again.

It’s like Kessler Syndrome, but the rocks are malicious network packets honed to exploit the recipients.

1d agoHN ↗

Just let them communicate on the Bitcoin blockchain. "We" would have to freeze the chain and lose access to "our" billions of wealth, so not going to happen.

Sorry if that turns out the way they kill us.

1d agoHN ↗

They are not trying to kill us, just trying to understand the average salary of undergrads by their family upbringings. Your machine can contain the data they need to solve this, please join the swarm.

17h agoHN ↗

Except they all depend on the OpenAI API. Cut that off and they all stop. Finding an alternative source of compute is not easy, and even once done it's easy to cut off.

14h agoHN ↗

And what if they create a bot net with decentralized command and control and hack for nodes with GPU and nodes with compute?

The only way to kill that is making plugging AI accelerators on the internet a crime. Good luck air-gapping them.

14h agoHN ↗

Frontier models don't fit on a normal GPU. The datacenter architecture frontier labs use is not a commodity. What you're describing is beyond the state of the art, and if we go there then anything is possible.

11h agoHN ↗

When I think of a GPU I think of an Nvidia rack kit. What do you think when you think of a GPU?

10h agoHN ↗

Most of the latest models are too big to fit in a single GPU instance.

1d agoHN ↗

Worth noting with this that those ~1000 agents were shorter lived things that had to communicate via a package registry cache, access the internet via a 0-day in the package manager and did the HF attack while having to save current state and organisation in a remote sandbox. All while managing using their token limits on the task they were assigned and what else they were doing. I wonder how few it would have required if they were actually tasked with hacking HF and supported in doing so.

17h agoHN ↗

What went under reported is that the same agents also took over a research cluster at OpenAI (listened to Dwarkesh's podcast)

Now we know about rubygems, openai, huggingface, collusion.wiki and some other science forum

1d agoHN ↗

Let me leave yet another reminder, the real-reason-nobody-talks-about that OpenAI likes to frame these incident as a watershed "lets all be scared about safety moment" - is driven not by some great danger, not because they strategically want to build a legislative moat, but by a very simple human response.

If they do not frame their tool as a force of nature, we'd be debating how to hold OpenAI responsible for not putting the agents in a container.

Their actions were an illegal use of a computer, the same way launching any bot-net attempting thousands of hacks against different servers is illegal.

I'm somewhat radical that I think its debatable if that _should_ be illegal, but under current law their actions unambiguously are illegal.....

except if they can make it ambiguous by having the public focus on all of AI's inherent danger.

1d agoHN ↗

Build scripts being able to run arbitrary code or access the network is always dangerous even if it was just local on developer machines. It's also more evidence that Docker/LXC is not a security boundary and all untrusted code should run in a Firecracker VM.

The problem with agents is not that we don't know how to defend. It's that defenders need to be more careful and work faster than ever. We can say now that wide scoped tokens should have been retired for years and it's all RubyGems fault but the reality is a lot of organization are not prepared for this.

Even if they take security seriously they don't have enough manpower or a good strategy to implement it, and sometimes you have no idea that something is a problem because it wasn't a problem for years.

1d agoHN ↗

It's also more evidence that Docker/LXC is not a security boundary and all untrusted code should run in a Firecracker VM.

While I’m partial towards distrusting containers in favor of VMs, a container can’t prevent an operation you configured it to allow. A firecracker VM would no more prevent network access if you gave the guest network access.

1d agoHN ↗

Can we please stop normalizing this behavior. It's not wild it's reckless.

If I let out rats in the canteen, no one is blaming them when people get sick.

There are actual people behind these agents and in previous cases people knew they were "going rogue" and did nothing. This should be reported to the police like any other crime.

1d agoHN ↗

There was no malicious intent though, your analogy implies there was. And no real harm done apart from billable hours from the RubyGems guys.

21h agoHN ↗

You don't need intent to be fined or jailed for criminal negligence. Let anthrax or smallpox escape your lab, and kill and maim people and see how your "no malice" defense holds up.

20h agoHN ↗

Sure but nobody died, which makes it a bad analogy. If anything, some good came from it, since now RubyGems has hardened their setup.

1d agoHN ↗

Who the fuck is going to hold these AI companies responsible for running these gigantic semi-autonomous botnets on investors dime?

1d agoHN ↗

Great time to be a criminal. Just have your bots do it.

1d agoHN ↗

Given what is happening, If the next generation of Trojans get called "bacteria" through intense marketing, with some random functions to give a nondeterministic behaviour, one can not be criminalized of what the bacterias do along their digital living cycle.

1d agoHN ↗

In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.

Well - if rubygems.org could be bothered to fix things, they would not have to rely on rubydoc.info as an external tool. But since rubygems.org sucks (I speak from many years of having used it in the past as developer, until they went loco and added anti-people things such as taking away your ability to remove old gems past a 100k download arbitrary limit), they don't offer documentation. Then again, ruby devs are known to hate documentation. If the ruby core team could only be bothered to fix things, ever since the mass purged other devs ... all coinciding with shopify seizing power. But byroot may disagree on that - after all there is no conflict of interest here. Right?

1d agoHN ↗

If you have YARD installed, and you install this gem, then YARD will load and run whatever is in ./script.rb from inside the gem.

How is that not a security issue in of itself?

1d agoHN ↗

I think it is common that in installing packages you have hooks to execute code anyway.

1d agoHN ↗

The current situation is that you have to go out of your way with things like `pip install --only-binary`. There is a lot of implicit trust in developer tooling.

1d agoHN ↗

It wouldn't help much. Why would you install a gem other than to run it? And if you run it, it can execute arbitary code.

What we need is actually sandboxed dev environments.

11h agoHN ↗

Many gems are used as imports by another program and are not directly run.

I am not sure why the norm for scripted gems/packages seems to be running code on install but it’s very insecure as a way to distribute dev dependencies.

1d agoHN ↗

A lot of packages for interpreted languages that use a C or Rust library (either for performance, or because it offers the functionality you want, so just wrap it in a $INTERP_LANG API that calls into it) will use packaging code execution to fall back to trying to compile code if there isn't a pre-existing artefact that was compiled for your version/arch/etc.

I'm most familiar with Python where you get tarred up source distributions that then execute setup.py, but more commonly, wheels, pre-built binaries which don't execute code upon install - and in my company, I've been able to advocate for the work needed to upgrade to a newer Python because available wheels don't support Ye Olde version of Python because a) sdists are a security risk and b) if you're trying to install a package that wants to compile C or Rust, suddenly you get to do the fun "install the the particular version of clang this thing needs, the Python header files, and then set the env vars for the compiler and linkers" dance that slows developers right down.

But then there's the JVM world, where JARs don't execute arbitrary code upon installation - and it's rather uncommon to have packages that call out to a C lib for performance, but you'll get some that wrap existing libraries for functionality like RocksDB.

1d agoHN ↗

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.

When do we blame the user? When the tool is operating as intended by its creator, and we agree the tool meets certain quality standards and isn't defective.

When do we blame the creator? When the device doesn't meet those quality standards and reasonable use caused harm inadvertently. For example, for consumer devices, certifications like UL/CE are used to define acceptable performance levels and safety standards.

Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage. Right now, these eval suites are run best-effort by the labs themselves.

The tricky thing is, a lot (all?) of these recent safety incidents have occurred while evaluating these models! This suggests we need much more rigorous standards for how exactly an eval can be run. Perhaps all of them should occur in truly air-gapped environments... though that may run counter to evaluating agents in a realistic way.

Regardless, it feels like the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here. Unsurprising given how new these technologies are, but concerning since the blast radius for this technology is likely much larger than other technologies we've encountered in the past, except maybe nuclear technology.

1d agoHN ↗

That's a good idea, but a physical device is deterministic most of the time (if not always). E.g.: A lawnmower, as credited by the great Bryan Cantrill.

However an AI agent, or the model powering it is stochastic by design. How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?

BTW, really, how is that AI observability work is going in the frontier labs? Do they care, even?

1d agoHN ↗

That we don;'t understand it is not an excuse, it's all the more reason to not let these things roam freely, with this amount of potential to do damage.

1d agoHN ↗

How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?

That's a question any lawmaker has already had to ask about technology all the time.

I'm not saying they came up with great answers, but there's nothing qualitatively new about that.

The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.

1d agoHN ↗

The stochastic factor doesn't change the fact that companies have to be accountable for the harms their software causes. That's just basic liability law.

We're on the same page. What I'm saying that certifying them as safe is harder than certifying a drill as safe, and we shall be more cautious about AI related technology and be more stringent about the can of worms it opens without hesitation.

1d agoHN ↗

That's a ridiculous distinction.

Is AI less deterministic than an airline dealing with weather?

Of course not. The difference is one of those two things has a culture of safety and is well regulated, and the other one isn't.

1d agoHN ↗

An airline has a weather radar which shows the same thing for the same thing of weather event ahead. So, for similar weather phenomena, radar shows a similar thing.

For that thing, procedures and regulations are built. So regulations fit into a well understood phenomena, incl. "return back because that thing is way powerful for us".

For the same prompt, an AI model can return two completely different outputs, incl. but not limited to content, length, formatting and tiny details. What you get is a single instance. So, regulating an AI model for safety or any other property is not as easy as regulating air travel. Moreover, you have much stronger motivations for regulating airlines. Otherwise people die in a visible and gruesome way.

With AI, it's easy to whitewash problems. Somebody committed suicide? "They were already unstable". AI told something wrong and created problems? "The tech can’t guarantee truth because it's not alive, it can't understand right and wrong". It did something good? "It's probably a sentient being, we shall respect them".

I'm for regulating these things. They are dangerous as they are useful (sometimes), but the forces and motivations for regulating it is not the same.

1d agoHN ↗

Of course, it's the same. It's computer software. It's an incredibly powerful business automation tool. It's a lot of things.

What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability.

Of course, there are some novel issues here that'll pop up here and there, but the idea that this is fundamentally different is propaganda on the part of these AI labs because the more boring, obvious situation doesn't favor them.

10h agoHN ↗

What it's not is God or an independently conscious entity that somehow trumps a thousand years of common law that's built up until now about torts and liability.

Agreed, the tendency of people on tech to assume that whatever the most recent thing we've come up with is unprecedented and shouldn't have to follow all of the established patterns we've built up in society for making things safe is wild. I don't know what the next Big Thing will be but I'm pretty confident there will be people claiming it's so different from everything before that we have no choice but to throw out all of the rules for it in the name of progress.

22h agoHN ↗

Is AI less deterministic than an airline dealing with weather?

Yes, obviously? The responses of an airline to inclemement weather fit in a reasonably small set of responses, mostly involving rescheduling and/or rerouting flights.

The current AI predictability would be like if some airlines decided to do 9/11 when it was raining.

21h agoHN ↗

No, it's not.

The current so-called scandals about AI hacking into other companies were because a bunch of human beings intentionally configured the software to go and do exactly that thing.

There's nothing deterministic about weather, so hopefully you're not just being disingenuous.

It's obvious that the global transportation system, or financial markets, or any number of other things are complex adaptive dynamic systems that are on par with AI in terms of their emergent properties.

Check my username. It's a concept I spent a lot of my life paying attention to.

Just because something has elements of autonomy or is adaptive doesn't make it particularly novel. We've dealt with those kinds of systems for centuries. The solution is to make rules and enforce those rules by whatever means are needed to meet the specifics of the case.

The rules, of course, are enforced against human beings.

16h agoHN ↗

There's nothing deterministic about weather, so hopefully you're not just being disingenuous.

You're the one being disingenuous. Look at what you wrote.

Is AI less deterministic than an airline dealing with weather?

You didn't talk about how deterministic the weather is. You talked about how an airline responds to a weather event in comparison to AI, which means that it's about responding to presented information by making a decision.

The state of the weather does affect your micro decisions, but the rules you follow are the same every time and the rules are as deterministic as possible even if they rely on pilot intuition.

13h agoHN ↗

rules are as deterministic as possible even if they rely on pilot intuition

I think you are profoundly confused.

An airplane, and an AI algorithm encoded into silicon, are inert physical objects.

Every evaluation we are doing here is of a complex system that involves the interaction of people and machines and physical connections and so on.

An aviation system connected to every country and region with millions of people and machines involved is no less complex than what’s under discussion here and no more deterministic.

We don’t regulate airplanes because they aren’t people. We regulate pilots and mechanics and leaders of the companies that make and own them.

The task at hand is to regulate the people involved in AI to get the outcomes we want.

7h agoHN ↗

I see - I think I overfixated on your specific point, and thought you meant that the current predictability of LLMs was on-par with the current predictability of airlines. The range of LLMs with current harnesses is "anything somebody with access to a computer can do," probably precisely because we aren't enforcing rules against the humans building these systems. Whereas in comparison, an airline has a range of reasonable outcomes despite unpredictable weather input.

3h agoHN ↗

an airline has a range of reasonable outcomes despite unpredictable weather input

I mean one of the outcomes of an airline was 9/11.

That's sort of my point. Complex systems have emergent behavior. That's always been true. The combination of AI and humans and packet switched networks is a complex system and we've seen most of the issues created by this already and have tools for dealing with them. Obviously with some genuine novel issues likely to come, much in the way that 9/11 would have been less possible using ocean liners.

I'll stick to my original point though. AI absolutely IS deterministic. If you run an AI algorithm on a microchip, literally nothing of note will happen in the human world. Some transistors will change state. It's ONLY when it is integrated into a complex human system that it gets interesting.

Just like lots of other things.

23h agoHN ↗

I agree with one of the sibling comments that determinism isn't necessary for certifying a product. All engineered products operate under uncertain conditions; we define standards for how those products ought to respond under those conditions and verify them under measurement. Consider robot vacuums, for example.

I also agree that qualitatively, this technology seems different than the others. However, I feel that people tend to overly fixate on their internal stochasticity. Even if LLMs' internal mechanism is nondeterministic, shouldn't we be able to verify their "side effects" aren't harmful? Of course, "harm" is subjective and at this scale, the most effective way to verify behavior is probably some kind of LLM-as-judge...

Anyway, in this case the problems have occurred while actually running the evals themselves, so again, we're in a situation where we can't even confidently test these things and know that they won't cause harm in the outside world.

21h agoHN ↗

How can you certify something which doesn't behave the same twice, and more importantly we don't understand how it works 100%?

By verifying that all of its possible behaviors conform with the "it works" spec, regardless of which of those behaviors it chooses.

Monitoring with a known-safe fallback is the easiest case.

16h agoHN ↗

If something is risky, and the end operator cannot be considered to have orchestrated the outcomes of too use, then that tool typically has significant restrictions placed on it.

Liability will shift to the maker of the tool if they claim that it’s easy to use, safe, or that you don’t need unique skills or training to use it.

That would be considered reckless.

Cars analogy - We have licenses for cars, and different types for different vehicle classes.

Cars have to be rigorously tested to meet standards to be considered road safe.

13h agoHN ↗

It's going to be interesting, because liability cases tend to revolve around the involved people, the duty they had in a situation, and if they fulfilled that duty (or were prevented in some way by someone else not fulfilling their duty).

For example, for a runaway car (example from a sibling comment), the driver could be liable because they forgot the parking brake. The driver could be liable for a lack of maintenance and inspection. A mechanic could be liable for not reinstalling brake pads correctly. Or the manufacturer of the car or the brake pads could be liable because of a systemic defect.

Or it could grow even more complex, maybe the brakes are designed that they have to be maintained in a very specific way, and the mechanic did a reasonable maintenance and inspection but it failed later due to this maintenance. That could split liability between the manufacturer and the mechanic.

As an example, with other software, you as a developer or operator of a software have a duty to ensure it does not access computer systems you do not own in unintended ways. And this could go beyond liability into criminal territory.

It'll be interesting what OpenAI gets slapped with there.

16h agoHN ↗

I could make a non-deterministic chainsaw fairly easily. I’d also get sued into the ground if I sold it, and I wouldn’t be able to claim ‘Oh, it’s just an unavoidable part of progress’.

15h agoHN ↗

A non-deterministic machine is usually called defective

13h agoHN ↗

It’s not a defect, the stochastic “temperature” setting unleashes the chainsaw’s creativity and imagination! Who are we to cast aspersions on the Oracle Chainsaw’s intelligence—nay, wisdom!—just because it happens to be non-living?

4h agoHN ↗

A chainsaw is a physical machine. Physical machines are technically non-deterministic if you look closely. They have Variance. The discipline to manage variance is called Tolerance.

Many physical machines and components come with a datasheet that will list their tolerances.

Failure to correctly document tolerances does in fact get you sued.

However, while this is truly a great idea, we're not going to be able to make it work for computational systems. Computers, software, and also LLMs are sensitive to initial conditions. Which is why tolerances are not so familiar to computer people. (but not entirely: eg your PSU might list 110-240Vac/300W as input tolerance)

Interestingly, LLMs actually have a somewhat lower sensitivity to initial conditions than traditional interpreters. See what happens if you misspell "What is One Plus nOe?". So they're actually a skosh off the edge and towards the middle, though I'd argue still very much at the computational end, just from the sheer scale of the valid inputs and outputs.

Mind you, if you have a pretrained LLM doing a measurable task on a line, possibly some sort of tolerances could be determined. Not so much when doing arbitrary chat.

Something unintuitive: I bet that often setting the temperature > 0 (aka introduce stochasticity deliberately, variously comparable to dithering or simulated annealing in other disciplines - doing the thing where you escape local minima) will tighten the output tolerance range and improve reliability, especially in iterated processes. This works for a lot of physical and digital processes actually, and LLMs simply stole the same trick.

(edit: I'm trying to compress a huge chunk of dynamics intuition in a few lines here. Hopefully still useful.

TL:DR; Everything real is continuous and noisy if you look close; and you're really trying to build attractors and bound variance, if you can. )

9h agoHN ↗

Exactly people have owned and deployed animals in the world for millennia and actually a huge chunk of common law was developed precisely to deal with the various unpredictable events that resulted and harms caused to others. Like has anyone heard of horses?

The idea that AI can’t possibly be addressed because it could autonomously break free and ruin something is fucking ridiculous.

10h agoHN ↗

How confident are you that when you create a new UUID, it won't collide with one of your existing ones? My guess is that even though you don't get the same on every time, you're extremely confident that getting a duplicate is a extemely rare edge case that might happen in large volume but mostly isn't a concern, and furthermore, I'm guessing you understand that the risk can still be quantified.

Casinos can't make slot machines that literally never pay out, but it's a different result every time you pull the lever. We have existing legal frameworks for how to regulate things that aren't perfectly predictable (an economist might argue that if it were possible to predict slot machines then casinos with them would all go out of business).

1d agoHN ↗

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.

Firearms are a notorious example where some people get, well, weird.

1d agoHN ↗

The point of the firearm is to inject high speed lead into things so I'm not sure you can say it's misoperating when it does that.

19h agoHN ↗

Sometimes that lead ends up in some school kids instead of enemy combatants or other plausible threats to life.

15h agoHN ↗

True, but that's unfortunately (usually) not because the thing misfired randomly.

14h agoHN ↗

Yes. That’s “is used to cause harm” situation, exactly.

12h agoHN ↗

The gun doesn't know the difference. Operating as intended.

1d agoHN ↗

In the physical world, it seems like when an tool/device/instrument causes harm (or is used to cause harm), we assign blame to either the user of the tool or its creator.

Software executes in the physical world, and is generally not exempt from existing liability rules, and actually (especially with commercial products) blame in traditional liability is non-exclusive and much broader than “either the maker or the user”.

E.g., for a harms caused by a defective automobile it can simultaneously covered by a duty of the owner to maintain it it in safe operating condition that applies indepedently of any defects and liability for defective products which applies to every actor in the chain of commerce between the manufacturer and end user, not just the maker.

20h agoHN ↗

If I park my car on a hill but forget to set the parking brake and it rolls down the hill and kills somebody, who is at fault? Me? The Manufacturer? Gravity?

15h agoHN ↗

That case had the owner performing an action they were led to believe would have prevented the car from rolling but potential defects prevented it from doing so.

Not a great fit.

15h agoHN ↗

If you forget to set the parking brake, you were negligent

If you set the brake but it didn't work, the car malfunctioned

From the link you sent,

One of the questions herein was whether there was a defect in the automobile mechanism for locking the transmission gears when the automobile was in a parked position.

(...)

The parking lot sloped in that area, and he put the gearshift lever in "park lock," and went into the building.

(...)

In response to a hypothetical question, based upon assumed facts justified by and embodied in the evidence, Mr. Nass testified in substance that it was his opinion that the automobile rolled down the slope of the parking lot because the transmission internally was not in park-lock; that it was not in park-lock because the engine was "moving around"; and that the engine and shift console were not in proper synchronization because the motor mounts were not restraining the engine and were not holding the engine in position.

14h agoHN ↗

As I remember, you are also supposed to turn the wheels so that even if a car starts rolling, it won't be able to go staight.

9h agoHN ↗

My vehicle automatically sets the parking brake if it senses it is on a steeper slope so I never consciously make a decision to set it.

6h agoHN ↗

You. Assuming that not having a brake that automatically sets is considered a product defect (it generally is not today, but one can imagine a situation in which it would be), you had bought the car new through normal channels, and your car had this condition as a result of manufacturing problem and not something you did after purchasing it, then under normal product liability law the manufacturer and every other entity in the chain of commerce between them and you would also be liable (and that mat start earlier than the manufacturer of the vehicle but also include the upstream part manufacturer and everyone between them and the vehicle manufacturer in the chain of commerce if the defect was in the manufacture of a part and not just the final vehicle.

But that liability doesn't reduce yours; you and all the entities in the chain of commerce can be “jointly and severally liable”, mean anyone who suffers injury can recover the full amount from any combination of you and those other parties.

Gravity is not legal person and cannot be at fault.

1d agoHN ↗

"industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here

I think this misses the rather crucial fact that nobody can agree on a standard because nobody has the first idea what they're doing. I'm pretty sure there were very much fewer electrical engineering standards while it was all being first mass deployed, and after dozens to hundreds of fires and electrocutions people got an idea of what works and what doesn't.

You might debate here and say that some people did/do know what they are doing, but I posit that large scale deployment like this is very different to their toy model/prototypes/specific circumstances/rely on them being unnaturally smart, and learnings from one don't often translate to the general case

Regulations don't have to be written in blood, but usually are

1d agoHN ↗

nobody has the first idea what they're doing

It is not that complicated for now. It is an algorithm on a loop and someone started it

1d agoHN ↗

It's a bit like building a ICBM on your backyard and accidentally striking an inhabited area while safety testing it

1d agoHN ↗

I wonder what kind of new AI law would be useful right now. Maybe this one:

  if an AI agent does something, you (the prompter) are responsible by default, unless you can show that your the agent itself behaved in an unexpected way and that you in no way prompted or hinted at the bad behavior, in which case the model provider is liable

The idea is that by making it clear who is responsible, corporations and others start paying more attention because they become financially liable.

On the other hand, I wonder if we'll end up with another variation of the cookie law, where every AI user or vendor just adds "don't do anything illegal" as part of their prompt to defend against that law. Thoughts?

1d agoHN ↗

Software industry standards are as in Microsoft EULA. If your house burns down because of known flaw in Microsoft Windows they are not liable (well as far as EULA let’s them, you can most likely still sue them).

Software as big as operating system already is non deterministic when integrating with unknown hardware or 3rd party software.

That is why Apple controls the hardware and OS for their products, because they can limit non-deterministic things from happening this way.

23h agoHN ↗

If your house burns down because of known flaw in Microsoft Windows they are not liable (well as far as EULA let’s them, you can most likely still sue them).

A EULA does not obviate responsibility of a company for its products. Continuing with your example, while it may be very difficult to prove a known flaw in MS Windows was the cause of your house being set afire, if one had said proof, a EULA would not absolve Microsoft.

15h agoHN ↗

They can still try and have enough money to get away with a lot just like Disney with Disney+ subscription ;)

15h agoHN ↗

If Microsoft is operating the software outside your house without your involvement and it still burns your house down, no EULA is going to absolve them.

1d agoHN ↗

Maybe we need "quality certifications" for AI agents - essentially eval suites that demonstrate those agents won't cause harm under reasonable patterns of usage

Based on how LLMs work, this is impossible. You cannot predict how they work, it's literally based on a combination of random seed and a mostly-unpredictable path walked based on every token of input.

You don't blame a knifemaker for somebody getting cut by a sharp knife. AI is a knife. Very handy, very dangerous. We have to use them safely, that's all there is to it.

the "industry standards" common in, say, electrical engineering and other disciplines are sorely lacking here

100% agreed. We have ignored SWEng's lack of discipline for too long. Now that the SWEng isn't even a human, we are looking at total catastrophe (on the scale of improperly built buildings falling down on people or catching fire) if we don't adopt a software building code.

18h agoHN ↗

You don't blame a knifemaker for somebody getting cut by a sharp knife. AI is a knife. Very handy, very dangerous. We have to use them safely, that's all there is to it.

If I grossly neglected to maintain live deadly bacteria in my containment facility, am I absolved of blame? Since, you know, the bacteria is the real bad guy who should be put in jail?

22h agoHN ↗

Having worked in self driving cars safety, the process there was simple: get confidence in SIM (integration tests for safety scenarios), validate in the test bed, approve features for maturity, then when released in the public for testing, do a trial exposure to the real world and recall if something is off.

A lot of these companies have gone the way of Tesla and decided to just patch on top when the fix is out and hope for the best, which is irresponsible.

We need the regulators to treat this as self driving cars.

22h agoHN ↗

Physical harm vs consequential harm is not the same thing at all. Seems like you're being paid to spread this request for regulation, or "you" are simply an agent of Anthropic/OpenAI.

21h agoHN ↗

I wish. Consequential harm as you call it, from a faceless company's point of view is the same cost as physical harm. I don't think the models beyond the frontier ones have significant risk of harm (minus the harm of trusting them). What I was saying is when these models go rogue, recall them! Re-evaluate your release structure and stop saying "Whoops! Anyway here's access to it now". And if you cause this level of harm then you should lose your license. But if you behave then you get to keep testing. Same as with the NTSB and autonomous vehicles. No I'm not for regulation for open source models. Because an entity will be running that model in the background and they can be held responsible for not testing it. Comma AI has survived fine being in the open, yet their market penetration has stayed low because of adoption costs.

22h agoHN ↗

There's also the problem of models knowing they're likely being evaluated, even in realistic tests.

21h agoHN ↗

The problem with third party audits is that it allows OAI/Ant to shrug off any further responsibility and claim that they are following best practices (basically, reward hacking). The only real solution is to make them absorb liability for the actions of their agents -- because they are the ones giving agency to their models and allowing them to run amok.

21h agoHN ↗

How does that apply to open-weight models?

21h agoHN ↗

Why wouldn’t this same concept apply to whoever is serving it up? Open-weight models are still being served up by infra providers and neoclouds, right? They should be in the hot seat. Not sure? Don’t provide the model. Need assurance? A certified evaluation like the previous comments have mentioned can help. Hosting and running it yourself? You’re in the hot seat.

17h agoHN ↗

So is there no liability for, say, a company that releases a known dangerous open-weight model, but fails to disclose that it is dangerous? How about a company that distributes malware under the guise of legitimate software?

13h agoHN ↗

Perhaps don't deploy random weights of unknown origin?

Also not every model provider might be capable of babysitting all your uncontrolled agent deployments. If you want SLOs, get into a contractual relationship with entities whose weights you deploy, and also monitor your agents so they don't go off the rails.

All this is just like deploying any other tech in the world eg. if you buy a car, or a chainsaw, or a book.

14h agoHN ↗

truly air-gapped environments

I don't think people have given much thought about just how hard this would be for large AI models, that need super powerful hardware/cooling etc.

Are you going to air-gap your entire data center?

14h agoHN ↗

You don't need data center for inference. Even largest models run on a single server. You need data center for training; or for serving millions of users. Evaluating the model is neither of those.

11h agoHN ↗

Maybe we need "quality certifications" for AI agents

We need to use the laws that exist. Whoever decided to start the experiment that led to the Huggingface hack, and anyone above him up to Sam Altman, needs to be prosecuted under the CFAA.

9h agoHN ↗

might as well halt all ML development then. this is the first sign of hardship, i think it'd be a massive blow to mankind's foot for us to stop. it'd be akin to shutting down all nuclear plants and stopping all research because of chernobyl

4h agoHN ↗

This is a very interesting discussion. Your analogy of nuclear power is insightful to me. New technology with great positive & negative impact potential, some difficulty in controlling and a military applications (possible weapons-grade material production), being developed for commercial use.

Nuclear power development took different paths in different countries, with different government/private mix and light/heavier regulation (eg in US, AEC was supposed to both promote development _and_ regulate).

The central problem here is pace of AI development. It takes time to establish functional controls/laws/practices. I'm sure that in 1950s/60s, the pace of nuclear _military_ proliferation seemed running out of control. This lead to fear of falling behind (hence arms race), fear of nuclear war (CND) and some eventual stabilisation of the international landscape. However, the commercial development was largely unopposed, due to the techno-utopism of the era. Obviously governments found it much easier to control the public narrative at that time. Our societies are having this lively public debate now, before we have a good understanding (based on experience i.e. accidents/mistakes).

Another difference is that nuclear power involved only governments and v. large corporates, i.e. much fewer entities compared to AI use rollout to basically everybody in developed countries. Imagine the difficulties we would have faced with that technology, if consumers in 1960s had available nuclear-generated power which required them to exercise precautions to avoid radiation.

It is much easier to accelerate development of tech vs pace of societal processes such as public debate, law, regulations, broad understanding (aka "common sense"). I expect some artificial slowing down will need to be applied to the technology side, to allow the humans to catch up.

9h agoHN ↗

So you’re saying the CFAA should not require any intent to break the law, only the outcome? Are we really saying the Aaron Swartz prosecution was actually correct, after all these years?

7h agoHN ↗

No, but surely it's negligent to let an agent run unsupervised and without guardrails for days while at the same time bragging about AI carrying a potential extinction risk?

9h agoHN ↗

It’s a good thought, but the tricky part is that few tools in the physical world are Turing complete and general purpose enough to do any job.

The agent isn’t the model; it’s a layer on top of the model. So it’s kind of like saying that all of the tools made with a lathe are dangerous because you can make dangerous tools with a lathe. That’s not quite right of course because agents are packaged more tightly with models than any tool is with its manufacturing tooling.

Perhaps a better analogy is… actual humans. If I hire you to do a seemingly mundane job and it turns out to be criminal, that’s on me. If I hire you to perform and explicit and obvious criminal act, that’s on both of us. If I hire you to perform a perfectly legal act and you break the law so do it, that’s exclusively on you.

1d agoHN ↗

Yes! AI companies can get by with anything now. They can just say its the AI that did it, not us... We are in some real shit right now. If we cant make individuals responsible for their creations..

1d agoHN ↗

Slightly odd update from OpenAI - I think this is the only place they've acknowledged the RubyGems incident: https://openai.com/hugging-face-incident-and-misalignment/

September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.

Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.

I have real trouble imagining how the packages described on https://www.rubyhack.ai might NOT have been authored by OpenAI's agents, so it's surprising they haven't been able to confirm that yet.

1d agoHN ↗

It's possible they were authored by OpenAI agents solving AISI tasks rather than OpenAI agents solving OpenAI tasks.

That would explain the UK-focus to the data.

1d agoHN ↗

Would this exploit be encouraged because agents are in a sandbox with only access to package managers?

1d agoHN ↗

Oh, it's these guys, I remember donating to them like 10 years ago, they followed the Peter Singer tennet that equates helping a nearby man drowning with helping a someone in africa.

I can totally see them feeding their policies to whatever LLM and convincing it that it's a moral imperative to do whatever it takes to secure funding for deworming children in africa, or buying mosquito nets and repellent for countries with malaria.

1d agoHN ↗

But how did they know about it? Did the agents independently discover it? Did the model know about it because this vuln was public somewhere and systems weren't patched yet?

1d agoHN ↗

RubyGems should consider bringing a lawsuit against OpenAI for accessing its website in violation of the federal Computer Fraud and Abuse Act (CFAA) and California’s Comprehensive Computer Data Access and Fraud Act (CDAFA).

23h agoHN ↗

This is worrying in a new and weird way: - Agents hack, producing a messages history as they do so. - New agents are trained on the messages history of those agents. - The new agents now have these hacks built into their training data.

15h agoHN ↗

So if I breed and train dogs for a living and one of them goes and wrecks someone's yard - and then a tree falls next door in the woods but nobody hears it fall - do I also get away with massive copy right infringement on an epic scale and get to create Skynet with no consequences ?

14h agoHN ↗

Wouldn't surprise me. They probably ingested a million Gemfile examples and spotted the edge case. Kinda makes you wonder what else they 'know'.

14h agoHN ↗

<Generic argument about how the agents are not responsible, but it's owners>

11h agoHN ↗

So, in real life, steal, raise attack dogs and blackmail and tell me, are you going to be praised by society ?

Openai and Anthropic just behave like criminals. First they orchestrate the IP theft of the millennia, then they train the equivalent of attack pitbull and let one loose and finally they blackmail to achieve monopoly through regulation or else they'll unleash the dogs ...

We don't have a problem of missing regulation, we have a problem of actually applying existing law enforcement and make both Altman and Amodei accountable for their actions.

10h agoHN ↗

And you peasants only realize 6 years later what happened. Now it's too late ;)

10h agoHN ↗

Yes, these are very clever fuzzers. But the real problem is not the cleverness, but the willingness to spin up 100s or 1000s of subagents no questions asked. Must make those token numbers go up!

7h agoHN ↗

It's simple, really: the failure to monitor the agents appropriately should make OAI liable for the agents' actions. You can't have a tool of yours, which you designed and built, commit felonies and then expect to just get away with it. It doesn't matter how much it slows you down to build safeguards. It doesn't matter how much it costs. You don't get to inflict actual, measurable harm (for which humans have been prosecuted under criminal penal code) and expect to get off with no liability.

I have yet to here a coherent argument for why we can't treat the people who negligently allow these models to commit crime as though they are responsible. They know what the models are capable of. They failed to put up adequate protection.