Even simpler: an agent is a simple while loop with tool calls that prompt an LLM continuously. That’s deterministic, standard software. You literally do not have to process tool calls in a way that will execute whatever the model generated. It’s a choice to process a tool call “run_bash” that provides an escape hatch with full execution permissions.
This is only one type of control, and it is certainly not infallible. Also, people will be incentivised to hook up AIs to real tools. But even if they don't, as long as people can interact with super-intelligent AIs without tool access, there are many potential dangers.
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Maybe I'm wrong, but your comment seems to suggest that you are skeptical of ASI. Do you have any arguments to support that?
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Unfortunately the article is behind a paywall. I would have been interested to see if he makes any substantial arguments. I don't see any reason to believe that AI being software entails that it can be controlled. Moreover, even if a measure of control is possible, giving that control to a handful of oligarchs seems undesirable.
"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
Everyone keeps talking about terminator scenarios or whatever, but the thing that scares me the most is what happens when people let agents run amuck in systems they should not be in, then the agent just starts doing random shit as the context overflows. We’ve all seen it. They just descend into madness, but what happens when they collapse with a hand on the wheel of, say, a backup generator at a hospital?
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
That would be 100% the fault of the people who set them loose on such things, and such things should not be on the open Internet for tons of reasons. This just adds a new reason that pathetic security around SCADA systems is dangerous. It was already dangerous before.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
...se that AI can easily become uncontrollable in the near future
Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.
That's not the point. The point is, regulation is required, and is coming.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
Held responsible? They're already being rewarded with vast fortunes and influence.
The point is, regulation is required, and is coming.
The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.
Yes, I fully believe that's why they're pushing on this. But as the general population, we don't need new laws to protect us from this AI-related problem. We just need enforcement of what is already on the books.
I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.
I think the other issue is that even if it's controllable, there's nobody representing us that is controlling it, except theoretically regulators who are facing an uphill battle to bring accountability and limits to these companies.
The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI models human behavior. Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
You guys know robots are coming, right? Like humanoid and all sorts of other robots too.
They're going to be running the infrastructure, self-improving, and will have human like power seeking ambitions and human like flaws. Because they're trained on human input.
And they will not require humans granting them money...
It's still software that runs on hardware someone owns. Whoever owns the hardware or the service can pull the plug, as long as they're willing to. That's the same situation as with legacy malware. Self-replicating worms have existed for decades and run without anyone controlling them, yet we don't call them uncontrollable. So what is the difference between your scenario and legacy malware?
as long as anyone anywhere is willing to make money by renting hardware to it.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
Yeah, sure, it's controllable. Because humanity has famously solved the crime accountability problem back in 1902, and no crime has gone unpunished since.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.
IMHO, if the model breaks a law, apply the law to the operator.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous?
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
A human is just a few chemical reactions, and fairly stable ones at that.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
So in the real world, what are the analogues to lithium and morphine we should feed to e.g. LLMS, how do we feed them, and how do we prove that it prevents unsafe behavior?
Software is trivially easy to control though. If you want to stop it hacking websites, you don't give it access to the internet. If you want to restrict it from connecting to arbitrary websites, you put in a whitelist. You can trivially sandbox applications these days to prevent them from accessing network or local resources
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
I believe that doomerism is a marketing gimmick to appeal to immature people who are attracted to danger, and get excited about it.
Unfortunately, a lot of decision makers with money fall into this category.
I heard it posited that the gimmick is on lawmakers, that they get to feel epic importance because they are writing historical laws that will save humanity from the machine gods... gives them tingles in the ugly bits
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
If AI is as powerful as an atomic bomb why could it only manage to kill a couple hundred Iranian school girls? Surely a truly society-shifting technology could manage to execute 1000 innocent children at least.
I've started to wonder why the big AI labs have no robotics yet, and I'm thinking now the reason why we don't see robots is because they want us (the general public) at all costs to not freak out. Imagine what would have happened if we had 6ft tall "friendly" robots marching through the streets, and suddenly the news came about of AI going rogue, like what we've seen from the recent hacks ... I bet it would have been the end of the story for AI labs right away (of course not for AI, because the department of defense would take over).
suddenly the news came about of AI going rogue, like what we've seen from the recent hacks
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
This is a website about making people freak out about how epic and all-powerful LLM code generators are. You should watch your tone unless you want to be replaced by AI.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
AI is a system of equations, used as a statistical automata. This should be able to be controlled by definition.
Even simpler: an agent is a simple while loop with tool calls that prompt an LLM continuously. That’s deterministic, standard software. You literally do not have to process tool calls in a way that will execute whatever the model generated. It’s a choice to process a tool call “run_bash” that provides an escape hatch with full execution permissions.
We do not have to do that!
This is only one type of control, and it is certainly not infallible. Also, people will be incentivised to hook up AIs to real tools. But even if they don't, as long as people can interact with super-intelligent AIs without tool access, there are many potential dangers.
Wish I could read but it's behind a paywall. Nice to see one person who doesn't think we're all doomed though
Obviously a bad businessman.
Many people don’t agree with the doomers.
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Maybe I'm wrong, but your comment seems to suggest that you are skeptical of ASI. Do you have any arguments to support that?
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
unwall.app does miracles
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Unfortunately the article is behind a paywall. I would have been interested to see if he makes any substantial arguments. I don't see any reason to believe that AI being software entails that it can be controlled. Moreover, even if a measure of control is possible, giving that control to a handful of oligarchs seems undesirable.
Seems like it is the uncontrollable aspect to me.
I'm not sure what you mean.
The oligarchs are under no controls, that’s what I understand from their comment
I think they are saying...
the rich and powerful may have so much sway over government that we may not be able to reach "Ai alignment" with society
... but not sure I agree, I hope this is not true anyhow
You think Gemini, Claude, ChatGPT are controlled by oligarchs?
In other words, that Alphabet, Anthropic, and OpenAI are run / owned by oligarchs?
People like Amodei and Altman are oligarchs? Come on.
Any link with no paywall?
ish...
https://cybernews.com/ai-news/arthur-mensch-apocalypse/
"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
That sounds still controllable.
Everyone keeps talking about terminator scenarios or whatever, but the thing that scares me the most is what happens when people let agents run amuck in systems they should not be in, then the agent just starts doing random shit as the context overflows. We’ve all seen it. They just descend into madness, but what happens when they collapse with a hand on the wheel of, say, a backup generator at a hospital?
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
That would be 100% the fault of the people who set them loose on such things, and such things should not be on the open Internet for tons of reasons. This just adds a new reason that pathetic security around SCADA systems is dangerous. It was already dangerous before.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
Sounds like motivated reasoning from someone worried about having their job stolen by wild monkeys.
Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.
That's not the point. The point is, regulation is required, and is coming.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
Held responsible? They're already being rewarded with vast fortunes and influence.
The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.
Why do we need new regulations? Hacking into hugging face is already illegal.
We need new regulations to impose costs on anyone who might compete with OpenAI and to outlaw open models to lock in cloud model oligopoly.
Yes, I fully believe that's why they're pushing on this. But as the general population, we don't need new laws to protect us from this AI-related problem. We just need enforcement of what is already on the books.
I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.
I think the other issue is that even if it's controllable, there's nobody representing us that is controlling it, except theoretically regulators who are facing an uphill battle to bring accountability and limits to these companies.
Implying there is some shared belief / morality / value system that unites all of _us_
Being "controllable" isn't determined by a system being deterministic.
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI models human behavior. Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
You just admitted it's controllable.
You guys know robots are coming, right? Like humanoid and all sorts of other robots too. They're going to be running the infrastructure, self-improving, and will have human like power seeking ambitions and human like flaws. Because they're trained on human input.
And they will not require humans granting them money...
It's still software that runs on hardware someone owns. Whoever owns the hardware or the service can pull the plug, as long as they're willing to. That's the same situation as with legacy malware. Self-replicating worms have existed for decades and run without anyone controlling them, yet we don't call them uncontrollable. So what is the difference between your scenario and legacy malware?
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
Yeah, sure, it's controllable. Because humanity has famously solved the crime accountability problem back in 1902, and no crime has gone unpunished since.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
We don't live in anarchy, or do we? We probably would still have slaves if it would not be so aggressively penalized.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.
IMHO, if the model breaks a law, apply the law to the operator.
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
So in the real world, what are the analogues to lithium and morphine we should feed to e.g. LLMS, how do we feed them, and how do we prove that it prevents unsafe behavior?
[delayed]
All it takes is one Harrison Bergeron…
a random number generator is software, so it can be controlled
Software is trivially easy to control though. If you want to stop it hacking websites, you don't give it access to the internet. If you want to restrict it from connecting to arbitrary websites, you put in a whitelist. You can trivially sandbox applications these days to prevent them from accessing network or local resources
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
Sounds like motivated reasoning from someone afraid of having their job stolen by rogue super-human AI. Keep crying, weakling.
People are animals. They can be controlled.
I believe that doomerism is a marketing gimmick to appeal to immature people who are attracted to danger, and get excited about it. Unfortunately, a lot of decision makers with money fall into this category.
I heard it posited that the gimmick is on lawmakers, that they get to feel epic importance because they are writing historical laws that will save humanity from the machine gods... gives them tingles in the ugly bits
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
Depends on your definition of "controlled". The ambiguity behind the term is doing a lot of heavy lifting here.
Boeing's MCAS system was also "just software". Which in principle can be "controlled", i.e. changed, updated, audited or whatnot.
But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
Literally the "AI is just computer" argument
I hope he didn't really say this: it is an incredibly basic take. More or less like atomic bomb is hardware, it can be controlled.
That is the point. It can be controlled by the operators if they want to.
If AI is as powerful as an atomic bomb why could it only manage to kill a couple hundred Iranian school girls? Surely a truly society-shifting technology could manage to execute 1000 innocent children at least.
Computer viruses are software too
I've started to wonder why the big AI labs have no robotics yet, and I'm thinking now the reason why we don't see robots is because they want us (the general public) at all costs to not freak out. Imagine what would have happened if we had 6ft tall "friendly" robots marching through the streets, and suddenly the news came about of AI going rogue, like what we've seen from the recent hacks ... I bet it would have been the end of the story for AI labs right away (of course not for AI, because the department of defense would take over).
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
This is a website about making people freak out about how epic and all-powerful LLM code generators are. You should watch your tone unless you want to be replaced by AI.
They will sell the company to Nvidia…
Cohere CEO said something similar: https://www.theglobeandmail.com/business/technology/article-...
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
That's all I needed to hear to completely disregard your motivated reasoning.
Did you make a new account just to shit post about AI?