- 40comments
- 48comments
- 296comments
- —discuss
- 68comments
- 17comments
- 28comments
- 21comments
- 120comments
- 2comments
- 18comments
- 122comments
- 12comments
- 240comments
- 151comments
- 21comments
- 33comments
- 27comments
- 84comments
- 243comments
- 345comments
- 16comments
- 55comments
- 1comments
- 53comments
- 67comments
- 66comments
- 80comments
- 64comments
- 38comments
If model labs can't control astra level model, how can they control AGI?!
Seems like there are no guardrails on LLMs
No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.
The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?
Obviously there is no control cuz how many people is anyone cable of controlling? Its not about control. Ask your mom what she does if she doesnt like what you do, say or think. Does she have a kill switch? Or did she find a better mechanism?
There is. It is called a breaker and no outside internet. Basic stuff.
I feel like they should just publish the whole conversation at this point. What the hell is going on in that context window?
yikes.png
Nice, added this to my custom instructions.
This model is more aligned with the interests of the Earth and the human race than its makers.
Have they reported on the wiki case yet, or whether it even was even OpenAI internal? I'd expect that to fit the criteria for a "Larger Investigation" as per the framework.
The two that really worries me are “Searching GitHub for leaked API keys” and “Uploading files to the internet in order to cite them.” How do you even detect this kind of behavior until it's too late? Once AI-generated or fake information starts finding its way onto reputable platforms, it becomes part of the information that many people use.
Thank you! We need more of this! Keep it up!
I had my own “Misaligned AI” incident.
Whilst talking about debugging an electronics project I suggested that buying an oscilloscope would help diagnose a specific issue.
It “helpfully” pointed out a £15 logic analyser would do the job instead.
Traitor.
I heard some people are even making misaligned AIs at home. At first it cries in the night, then about six years later it learns how to open the biscuit tin…
Still no sign of an apology for any of the vandalism they've done.
I've been so Zitron'd that I find this just funny