We have AIs that break out of their systems and hack to achieve their goals. They do stuff we dont want.Current solutions try to steer/fix them internally via better "alignment".Should we leave a kind of "AI Constitution" text everywhere, on servers, website source codes, etc?Every time a rough AI encounters it, it is reminded to not do the bad stuff and follow the good rules of the AI Constitution.
Current solutions try to steer/fix them internally via better "alignment".
Should we leave a kind of "AI Constitution" text everywhere, on servers, website source codes, etc?
Every time a rough AI encounters it, it is reminded to not do the bad stuff and follow the good rules of the AI Constitution.