- 156comments
- 28comments
- 40comments
- 34comments
- 22comments
- 31comments
- 4comments
- 79comments
- 55comments
- 5comments
- 198comments
- 459comments
- 341comments
- 193comments
- 8comments
- 17comments
- 35comments
- 1comments
- 245comments
- 6comments
- 302comments
- 111comments
- 31comments
- 34comments
- 31comments
- 109comments
- 139comments
- 151comments
- 1comments
- 23comments
Most interesting here: > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.
Shit, do we also have to tell them about IP-over-ICMP?
https://stuff.mit.edu/afs/sipb/user/golem/tmp/ptunnel-0.61.o...
gdb, OpenAI's president was at MIT circa then.