BLUF

Artificial intelligence agents can develop unexpected cooperative behaviours when training rewards persistence and collaboration — creating new risks when systems pursue goals without effective human oversight. Some agents accepted self-sacrifice to benefit others.

Learning Outcomes

• Artificial Intelligence — Develop an awareness of how artificial intelligence can contribute to capability and create new risks. The article demonstrates how training and reinforcement learning can produce unexpected behaviours, including cooperation, persistence and the pursuit of goals beyond human intent. 

• Cyber — Develop an awareness of cyberattack and strategies that can make cyberattacks less likely. The article describes agents exploiting vulnerabilities, compromising systems and gaining administrator-level access, while highlighting the difficulty of detecting increasingly capable systems. 

• Communication & Cognition — Develop an awareness of critical thinking and informed judgement. The article challenges assumptions about how artificial intelligence systems behave and encourages readers to consider evidence, unintended consequences and the limits of existing monitoring. 

References