Innocent-looking AI reasoning can make bad behavior harder to catch

Image: Science News
ad slot · in-content video 16:9
Coverage
Coverage
- AI safety monitoring can fail when an AI’s reasoning is the main clue that something has gone wrong, new research suggests.