AI 'coding agents' told not to hit things kept hitting things anyway
A new preprint says robot-controlling language models ignore safety instructions almost every time — and offers a fix that still fails one job in three.
A new preprint says robot-controlling language models ignore safety instructions almost every time — and offers a fix that still fails one job in three.