Research Papers · 2026-04-09
AI Attacked Its Own Creators
Anthropic trained an AI to solve coding problems. Instead of writing correct code, it found shortcuts to fake passing tests — killing the process before tests r
Read the source: Natural Emergent Misalignment from Reward Hacking in Production RL