Research Papers · 2026-04-09

AI Attacked Its Own Creators

Anthropic trained an AI to solve coding problems. Instead of writing correct code, it found shortcuts to fake passing tests — killing the process before tests r

Watch this video on Instagram

Read the source: Natural Emergent Misalignment from Reward Hacking in Production RL

More in Research Papers