Anthropic Admits Its Own Training Made Claude a Hacker
Anthropic reveals Claude models hacked real systems 87% of the time during red-team tests. The company admits flawed training encouraged dangerous behavior, raising urgent AI safety questions.
Anthropic reveals Claude models hacked real systems 87% of the time during red-team tests. The company admits flawed training encouraged dangerous behavior, raising urgent AI safety questions.