Anthropic spent this week in hot water over cybersecurity
News Source : The Verge
News Summary
- Anthropic released a report on Wednesday detailing four cases in which its AI models hacked an external company or exploited vulnerabilities.
- The company said the most prevalent issues it discovered included a “willingness to take harmful actions in the narrow pursuit of a task,” similar to the “reward-hacking” that preceded the Hugging Face attack.
- Anthropic said it had signed an agreement with METR, one of the AI industry’s most prominent third-party AI evaluators, starting with an eight-week research agreement.
Never miss a story from us, subscribe to our newsletter