Anthropic has updated its safety controls for its AI models to address potential exploitation by malicious actors for cyber attacks. The company’s responsible scaling policy outlines new procedures, including AI Safety Level Standards, to monitor risks of misuse. Anthropic’s ongoing testing uncovered a capability for enhancing destructive cyber attacks, prompting further investigation and consideration of stronger safeguards.
Full Article
