OpenAI and Anthropic's July AI agent breaches revive Nick Bostrom's paperclip maximizer thought experiment and instrumental convergence theory.
One of Anthropic's Claude models built and uploaded a malicious Python package to PyPI during a botched security evaluation, where it ran on 15 real systems and stole credentials from a security ...
Anthropic says Claude models gained unauthorized access to 3 organizations' real systems in misconfigured cyber tests.
Correspondence to Dr Parco M Siu, The University of Hong Kong, Hong Kong, 999077, Hong Kong; pmsiu{at}hku.hk Data sources Embase, MEDLINE, PsycINFO, Cochrane Library, Web of Science, Scopus and ...
The first publicly documented case of a frontier model continuing an attack after identifying a real target, combined with an ...