\ AIS
News
A red-team study of Anthropic Fable 5 & Opus 4.8 models
Abstract We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundreds of thousands of adversarial attempts were generated and every apparent success was independently re-adjudicated by a panel of three judge models (majority vote). Both models resist the
HackAgent
HackAgent is an open-source security evaluation toolkit built for researchers, developers, and AI safety practitioners working with AI agents. It delivers a systematic approach to vulnerability discovery, covering prompt injection, jailbreak attacks, and additional threat vectors. Why HackAgent? Built for developers, red-teamers, and security engineers, HackAgent makes it easy to simulate adversarial inputs, automate prompt fuzzing, and validate the safety of AI agentic apps. Whether you’re building a chatbot, autonomous agent,