About this talk
This talk, presented by Giulio Zizzo from IBM Research, explores the evaluation and enhancement of the robustness of large language models using scalable red teaming methodologies. It addresses the challenges faced as these models evolve from isolated chatbot applications into critical business systems, highlighting the expanded attack surfaces that arise when models interact with various tools and workflows, thereby increasing vulnerabilities.
The session introduces ARES, the AI Robustness Evaluation System, an open-source framework that supports systematic exploration of attack objectives and strategies, including multi-turn interactions and obfuscation techniques. ARES features a modular architecture that allows users to customize evaluation methods and attack scenarios, facilitating the automation of adversarial testing workflows while mapping vulnerabilities to threat models for effective risk prioritization and mitigation design.