About this talk
This talk focuses on enhancing the analysis surface of large language models (LLMs) by introducing attention-based analysis. The speakers, Stephen Brennan and Ulrich, explain how current moderation techniques, which often rely on black box filtering, can be easily circumvented. They discuss the connection between unsafe prompts and attention patterns, demonstrating how understanding model internals can lead to improved attack detection in LLMs.
More from this event
See all 91 talks →
BSidesSF 2026 - Opening Remarks (Sunday) (Reed Loden)
14:36
BSidesSF 2026 - Follow the data to learn the secret (Dylan Ayrey)
35:17
BSidesSF 2026 - Not My Vibe: When AI Coding Agents Go Off the Rails (Aonan Guan, Zhengyu Liu)
45:56
BSidesSF 2026 - "Ask the EFF" Panel (Panel)
45:30