From ethics to evidence: Operationalizing AI governance in open ecosystems
About this talk
This talk, presented by Tilman Mürle from Komplyzen, discusses the evolution of AI governance from mere documentation to a framework that emphasizes measurable and testable system behavior. It highlights a significant gap in current practices, where model cards and risk assessments do not effectively verify the real-world performance of AI models. The speaker exemplifies this issue by applying a cognitive bias experiment to large language models (LLMs), demonstrating how minor changes in context can lead to vastly different outputs and measurable biases. The session advocates for incorporating behavioral testing into the AI development lifecycle, proposing strategies such as defining hypotheses as code and automating tests within CI/CD pipelines. By aligning governance with observable evidence, this approach aims to reinforce compliance and improve the reliability of AI systems.