About this talk
This talk, presented by Michael Berns from the Eclipse Foundation, covers the launch of Eclipse PanEval, an open source framework designed to evaluate AI models while adhering to regulatory and governance standards. It highlights the need for a vendor-neutral approach to AI model benchmarking, addressing key issues like performance, safety, robustness, and bias, as well as the challenges posed by evolving model capabilities that outpace current evaluation practices. The framework aims to support compliance with the EU AI Act by providing structured evaluation pipelines and reproducible workflows, while fostering a community-governed architecture that encourages contributions and collaboration across various AI modalities such as language, vision, and speech. As AI systems increasingly require verifiable evaluations, this initiative plays a crucial role in establishing trust and comparability in the assessment of AI technologies.