Open Community Experience (OCX)

How to see in the black box: Uncovering AI alignment

14:03 · 21 Apr 2026 – 23 Apr 2026 · YouTube

About this talk

This talk, presented by Ji Darwish from Lunatech, focuses on AI alignment challenges in large language models. It examines the risks of misalignment and deceptive behavior, highlighting how these issues arise from the training process rather than explicit programming. The speaker discusses the interpretability challenges posed by neural networks, the complexities of emergent behaviors, and the gap between model evaluation and real-world deployment, complicating risk assessment. This session sheds light on the implications for AI safety frameworks and the regulatory landscape, making it essential for understanding the integration of AI systems into production environments.