About this talk
This talk, presented by Ji Darwish from Lunatech, focuses on AI alignment challenges in large language models. It examines the risks of misalignment and deceptive behavior, highlighting how these issues arise from the training process rather than explicit programming. The speaker discusses the interpretability challenges posed by neural networks, the complexities of emergent behaviors, and the gap between model evaluation and real-world deployment, complicating risk assessment. This session sheds light on the implications for AI safety frameworks and the regulatory landscape, making it essential for understanding the integration of AI systems into production environments.