PyTorch Conference Europe 2026

Lightning Talk: Enabling the Audio Modality for Language Models - Eustache Le Bihan, Hugging Face

14:03 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk features Eustache Le Bihan from Hugging Face, who discusses the integration of audio into large language models within the `transformers` library. The session offers an overview of the current landscape of Audio Language Models, highlighting trends in incorporating audio into pretrained text backbones. The speaker examines the convergence of architectural choices inspired by Vision Language Models and introduces concepts like audio tokenization and streaming. Key insights include the differences between audio encoders and audio tokenizers, their advantages and limitations, and how these innovations are implemented in PyTorch to standardize audio as a modality in the open-source community.