Lightning Talk: Enabling the Audio Modality for Language Models - Eustache Le Bihan, Hugging Face
About this talk
This talk features Eustache Le Bihan from Hugging Face, who discusses the integration of audio into large language models within the `transformers` library. The session offers an overview of the current landscape of Audio Language Models, highlighting trends in incorporating audio into pretrained text backbones. The speaker examines the convergence of architectural choices inspired by Vision Language Models and introduces concepts like audio tokenization and streaming. Key insights include the differences between audio encoders and audio tokenizers, their advantages and limitations, and how these innovations are implemented in PyTorch to standardize audio as a modality in the open-source community.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17