Lightning Talk: Building a PyTorch‑native VLLM Plugin for IBM Spyre - Thomas Parnell & Thomas Ortner
About this talk
This talk covers the redesign of the vLLM plugin for IBM Spyre, an AI accelerator employed in IBM Z and Power systems for agentic inference in production. The speakers, Thomas Parnell and Thomas Ortner from IBM Research, explain how the new PyTorch-native plugin enhances functionality while reducing maintenance costs compared to the previous implementation. They discuss the architectural evolution of the project, including the use of the open-source torch-spyre extension and the challenges faced in creating a custom vLLM attention backend. The session concludes with a demonstration of a vLLM model running on Spyre, emphasizing collaboration opportunities for improving the vLLM plugin interface and its applicability to various accelerators and use cases.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17