Portable High‑Performance LLM Serving: A Triton Backend for... Burkhard Ringlein & Jan van Lunteren
About this talk
This talk covers the introduction of a Triton backend for vLLM, which is becoming the industry standard for serving Large Language Models in production. The speakers, Burkhard Ringlein and Jan van Lunteren from IBM Research, explain how traditional performance reliant on hand-written CUDA or HIP kernels can limit portability across different hardware. They showcase how the Triton attention backend offers competitive performance across GPU platforms using a single code base, eliminating the need for specialized kernels. The session includes insights into the engineering, system aspects, kernel enhancements, and optimizations that enable this Triton-only solution to deliver consistent high performance on both NVIDIA and AMD GPUs.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17