PyTorch Conference Europe 2026

Lightning Talk: Slash LLM Cold-Start Times by Pre-distributing GPU... Billy McFall & Maryam Tahhan

6:47 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk focuses on optimizing Large Language Model (LLM) deployments by addressing GPU cold-start times and reducing startup latency. The speakers, Billy McFall and Maryam Tahhan from Red Hat, explore how distributed inference at scale can be more efficient by pre-distributing GPU kernel caches across inference nodes using KServe, a Kubernetes-based model inference runtime. They provide insights into the technical implementation of signing, verifying, and mounting cache images, ensuring supply-chain security throughout clusters. Attendees will gain a practical blueprint to enhance the performance of GPU-heavy workloads in production environments.