Lightning Talk: Slash LLM Cold-Start Times by Pre-distributing GPU... Billy McFall & Maryam Tahhan
About this talk
This talk focuses on optimizing Large Language Model (LLM) deployments by addressing GPU cold-start times and reducing startup latency. The speakers, Billy McFall and Maryam Tahhan from Red Hat, explore how distributed inference at scale can be more efficient by pre-distributing GPU kernel caches across inference nodes using KServe, a Kubernetes-based model inference runtime. They provide insights into the technical implementation of signing, verifying, and mounting cache images, ensuring supply-chain security throughout clusters. Attendees will gain a practical blueprint to enhance the performance of GPU-heavy workloads in production environments.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17