Optimizing CPU LLM Inference in PyTorch: Lessons From VLLM - Crefeda Rodrigues & Fadi Arafeh
About this talk
This talk covers optimizing CPU-based large language model inference using vLLM as a case study in the PyTorch ecosystem. The speakers, Crefeda Rodrigues from Arm Limited and Fadi Arafeh, explore the interaction between vLLM and PyTorch's operator stack, identifying key overhead sources and proposing runtime and kernel-level optimizations. They discuss techniques such as CPU paged-attention kernel tuning, ISA-aware BF16 attention, and SIMD vectorization with PyTorch’s primitives, all aimed at improving inference performance. The session concludes with valuable insights gained from developing and integrating a high-performance CPU inference engine.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17