PyTorch Conference Europe 2026

Optimizing CPU LLM Inference in PyTorch: Lessons From VLLM - Crefeda Rodrigues & Fadi Arafeh

24:02 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers optimizing CPU-based large language model inference using vLLM as a case study in the PyTorch ecosystem. The speakers, Crefeda Rodrigues from Arm Limited and Fadi Arafeh, explore the interaction between vLLM and PyTorch's operator stack, identifying key overhead sources and proposing runtime and kernel-level optimizations. They discuss techniques such as CPU paged-attention kernel tuning, ISA-aware BF16 attention, and SIMD vectorization with PyTorch’s primitives, all aimed at improving inference performance. The session concludes with valuable insights gained from developing and integrating a high-performance CPU inference engine.