PyTorch Conference Europe 2026

Lightning Talk: Training Embedding Model Resiliently for Multimodal M... Huamin Chen & Haichen Zhang

12:48 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk explores the training of embedding and classification models for efficient routing in Large Language Model (LLM) systems, focusing on balancing cost, latency, and quality. The speakers, Huamin Chen from Red Hat and Haichen Zhang from AMD, present the vLLM Semantic Router, highlighting its reliance on fast classifiers for the Mixture-of-Multimodal Models architecture. They discuss their end-to-end approach using native PyTorch on AMD GPUs, showcasing a multilingual text embedding model and multimodal embedding models enhanced for high GPU utilization through distributed training optimizations. Attendees will gain insights on training efficient classifiers for LLM routing systems and how to integrate these models into production inference pipelines.