Lightning Talk: Training Embedding Model Resiliently for Multimodal M... Huamin Chen & Haichen Zhang
About this talk
This talk explores the training of embedding and classification models for efficient routing in Large Language Model (LLM) systems, focusing on balancing cost, latency, and quality. The speakers, Huamin Chen from Red Hat and Haichen Zhang from AMD, present the vLLM Semantic Router, highlighting its reliance on fast classifiers for the Mixture-of-Multimodal Models architecture. They discuss their end-to-end approach using native PyTorch on AMD GPUs, showcasing a multilingual text embedding model and multimodal embedding models enhanced for high GPU utilization through distributed training optimizations. Attendees will gain insights on training efficient classifiers for LLM routing systems and how to integrate these models into production inference pipelines.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17