Lightning Talk: From Hugging Face To Handheld: Scaling LLM Deployment W... Cormac Brick & Weiyi Wang
About this talk
This session explores the end-to-end journey of deploying custom PyTorch-based Open Source Large Language Models (LLMs) on cross-platform devices using LiteRT. The speakers, Cormac Brick and Weiyi Wang from Google, demonstrate how to convert a Hugging Face Transformers checkpoint for on-device execution, covering essential steps from conversion to deployment. They discuss automated optimization techniques, seamless fine-tuning integration from an Unsloth session to a TorchAO-quantized model, and how they enabled the QWEN0.6 model in just 20 minutes. Additionally, the talk includes interactive validation methods for verifying numerical correctness before device deployment, ensuring a streamlined fine-tune-to-deployment process within the PyTorch and Hugging Face ecosystem.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17