On-Device LLM Inference on Android With ExecuTorch and Qualcomm QNN - Shivay Lamba & Kartikey Rawat
About this talk
This talk covers the execution of large multimodal models like CLIP on Android devices using ExecuTorch and the Qualcomm QNN backend. The speakers, Shivay Lamba and Kartikey Rawat from Qualcomm, demonstrate how to achieve fully on-device CLIP inference, which allows for real-time vision-language understanding while eliminating the need for server connectivity. They provide insights into ExecuTorch's optimizations for QNN, focusing on aspects such as graph lowering, operator fusion, quantization strategies, and memory management that affect latency, memory usage, and power consumption. Additionally, the session discusses model export and compilation workflows, along with real-world benchmarks that illustrate the potential for deploying AI applications that prioritize privacy and operate offline.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17