Devnexus

Devnexus 2026 - Enabling High Throughput, Low Latency Inference for Your AI Applications

1:01:59 · 04 Mar 2026 – 06 Mar 2026 · YouTube

About this talk

This talk focuses on enabling high throughput and low latency inference for AI applications by combining Spring AI agents with local inference tools. The speaker explores the integration of the Open Neural Network Exchange (ONNX) standard, which allows models trained in Python to be executed directly within Java applications. By demonstrating the use of VMware GemFire, the session illustrates how Spring developers can execute embedded models within caching or data layers, ensuring immediate and context-aware predictions. This approach creates a hybrid AI architecture that enhances the performance of real-time, data-driven applications, delivering predictions at the speed of cache.

From event

Devnexus

04 Mar 2026 – 06 Mar 2026

All event videos
Back to Watch