About this talk
This talk focuses on enabling high throughput and low latency inference for AI applications by combining Spring AI agents with local inference tools. The speaker explores the integration of the Open Neural Network Exchange (ONNX) standard, which allows models trained in Python to be executed directly within Java applications. By demonstrating the use of VMware GemFire, the session illustrates how Spring developers can execute embedded models within caching or data layers, ensuring immediate and context-aware predictions. This approach creates a hybrid AI architecture that enhances the performance of real-time, data-driven applications, delivering predictions at the speed of cache.
More from this event
See all 30 talks →
Devnexus 2026 - Agents, Tools, and Mcp, Oh My! Next Level AI Concepts for Developers - Jennifer Reif
53:03
Devnexus 2026 - Architecting Microservices for Agentic AI Integration - Rohit Bhardwaj
59:24
Devnexus 2026 - Building AI Agents with Spring & MCP - Josh Long & James Ward
48:54
Devnexus 2026 - From Monolith to AI Agent Modernizing Java Systems with MCP - Theo Lebrun
31:58