How Many Spark Applications Can Your Etcd Really Handle? - João Soares & João Azevedo, Feedzai
About this talk
This talk examines the effective integration of Apache Spark on Kubernetes, focusing on the development of a self-service Spark platform. The speakers, João Soares and João Azevedo from Feedzai, discuss their journey transitioning from legacy YARN clusters and the challenges faced, particularly with managing Spark custom resources that can lead to instability in etcd under high workloads. They highlight the significance of open-source tools like Volcano and the Spark Operator from Kubeflow, as well as propose innovative solutions, including the Kubernetes API Aggregation Layer, to enhance storage management in large-scale Spark deployments.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32