18 Bluetooth Controllers Walk into a Bar: Observability & R... Simon Schrottner & Manuel Timelthaler
About this talk
This talk explores the integration of observability and runtime configuration using CNCF tools in a real-time party game. The speaker discusses the challenges of monitoring an open-source game, which runs on a Raspberry Pi and uses PlayStation Move controllers. They emphasize the importance of observability tools like OpenTelemetry and feature flagging to gain insights into game performance. The session illustrates how to enhance metrics collection and manage data effectively, especially in scenarios involving rapid accelerometer inputs. The speakers share their experiences in transforming the codebase into a microservice architecture and how it opened up new possibilities for monitoring and analytics. They also highlight the need for documentation on using these tools beyond traditional cloud-native applications, specifically for IoT and real-time systems.
Full transcript
Hello, everyone. Hello. We're super glad that so many people showed up to the last talk of this year's KubeCon. Thank you for being here. Thank you. We're honored. We tried to make this entertaining for you, so you won't regret it. 18 Bluetooth controllers walk into a bar. We'll be talking about observability and runtime configuration with CNCF tools for a party game, for a real-time party game. And
for that, of course, you first have to understand what is this game we're talking about. Yeah, and that actually brings us to those PlayStation Move controllers. Actually, you can imagine this is an acceleration-based game where it's where those controllers are like a spoon with egg on it, and you try to protect your egg while you while you try to make the other lose their egg. Just to
give you a little fun demo, let's see who wins. Basically, when the game is starting, it's like this, and then you try to to Yeah, you try to push the other around, and that's the whole sense of the game. And we try to this game is really fun. It's a really nice icebreaker. I love to bring this to conferences because, as I said, you you're finding people,
you talk a little bit, is a really good starting point. But from now and then, when I brought this game to conferences, there was always this this moment when suddenly something was not happening. It's an open-source game running on a Raspberry without any kind of monitor, and it was really really frustrating. But Simon, don't you work at an observability company? Yes, I'm working actually at an observability
company, and I should know how this is how how I could get more insights into that. I'm working at Dynatrace, but in the end, the thing was, well, for me, this does not look like my usual cloud-native workflows. This looks for me like a little bit of IoT, and I thought, well, I know somebody who works at an IoT company, and that's why I asked Manu if
he can help me with that. >> Hi, I'm Manuel. I work at Tracktive. We're one of the global leaders in health and location tracking for pets. And while we, as a company, do work with hardware a lot, I don't. I work on the web services. But hey, sounds interesting, I mean. So, we asked ourselves basically the question, how can we make this little tool a little How
can we know what's going on? How can we get a little bit of observability into it? I mean, we already have in the CNCF this really cool technology which is like Open Telemetry. And as I'm an Open Feature maintainer, I thought, well, let's sprinkle a little bit of feature flagging also on top of that. And that brings us to our experiment. Can we use the tools in
Yeah, their default state? What works? What breaks? What do we have to tweak to use for a real-time game? Yeah, and we want to take you with us on our learning journey. We've selected six learnings that might be interesting for you, I hope. We had many more learnings. It's usually not the ones that you expect. But yeah. Okay. But first of all, to understand what how this
game looked like, we need to take a little bit bit a little look at the code base because that's an open source game written in Python uh with multiple processes, we IPC communication, starting own sub-processes and things like that. Really tedious to work with. And we have one or two instrumented. So, we tried to understand the code base, tried to add Open Telemetry to it, but it
was really, really tedious. So, Yeah, that was actually quite what did we do? Yeah, we we we we we do the did the obvious thing like everybody would do. We uh refactored it and we over-engineered it heavily. Uh we turned it into a microservice architecture with GRPCs. So, getting a little bit more into the cloud-native space. But this opened up a a of cool tools already. So,
with auto-instrumentation, we had already all the insights we wanted to have there. And it enabled us to see way way more than we expected. Onto the Raspberry Pi, we also added a uh we added a little bit of an observability stack to it. Of course, we need to have our collector running where we get uh which forwarded all our data to our observability backend. And we had
this little fun tool called Flectory in the middle for our feature flags, our feature flag proxy. So, what do we get now? Yeah. What do we get? We get the obvious thing out of it. We get spans and traces with what everybody wants to see. But in this case, it's a little bit different because we are actually in a game. So, for us uh it's not just
about a certain request or things like that. We can enhance them with way more information to have way more insights into the game like how long does a player live and those kind of things. And we can already push it to different backends, for example, also Dynatrace. But the thing is, we now have the microservices running on the Raspberry Pi for our game. And for example, if
you're using Dynatrace or any other solution for that, all of the data is in the cloud. And we were sure that we have to do this. But But yeah, we took a look at our Raspberry Pi and that's quite funny. So, we thought we we over-engineered a lot of stuff and thought, well, having the observability stack on the same machine, that would definitely not work. And it
would definitely be a problem. But it turns out the Raspberry Pi is a quite powerful machine. We even tried to run this game with 32 controllers and more on 60 Hz on 60 FPS. It was no problem at all. Pulling the data, using the data, and the adding observability data was no issue at all. So, we were had the possibility to actually call combine everything into this
nice little over-engineered architecture on one small little device. And now we can use it off-grid, which which was really cool, especially for a party game. Yeah, and also cool for conferences because Wi-Fi isn't always that great and we can just, yeah, play it because everything is on the Raspberry Pi. Yeah, so we have logs and traces. Of course, we also want metrics. And for that first, you
have to understand all the game input is basically from the motion sensors. Everything is based on the acceleration. So, how fast you either move or how fast you get moved. So, this is our main metrics. the games are short. I mean, with 36 people, it's a bit longer, but if you play a game with four people, it could it could be over in in 10 seconds. And
we were using Prometheus for getting, yeah, the metrics data uh somewhere and Prometheus primarily is a pool-based system or the majority of people use it with scraping and the default is 60 seconds. So, it gets data every 60 seconds from the game. We can fit, I don't know, four, five, six games into this time frame. And we tuned it a bit to 10 seconds, tried to get
a bit lower and then it stopped being reliable in our setup. And the problem is if you get a if you have a pool drop, that is when you're missing some data like I as a player have the feeling, I don't it's it's not responding to what I do. And the game will just silently ignore the data when it can't fetch it in that interval and it
just goes on. You don't you don't see that. You don't get any feedback. So, that's an issue. How does this look like, Manu? Oh, yeah, glad you asked. Oh, that doesn't look too interesting. It looks like this square wave. It does not look like there's a game been happening. This doesn't tell us anything. Okay. Can we do better, Manu, than that? Of course, you can do. The
thing is out of the box, Prometheus also supports pushing metrics. It's something that not many people do, but for such a scenario where the data is only collected for a short time, only pushed for a short time, it's a cool use case. we're now pushing data from the controller manager. We're aggregating it every 100 milliseconds to Prometheus. With more less often for other services. And that can
get us down to three, four, five hundred milliseconds reliable with all of the overhead and everything that's happening. So, it's already near real time. How does this look like, Manu? Can you show us a picture? It looks completely different. Now, I have played this game very often, but when I looked at this blue graph, I can imagine how the game went. At the beginning, you can see,
yeah, the people might have been very cautious in the beginning. There's not a lot of action, not a lot of acceleration that you can see it, and it gets more hectic. Then you have a couple of spikes where there's where there's probably the jostling going on, and then some players died, and one player won. But this still looks like very choppy, Manu. Can we get even more
data in there? >> choppy, and we also tried tweaking Prometheus to get more data, but then we also tried different backends. And, yeah. You can see then if we get down to 100 milliseconds, you also see that some of the spikes don't really look like in the picture before. So, higher definition data tell you more about the So, yeah. Yeah. That brings me to the topic of
the cardinality. We wrote in our proposal that we talk about high cardinality, but when we thought about it, so we didn't take a look at the game and did not thought what data we want to have in there. And in fact, we do not have high cardinality, and in fact, it's pretty low. I mean, the acceleration is just three axes, and that's really important for the game,
and that's what we wanted to track. The problem was more the high density. How to get in that much data into our projects. sending this all the time or pushing its metrics one after each other is quite is really really uh resource intensive. And luckily, all the open telemetry is there quite helpful because you can define your buffers on the SDK, you can define your batching on
the hotel collector, and this works really really well. there's also a couple of cases where it will fall apart if you were to implement Exactly. So, for example, we thought about adding just one ID to the as a one additional label to our metrics, the game ID. Suddenly, with 32 controllers, when we we would end up instead of 1,200 series in Prometheus with 18,000 series, which is
theoretically also from the writing side not a problem, but we wanted to display this in real time or in faster times with a timeout on Grafana with 500 milliseconds. And when you take a look at it at the metrics, we wrote a test for that. We already reached 400 milliseconds, which is with the computation in the browser, etc., it's not really suitable. For example, if you take
Victoria metrics, which is a drop-in replacement, you can have steady times for that. So, there is there are tools with what you with whom you can achieve this basically. You need to just be very careful with every label you are taking. And it's not just about the observability aspect. We also want to react to things in the game in real time. So, the faster we can get
the data there, the faster we can do something in the game with feature flags react to what's happening. I'm pretty sure you want to show some actual data of the game. we want to show you how all of this looks. And one thing this game or any game is very difficult to test because you want to have all these faults. And this morning, Simon even tried to
introduce some disconnects and port drops and he he from one edge of the conference to the other edge and apparently when there's no one in the hall, Bluetooth connections are pretty good. So, to make this easier, we introduced a couple of flags to help us inject faults into the game. in the slides, if I have a game running or even if I don't have it running, I
can select a couple of controllers where I want to have port drops or disconnects or acceleration spikes to help us with testing. And if you've not played the game before, it's also difficult to understand what this would mean for the players. Um so, everyone gets a color and when the light goes off, you lose, you're out. And if you have for example LED flickers, you get confused
as a player because you don't know what's happening. Have I died now? Am I Am I back in? What has happened? I don't know. Could also be that you have a full disconnect like old broken controller, batteries running out and so on. And there could also be acceleration spikes. Your accelerometer or a motion sensor could be badly calibrated, it could have other issues and whatnot. there you
can also see a super quick game because the computer plays for itself. Uh yeah. And the fourth thing, port drops, you don't see. This game has no UI. You have a bit of feedback via the vibration and the color and you have a bit of feedback via the audio. But if port drops are happening, you don't know that. But we see that in our dashboards then. I'll
first maybe go once again to Yian and to the traces. Yeah, the demo got a refresh. I'm still mirroring my screen. Are you still in the full screen mode of futures? That's a slide there. No, I see my screen. Okay. So, what we could see now and that's the that's more the important part. What we achieved is I said we achieved actually what we wanted to achieve.
But now we seen in in in the back end in our Grafana dashboards which Bluetooth controller is is working on which adapter. We can inspect actually our tooling. What is where we have issues. We currently also run can more or less run analysis if one controller has more connection issues than usual. Is it working? No. I've stopped the mirroring and added it again. Maybe we just continue
with the presentation and then we can play a little bit instead of showing the demo. Yeah, I cannot >> us and cannot continue. >> part is it actually works. So, what is going on? I plug it in again. Let's see. Yeah! Oh, I'm back. Cool. And I'm in the browser. Why don't I see? What actually has happened here? Oh, yeah. I'm back. That's good. So, do you
want to continue or should I? I will continue. What you can see here is the full game and you see also the actions and traces for each player. All of them receive for from our system just a random name like Gold Eagle and you can see all of the warnings when someone was very Yeah. Did a lot of action. You can also see how how long they
were in the game. And you can also see errors in here. So, for example, we here have a span with all of the poll drops that we injected before. So, we know for this player the game might have been different. Of course, this is a very small and short game example, but you can imagine how this would work for a larger game. And we have lots of
dashboards that tell us the connection quality, the signal strength, the battery history. Can also see the poll drops over time. We have a couple of dashboards again for the acceleration data. And what you can see here, for example, are also a lots of spikes that happened. These were the ones that we injected. It's more easy to see here, probably. You can see there's some things out of
bounds. This is not you as a player hitting someone with a force of 10 G. That's not very realistic. Or I don't know. I don't want to play against you if you do that. And you can also see disconnects where there's where there's no data. For example, here's here's a huge gap. Right. Is my clicking not working? Cool. So, so such a talk, I mean, we did
not do any kind of rocket science here. I mean, that's I have known for most of you. We're in the cloud native space. Everybody knows what we are doing. Most of you at least know what we are doing there with spans and traces and instrumentation. But it would would not be a good good talk when we would not talk about how we move forward from this presentation.
What can we gain out of that? And that brings us actually what we what we can do as a community. Because we feel like we have really cool tools which work really really well in a lot of use cases. But our documentation are only made for the cloud native world. When you want to start with that, when you say you have IoT, you have high frequency data,
and you want to make this work, where do you start? The hotel demo is a really, really cool place to start with. You have a lot of information, but it's just web services, web application in a way that you normally not use in a game or with IoT So, our call to action is do what we've done. Try these tools also on real-time systems. There's lots of
them. And document what works and contribute back to the community. And something that we also realized the past couple of days here at the conference, we're talking about real-time system. There's one thing missing and that's also as a side note or maybe for as an inspiration for your next talk. We have to throw in that now because everything is everything is about AI. If you want to
observe millions of AI agents and doing and seeing what they do, we'll also probably need fast observability and we have to see if we want to tweak our tools to also be able to do that. And yeah, that leaves us with a shout-out to the community and also to the creators of this game. This is a very old game. It was originally on the PlayStation 2 with
these controllers and it only supported two and then four players, I think. Yeah. And that was done by the good old fabric and there's a Python version, an open source version that we've used that you can tweak to have many, many more controllers if you want to do a daisy chain of Bluetooth adapters. And yeah, you can fork it. You can break it even more if you
want to. And please give us feedback. If you try it out and you see something, maybe maybe we can use this basically as the next IoT hotel demo because that's something everybody can easily use at home. It's a Raspberry Pi, it's a Bluetooth adapter, and PlayStation Move controllers. That's not something hard to get by. Generally, I think I thank you basically for being here with us because
it has been a a long week for most of us. Uh big shout-outs to you. Big thank you. Also, uh sorry. Also, thank you Manu to you for joining me on this journey with this little adventure. Thank you for inviting me. It was fun. I hope you enjoyed it. I mean, it's just fun little end talk basically. Uh if you have any questions, feel free to ask
questions. There's also our social media information. You can also find us on the CNCF Slack. We're more than happy. And as we are really, really fast this time, if there are no question, we also can play with 80 with 20 people. 22 people, I think. I did not count the controllers. Thank you very much. Are there any questions? So, who wants to play? Come on. Yeah, come
here to the stage if you really want to play. It's a Do you want to drive it for me? Cool. Thank you. Have Have a great way back
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32