DEVWorld 2026

Srini Srinivasan - Engineering Productivity Starts with Predictable Behaviour

26:32 · 07 May 2026 – 08 May 2026 · YouTube

About this talk

This talk by Sreeni Srinivasan, founder and CTO of Aerospike, focuses on the challenges of engineering productivity, particularly in the context of building and maintaining consumer-scale applications. He discusses his experiences at Yahoo, where scaling mobile applications led to unexpected operational difficulties and declining team productivity. The speaker emphasizes that sustaining high productivity involves not just developing features but also ensuring that systems operate predictably despite the volatile nature of user loads and market conditions. He reflects on the complexity of modern applications, including the need for optimal database performance and the impact of caching strategies. Sreeni also highlights real-world examples, including the TIPS instant payment system, to illustrate the importance of building resilient infrastructures that can handle unpredictable demands while maintaining efficiency.

Full transcript

Without further ado, big big round of applause for for Sreeni. Um >> [clears throat] >> Thank you, Alex. I can't really see [snorts] the people here, so the lights are fairly bright, but uh good morning. Um it's a pleasure to be here uh and speaking to developers. Um just a little bit of a uh introduction to myself, you know, my name is Sreeni Srinivasan. I am founder

and CTO of Aerospike. Uh it's a NoSQL database. I'm not really going to talk about Aerospike much today. I'm going to talk about the problem of productivity of um building products. You know, I'm an engineer, you know, an original uh kernel programmer, C programmer, basically, an art which is probably dying out, but with Rust maybe some of it will be um reinvented. Uh But fundamentally, I'm going

to talk about uh productivity and how um and that's the work we have done at Aerospike. We have tried to build predictable systems which uh serve uh a number of consumer scale applications. I'm pretty sure uh many of you in your daily kind of life use um applications, consumer apps under which Aerospike is used. I won't go into that today, though, but come talk to us, uh

myself and Dirk, um he's, you know, uh after the talk if you want to learn more about us. So, let's get to Okay, now this is not working. Okay. first I'm going to give you an anecdote of what actually happened uh to me uh when I was working at Yahoo, you know, uh I'm a very early um I guess uh adopter of mobile. We built a lot

of the mobile technologies all of you take for granted today. We built those in the early 2000s, you know, 2001 to 2005. And then I ended up at Yahoo, and um, we ran into a problem. So, we were just, uh, scaling up mobile apps, and I inherited a team which was building all kinds of interesting mobile But, what happened was, uh, a lot of the problems that

we talk about here, right? I mean, we we we when you talk about productivity, you talk about how many people are there in the team, how many engineers, what are you coding, how are you adding features, things like that. But, what we ended up doing was we would release a, you know, a particular version of the product, and then after we released it, uh, because of the

scale, we were just learning to deal with scale at the time. So, we didn't know what the scale would even be, you know, this is the very early times in mobile. So, a couple of engineers will be taken out, and they'll be doing operational work. And then, so now you have a five-member team, you end up with like, after the first release of the product, you have

three people left. And then, after the second release of the product, you have one person left. So, this is kind of how things used to happen. So, productivity was really affected by the fact that we had to operate our own systems at scale, and it was all unpredictable in terms of what's happening. So, when you talk about productivity, you and especially I'm talking about productivity for products

which are running at consumer scale. Think of all the people uh, in Europe uh, using a product on a daily basis, you know, it could be tips, I will talk a little bit about the tips project later. Uh, that's that's in Europe, but we have we have running things worldwide, you know, like Barclays uses us for um, fraud detection, for example, and so on. essentially, the main

thesis is engineering productivity doesn't come only by developing products, it also comes by running these products well in the presence of volatile situations. How do you produce products which provide this kind of predictability of performance across various things that happen over over time? So, and and you can basically come up with Here's an example of I'm talking about millions of consumers interacting with these products, you know,

on your phone, on your apps, and so on, right? And when you do that, there are a whole bunch of components to these systems. These are complex systems. You know, you end up having a database. Sometimes the database is not fast enough, so you put a cache in front of it. And then you have an application which is also caching. And there are, you know, as you

know, you know, I'm a database guy. You know, I have a PhD from like, I don't know, a long time before many of you were even born. But the point is what I have done over the years is learned that with the consumer scale that comes with it, databases need to be transformed. And you know that there are a lot of databases which have been invented in

the last 20 years, and many of them you probably use, like Cassandra, Scylla, Redis, and hopefully also Aerospike. But the point is that means that we are solving problems which are evolving in terms of the workloads, in terms of people using them more frequently because of mobile you know, device adoption worldwide, and so on. So, there are a lot of these different components which come in, and

consumers essentially are affected. But what actually happens is a lot of unpredictable things happen in the system, and things keep changing. And how do you keep your system to be predictable in the presence of things which are not under the system's control, okay? And what happens in those things is workload can change. You know, take you know, we we have customers who use brokerage, for example, you

know, where you do trading, and various things happen to the stock market, especially with the the recent situations of the oil prices, and so on. It's like every week you have a crazy day in the stock market, right? So, if you want to if you have a company which is delivering, you know, if you're developing, for example, uh a margin loan application, now how do you handle

margin loans in the presence of this kind of volatility? You need to do real-time computation of these risks, for example. You know, and then so the workload changes a lot, you know. Uh the cache itself becomes less effective because um caching fundamentally is a losing game, you know, because I am actually a database person. If you can build a database which is as fast as a cache,

you're better off. And that is actually the project I've been personally involved in for over 17 years now, and that's what Aerospike is. And we're not talking about that, but fundamentally caching is a losing game because you could lose data if the cache fails, and there is, you know, and so and you and users depend on the, uh you know, more and more on the real-time SLAs.

So so you but but caching becomes necessary with a lot of databases out there, as as you very well know. And then uh sometime a background process starts, you know, compaction of the database happens sometimes. Um people, you know, refresh uh operating systems a lot. Many of the Many companies have the requirement that you have to refresh your operating system once a month, which means you're running

a system 24/7. System can never go down. It has to be up all the time, you know, at any point in time if you want to transfer money to somebody, if you basically want to book an airline ticket, or if you want to do something else like get an Uber, you expect things to happen, no matter what the time is. And that's where a lot of these

technologies There's a lot of unexpected behavior, right? I mean, I don't know um uh you know, I travel a lot, and very recently I was stuck in a particular situation where some flights got delayed entering into the airport I was going in, and then everybody arrived at Uber at the same time. It and and we we had this enormous traffic jam, and essentially it took me, this

was in Las Vegas, it took me like an hour and a half for a 5-minute uh Uber ride to get out of the thing, you know. So things like that happen in real time, and and and we have to deal with it. Um and then there are simple failures like a node dies, dependency, you know, like for example, when you're doing a complex application, for example, ad

tech if some of you are familiar with ad tech, you know, typically there are these big systems like Google and Meta which which an ad tech company like Adjust for example, which is based close by or ad form, they actually work with these large providers. So one of those providers could slow down, which means your ads won't show up, which means they don't make money. So these

are things which you have to deal with. And the other one is, you know, traffic always changes. We already talked about it. You know, it could even based on today's oil price, a lot of traffic can change because of some other event that happened. So So the main thing you want to think about it, what are these operational issues? And then why do they why do you

lose productivity? You know, what actually happened in my case when I was at Yahoo is is a situation like what is shown up here. You know, suddenly there's some latency. In this case, this was before my work at Aerospike. So I was using relational databases and various other systems which are pretty standard in terms of how they run, but they did not scale up. This was before

the days of Cassandra and Redis and all that. So we had to put caches in front of it. So essentially what you end up with is like you know, latencies go up. And and the worst part is when a system goes down, this is how Yahoo used to work in those days. I'm sure every other company works that way. If it was the system was down for

15 minutes, it it goes escalating up the chain. So if if your system has been down for about an hour or two, the CEO gets to know. And then they're all in and I'm a developer trying to keep the system up and then you have enormous pressure, right? And and you need in order to do that, what you end up doing is you end up fighting fires

all the time. And you the goal is to avoid it, you know? And and and in order to avoid it, you have to design systems properly. Um but fundamentally, there's a whole bunch of things that happen, and this could this results in lost productivity. The end result is the team which is building the system also has to maintain it. It's an important principle. And then, if you

spend all this time maintaining the system, you don't get a chance to innovate and compete in the marketplace. So, your users essentially will have a really, really bad experience. Uh I'll give you an example, right? Um if you have a failure, which is rare, like a node going down or a data center dying, right? And then what happens is you end up having to affect users. Let's

just take a very simple example, okay? 2% Now, this actually happened to us in PayPal. We had a use case in PayPal where PayPal um computes the fraud score for essentially um every transaction before the transaction is committed. These are millions of transactions that are happening, and the time you have to compute the fraud score is about 100 to 150 milliseconds. If you don't have a fraud

score computed, you still have to let the transaction through because there is no reason to stop a transaction because you were not able to compute the fraud score. Now, when that happens, if you've missed the fraud score for 1.5%, that's basically all they were missing. 98.5% of the time they were able to get the fraud score. But if you have millions of users using the system, thousands

of users every minute or even more even more every minute will be facing this kind of behavior where you're not getting the best user experience. This results essentially in a lot of complaints on X and various other platforms, and then you end up in the news, okay? So, it's not like, you know, uh you 1.5% is too much is what I'm trying to say when you have

millions of users using it. Okay? And that's fundamentally something for you to keep in mind. Um that that's kind of And therefore, uh you can't just assume that oh, it's a very rare thing, you know, a node fails, you know, my timeout is just 1.5% it's okay. It's not okay. Okay? And then, there's a gap between what system metrics indicate and what users actually feel. I mean,

this is essentially really important to keep a pulse on the end-to-end interaction of users. If users are not able to uh uniformly have a great And And there's actually a way I used to say it, you know, uh about 10 years ago, either first user and the millionth or the 10 millionth user or the 100 millionth user need to have the same user experience. In consumer-oriented deployments,

that is important. And you have to do it from day one. You can't actually add to this later. Okay? That is impossible to do. And that's fundamentally the lesson that And you have to always think about the end-to-end interaction. As As a developer, when you're writing code And the same thing happened to me when we wrote Aerospike, we had to keep in mind the fact that we

do not And this has actually been one of our successes is as the system scales, you know, we run systems in Europe, in North America, um in India, in Asia, and so on, uh where there are potentially hundreds of millions of users using these systems. Uh and we are not actually the one running it. Our customers are the one using our technology to run it. But that's

These are the problems that we actually work, uh you know, uh we've been actually worrying about and figuring out over the years. Um the other problem with what Aerospike says is if you don't actually think about the hard problems at scale when you start building a system for consumer scale applications, it's really important. Okay, there are And And even enterprise grade applications now are getting to consumer

scale because of the use of agents. Because each of us can have 100 agents, which means that uh 100,000 or 10,000 enterprise 10,000 people enterprises can potentially have a million agents. You know, we have some use cases we've already seen uh in some of the pioneering companies where enterprise workloads are requiring operational databases, for example. But, the point here is the workarounds that we build because we

didn't think about it right and build the systems well, create many more problems long-term and it ends up in your entire team essentially maintaining work around work around as an architecture. And that is a real disaster because what's happening is if you can't design things right from day one, you end up in all these compromises and inventing things on the fly, you know, I I've done this

myself. So, it's not like a criticism of anything. It it happens because there are business reasons. There are you know, you you launch something, your app, you know, it could be a game, it could be something else, it takes off and now uh now you need to basically make sure that all of this actually works out. So, the hidden cost of engineers, right? I mean, essentially what

happens is all of this means that um you basically have your best engineers working on the worst problems, basically. And and that is fundamentally not something we want to essentially focus on, okay? Firefighting is great, but the most important thing engineers should be able to do is to work on problems that they that they can solve for all their customers, add features, and so on. But, if

that's not what you do because you're actually having an architecture which essentially doesn't work really well, that is fundamentally not a great issue, you know, to be focusing on. Defensive engineering is something we people do a lot and I you know, I've even done it quite a bit and I'm sure many of you have done it. It is not actually a long-term win in any way because

things system changes really really well, for example. Now, I'm going to take you through a few small areas of performance, actual performance, and there are some QR codes, okay? So, you can actually um go download these things. Now, this is this is just to illustrate what is going on. The first thing is when you have unpredictability, right? I talked about workloads and unpredictability. This particular uh benchmark

is about Cassandra. You know, there's some details here, but fundamentally it's a cluster running about 50,000 reads per second uh and 40,000 writes per second. And that's the other point I want you to remember is it's not all read-only, okay? Or read mostly. It is a mixed workload with 50/50 read-write workloads. And this means that what you have is uh when the workload goes up, you see

that the latency The latency is shown on the right. It's It's It's a 10 milliseconds. And then it ended up at 15, you know, it it and it changed forever. Once the load went up, the system's uh behavior uh it is kind of predictable, but it changed, okay? So, as as this is still available, right? I mean, there's a CAP theorem. What is availability, right? Complete unavailability

um is basically the system is not even available. But then, if you're available at a slower pace, you lost some availability, also. So, I think I think this is this is a benchmark which is worth looking at. And then what happens is uh that change happened when the load changed. And then it does not, you know, come back to the mean. It just stuck in the new

situation where the And also, it's all scattered, the the whole thing that's happening. This one is ScyllaDB is a more um is a more efficient Cassandra. It still has the same architecture, uh but you can seamlessly it's written in C like Aerospike, for example, you know? But uh but then you run up, you know, and and it actually looks healthy. And and essentially what happens is over

time, though, you see this particular workload is slowly deteriorating over time. You can measure this, and the end result is if if that is actually happening the headroom is actually disappearing. You're going to pay for it at some point. It's important to realize that early and and and make sure that you don't work on. Again, these benchmarks you can download from with a QR code and then

you can see you know for yourself more details. You know, I don't have time to go through all the details here but there's a lot of deep technical work that we have done over the years. Um and to give you an illustration, okay, this is the one thing where I can illustrate why Aerospike even exists as a company and what we do. If you see the original

this is this is coming from uh essentially um the use case of company um called LexisNexis. It used to be called ThreatMatrix. But what they do is every time you many of you log into a bank account uh the bank reaches out to a service like LexisNexis and asks them, "What is your level of confidence that the person who is logging in is exactly who they say

they are?" This this answer has to come with a very short amount of time, 50 ms or so, otherwise your login is going to be blocked. And in fact it's actually not going to be blocked. If if if it says no, it's it's not the confidence is low, your login will be blocked. If not, uh essentially the company which is being asked, if they don't respond with

a proper answer within 50 ms, they don't make any revenue. They don't get paid. So, the fundamental thing here is uh when they're running in Cassandra they have this thing that they can miss their SLA if the load goes up. You know, and that is not acceptable. And then what you see below is a system which is basically based on Aerospike which they ended up launching. You

see how predictable it is and how low the latency is. It's the same workload, same amount of data. And again, I'm not going to go into the details of Aerospike technology. Come talk to us after this and I can I can tell you what it is that we do. But there are a lot of technical innovations which go into maintaining predictable system. And that will increase productivity

for the people at LexisNexis or ThreatMetrix because they don't have to worry about you know the kinds of things that could happen where the load goes up and they have to deal with this the running system. But the predictable behavior is super important. Now, if you look at the the other one we looked at right the other you know the ScyllaDB one ScyllaDB is pretty predictable. Okay?

But essentially if you run the same kind of system with Aerospike you can do it with fewer nodes. And again that's because of the technology and the way way we leverage SSDs. But the point is what you end up having is a million transactions per second as opposed to 400,000. The capacity is higher in in a more predictable system. And that's fundamentally really really to kind of

start summarizing when systems are predictable an engineer can be more confident of how the system will behave. Which means they know what to expect. They can make changes with confidence. This is a big thing. Upgrading systems in such a way that you can actually run the system you know and keep upgrading it with your fixes and features is really important. systems which are actually I would say

the short-term things that I talked about where they can't actually upgrade the system because the system isn't working well enough. It's not operating properly. And that's actually a huge problem. And and the fundamental thing is you know you need to know if every change is risky or not. And if you run a system which is predictable in the presence of various volatile changes that happen like failures

and so on then you can actually I mean I'll just sum it up. You can basically sleep at night. There have been people which have used Aerospike over the years who have come back to me and thanked me that they are able to sleep at night because running a system 24/7 for your customers can be really really traumatic for the engineers if it is not architected properly

and that is really important to This is I have a couple more flight slides but I want to spend a little bit of time on this one. This is a real use case which is being used in Europe since November 2018. Uh there is an application called TIPS. You know, essentially this is I just took this from the website yesterday. Uh it is basically the target instant

payment system settlement and what it does is it allows an instant payment originator anywhere in Europe. You know, it will allow you know, payments in euro and the Swedish and Danish currencies. Um basically user a person with a bank account somewhere in Europe uh can send money to else in Europe uh in whatever small denominations you want to send. This has been going on since 2018 and

the reason this is interesting is you need to reduce the transaction cost. You can do all of this with um other systems, okay? You can build a system which is super inefficient and unpredictable and build stuff around it. You know, it is actually fairly good job security for people like us who are developers but that is not actually doing anybody good service. Whereas they ended up picking

Aerospike for example. Why? Because they were able to do the TIPS is in the middle. It is able to one once the payment is initiated from one side, it manages the transaction for the next second or so until you the money itself is transferred. Okay? And that's fundamentally something um which all of the principles I talked about uh that's been used. There are many other use cases

uh worldwide and also in Europe but this I thought would be interesting in case any of you do instant transfers and you can go to the TIPS webs web page using the QR code there, you know. So fundamentally, right? Um um developing products fast doesn't really help, especially for consumer scale applications, unless you are able to run it predictably and continue to enhance it. And that's the

other thing. There are There are two things here. One is there is a real-time thing, like when you're transferring money, you want to make sure it transfers properly, and you're able to get um the answer, you know, using SMS or whatever, immediately, so everybody's happy. There's another thing. When you add applications like this at consumer scale, you want to be able to add more and more applications

because that's what consumers expect. Like it or not, you know, I'm I've become a consumer myself, you know, since 20-plus years ago when I started working on mobile. Um my level of patience is very, very low for any of these things that happen in real time. I mean, even getting an Uber, you know, you just switch it around. I mean, if it doesn't show up, you know,

you take a taxi. I mean, it's like nobody is actually patient anymore. Okay? And And that's not a bad uh it's just because we can spend our time on other more useful things. Uh that's That's the way I look at it, at least. So, so the So, the way to protect your productivity is to make sure your system's running at scale for consumers work really, really well.

And uh and are predictable. Finally, right? Uh I want to stop here. Um I think okay. I think we're coming to the end of it. So, what do we have here is uh okay, I'm wearing this T-shirt, you know? So, we we try to think of us as dragons dragon slayers, and I'll just I'll conclude with this story from one of our customers. This actually happened uh

in 2014. There's a company called AppLovin, um based out of the US, started using Aerospike. And at that point, they were actually using Cassandra. And Cassandra is a great product, actually. It scales really, really well. And I have spoken to the people who wrote Cassandra and so on. But, the point is it is not a real-time system. And they were using ad tech. And they And because

of that, what happened was when the load went through the roof uh and the way this person's name is John Cristnak, he was the CTO of uh um Applovin. He came and told me, "Srinny, then they switched to Aerospike and then he said, 'Hey, you you killed the dragon in our data center.'" I said like, "What are you talking about? What dragon? What data center?" He said,

"No, when when the load went up uh he he this particular database uh would just breathe fire in his data center and use up all the resources." So, he kind of and said like and and then he switched to Aerospike and because of our technology, we were able to with the same load that they were growing because the whole point is it's a growth, okay? It is

not that they had a load which is predictable. The load itself is unpredictable. The system has to be predictable. And that's that means our system made it more predictable and they started basically reducing the resources they needed in the data center to handle the same load that they were doing with a more inefficient system. And that's what the slaying the dragon meant is basically the dragon is

the unpredictable load, but the predictability is what slays it. Okay? And then we have a lot more of this and if any of you have been slaying dragons, come come to us, right? I mean, both Durk and I are here um for, you know, until well after lunch today. So, come find us and uh if you basically give us a nice dragon story, Aerospike, uh we will

arrange with our marketing team to get you one of these t-shirts. So, uh if you're interested, okay? So, thank you for listening to me and coming so early after the late party last night. I appreciate um the the opportunity here. Thank you. >> [applause]

From event

DEVWorld 2026

07 May 2026 – 08 May 2026

All event videos
Back to Watch