Great International Developer Summit (GIDS)

Scaling Systems for a Billion Users: What Breaks, What Survives - Sarath Chandra Kummamuru

16:18 · 21 Apr 2026 – 24 Apr 2026 · YouTube

About this talk

This talk focuses on the Bharat Bill Payment System, a platform that facilitates bill payments in India. The speaker discusses the scale of operations, noting that the platform processes nearly 3 billion transactions yearly, highlighting the importance of interoperability and collaboration with various partners. Key challenges at scale include ensuring system resilience, handling failures gracefully, and maintaining robust observability across the ecosystem of billers and customers. The speaker emphasizes the role of AI and event-driven architecture in enhancing user experience and simplifies integrations. The talk concludes with a vision for future development focusing on open, interoperable systems and the commitment to earning user trust through reliable services.

Full transcript

Excited to be here. Humble to be here. I've been an engineer. I've been a product person for all my life. Been an agile coach. And learned a lot in the last 20 odd years that I've been in business. I right now I'm with NBB We are in Bharat Bill Pay System. Uh we have a few platforms that you must be using. So before I talk more about

today's session, how many of you pay bills? Raise of hands. How many of you do them online? Most of us. And I would believe that a lot of you about 10 15 years back weren't doing it online, right? Weren't paying the bills online. That's where we come in. We are India's biggest bill payment back own. We are a uh DPI for India. We are an NPCI company.

On a lot of the learnings that NPCI has learned with UPI, IMPS, and many other payment DPIs, we've picked up the use case of bill payments and that's what we serve. Over the last many years, we're very proud to see that you have adapted us very well. Look at the numbers. Very powerful numbers to look at. Almost about 3 billion transactions in the year. And a few

trillions in rupees that are being passed through our platform. We are a DPI for the citizens. That's what we are for all of you. Right? What is our vision? Open interoperable platforms. We believe when we invest in open interoperable platforms, bring partners onto the platform, they do the magic. The magic doesn't happen with us, but with companies and developers like you, engaging on our platform, solving customer

problems. That's what our this thing is. Now, why are we doing all of this? How do we make sure that we are scaling? And look at the scale, right? We possibly are ready for about 10,000 TPS on our systems. How does that happen? And what are the various infrastructural or architectural aspects we have uh experienced? By the way, how many of you noticed this on your newspaper,

the bill payment times? It was on the newspapers a couple of uh days back. And this is something that we're very proud of. Whether you are in uh any city in India, whether you're paying an electricity bill or you're paying your loan EMIs or credit card bills, we are enabling that along with a lot of the partners that we have on our In a very simple ecosystem,

but have a lot of uh partners who are participating with us. Overall, we are Bharat Connect, the bottom layer which is supporting all of our uh partners. We have consumers, all of us on one side, and we have billers on the other side. That's basically what we do. And there are something called we customer operating units and biller operating units who participate, connect with us, and work

with this. We have I'll just remind you some numbers. We have almost about 25,000 billers on our system. We have about 7,000 apps. Various apps that you're using are on our platform and using our platform when you pay your bills. And then we have a few uh millions of customers on a regular basis. What's the challenge? And this is what we focus on. The scale. It's one

platform, thousands of partners, billions of customers. We need to make sure that all the entities in our listing are ensuring that they're delivering they're getting critical value from the platform. Otherwise, why should they be on the platform? And we also ensure it is resilient. So, the question is if so many people trust us, so many entities run their businesses on our platform, how can we ensure that

under real uh unforgiving traffic population scale, how do we sustain it? And that's what I'd like to talk about today. Challenges at scale. Uh most of you are developers, and you must have faced it in your own enterprises. A lot of enterprises run at scale. We are saying we're bringing multiple such entities onto our platform, and hence the scale multiplies. So, how do we What are some

of the key things that we've seen? Every API call that's happening, whether it is a financial API or whether it is a non-financial API, coming in from all these billers, BOUs, going to the millions of customers, everyone is generating data. That's a lot of load and something that we need to manage reasonably well. Timeouts. No system is perfect. I hope all of us as developers know that

any system is going to fail. And how does the system or the service ensure that failures are handled gracefully and recovery happens is extremely important. When one of my partners times out, I cannot say the whole platform and the ecosystem is down. How do we handle that? Cascading failures. One of the partners, one link somewhere in the ecosystem fails, we cannot have a domino failing across the

system, right? And then, lot of us invest in observability. How many believe that observability is our great savior? That's for us in individual enterprises. Now, what we're talking about is at an ecosystem level with 22,000 billers, 25,000 billers, possibly about hundreds of BOUs and 700 apps. Any failure, if you're not able to observe, I am not in a position to inform my partner that there is a

challenge and there is a customer impact. Hence, ecosystem level observability and end-to-end observability is something that we believe is extremely critical. And these are no longer nice to haves. And neither is performance an engineering afterthought. Our product managers have to think about how this will scale. Our operations teams have to think about how are we going to manage scale on a regular basis. So, neither of these

are nice to haves or afterthoughts. Some architectural principles, I think some of you yesterday would have heard one of my colleagues talk about trust and security. As a DPI, as a critical infrastructure for India, we need to ensure that everything is safe. One big failure and all of you as customers will stop trusting us. That's a challenge. Security. We cannot be an insecure system. We cannot allow

our partners to be insecure in any way. Safety and compliance are important. While these are our core architectural principles, we have some others. Like all of our architecture is composable, pluggable. We have a lot of context. That context defines the configurations and how our systems behave. We are observable, like I said, end-to-end. And then we have to be performant. With these architectural we still believe that keeping

things simple is what makes us effective and successful. And all of us know, keep simple, keep things simple and uh stupid is one of the first principles we've always learned. And that continues even at For example, our APIs, we have various resources or entities that we expose through our APIs. All of them are asynchronous. They have a request and a response system. That's all that any of

my partners need to understand. Keeping it simple that way ensures integrations are easy and developer experience is much better. Along with this, we believe that a domain model and clarity about what is the domain model for bills as an ecosystem, as a space is extremely critical and we have standardized that. So, when a partner talks to us, a consumer talks to us, our own engineers and products

talks to each other, we're talking the same language and that I think is very powerful. Interfaces uh are become the future APIs and then clear domain boundaries. Event-driven architecture, I don't think I need to sell this to you guys. Event-driven architecture has been something that is a saving uh light for me for the last decade and a half. And I love how some of the distributed architectures

have helped. And this is really, really is deeply entrenched into us. All of our APIs and systems are event-driven in that sense. A quick other thing, build for developer experience. Keep things simple. Make sure the developers are able to build on top of our infrastructure. Whether it is our internal developers or whether it is our partner developers. I'll call out one example where we want to be

able to go deeper. All of us in enterprises would have built our APIs. We expose client SDKs. Yes or no? Are client SDKs good enough? And what we believe is while client SDKs are good, they're only partially helpful. Once I give a client SDK, my customers or partners need to integrate these client SDKs within their own ecosystems. That takes a couple of weeks. That takes a month,

depending on the kind of partner we are talking about. How can we reduce that even further? And that's something that we work very closely on. And possibly, this is a place where AI can help us. Can we build tools that our partners are able to use which leverage our SDKs and understand their domain and help their developers build faster. Talked about end-to-end observability multiple times, but there's

something that I wanted to call out here for a couple of minutes. In our system, we define what we call observables and observers. There is one goal for my engineering organization support organization. Can we catch issues before a customer does? And that's the only thing that I am concerned as an engineering leader for my support and operations teams. Systems will fail. Issues will happen. But, if we

are able to catch them early in the cycle, before a customer can complain, I think we are doing pretty well. And that's how we're looking at observability. What are the things we can observe? We look at system metrics, service metrics, traffic, and then business metrics. So, before a business metric fails, there are a lot of indicators within the system telling us there is something that's going wrong,

possibly going to go wrong. Observe those. Make sure that there is enough tools and availability of various people with SOPs to understand how they should observe what needs to be observed. And this is extremely critical for us. We need to talk about AI. I think AI is an amazing disrupter. I love it for what can be enabled with AI. Michael was talking about a lot of the

stuff. I believe very deeply that what he's talking about is very powerful. If you have a 100 people, they're doing something that they're doing. With AI, with the right first principles, and ensuring that the mental model and the control is with us, we can use AI to build great things. I think this is a very powerful enabler. One for the developers, but here I'm looking at a

different use case. To pay bills, we use mobile applications. There are these 700 different apps through which you can pay. But a rural Indian might not be able to look at uh English application. While we've talked about making sure that applications are multilingual, the opportunity I believe with AI is language and voice. So, if like you, all of you are paying your online bills with some app,

if a deep rural person is able to pay their loan EMIs just by talking and saying, "Hey, how how much is the EMI this month? And let's go pay it." Just with voice, that's a huge distribution enabler for It's obviously a lot of agentic opportunities exist. And also, simplified integrations like I talked in the other testing. Does it mean that this is only for distribution and other

business use cases? No. I believe this is very powerful for us and our ecosystem developers. We look at it to improve our product and research. We use it for development and testing, especially with agents that can help. And then operations and support I believe is another area we're very deeply involved with leveraging AI. Again, it's good to see that Michael was talking about us developers understanding the

mental model, the context, and the domain. And using AI. And I would love to look at this serendipity that this is something that I have also called out. And I believe this is how we're looking at AI. And this is a very powerful outcome for us. Quickly to close, the future we believe is open, interoperable, and scalable systems. All of this is not magic. It needs disciplined

system design. That doesn't go away. What we have learned over the last couple of decades or more about how to build systems doesn't go Performance needs to be our first culture. We need to think everything how about how do we build for performance? Resilience doesn't come outside of the architecture. We need to start thinking about an architecture that's inherently And then we believe open, consistent, and interoperable

this thing is the future. Something that I want to leave all of you with. We believe that if we are doing our jobs well, a billion people might not know our names, but they will trust the systems that we build. And that in Hindi, Aap Humein Mante Ho Aur Jante Nahi Ho Magar Mante Ho is a video campaign that brand campaign that we've been running, received very

well, which indicates that while a lot of you are using our systems inherently, we don't we're not upfront saying this is us who's helping us. We enable the apps which you recognize, and then allow you to pay bills. So, while you might not know us directly, we believe you trust us deeply because of all the numbers that we're seeing, and I'd like to thank all of you

and your families and your friends for using deliver much more for you guys. Again, for anybody who doesn't understand Hindi, you may not know us, but you do trust us. is always something that we are pushing ourselves with. This fact that this country trusts us, and we want to deliver more for all of you. As engineers, I'm looking forward to anybody who is interested in building for

India. I believe the next two decades belong to India if you do the right things and all of you are welcome to talk to me, talk to my team, and engage in any way. We'd like to hear feedback, we'd love to see how we could leverage some of your knowledge better, and if you could join us, even more happier. But, thank you. Sharath Chandra here again, and

wonderful to be here. >> [music]