About this talk
This talk explores the journey and current state of Bharat Connect, a platform developed by NBBL, a subsidiary of NPCI, which facilitates bill payments in India. The speaker details the architecture of Bharat Connect, highlighting its role as a switch that connects over 20,000 billers and allows for transactions through various channels, including mobile and web. The platform handles up to 170 million transactions daily with a focus on maintaining high availability and low latency. The complexities of managing different billers, ensuring seamless service delivery, and the technological challenges encountered over the past decade are discussed. The speaker emphasizes the importance of domain-driven design, event-driven architecture, and robust reconciliation processes to support a scalable payment ecosystem. Additionally, the talk outlines the strategic use of open-source technologies like Kafka and Postgres to enhance operational efficiency.
Full transcript
I I hope many of you, I mean all of you in this room would be using UPI. Right? We do have a platform called BP S which is which had been rebranded as Bharat Connect which uses very similar kind of architecture but slightly more complex. Okay? So, I'm going to talk about our journey where we are right now a bit and what worked for us. How we
have we are able to build this particular is what I'm going to cover in this off an hour. Okay? So, we do have a three products and from NBBL NBBL is a subsidiary of NPCI. So, we do have three main products Bharat Connect which is a bill pay system. That's what I'm going to talk about. We do have Bharat Connect for business specifically for businesses which do
integrate with quite a lot of ERP products facilitates in information invoice exchanges and payments and we do have banking connect which is something that is an upcoming product but equivalent to IMPS but targeted different audience. Okay? So, getting into our product, right? So, we do have I think most of you would have done bill payment. You would have done gas booking. Recently IVR systems went for a
toss. People who had used Amazon GPay they would be able they would have been able to book. Right? Have you noticed when you make a payment there is an ID that gets generated which informs that there is a bill reference ID and this is from the system and that backbone is us. Uh the bill pay system we do have is a switch that works in the background
that connects over 20,000 billers. Uh quite a lot of agents, crores of agents, foot soldiers who go to India three, the villages, collects payments and uh uses our system to settle. And we do have as I said multiple channels, multiple agents. You use your phone, banking channels, web, quite a lot of channels are available for us. We work 24 cross 7 obviously and our latencies are in
milliseconds. So, what we have done? Have we been here from day one? No. We had learned, we had made our mistakes, we had major catastrophic failures as well before we upgraded our technology. Uh we had those journeys in in over a decade uh that we had been in existence. Uh we had all all those journeys and the learning that we had uh they are not I I
won't call it as transformative or path-breaking. They're all small things that works. All right? All of them are needed for this system of this scale to work. We process about anywhere between 130 million to 170 million per day. Okay? That's the amount of transactions that we do. we are probably about 1/5 of what UPI does and we will probably be also be in top 10, top 15
in the world. Even though not many of us not many in this room will know about us. Right? Uh so, that's one reason why I thought we will come to this and talk about the system that we have built, which is very interesting, which is probably useful for everyone to understand and probably build Yukon that way. Yeah. So, what we had done uh It It It looks
very simple when you go use your GPA Amazon, say I want to book a gas, it looks very simple. But, there are multiple legs here, multiple places where it can fail. Okay? As we had said, it is deceptively simple, but really really hard for us, Lot of payloads, lot of categories, it needs to work seamlessly. When a new biller comes in, the existing system cannot break. It
needs to be flexible. These are all problems that have been solved, but it needs to be solved at scale. That is where we had the challenge. Uh money in motion, 24 cross 7 sections keeps happening. I I cannot declare a downtime anytime, Scale with consistency. We do operate at 99.999 availability, uh consistency, uh our system don't fail. We might have problem with our ecosystem partners, and I
have to help them as well. That's what we are looking at, right? very simple, nothing fancy, I When I mean nothing fancy, nothing new that I'm introducing here. This is what we had done. It is not where we started. This is where we are right now. Okay? So, we have domain driven design, event driven architecture, and there are lessons that we had learned and data platforms. Okay?
One of the main problem for us is also we scale. And we are technologically advanced, but the partners that we work with it there are banks, cooperative banks who are very new. There are technology players who are coming in new, who don't have the scale, who don't have the money, uh who don't have the techno-how to actually scale, right? And we have to support them. And that
is one of our core principles right here. Okay. So, what we had done with respect to domain-driven design, uh we had classified it as very simple six basic domains. How do I onboard an entity? There are multiple entities in in our system, the banking partner with whom we will do the settlement, uh Amazon, Google Pay who works with the banking partner, the agents who gets onboarded, the
foot soldiers. Uh there are banks who connects with the biller uh for them to enable billing, settlements, and the biller themselves. So, there are multiple players here. All of them need to confirm to the contract. All of them need to understand the language so that they can seamlessly get added. So, onboarding is a very critical area for us for them to confirm so that they don't end
up breaking the system, right? We do have a fetch very simple. When you go look for your electricity bill, it gives you the exact amount. When you do a Fastag, you know the balance before top up, right? So, you need to fetch. Then you you make the payment. It looks very simple, but there is a context. Not all billers operate accept any amount. For a Fastag, there
is just a simple principle that you have to do it in multiple of 100 or a minimum amount of so much. Whereas for a credit card, some credit card biller will come in and say, "You can pay a minimum amount or the full amount, not in between." But, some people will say, "Hey, you can pay any amount." Even if the bill is not generated, you can go
out and make a payment. So, each biller have their own logic. And as a switch, I need to be able to accept that, handle that, and pass it on the apps of your choice to be able to understand that language and pass it on. Right? So, we do have to do that. And settlement is the most critical part. Billers come to us. Customers come to us comes
to the platform the GP or Amazon because they know if a payment is made, the service is available for them. The payment is successful, the service should be enabled. If for some reason there is a failure, which can be 0.001% or even lower, the money gets refunded quickly. The challenge that we also have is we work with multiple categories. Say, for example, electricity will have their own
dispute resolution mechanism. They also have a regulator and they have to stick to it. Credit cards will have its own regulatory framework. We have to have a model which works with all of them. And that's where the dispute also comes into the picture. You will need to confirm to them and be able to help individual categories, individual fields in which they operate to be able to come
to our system. Uh by the way, can you guess how much of what percentage of electricity bills we process through for the whole country? Any guess? Seven. Nine. Sorry? What percentage? We do about 40%. A lot of queries do happen, but think about it. India 3, the villages still are not paying through the phone. Which means we are very much there in India 1 and 2. Most
of the urban areas do use any one of the apps to make the payments. So, we are someone who needs to make those high volume, high value and if you look at it, many of the people will make payment on the last day. Right? And I will have to have a strong reconciliation mechanism to make sure the service gets delivered. this is where we will have to
make sure everything gets delivered fast and it is settled [snorts] and the billers are able to understand the service had been delivered. Right? Or the customers understand the service had been delivered for them. What we had done, again very simple. These are all the contacts that we have. We do have a service equivalent to it and we do have teams equivalent to it. There are certain requirements
as well. If you look at it, ops dispute, it is also a regulatory requirement that we will have to have a separate team which takes care of it. Operations needs to be there, dispute resolution needs to be there. Many of them are mandated or right or regulated. For example, the bill payment, we will need to keep records of it for about 10 years. For anything that comes
in. Regulator can come in and ask data for any amount of right. But it cannot go out of India obviously. That's again something that we have to do and we'll have to keep it encrypted. When they request with proper authorization, we have to decrypt and give it to them. Even my production team cannot access individual customers data. So, all those things are something that we will have
to take care. Right? What we had done is since we do have multiple players, it has to be even driven at every stage. A request comes in, we do basic validation. That service doesn't go to the biller, the next stage. It just puts it into a queue, replies. My system, a single service can actually handle 5,000 requests per second. Okay, more than that. Actually, I'm talking about
5,000 transactions, not just request. A transaction can last up 3 minutes for us. So, within that time, so many number of transactions can come Right? So, it has to be even driven. It has to hand off to the next service pretty quickly, and that is something that we do very well. We do have a trace ID with which we can pass on the context, and there is
also correlation. When you do a payment, you will need to go back to your fetch for few of the cases. What is the bill value? Is he paying the amount that had been mentioned? And that is specific to a bill. So, it cannot be stand-alone. I'm not picking up a request, I'm just passing it on. No, it is not just routing job. Right? So, I also need
to maintain the audit trail. We work at such a scale. I do have a blip for couple of seconds. I do operate multiple data centers. Cross communication, something goes off, yes, the van kicks in in a minute. Within a minute, I would have had 1,000 transactions, 10,000 transactions based on the time. Right? 10,000 transactions, all of them goes to GPay, PhonePe, Amazon, they come to us. All
it happened was 1-minute blip. Right? So, we will have to have a clear audit trail, have to reconcile, have to go back quickly and do. This is real-world problem I'm talking about. A network having a blip and having time to recover, understand that I will have to switch to my backup device. Does take some time. And I cannot avoid that. Right? So, all those things happen. So,
replay-ability, we do use Kafka, we do use Postgres, we do use Cassandra, we do use ClickHouse. All open source system that you talk about, most of them we would have tried. We do have cash at different layers. In transit systems, we have in memory cash, Redis, KDB. We do store it in Kafka so that it can be processed. We have to do that. At the scale that
we are operating at, we will have to do that. And it has to be decoupled. So, each stage can operate, go back quickly to arrive at the scale that we are able to do. What we do? Simple. Actually, if you look at it here, I had gone with a JSON, but it's sort of mixed. We started so long ago. Uh we do use XML. And there are
few APIs, latest APIs, where we had started moving to JSON as Right? So, but it is hierarchical. Where category specific data are are added. Because if you look at it, when you search for it, you can use phone number. Sometime ago for LPG, you will have to give your consumer number. for biller, it will vary. So, you will have to go with specific things to find Right?
So, there is a tree model. There are category based handlers. So, a specific category will have certain specific things. And actually biller specific data will also be there. And on top of that, I need to make sure the data is encrypted at each stage. If there is a PII data, even within my log, I cannot have them recorded. It needs to be hashed or encrypted. Okay? Obviously,
there needs to be backward compatible. I have a new biller. So, if you look at it, these are not challenges that I'm looking at AI to solve. Right? These are my domain problems. These are my practical problems. I tell AI to develop a code for it, it will do. But I'll have to specify this is what what what I want. That's the reason why I had taken
this topic, this approach. Hey, yes, AI will do it for you, but you will have to tell exactly what you want. And how do you approach? This is what we had taken. Right? what we tell ourself on day, right? One contract before code. Obviously, we are working in payment domain. We will need to have proper contract, right? Whether it is API or stamp duty-based, stamp paper-based contract.
I'm talking about both here. idempotency, I can probably make, I mean, the same because I process at multiple places. Some can say disposable, very rare case. Same ID could be passed. app could have posted it multiple times for some reason. I'll need to be fault tolerant. My builder need to be fault tolerant. So, we do have those contract. So, those things we will have to take design
for failures. There will be failures. There will be milliseconds failures, seconds-based failures, couple of minutes failures, and I will have to recover from it. I'll need to have a mechanism for it. That's where the events come into the picture. How do we store? How do we take it? All those things matters. Small events, rich context. What we do is when an event is produced, it will be
small. For analytic purpose, we will enrich it. It will happen in the pipeline. Not when it gets produced. Right? We will have to be very prudent in that. Right? To be frank, we use lot of boring tech. But it works. Uh doesn't mean we don't use exciting ones, but boring tech does solve lot of problem. So, we don't have to run after new fancy technology before it
matures for me, it is more important to deliver those 5,000 odd TPS per service per microservice then we otherwise I will end up scaling. I need I'll occupy the whole data center for the skin to feed delivery, right? So, tech partners as first class. We have to focus on tech. See, one of the things that we are when we are looking at AI, right? Internally plat- as
a ecosystem player, right? As a switch, I can use AI to develop code faster. there had been partners who had struggled to adopt to some of the standards specifications that we can develop. Here exposed things that even after a year, they were not [clears throat] able to come back. Obviously, there are regulatory requirement that takes precedence over new feature development for some of the players. How can
I help? Can I have a A to A which when I give a specification, when I have those specification, it goes to them, helps them develop based on their context. I may not know their code base, but communicates with their agent, develops a code, and comes to my UAT environment, does the testing. We are able to confirm. And it gives that and it goes to their governance,
goes to my governance board, and we can say, "Yes, this looks fine. Let's take it live." This could happen in days, all right? This is where we wanted to be. We want to go. This is where we are investing right now with respect to AI. For all of this, the context that we have, the domain clear domain, the events that we have, how do we do reconciliation?
All of them are the fundamental. Unless I have this, I cannot go to the place where my ecosystem player can adopt, can come in fast. Right? What do we do? Our data platform, we do have services, Kafka, CDCs. We do use streaming, batching, nothing fancy. You would have heard all of them, right? We do do hot, warm, cold at different layers. When the request comes in, I
do have hot, warm, cold. When it gets processed, when it gets offloaded, when I store at different layers, I do have hot, warm, cold. If not hot, warm, cold, at least hot and cold, right? I'll need to store huge amount of data for regulatory requirement, and I do have a cold system, which needs to store it for longer duration, right? Obviously, we use BA or ML regulatory
related needs. When RBI comes to us for auditing, we will have to have all the data. We will have to allow them to look at all our data, right? So, all of them are there, right? Again, one of the things I told talked about is our ecosystem players may not be strong enough, One simple case, right? A biller doesn't directly connect to us. He comes through an
intermediate system, what we call as biller operating units. And if a biller goes down, right? How do I directly know? We do have an API, right? For the BOU to tell us the biller is down, because they know. But the complexity we have is the same biller can be connected by multiple if if they are large billers, say for an example, an electricity biller, a large one,
Karnataka or Tamil Nadu, or Maharashtra, they connect to two different, uh, BOUs, what we call as biller operating So, it might be down with one guy, there is a link problem, but the other day it might be up. I will have to obligate, I will have to inform, I will have to make sure we operate in a way, and any new request comes in, I don't unnecessarily
increase the failure rates for the customer, right? That is something that I will have to do, and I will also help the biller saying, "Hey, you seems to be having problem with this. Do you want to look at this?" They may not have a strong team. They may not have all the tech that is required. You know, I had spoken with billers in Bangalore, the Silicon Valley,
and they talk about, "I operate with single box." Right? And just a single box. Maybe they will have a database, which And some of the billers prepare an Excel sheet, give it to the BOU, and say, "Hey, this is the this month's bill date. You collect all the amount, you let us know." That's one reason why you might be receiving your SMS little late, right? In few
cases. So, we have to deal with all of them. We will have to call them, and tell them, "Boss, your system seems to be down. Please do something." Right? All some of Sometimes all they do is press the restart button. Right? But, we will have to tell them. So, five things to carry home, obviously we had talked about domain, that is the key, right? Uh, when I
talk about domain, you have a proper ontology, you do have a proper specification, uh, with which any of agents can actually operate on, right? Then, you do have events for you to be able to recover what happens in the system, for you to be able to replay at a later time. You'll will to have events and need to log it properly, simple, fast, that can be processed,
right? You can enrich it at later. Design contracts but don't look at payloads. It needs to be flexible, right? Ours is pre-base structure based on category, based on billers. It can change and it cannot break. It needs to be backward compatible as well, right? Take data as a product. Yes, we are, right? Because if you look at it, a biller is down or not is a data
for us. Right? Uh but can we sell it? No. As a switch, we have we are assisting our ecosystem. That's where the product is for us, right? Observe the ecosystem and help them. Because without them, without a biller, you are not going to use Amazon to pay the bill, right? If there is if they are not there, you are probably not going to use your Amazon or
GPay to So, they are the most critical part and we will have them. Whether they are tech-savvy or a simple right. So, we are at booth 20. Okay, I welcome you to join us. We can discuss. We do have quite a lot of openings as well, opportunities. We operate from Chennai, Hyderabad and Mumbai. Obviously, unfortunately, we don't have a center in Bangalore, but all are welcome. Thank
you. >> [music]
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59