Great International Developer Summit (GIDS)

AI Inference at Scale: Reliability, Observability, Cost, and Sustainability - Rohit Bhardwaj

1:00:41 · 21 Apr 2026 – 24 Apr 2026 · YouTube

About this talk

This talk delves into the complexities of building reliable AI pipelines, focusing on the ROCS loop framework, which emphasizes reliability, observability, cost awareness, and sustainability. The speaker discusses the challenges faced by engineering teams during unexpected traffic surges and how to address issues such as latency spikes and resource utilization in AI systems. Key concepts such as cold starts, AI agent orchestration, and the importance of caching and micro batching are highlighted. The speaker also introduces techniques for managing costs in AI operations through model selection and intelligent resource allocation. Overall, the session underscores the need for robust systems to anticipate and mitigate potential failures and optimize operational efficiency.

Full transcript

So, welcome everyone. So, this talk actually came along from the inference perspective. As you can see here, there are two links which are important for you for later use, you know, you can connect with me on LinkedIn, find out uh if you have any questions, you can ask me also. Uh you know, we'll be able to answer later on. I'll be available for you to answer that.

But, there are two links here. Number one is the link where the slide deck is. Okay, people ask me this question. Second is the agent. So, I created an agent for this talk only. Because you're going to talk to this agent to find out your specific project or your specific challenge. So, I may not be able to answer all your questions. That's the reason I created my

clone because it makes it easier. If that does not answer the question, then definitely let me know. I'll be able to um you know, provide the details afterwards. So, this is where these are the two links for you. But, let's talk about why this actually evolution of this talk came along. It came along from something called as ROCS loop. This is coming in my book uh for

system design with AI and with Apress. It's coming out in 3 months or so. So, we'll be able to uh talk on that. Uh so, so that's that's uh that's possible. So, as you can see, there are few elements to this. You want to build uncontrollable AI pipelines. Who wants to build uncontrollable AI pipeline? No one, yes? I don't think so. So, what you want to do

is you want to build with the reliability, observability, cost awareness where you're having resource utilization, sustainability. So, now if you were to have carbon footprint to sustain this. So, this is the framework I created and the agent I have created is going to solve for this, okay? So, you can ask the question for your toughest challenge, it should be able to solve for it. If it does

not solve, you know, LinkedIn me and I will put that in as part of the challenge for myself, all right? So, that said, let's move ahead with with our first thing. So, what's happening is the weekend is starting. Now, this this is a quiet weekend, no alert, dashboard is green, traffic is stable, undetected issue. How many love that? No issue. It really I like this, you know,

this is so good. Like, you know, if you do this, your boss is happy. That used to be the case 5 years ago. Not anymore. That's That's what we are here for. So, what's happening Monday, there is a discovery phase which is happening. Now, engineering teams are reviewing what's what happened. There is a P99 latency doubled. Uh finance is saying $60,000 bill has come in. Now, the

CFO is saying, "Oops, there is a bill coming in. What do you want me to do here?" And now, we have the ESG team is there, carbon spike detected and that is causing these all these problems are coming in for us to work with. So, let's see what really happens behind the scene when this this kind of thing takes place. Okay? So, what I'm going to do

is Let's just put it here, okay? Great, thanks. So, there's a quiet weekend is there which we are kind of working on right now and this problem is given so you can actually see this is what happens actually in any project. You're working on a project, this is exactly what happens. Everything is fine. You need to understand what are the problems which can come in. Cold starts

can come in. Cash delay decay can come in. Indexing can be having Now, these are all the terminology you need to know. You ask me skills, what skills do you need? You need to understand what is cold start means, you know. How many here might know what cold start means? Not cold shower, yeah? Yeah, go ahead. Anyone? But but like when you start means like your application

you start that it didn't used to handle the traffic when people used to do it. Yeah, go ahead. He's coming, yeah. He is It's good to discuss this, yeah. Mhm. Yeah, go ahead. So, like when your application starts on a weekend, it starts at a scale which is not ready to handle the traffic. So, when the traffic suddenly comes in, it is not able to scale uh

immediately and uh start generating issues. So, Absolutely. As you are absolutely correct. So, the problem here is not just this. The problem is that, you know, you have to go through certain phases, which is the things are happening. And we'll go over in detail what what really happens is system is looking stable for yourself, but really speaking uh you know, Monday morning you got getting the failures

on latency is increased for any of the solution coming to you. And we will be having fixes coming in to fix those issues, okay? So, this is the architectural insight we'll be talking about how that really works. Okay, so let's talk about this. I'm going to simulate this with my agent for uh you know, uh let's do this for for for our baseline GPU utilization, okay? So,

we're going to do this one first. And that's the first thing we're going to try it out. Okay, so this is my agent, okay? So, I'm just going to go in here. So, I'm simulating this problem. Warm caches are coming in. Uh I'm not doing everything, but you get an idea like, you know, this is just like you have to play around with a lot of things

happening here behind the scene. But we're just starting on on Monday morning or anybody who has sale coming out, you know, you were preparing for the sale. How do you prepare for the sale in an AI land? That's what I'm trying to discuss in in in this case, right? So, it's just going to Okay. Okay, let's see here. This is the reliability test. Monday morning test, you

know? So, I'm just preparing us for this kind of test right now. Okay, it started again. That's good news. All right. So, our RCS loop, oh, what was that? Yeah, okay. So, that's coming in the quiet quiet simulation with a healthy baseline coming in. Now, this is something called as queries per second. Low queries per second, you know, these are all our assumptions which are there. But,

now what happened is that whenever any request is coming in, now this is a steady thing which is coming in. Now, in this case, you're trying to somebody is asking for a question. Hey, what is your policy looks like? Based on this policy, it's asking question like, you know, hey, what really happened in that case? And there is a cache hit rate is there which is really

good. Now, whenever there is a reliability coming in, there are issues in the reliability. There is an issue in observability. There is an issue in the cost. There is an issue in sustainability. That's the real problem which is happening right now in this case. If you if the cache is not hit, you're going to the database and getting that information from the database when you're producing that

result. So, we'll be discussing this in more detail today. All right? So, I'm just preparing you for that. So, what does that mean? It means that there is a lack of alerts. There are no alerts coming out from our system. Immediately, the warning is not there. They're green dashboard. So, we need to we need some way of catastrophic failures are coming in and we need a stable

traffic for us to work with. That's what we'll be discussing today. So, let's talk about one at a time. Black Friday P99 blow up to place. E-commerce lead is saying here, why this failure is coming in? Based on this failure, now patient patient is waiting. You know, architecture decision needs to be made. Now, this is for health care. You don't want to wait for the decision optimization.

You want patient to get this information as soon as possible. Now, you can think about health care, uh you know, financial industry, everyone faces the same problem. Like, you know, so this is this is where uncleared GPU spend. How do you make sure what you need to do? Something called a FinOps lead. Now, this used to be not someone doing it. FinOps was not there, but now

FinOps is a big thing. Cost You somebody talked about the cost. So, you need to have a observability with monitor the resources and align on the goals what you want to achieve. And how do we do that? We do through certain things here. There are four stages of AI outages which are there. Four total stages are there. Number one stage is Number one stage is tail latency.

Tail latency means and we'll talk about more this slow response is coming in. When I say slow slow response is coming in, that mean that it's 99% of the time it's working fine. It's only 1% it's not working. And 5 years ago, and I'm talking about 5 years ago, I used to say ignore 1%. Yes, we do that, yes? Not anymore. Because that 1% amplifies. And now

what happens because that amplify, the queues increase and your data is not coming to you. That's the problem. Observability There's no observability there. Cost amplification is taking place and carbon uh compounding is also there. So, these are the inferences when we get the reliability, observability, cost, and sustainability. When you provide these things, now these are the thing which we need to come in and fix them with,

okay? So, we don't know why the P99 spikes. You know, why my bill is increasing? Nobody knows it. Like, you know, you are people are asking you, "Hey, why this increased?" And if somebody asks, you find out and tell that answer. How do we do that? Let's discuss it further. So, e-commerce liability cost spikes. So, the flash sale took place Now, when the flash sale sales took

place. Now, think about a flash sale Mother's Day coming, Mother's Day coming, you have to give gift to mother, yes? It's a must, yeah? So, uh that's the retrievers are now overloaded. Now, because they're overloaded, this is causing causing the GPU spikes happening. This is actually adding the queuing delays. P99 5-second delay happened, right? Because of that reason. Because that happens, the GPU bills skyrocket right away.

What we could have done? And we'll talk about what we could have done to fix that later on, but I'll just give you one example. What we What are you going to do to fix this? Anyone? What you can do to fix this problem? Anyone? Rate limiting. Rate limiting, you can do rate limiting, but rate limiting is not the only thing. You can plan for the flash

sale, yes? For the flash sale, what utterances the user are going to have? Can you feed that in first so those are all cached? Can you do that? We'll talk about in my next talk, but like, you know, but but that's what is called as Rag orchestration, where you're caching the Rag value before you are So, you're not even hitting hitting hitting the load, you know? So,

that's one thing you can do here. There are a number of things you can do to fix these problems. Healthcare, the same problem with the radiology data is coming in. Multi-agent inferences are there. Now, with the vector cold partition detected. When I say vector cold partition, that means all the vector database is doing is it's a database behind the scene. Okay? But it is a cache, Redis

cache. Okay? We all know about the Redis cache. It has a cold partition attached to it. Now, even if you're not doing it on your side, you're calling vector database and if it is not there, it will take time to infer the data and diagnosis will be not that fast. That is what is the real problem which is there. Digital twin Now, did I mean did you

hear from manufacturing industry? You know? If you are from manufacturing industry, like this is for manufacturing, the vector search is done, the influence pipeline is created based on this influence pipeline, control loop is there which which we are trying to do. Based on this Now, it's sending it out as a digital twin to really find out what's going on. So, you basically you're creating a digital twin

control loop. Digital twin of your whole system. And you can understand it differently. So, what happened used to what used to happen is that like you know, I used to go to a restaurant and you eat anything. You know? If a digital twin of mine is there, yeah, my spouse is calling me, "Hey, you ate too much food today. Stop eating more food, you know?" Can she

do that with digital twin? Yeah, she can do that. And I'll be careful like you know, what I think now, my wife controlling what I think in next 5 minutes. Now, she's controlling what I eat also in next 5 minutes. And that's products wants. They want that data to be there to be able to process it very very fast. Predictability He was asking pretty That's what the

pretty digital twin If you build digital twin, now you have a better answer for yourself. It's there for every industry. Financial industry, you can build build digital twin and you can do all of them. Okay? So, what do that means? Why this tall exists? I think we already discussed the SRE metrics and all these things we discussed. I'm going to skip some of these things. Initial understanding

And detecting hidden loops and vector store and what are the failures coming in? I'll discuss that in more detail. So, unveiling the inferences is the first thing. Whenever something happens, explosion. Now, now, after explosion happens, everybody know, this is volcano, yeah? Explosion you don't want to wait for it, yes? That's why you are here in this talk. Why you are in this in this talk? You don't

want to listen, "Hey, something happened." Okay, something happened, so what? Like, what what can you do about it? But, can you act on it before it happened? Like, just like, you know, my wife can't control what I think, you know? Uh uh so, that's that's one thing which is there. So, revealing the inference mechanism, you know, that's where the mechanism is coming in. You're influencing the data

which is there, and apply these forces to really build build the solution. So, you build a clear understanding how to do this, okay? So, let's talk about this reflection mode on this one. So, what does that mean, the reflection mode here? So, e-commerce, we are trying to build the e-commerce solution, and we're trying to do semantic search for it, you know? When we are doing that, there's

a flash sale which took place, GPU cost, uh you know, in 2 days were, you know, three 3.5 it increased. Gateway it you know, LLM increased. So, your task is to make sure that you create a planner, and you're able to inference the API, which is pricing API, and other APIs are there, to really build the full-proof system using P99 100 900 ms become 4.5 seconds. You

know, that's the real problem which is there. So, let me talk about what does that really means. When the user is making a call, you know, it's calling API gateway first, yes? After the API gateway, it's making a call to you're doing a asynchronous queue. How many here love asynchronous queue, you know? Everybody say async queue is really good to do, which is which is a good

thing to do, you know, you're getting search requests are coming in from there. You're searching the rag, you're finding out similar products and the policies. Retrieve the data from the database or vector database, you're retrieving that, re-ranking the data, and then LLM is now sending this data back to us and we are getting this data back and there are sub-agents which are kind of looking at this

data and then critiquing the evaluator agent. See this? This is where This is where our ROC starts. You have an evaluator which is saying that, "Hey, who's the monitoring Who's monitoring how the agent is working?" That's the first thing you can do. Rewriter agent is there, improve the tone. How can you improve the token tone on like to help help the system out? Based on that, now

you we created the reliability loops for us to work with and once we know the reliability loops are there, you're able to work on this. Okay, I'm not going to go through all of them because we'll discuss in more detail later on, okay? Now, this is inference is not one single call. This is very important to understand. It's a dag of hidden loops. Okay, let's talk about

this for a second. Let me make it like this, yeah. All right. So, I'm going to make it a little bit bigger so we can actually see this through what's going on. So, we all know about this one, yes? Everybody know user is making a call. When the user is making a call, you create a async queue, yes? You create an async queue and then if you're

trying to have a decision planner is is the agent is there which is deciding what to do. What should I call next, you know? It's trying to decide what action should I call next. When it is calling this action, now what's happening is there is a sub-agent is there which is now looking at It's picking the sub-agent and finding out what actions can I perform for this

one. This sub-agent is now going in to retriever with the retriever logic and retriever logic is now going to the vector store. Yes? When it's going to the vector store, it's going to go to the cold partition. We talked on this before, yeah? Cold partition means if if that vector store you have already warmed up, hey, somebody will be asking for store opening. Somebody will be asking

for uh you know, flash sale. And and the code for flash sale, which is coming up, yeah? So, that you can warm it and ready to go. And if it's if it's if it's a hot partition, you get the data and then you now send this to LLM to process this. Okay? Now, if you can think about this, what's really happening here? There is always like you

know, there is always a knowledge which you're trying to work with, detecting any problem. This is all observability side of it. Now, based on that, now you're going to LLM gateway with all this information coming in with the governance layer. Now, you're going to the LLM, making sure you're finding, fetching the right data, tokenizing it, and returning back to the user. Now, can you guess when you're

doing all these things, you need to monitor cost and other things like this? So, this is not queries per second. Everybody with me? What is this? Huh? Hops per second. So, what are we trying to do? We are trying to build this as a hops per This is not one one query. This is hops per second which is coming in for us. Now, when this comes in,

now the inference DAG which is coming in, now we are trying to come up and find out the cold partition and then find out what the issues are. That is where the reliability failures are coming in. And what do we need to do? Tail latency. So, what is tail latency means? That when I'm trying to retrieve the data, it is not there. Now, I will be making

a call to the other system to get that data. So, what happens is and uh and let's go through this math for a second. I have to see this here. All right. So, what's really happening is that whenever somebody's calling, PII and latency is there. Whenever anybody is making a call, you are trying to uh reflection question is there, you know, you're asking a question and if

the other API is not there, if the other API there is a problem going on because the network fails, yeah? When the network fails, what are you going to do? You're going to retry that network. Then this P99 latency will now, you know, it actually collapse this pipeline to 3 to 5 seconds latency instead of P99 latency. And because why it is doing that? Because the queue

depth increased, yes? Because the queue depth increased, what are you going to do now? You're going to add more worker threads to solve that problem. Now your cost increases. So it's an explosion not only just just by doing small effort, but the but then the vector store also a lot of people are making a call to and they are they are they don't measure P99 per hop.

So the first thing from this slide if you want to get is P99 you need to measure per hop. Okay, you're hopping from one place to another place, like you know, you're going from user to LNN, that's one hop. From from to vector store is another hop. So you need to do P99 latency for all the hop and then trace all the critical paths. Just doing these

two things from this diagram, now you are not also managing the GPU queue. Queue, how much queue you have got? And if the queue increases, then you throttle. Yeah? That's something which we can do All right, so that's good. Now now this is another one and so this is another one white tail and the explode. I think we already talked on this like GPU try to increase

and this is basically increasing the increasing the value. So when I apply this loop and GPU queue is happening here, so this is this is what I was talking about before. The queue increases, that's what increases spike for us. Right? So what happens when I'm trying to make a call here, I'm doing batch of eight and there's a delay and there's a batch of eight is also

there to compute whatever is coming for us from this side. So, what can we do to fix this? Number one is batching. You know, micro batching. So, all the requests coming in, don't send one request at a time. Okay? You do micro batching. The simple one tool from this diagram, okay, everything else is doing micro add micro batch What is micro batching? Anyone? What does that mean?

10 users are asking the similar kind of questions, send them all together. Okay? Don't send one request at a time. If you send through API one request at a time, it's not optimized to make a decision well. Makes sense? So, that's that's what micro batching is. So, this diagram was just for the micro batching purpose. The vector store for tail behavior. Now, what happened? This is the

vector storage is there, the hot partition is there. What can we do really? Forget about everything else. There's the instant spike which is coming in. Now, what can we do to fix that is what we'll be discussing, you know, after sometime we'll be discussing how can we fix this through warming what the requests are coming for us. Okay? There are number of tools to do that. Now,

this is where the observability comes in play. You need to observe. Reliability we talked on what are the different ways reliability is there. We can increase the reliability by what? Different ways. Micro batching. Okay, what all thing we discussed so far? We discussed micro batching. It's very easy to do micro batching. Well, if I do micro batching now, I'm not with jitter I'm doing it. Now, this

will create an easy way to handle it behind the scene and and I am also now able to do 10 to 38 millisecond to solve that problem. Okay? So, that is one thing you can do. Another is vector store. If I have a vector store is there, what you can do is you can inference for that vector store and then basically what you can do is you

can you can try to see if you can warm partition that. That's number two which we can get from And number three is observability. Now, observability is important because you need to understand the diagnostics, prognostics, and you know, the prolonged down down time and reduced reliability is coming in because of that reason. So, let's take a look at uh this one. So, invisible pipeline. No, this is

the invisible pipeline. So, when I say invisible pipeline, you don't see this actually. You don't see this. You're just making a call. Planner is making a call. Now, what is a planner, anyone? The planner is where which is deciding what to do. What should I call at this time? Let me show you a planner just to get an understanding what I'm saying here. So, what I'm going

to do is I'm just going to go in here to an agent. So, I'm just going to an agent now. As is the reliability example. So, this is what happens. Like when you're trying to make a call to an agent to do something, you know, it's just sometime works, sometime does not work. So, it Yeah, it's coming up. That's a good news, yeah. Agent. So, I'm just

going to look at one of the agents which is already created. Um and and the reason I'm showing you is this will help us understand. Okay. So, I just you can this agent. This is an agent you're working as you're working as a orchestrator agent and your role is customer service representative to solve the problem. Now, tell me what customer service does. Anyone? What the customer service

do? Come on, quickly. What all things customer service do? If you call customer service, "Hey, I have changed my address, yes? Changed my, you know, uh detail. Email address is different. Uh I'm having problem with password reset, yeah?" Things like that, yeah? So, if I need to do that, what are these? Micro agent, sub-agents, yes? So, whenever any inference is coming for me, what am I going

to do? I need to understand what happened there, you know? What inference are you trying to make here? So, based on that agent, I will be able to work on this agent. Everybody with me? Now, what I'm trying to do right now is I'm going to go in and I'm asking the question here saying that, "Hey, I am this I want to find out because this is

a resort. I'm going to go to a resort and find out this resources." So, I'm just going to go in there and then look at this resource. So, I'm this person who's going to go to this resort. And by the way, some people you some of you might have seen yesterday. I was if if somebody were there in uh in the meeting before. So, user prompt is

coming in. Based on the user prompt, what I'm doing is I'm going to go in and then find out what sub-agent I need to make a call to. Sub-agent selection is done here based on the instruction, okay? Based on the instruction, sub-agent selection is done and then create expe- These are all the actions I can perform for this agent, okay? And and these are the agents I

can perform. It's asking me for a username and password. That's exactly what it's doing. So, I can perform the username and password now. And ob- I'll show you observability in just a second how we do that behind the scene here. Now, how do I debug this? Some people ask the question, "Hey, how do I debug this thing? What happened if if there is a cost implication for

this? How do I solve for the cost implication?" All those things we'll be discussing now. Now, I put in the number. So, I 1 2 3 4 5. Uh that's not the right number, but I think it takes it here. Uh yes, I want to go next weekend. And then I want to book for four people for that one. So, it's able to use, you know, 25th

April. Yes, two people. Yep. So, that's what we are trying to do here. So, behind the scene it's actually doing doing the analysis and autonomously be able to decide what to do, okay? While it is deciding that, this is the planner, which is going to go through all these these levels. Your monitoring shows 20% of the system. That's the real problem what I'm saying here. It's not

monitoring uh, you know, retrieval. It's not monitoring LLM calls. If you start to monitor all of them, now we are trying to do the retry storms, vector partition, embedding embedding data, GPQ queuing, and and these are all the problems which can happen, you know? And tool calls can also happen behind the scene. Now, what all that means? It means to be architecture. So, when I need to

build the proper architecture, how am I going to do that? Uh, we're going to add observability, reliability, and other things here. So, I'm going to skip this one for now. Financial industry. Let's take a look at from the financial industry perspective. So, whenever uh, whenever any uh, transaction event is coming in, I want to make sure that the fraud agent is doing the LLM scoring. And then

when it does the scoring, it needs to find out what things are coming in here. Now, this is not one call. This is not one call. What you're trying to do is whatever the user utterance is there, you're going to increase the user utterance, you will increase the user utterance to uh, they'll be like three times the token will be there to really solve this. So, you're

going to embed more things to to basically give the context engineering to the LLM before you can call that. Now, what happens whenever the agent is making a call, it's going to do retries. Few retries it's going to do because because network is not there. And then if that happens, you're going to do retry two and retry three. When I'm doing these retries with fraud coming in,

what happened to the fraud agent? Now, what could happen to the agency behind the scene? What could happen to the agency behind The agency where this user is getting called, they're also thinking somebody's calling them they are also fraud. You see what I'm saying? They You're calling so many number of times to them. That's another problem which comes in when you're kind of working on financial industry

problems are there which which can come in through. So, what does that really means? It means, and let's take a look at financial industry. So, this is for the financial industry risk multiplier which comes in with the ROCS loops, okay? And by the way, after this, I'm going to go through each loop. We have not done that. First half an hour is to understand what the real

problems are. Second half an hour will be I'm going to go to deep dive into each loop. What you can do to fix that. Everybody with me? So, this is the first half I'm going to go through right now just to make sure we are able to understand this, okay? So, this is what we need to do. We need to do the financial retry cascade you want

to do here. Now, when I do the financial retry cascade, what really happens and I'm going to copy this. By the way, this is a homework you can do These are all the prompts ready to go. You can try this thing out, okay? And when you try this thing out, it's going to it's going to do the mimicking of what really happens in the financial risk scoring

agent, you know? So, so what happens if the customer transaction took place? Initial timeout is a 200 ms SLA is there. Three automatic retries we are trying to do in this case. And we are calling risk engine, you know, risk engine timed out. That's the problem which is there in this. Now, when the risk engine timed out, the number of retry attempts are three. Total executions are

four, okay? Because the total executions are four, the carbon footprint increase here. See this Carbon footprint really increase. Because that's what will happen. Now, for each hop, you can think about original request, you're doing vector query, embedding, little and call, and then it's taking 310 milliseconds. And it timed out after that. You see? Another retry you did, another retry you did. Now, these were successful in because

they're successful, duplicate successes are created. And this is what's happening behind the scene. It's not in your control. You know, these are the things which are happening behind the scene for you. Now, this is where the telemetry comes in. Using the telemetry log, you are able to now find out, hey, how many attempts were made and what what's the how many seconds it took to do that.

And then reliability collapsed because it timed out. No retry budgets are there. Now, what you need to do is you need to do the retry budgets. How much you should retry? Oh, I'm going to only retry one time. Okay? Why one time? Because it will user will send it again. The other agent which is calling your agent What's wrong in that? Why you need to solve for

world hunger? Everybody with Yeah, just say why or yes. You know? Yes, it makes sense? Yeah? So, what I just what I just said don't try to solve for hey, I'll do three times retry. No, that's you're increasing your cost. That's something you don't want to do. Uh and then and then and then duplicate execution after success. That's the real problem which is there. And observability, retries

are not are not trace linked. You're not linking with the tracing which is going on. No visibility behind the scene how the retries are taking place. So, these are actually four successful inferences. Are you saying one request, three hidden retries? Are you saying that in your call? Who is working on the observability? I think you are working on it. Are you doing this? No, okay. Try to

do this one. Because because what's happening is that you are thinking everything is coming in, but which one is a retry, which one is not a retry? If you are able to do that, now you're able to solve this problem for for for root cause of the real problem. Once you do that, the cost explosion taking place, sustainability problem is coming in. And this is where the

retry bag is broken. This is what we are trying to say here. And there are silent retries and no item potency is there. Now, item potency is important, like we'll discuss more on the API side tomorrow, but really speaking, item potency is making sure that you have for every API, how many here have for your every API, somebody calling post method to save an order? How many

here send out item potency key? You guys are doing that? You're doing awesome. Anybody who's not doing that, it's a great time you say that I I have learned this thing is a important thing for me to do, go back and do it. You know, you're adding your value or job security from your side, you know, just by doing that activity. Okay? And full execution latency is

there, so that's the real problem which we are getting out of it. E-commerce is the same thing, you know. What is e-commerce doing here? Take a look. E-commerce, the user prompt is coming in, 200 tokens are there. Now Now, what happens is personalization is kicking in. Oh, what does that mean, you know? Show me Show me breathable running shoes under $100, compare them to my last purchase,

and highlight comfort features. Oh. So, you What is user sending it? Very small tokens. Is it really small token, anyone? It's doing what? Table scan. You know, in in database terms, what is this? Table scan, yeah? Is it a good thing? Should we allow them, "Hey, go for like a last purchase from past 10 years?" It's even worse, you know? So, you think about this personalization what

is added top K notes which is coming from all the notes which are there. Now, you are coming in here. Tool expansion LLM is calling so many two 2,600 tokens are there which is coming in. Cost latency is there. Now, if in this network fails, God bless you. >> [laughter] >> It's already overloaded. God bless you. Like, you know, what can you do? Like, this is what

the real problem This I'm discussing real facts. This is in production. This is what is happening right now. And we need to someone fix that problem, okay? And that is where the cost comes in. Money, money, money, money, you know? You want to make sure that if there is a cost there, fix the cost, you know? Why are we are trying to work on solution? The Black

Friday cost explosion I think we already discussed this one. So, I'm going to skip this. Uh cost is 10x more token drifts which are happening, we are adding more tokens, auto scaling lag, you are adding the queues are there. You're adding more queues to this solution. Now, this is where the cost curve is increasing, cost explosion increasing, and that is causing more money to be flown. But,

what do I do? Rohit, this is something which we uh before and sustainability failures are there. A carbon compounding in manufacturing, that's another thing which is happening. Now, you're doing predictive maintenance. This is your digital twin. Based on the digital twin, you create for the digital twin another place where you have all exact same parameters are added. Now, you are able to find out the anomaly detection

and you can work on it. And this is the way to really control, you know, control the problem which are coming in. You know, this is called a predictive agent, you know? We can we can call it as a predictive agent in manufacturing. Is it only for manufacturing, by the way? No, health care I told you my spouse can control what I do, you know? And same

thing is true. No, but basically it's really good for health care because if my mother is having a blood sugar spike, I come to know before her, you know? And that is really good. I I mean I would love to have that, you know, from from my side to happen. All right, so let's take a look at ROCS loop. Now, this is our second part of the

journey we're going to do it together in 25 minutes. Everybody ready? Say ready. We are ready for this. Okay? We are all ready for this. Okay, I don't want anybody to sleep yet because oh, it's it's the first task, so nobody should. My last talk, maybe, you know, >> somebody is dozing out. So, what's really happening is that in this one which we are talking about, you

know, few aspects of it. Implement control loops. When I say control loop, that you have an unstable system. Are you providing rate limiting? Rate limiting is not only for for one thing. Rate limit for everything. You know, what are the different rate limits you can have? Anyone? Agent level? Request level? Which? LLM level? Yeah? Token level? How much tokens are coming? So, there are so many rate

rate limit. Tracing. Monitor every system performance. Budget. Have a budget control. And carbon visibility should be there. So, we'll talk about all these implementation loops through our reliability loop. So, now this is what is called the loops. So, I call it loop because you have to continuously improve it. Reliability loop, observability loop, cost and sustainability loop. In the agent, you can ask for each one of them,

it'll give you perfect answer for that. Okay? Now, we're going to talk about the reliability loop. So, Rohit, you discussed so many problems with me. So, let's solve for this problem now. Okay. So, this is the reliability loop from the reliability perspective. Now, what I'm going to do is unreliable loop what what I'm going to add here is stabilize, reduce the system variability. How much stability variability

absorb handle any unexpected traffic which is coming to you? If I reduce the stabilize unexpected traffic and route efficiently. Oh, efficient routing is possible, you know? Why do you need to go to the same rack? Like, okay, somebody ask a question, you know? And somebody ask a question, a simple question, what time is your store open? Does it have to go through a complicated rack or data

to come and get that value? No. So, are you doing that optimization? Which one should I call, you know? If you do that, you save money and you save the efficient route and regulate, control the resource allocation. What you're trying to do and recover. If some failure happen, you recover from that failure. If you apply this loop, now you're able to work on it. So, we're going

to start with the reliability for the Black Friday. Spike data is there, traffic burst happen, unpredictable P99 is happening here. So, pattern number one, async queuing. That's number one pattern. So, what that means? Whenever any request is coming in, you asynchronously create the worker threads to solve for this, you know? And this is where when we are doing this, also do something priority queue. Okay? What is

priority queue means? If you have somebody doing the checkout, you know, get that get that more preference because checkout means money coming. If somebody is doing search, hey, you know, give it important but not that importance as a checkout. So, that is where priority queue plus async queue is coming in. Now, this is where now you can batch this also. That means when the request is coming

in, I'm going to put five of them and send it together. So, instead of having one worker thread just look at one request, yes? You're going to look at five requests at the same time and then process and return back. That way you can micro batch it and make it faster. Right? That's one async queuing is there. If you don't buffer the traffic, your GPU will buffer

Is that true? Yeah. GPU will say, "How many here like Nvidia stock?" You know? Why is it going up? Maybe GPU. And he's so happy, you know, he's using, "Hey, I'm happy people are saying thank you to ChatGPT." They're not realizing the cost coming out of it. Simple. If you do this buffering of the traffic with async queuing with priority queue, you are able to take care

of this. Awesome. Second pattern is back pressure. Now, back pressure is a protecting the system from unbounded work which is coming in. What do you mean by unbounded work? That means now Now, token caps don't provide more than certain amount of tokens for you, okay? Second, always implement your agents as bag. Directed acyclic graphs. What does that means? Do not have agentic loops. Now, is it possible

that, you know, one agent is calling another agent is calling another agent calling back to this agent? Is it possible? Same request, by the way. It's possible. It happens all the time, you know? Because I'm calling sub agent which is calling another agent. Now, if that happens, like, you know, then we need to predict it and fix that. And that is where where agentic loop is there.

I think if go and watch my talk from yesterday, I talked about that with graphs. You know, we discuss graph and how do we detect that anybody joined from that one? Okay? Okay, how do you detect this? Detect the loops. What do you Depth first or breadth first search? Depth first search? Yes. Depth first search we use to detect the cycles which and then and then build

the solution for us. And then the third thing is query rate limits. Vector query rate limits also we need to apply. If I do the back pressure these three things, I'm golden. So, what do that means? That's the pattern number two. Okay, async, back pressure, number three pattern is GPU pooling. GPU pooling auto scaling we you know auto scaling sounds good until until instantly burst comes in.

Now, GPU warm the GPU pools ready to go. Predictable latency, predictable cost and no cold in San Francisco style style is very important. That means if you know Mother's Day is coming, you know Christmas is coming, have the GPUs ready to go. You know, why are you waiting? So, this is what is called a predictive analytics. One way trying to do that. Cold auto scaling is really

bad like you know. This is the cold auto scaling. Cold cold cold. So, when this become cold like you know, how many here like to do cold shower? Until unless you are a Wim Hof fan, you're going to say no. Anybody here Wim Hof fan? Yeah, search online. You know, it's it's really good. Yeah. But I do that but not that much, you know. It's like too

cold for me. But warm is really very good. Like now I have microservices also warm, predictable latency is also warm. Now, once we do that, now we are able to do the GPU pooling and that's number three. Okay? Dynamic batching is another thing you can add to this. Now, what is dynamic batching is? Whenever anything comes in, you micro batch every 5 to 20 Did actually smoothing

what? Chaos is coming in. You are smooth the chaos just by doing that. Micro batch 5 to 10 sec milliseconds you're doing that is actually stabilizing the tail latency. Now, whenever any request comes, you put it in a batch and you process these as micro batch. So, these requests are coming in 1 2 3 4, they're batch, they're batch, they're batch, they're going to LLM and returning

the value and you're able to solve that problem. All right? Simple technique. Very powerful technique. How many here are doing this? Something which you can do when you're working on it. So, this is something uh add to from your side. Pattern number five is model tearing flow. Now, somebody say that, "Hey, you know, what's happening is that predictable latencies can Now, what's happening is model router is

there. How many here have built a model router? I don't think anybody This some of this is coming in 2026, you know, so you're going to see this. You're at right place because you know this coming. You prepare yourself for this one. So, what's really happening is you're routing your traffic for the search. Hey, somebody just searching something. Why are you going to LLM? For this product,

give me the data. They're telling you the product name, you're going to LLM for what reason? And that is where there no need to go to LLM for expert tier is there. But now you're saying reasoning. You reason and then say trend analysis I want to do. No, you need to go to complicated, I understand. Q&A normal tier and fallback tier is also there. Now, you're adding

this One just by doing this model tiering is done with routing, you're able to solve money also and you're able to solve the complexity with this one. All right? So, that's the benefit with this. And pattern number six is multi-level caching pyramid can also be added. So, caching is not just done in one place. Oh, let me cache uh you one place. No, it's multi-level cache is

there. You're going to cache at the prompt level, prompt cache. Somebody is asking a similar question, you prompt there and then if TTL that prompt, this is easy to do. Embedding cache, embedding data is coming in now. Embedding means your vector store data, which is your company data, can you cache the data? You can cache the data also. Because you're caching data, you're able to work on

this. Vector semantic cache can also be added. Vector semantic means for this or these kind of products, you have a cache associated with that product. That also you can do here. And final answer cache can also be added. As you do that, you're able to save the money and also provide the reliability for your solution. Makes sense? That's said, that we completed one loop. Awesome. Observability style.

This is the observability cycle. We're going to go through this. Now, observability has few elements to this. It starts from measure, trace, explain, correlate, and attribute. So, these are the things which are important to do. So, the first thing pattern is prompt level tracing. So, we need to do prompt level tracing. Prompt are expanding unex- unexpectedly they're expanding. Now, token spike increases. So, you Now, you have

a silent retry happening here. Prompt template is calling one call, then the second call, and this is all are happening here. So, prompt flow timeline can be there. And once that is added, what will happen What will happen to the prompt level trace? Now, you have more tokens at See, all of these are explo- explosion which are happening here. So, we need to do product, you know,

comparison assistance can be added. So, you can you can do the tracing for them. Trace for what? Prompt expansion, trace for token generation, how many retries you're doing, and template versions. So, are you going to have only one version of prompt template in your whole life? No, you're not going to have it. You're going to have a different version because the new version has come out. You

want to make sure that the transition from the old version to the new version takes place without any problem. You need to always do the versioning also with this one. All right? Now, that's number one. You know, caching the prompt level. Vector query telemetry. So, vector query telemetry means whenever any request is coming in from the vector query, you need to look at, you know, embedding ANN

search time. That means nearest neighbor search time, you know, partitioning loop, cold cold and hot partition. So, you need to find out all these parameter which are vector parameters are there. If you do that, you are able to get the observability for all of them. Something is good to know how how how is this going for for you. And then you do the GPU telemetry, see the

GPU queue, not just the utilization. CPU utilization used to be the case before. Now, it's called GPU utilization. How much is the GPU utilized? And that is actually determining the utilization for us. So, that's the important thing for us to understand. Q Q depth is spiking, batch size is high, and that is where the real problem is. GPU at 70% looks fine. But inside, the queue is

full. The batch size are all over the place, and kernel is is slowing down, and memory is heavily fragmented. So, these are the real problems for the for the GPU perspective. So, we need to watch the GPU how this is flowing. Multi-agent span, you know, correlation. Now, what happens is that you're never going to have one agent working on it. Whenever any agent is working on it,

now what happens when you have multi-agent which is there? Multi-agent means if you have multi-agent, are you using the same rag for both the agents? Hopefully, yes. It's a good thing to do, yeah. If not, then it's like something you have to see. Because it should give the same answer for you whenever I'm working on it. So, you need to control control the span correlation with this.

So, that's the reason like you know, you know, which of these three running shoes is the best for trail running? Now, I need to have a planner agent, retrieval agent, vector database, reasoning engine, evaluate pricing service, and all these things are added. Now, when I just said three running shoes is best, which of these three running shoes are best? When I say that, it's doing analysis behind

the scene. Step by step what to do, and then processing that request. Okay? So, we need to observe this and then build the perspective for us. All right. So, what do that means? So, multi-agent is also good. We need to retrieve the data, evaluate what's going on, and then be able to solve the puzzle Next one is prompt logging, structured prompt logging. So, whenever any request we

are able to normalize the prompt. After we do the normalization of the prompt, now we are logging that prompt, and then based on that, we are we are able to process process the request which is coming to us. So, that means we are able to see that, hey, if I don't do it, then it will be token expansion problem. Uh don't have ceilings. Don't hit the limit

for the ceiling. That's very important. If you just do this, you're awesome. You're you made the right choice for yourself. All right. That's the that's the important thing for us to look at. All right. Now, we completed two of them. The next next one is cost, money. How do I control the money? A lot of people are asking questions on the cost. Let's take a look at

cost attribution in AI system has these parts to it. Everything should have a cost for the pipeline, whatever is coming in, we need to be able to to cost cost attribution. So, the first thing which we need to do in this is pattern number one is pattern number one is cost attribution So, whenever any request is coming in, CFO asks, "Why this money is coming in?" I

should be able to go and find out what? Agent level cost, model level cost, prompt level cost, vector level cost, and pipeline level cost. If I do all three all these four of them, now I can go back and say that, "Hey, these are all the cost level perspective is there." And then GPU cost increase because of that reason. And we can say that you why this

cost really Pattern number two is this distillation. Now, distillation in the e-commerce solution. Now, fast specialized student models for large teacher. Now, there is a teacher is there. Now, teacher is using 70 billion parameters to solve a Do you need 70 billion parameters to solve like, you know, if you are working on a particular math problem? Do you need 70 billion? No, you need only maths. Yeah,

integral, a little bit more complex physics maths should be there. That should be good enough for the digitalization. This is where these things are getting converted into digitalization. And if I do that, the student model is returning, which is 1 billion, is less than 20 milliseconds response time. Okay, so that is what the advantage is because it's less than that. It's not taking that much amount of

time to do that. So, this is another thing people are actually coming up with to really save 80% of the workload. Don't need deep reasoning. Do not send to the deep reading reasoning. Send it to a smaller model, like, you know, save money. Just simple simple simple trick doing that, you're going to be able to do that. Then the next thing you can do is quantization. So,

what is quantization means? Like, whenever any request is coming in. Now, if I quantize the data. Now, whenever any data is coming in. now, how does LLM work? LLM stores all these parameters in dimensions, you know? Uh what does a dimension means? That means it's like, you know, we have X, Y, and Z dimension is there. So, similar to that, there is a dimension for everyone who

plays football. There is a dimension for everyone who who who is endorsing the shoes. There is a dimension for companies who are targeting footballers or uh you know, basketball players, you know, Michael Jordan, for example. Yeah, they they are doing that. These are different different dimensions are there, and these dimensions are stored in what? Huh? In vectors. And they're stored as what? Floating floating point numbers, yes?

So, instead of floating point numbers, can I store that in eight eight eight bit? And four bits. And that is if you do that, it basically helps you optimize the model with eight or four bits. Now, this is cost per call will decrease. Queries per per second will increase. And model size will decrease, and throughput will Just by doing this exercise. So, simple technique, but very powerful

technique you can apply. Now, somebody wants to reduce the cost, this is one way to do Now, pattern number three is right model sizing. That means don't give everything high model cost, low model cost, depending upon use. You're going to pick the right model to solve a particular problem. You know, small model, you can create a small language model to solve Now, another thing is token budget.

We talked on this before, yeah? So, max out the token budget. Uh you know, max out the token input token should be maxed out. Output token should be maxed out. Oh, it's not just input, output also, because LLM will throw you everything in the world, you know? It's like it's going to throw you everything in this case. And retry caps, how many times you're doing the context

window limit? Now, there is something called context window. Now, you need to work within the context window given to you. That means when you're sending me sending it to LLM, LLM can take, you know, how many you like 1 million parameters, yeah? Somebody was saying that, yeah? It can do that. But, you know what that will means? Increase the cost. You know, what are you here for

to use the cost? So, you need to distill what the context is saying and then work on top of it. So, main job is to really understand what the user is looking for and then work on I'll talk 1 hour just on this. Just one sentence I just said here, how do you interpret what the user is saying and create the orchestration for it in my talk

in the afternoon. Just letting you know, if you're interested in that detail, I can go over and explain you how to do that. So, context window is a big topic. You need to understand the context. If you don't understand the context, garbage in garbage out. Like, you know, that's that's that's what the problem is. Okay, pattern number three, spot. You know, preemptive GPU GPUs which are coming

in the AI pipeline. So, now this is where the spot of preemptive GPUs which are there. Now, spot GPUs are there. Why why do you have spot GPU GPUs? So, huh? Why? Spot instances are there, yeah? Cheaper, cheap. You know, it's like dirt cheap. It's like 90% cheap data is coming to you. If you just do that, you're able to solve it. So, let's say you're running

a batch processing. Can you do this? Yeah. It's doing at night time. You picking the spot and running it. Easy easy to save the money. And I'm betting now 70 to 90% less cost. Wow, I love that. You know, I love the money saving, you know? Why why don't we give it to Say, I'll tell my manager, give that money to me in my bank account, you

know, instead of giving it to someone else. I worked very hard, you know, just to do this. Yeah. All right. So, that is what is important. Then the last part is sustainable. We have like 3 minutes left. Okay. Sustainability have like few things in it. Now, we need to make sure the sustainability loops are created, which is carbon aware. Loops are created. And once we create this

carbon aware loops, we are able to create SCI. SCI software carbon intensity measure. Once you apply this measure, which is also sometimes carbon aware scheduling. You can schedule at night. Low carbon region you can pick. And you can also do intelligent warm pool also to solve this problem. Another thing you can do is hardware efficiency for e-commerce logistics. Now, if you are doing e-commerce logistics, can you

use quantized model instead of doing the actual high high high model to use that? You can use the quantized model. Now, hardware efficient model can be created to solve this problem. All right? And hardware efficiency is added to solve this and do less work, you know, that's something after after optimization to solve a particular problem. You can also do, you know, when we are working on the

So, that's another thing which can be done. So, these are all the patterns. The agent is also there. I will I'll share the agent with you later on. But we covered most of these things. See, take a look at this one. Here. Reliability, async queuing. Back pressure, GPU pooling, dynamic batching, model tearing, and these ones. Observability, prompt level, vector query, and all the If you apply these

principles, now you are in the nirvana state of consciousness. Yes? You feel so good like, you know, these things can really help you improve your project project perspective. So, you take a photo if you want to take a photo on this one. This is the important part. Once you do that, once you understand this aspect of it you apply the agent which I have got, you can

apply it to your project and start using it right away in your project. My book is coming out on system design. We I have discussed much more in detail, uh you know, on this one for all all the projects perspective coming. Okay? So, these are the ROCS loops which we discussed. Reliability loops are there and then, uh now, I'm going to share the link with you, the

previous link. Uh just 1 second. Yeah. Yeah. Yeah. I'm just putting it right here. All right. So, these are the two links which are there. I just take a photo if you want to take a photo. I really enjoyed talking to you. Hope you join another session and have a wonderful time here. Thank you. >> [music]