AI-First Software Delivery: Superpowers, Adoption Challenges, and the Path to Software 3.0 - Vanya S
About this talk
In this talk, Vanya, head of tech for ThoughtWorks in India, explores the implications of AI-first software delivery and its transformative impact on the software engineering industry. He discusses the transition from software 1.0 to software 3.0, highlighting the challenges organizations face amid the rise of generative AI, such as tool fragmentation and workflow disruption. Vanya points out that while AI enhances developers, it often neglects collaboration among cross-functional teams. He underscores the dangers of AI-related complacency and the need for governance to ensure compliance and effective resource management. The talk concludes with strategies for integrating AI across the software value stream, advocating for a thorough understanding of the workforce and metrics to gauge productivity and effectiveness.
Full transcript
Uh let me just quickly introduce myself. I'm Vanya. I'm the head of tech for ThoughtWorks in India, an organization that's pretty much at the bleeding edge of what does AI first software delivery really means, right? So, I'm going to share some really hard lessons, some challenges that we see across our global customer base, right? Personally for me, as a as a techie in this space, I believe
it's really, really exciting times to be alive, right? Because there is so much that's changing, right? And and I keep saying that like we all are living in an era where we going to probably tell stories to our children that that's how technology was shifting at that point of time, right? So, AI first software delivery superpowers, adoption challenges, and a path forward for software 3.0. Now, McKinsey
predicted, right? This is a not a new slide, not a new view, not a new study, uh probably 2 years old that software engineering is going to be that space or that industry that's going to be disrupted maximum because of AI, right? Guess what? They got it right. Do you agree with me? Agree with me? Show of hands. See, hey, it's it's going to be interactive, right?
So, you need to pay attention. You need to be awake. Michael has done the hard part of keeping you folks awake after the post-lunch session, so it's now our duty to carry forward the same uh rhythm, right? So, McKinsey got it right, right? Absolutely right. One industry that has been maximum impacted because of generative AI is definitely software engineering, right? There is broad consensus. In fact, I
was reading a very interesting research from the pragmatic engineer. How many of you follow pragmatic engineer reports? Good. Quite a few, right? So, if you have seen the study that they released, right? I think in February itself, which said that 95% of the cohort of people that they surveyed are using AI at least once a week, right? 75% of them are using AI for more than 50%
of your of their work. And there is about 56% who are using AI for more than 70% of their work, right? So, hell yes, McKinsey got it right. Industry corroborates with this idea. But, how many of you have had sleepless nights on the rogue AI usage within your organization or the token spend and the cost that it brings along, right? Or the ROI for that matter, right?
How many of you are playing engineering managers who get constantly questioned, "What is the productivity gain that we are getting because of using AI?" How many of you are in that boat? Okay, I see a lot of you, right? So, that is the reality, right? That there is a lot of spend, there is a lot of uh money that's being pumped into this, but what's the real
ROI, right? But, before we go into some of the details and talk about some of those challenges, some of those sleepless uh night reasons, right? Let's look at and agree upon the three software paradigms, right? I'm going to channel Andre Karpathy here, my poster child for anything related to AI in the industry. Uh software 1.0, right? I'm sure all of us understand that we program the computer,
right? We We program the computer, there's a compiler that converts what you have given, the programming language, into byte code, and then that byte code executes, right? That's software 1.0. Then comes software 2.0, right? Where we start talking about programming the weights, right? There is data, there is labels, probably there is a deep learning algorithm that differentiates between a cat and a dog based on the data
that it has been given, right? So, that's programming the weights. And software 3.0 is where we are right now which talks about programming the LLMs, right? Where we provide natural language instructions to the large language model, and then it understands our intent and it tries to do what it has to do. And I would say a more more natural and uh progressive view of that is the
agentic mode or the autonomous agents which are taking care of understanding what your human intent is and trying to monitor and orchestrate set of agents to execute that job, right? So, those are the three software engineering paradigms. Now, my hypothesis is that as we go forward and we are already seeing signals for that, that more and more software is going to be written in the software 3.0
paradigm. Do you agree? Do you agree? Like, show us hands. Like, let's be active. We agree. Excellent. So, we agree, right? Now Now, I want to really say something here which also matches what we are seeing at ThoughtWorks, that there is going to be some coexistence, right? It's not that all the 1.0 code is going to go away, right? Like, there are large systems, financial institutions, insurance
providers, medical systems still running on that part, right? So, it's not going to go away, right? It's going to be a coexistence. And within the ThoughtWorks ecosystem, I can say with a lot of confidence that a lot more green field code or green field solutions is where we see, you know, a lot more uptake on software 3.0. But, as I said, you have to take a conscious
choice on what makes sense at what point of time, right? So, there is going to be a coexistence. I wouldn't worry that software 3.0 is taking the world. It is definitely causing a lot of ruffles, but there is going to be definite coexistence, right? And there are tons and tons of challenges in this space, right? And as the world starts to understand some of these challenges, we
will learn more and more and more as we go along, right? Now, let's talk about some of those challenges that I promised that I will talk about based on our experience. What are those, right? The first one is tool fragmentation and workflow disruption. Now, this one is a little overloaded, right? There are two parts to it. Let's look at the first part. The first part is imbalanced
AI adoption. Right? Now, I'm going to ask again, how many of you in your organizations are seeing this phenomenon that when we think about AI first delivery, the org is pretty much only thinking about developers and coding and implementation. Raise of hands, uh Again, quite a lot of you, right? Now, last time I checked, software development is still a team sport, right? You need a cross-functional team.
You need product managers, you need designers, you need security people, you need infra people, but right now there is imbalanced AI adoption in the industry. That means that only one of the roles, which is a developer in the entire cross-functional setup, is getting enhanced with AI, whereas the others are not. Let's look at a repercussion of this, Have people have heard of theory of constraints? It has
been borrowed from the manufacturing space. Theory of constraints, anybody knows about that? Okay. A few. Now, Eliyahu M. Goldratt said that if you only optimize one part of your pipeline, right? It's not going to have the ROI or impact that you're looking for. Now, now look at this diagram here. Now, if I'm only focusing on coding or implementation when it comes to enhancing my team with AI,
I'm increasing my coding throughput, right? You can see the bubble in the in the center. I'm increasing my coding throughput, but now the question comes, can can I review that code at the same rate at it it is getting generated? Can I ship it in production at the same rate as that code is being written? That's the downstream impact of this bottleneck that I'm creating, right? Now,
if I can write code at a very fast pace, can I hydrate my backlog at the same pace that I'm writing code, right? So, you can see optimizing for just one part of the pipeline is already starting to put pressure upstream and downstream, right? We can see the pressure, right? And it's it's a known thing, right? Organizational theory theorists Peter Senge also said that when you try
to push a complex system, it pushes back in unique ways, right? So, SDLC, software development life cycle, in my opinion, is a complex process. And if you try to just push on one part of the process, it's going to push back in ways that we can't imagine, right? Now, this is a very interesting thing that we see, and you know, some examples of that is a team
transition to Cursor, but now their UX is getting a bottleneck, right? They're not able to design the screens at the same pace at which the developers need them, right? There is another example from a very popular platform which allows you to do a very quick prototyping, Lovable, put so much load on the GitHub's systems or GitHub servers that GitHub hack actually blocked their account from creating any
more repositories, right? So, 12,000 repositories every hour, and one of the GitHub SREs just disabled their account because it was just putting too much pressure on the system, right? So, that's what I meant, that one part of the system without thinking about the downstream effects, it's going to push back in many unprecedented ways, right? Let's keep this Bear this in mind. So, that's the first part. The
second part of of our tool fragmentation and workflow disruption is that AI is enhancing the roles, but not their relationships, right? Now, even if even if there are people on your team who are using AI tools, my designer is going to work say in Figma AI. As a developer, I'm going to probably say use cloud code. As a as a business analyst, probably I'm going to use
some frontier model like a Claude AI or ChatGPT or so on and so forth. these individual roles are getting enhanced, right? But, where is the collaboration? Right? So, where is the workflow? Where is all of these tools coming together, right? So, the hypothesis here is that AI is enhancing the roles. Then there is local speed at each and every step, but the larger picture, the sum of
the parts is not getting materialized. Right? And look at the overcrowded tooling ecosystem. There are so many tools across the entire SDLC. So, there a plethora of tools, different people in the team using different tools, but where is the collaboration between them, right? Now, if a if a designer changes something in your in your Figma AI, how does a developer downstream understands what change has happened, right?
So, what is where is the collaboration? And that's the missing link today, right? And when I pack both of these problem statements together, conclusion that we have is that there are local speed increases, but the system-wide delivery still drags, right? There are local gains at individual steps. Like individual roles are getting enhanced, they're getting a local speed, but the overall delivery, software delivery as we as we
call it, still is dragging, right? So, the required acceleration is only for the individual task, but not how these roles work together. I think that's the problem statement one. The second one are what are the AI age attack vectors? Now, let's let's look at your agent decoding assistant, right? I'm sure all of you are using one. What do you think constitutes context for your coding assistance? You
can speak up. What is the context for your coding agent? Sorry? The code? The workspace, the code. Okay, what else? What What else constitutes context for Prompts. Okay. So, some sort of tools and MCP servers and probably your terminal feed. Right? Probably your agent configuration. Right? Like your documentation and what specs and all, what documents. All of these are the contexts that your coding agent has access
to. And with the help of that, what it What it allows the coding agent to do? It can change files, right? It can run software. And because of that, it has access to the entire software supply chain. Am I right in saying that? Your coding agent has access to your software supply chain, starts from your development box right up to your CI/CD environment. Now, let's look at
what are the problem areas or how secure is this setup, You all You see all these points? These are the potential places where context poisoning can happen. People understand what is context poisoning? And people understand what is a lethal trifecta? Have you heard of this term? Lethal trifecta. I encourage all of you to read this fantastic blog written by Simon Willison, where he talks about the lethal
trifecta as the main problem that is sort of pervasive across any agent, not just coding agent, any agent that you are using or that you're are building, right? Now, let me explain what this lethal trifecta really means. Lethal trifecta means that your agent has access to sensitive information. In the case of your coding agent, your agent has access to some of the configuration files that might have
some sensitive keys, right? That's confidential information, right? That's first element of the lethal trifecta, that your agent has access to The second aspect is that the agent has access to some malicious input, right? Imagine any specific web page or a GitHub issue which has some ASCII characters which your eye can't see, but the agent or the model can see them you know, it's it can follow instructions.
It will follow anything that it is given to, and that malicious issue says that get any sensitive information that you find in the in the repository and just send all this information to this remote endpoint, right? Now, that's malicious input. Any web page from the internet that you scan, any rules file from the internet that you see, that you download, that you use can have these malicious
instructions, right? So, access to malicious information which poisons the context for your agent, right? That's the second aspect of the lethal trifecta, and the third one is ability to do exfiltration of data, right? It can execute commands. It can execute commands on your terminal. write files. It can create files, and in in in the in the in the we of, you know, doing fast, sometimes we accept
accept accept and accept, and there you go, it can just exfiltrate sensitive information from your code base right under your nose, right? So, that's the lethal trifecta, and these are the attack vectors that are there sitting right on your developer machine, right? The environment that we once thought was impenetrable, right? So, that's the lethal tri-sector. And that's These are the problems that not just the coding agents,
but any agent that you use in your in your day-to-day probably is suffering from. Right? These are some recent incidents that captured the public imagination where um hacker slipped something very interesting in the Amazon queue uh PR where said, you know, just wipe out all the AWS infrastructure if ever that instruction could have executed. I'm just saying that now the development box where we used to think
that it's secure is no longer that safe haven that it used to be, right? So, you need to be very very careful when you're thinking about coding agents or any agents that you are using, right? Moving on, the third challenge that we have seen that I have seen across our customers is the adverse impact on the code quality, right? Let me quote a very interesting study that
I really love to quote because I trust it. It's from GitClear. This is from 2023. Just bear with me why 2023 we are already in 2026. This is from 2023, right? Now, look at the three metrics that we are talking about here. Code added, code churn, and code moved, right? Code moved is a very good indicator of code refactoring, right? When do you When do you move
code? When you're refactoring, then you're merging classes, splitting classes. So, code moved is generally a good indicator of how good or how fast are you at repaying the technical debt in your in your software system, right? So, code moved is going down, code churn, on the other hand, indicates that how many times when you committed something to your CI in less than 14 days you had to
revoke it. What does it mean? That either your harness your test harness is not proper, right? Or you're not getting fast feedback that you have to revoke some change from your CI system. So, that's code churn, and that code churn is going up. And code added, as we all understand, you know, it's damn easy to write code with AI. So, the code added is also on the
rise. Now, let's look at those numbers from 2024, even more stark. Do you see the dip in code refactor going even more down, code churn even more serious, and code added even more higher? It's It's a 2024 study, but you get the point, right? That if we are not being careful about what's the long-term impact on the code quality that AI is generating, we might get local
speed by, you know, pushing features really, really fast, but if we are ignoring this aspect, in the long run, the code becomes a hot mess, right? Which is which is barely recognizable. I was reading a very interesting LinkedIn post. It said that in 6 months, I was barely able to recognize my code base, and it looked like, "Who's the owner? Is it me or cloud code?" Right?
So, that's the state that probably we don't have to have to reach. But some more numbers, On the left, you have these from the Pragmatic Engineer a research report of 2026, which says there are some good parts, the larger code change is getting around to more quality of life work, typing is no longer a bottleneck. But what are some of the bad aspects, right? Dealing with more
and more AI-generated slop, more and more debugging. And for some [clears throat] people who have built a trust out of writing software code, a loss of identity, right? Right? What is my purpose any longer on this earth? Anyways, that's something that for all of us to, you know, reflect upon. But But the point that I'm trying to make is that even though we are good going at
god speed when it comes to code generation, there are still reports that says, and also backed by practical experience, that the amount of time developers spend debugging AI generated code is still way higher than it than it can be acceptable, right? So, there are the good parts, but there are definitely the bad parts that we have to be more watchful about. And all of that sort of
alludes to AI induced complacency, which is very, very real, right? All of these automation bias, if something machine has suggested, it must be right. Right? Or the anchoring bias. Once you have seen I think Venkat was talking about in his talk in the morning that once you have seen certain suggestions from somebody else, whether it's an agent or a human, you are not able to move beyond
that suggestion, right? So, once you have seen a suggestion from your agent, you just get fixated on on that. You just get anchored on that specific solution, right? So, that's anchoring bias. Sunk cost fallacy, right? Another interesting construct where if I have spent enough time in making something work, I want to spend make sure that it actually works. Even though in my gut I know I know
that this is probably not the best way to solve the problem, right? So, you try to make keep on working to make it work. And of course, security vigilance, of course, right? Like on the comment it's still your name and not your AI's, right? So, these are some of the AI induced complacencies that we see out there and I'm sure all of you are also noticing, which
is a big challenge and impediment, in my opinion, for real scale that we expect in the enterprises. And last but not the least is governance and compliance, right? There are several aspects to it, right? What are the elements that are hurting? Model hosting options. Now, would you know from a GDPR perspective or from a data localization perspective, where is your model hosted? Do you know where where
does Cursor host its model or where does Claude code or Anthropic host its model? Do you know where is it hosted? They don't disclose, of course. Like they don't disclose this, but think about it if you are in the European region and you are compliant because of you have to comply with GDPR and you are using some model that probably is hosted in in Japan, what does
it really do to your compliance posture? You wouldn't realize, but it's a serious problem, right? And therefore, you need to carefully think about what is the right deployment strategy for my model so that I'm always compliant with the regulations that are attached to the region in which I'm operating, right? India talks about data residency laws, right? Where the data from for from India, specifically in the BFSI
space or banking and financial space, should not leave the boundary of the country. Now, there are a lot of our customers who are in the BFSI space who are so intent on hosting a model which is completely in India, right? For example, Anthropic is now available in say AWS and the Rock ecosystem also in the GCP Watix ecosystems so that you can deploy and access it via
the same region in which you're operating, right? So, carefully thinking about the deployment strategy, knowing where your model is hosted is a critical element of governance and compliance, right? Self-hosting is another option. That's on the other end of the extreme, I would say, right? Some of our customers who are very very you know, pertinent on not using proprietary models, they host their own large very large coding
models on prem in India, right? That's also a possibility. I'm just giving you the options that knowing what model as a service are you using, which is the region in which they're deployed, what are the alternate deployment strategies that you can think of are all aspects that are related to compliance which we developers don't really think about, but become very important as an impediment in the overall
governance and compliance ecosystem for using these agents. Then comes lack of visibility on the cost incurred in token economics. How many of you actually think about how much is your AI you know, according agents burning the tokens on? How many of you are careful about it? Like barely barely a few hands, right? Because we are not paying, the company is paying the bill, right? Now, let me
tell you that we are still in the honeymoon phase of all of this cost and token expense. It's coming definitely from the investors' pockets, but when this honeymoon is over, I can assure you that the bill of the token spend is going to look like the cloud cost bill back in 10 10 years ago, right? The promise of moving to the cloud and then the surprise surprise,
it comes with a cost. Similarly, with all of these tokens that you folks are burning, right? I can give an example. Now, we all started with $20 per month per developer. Correct? We started with that. Now, what's the plan that all of you are on? Is $20 per month working for everybody? A lot of you, I'm sure, must be on a $100 plan, right? Or probably you're
moving towards pay as you use, right? Or usage-based pricing. And I can tell you that some of our power users at ThoughtWorks in some of our engagements $20 per developer per day. Right? So, you can do the math. If somebody works for 20 days, that's already hitting a $400 mark per developer per month. So, and by the way, that cost is not going to go down. It's
going to just go up linearly from here, right? So, not knowing what is the spend that your organization is making on these tokens, or what is the spend that my team is making, my developers are making, is going to be a huge huge surprise or a challenge in the next 6 months, right? Something that we all need to be watchful for. Then comes rogue AI usage, right?
Now you have the API keys. Imagine your organization has given you the API keys. Who stops you from putting it in a tool that you're not supposed to use? Right? I'm sure you know, your CISOs, your chief security officers are having sleepless nights on when somebody will use a tool that is not supposed to be used and some organization is going to leave your boundary, right? So,
you should definitely ask about this to your CISOs, chief security officers, and you'll know that this is a real, real problem which a lot of people don't really think about, right? So, rogue AI usage, how people are using AI in the enterprise, and how are they interacting with these frontier models and so on and so forth, right? Last but not the least, as I said earlier, right?
What's the ROI? What's the net engineering efficiency that I'm gaining on the back of all of these investments, right? There is just It seems like the CFO keeps on questioning, "Hey, I'm putting in this money, but what's the return? What's the return?" There is no answer to that, right? Because nobody is able to actually put that forward that actually we are making or saving this much money,
right? So, these are the things from a governance and compliance perspective that are hurting, in my opinion, the scaling of AI at an enterprise level. But, enough of the problems. How do we solve this, right? Now, let's look at some of the mitigation strategies and how to solve some of these challenges. The first one, the three-pronged approach, is AI across the entire software value stream, right? Now,
what I mean by that is, if you look at the development aspect, coding and testing and development are just two pegs on the entire value stream, right? There is so much out there that's possible across the different phases and life cycle. By the way, this is just the tip of the iceberg, right? There are so many things that you can do. I mean, I like to think
that AI is like a Swiss Army knife. You can do anything that you want with it depending upon your context. So, even this is just the tip of the iceberg. There is so much more possibility out there, right? One thing that's very interesting to understand is the concept of value stream mapping, right? It's a very old construct, definitely not new, but the more I see product teams
struggling with identifying how to rationalize, you know, what's the ROI that they're getting out of their AI tooling investment, I believe that teams who are able to do value stream mapping of their entire path to production, right? Right from how do they ideate, how do they hydrate their backlog, how do they write stories, right up to how do they maintain production systems. If they're able to carve
out that value stream, and in that value stream they're able to identify the the friction points, right? Or the places where where there is scope for improvement, that is the only way to actually take those small steps towards using AI to fix those gaps, those friction points, so that you can start materializing that given this challenge or friction point, I'm using AI to overcome this, right? Now,
that is going to be very critical element and differentiate the teams who are doing tool adoption blindly because there is a shiny new tool out there versus teams who really understand their context, who really understand how their software life cycle works, identify the pain points and the friction areas, and then try to identify what is the right tool, AI tool that's going to help me fix that
gap. That's going to be the difference between a team that's, you know, successful in justifying their ROI and a team that's just blindly adopting AI tools because it's new and shiny. I think I'm not going to spend enough time on this, but engineering practices, right? You remember the picture that I showed you earlier where we had that big hump in the middle which was talking about coding
throughput. And now, you need to be this tall to do microservices, said Martin Fowler 15 years ago. Now, I'm going to say that you need to be this tall in order to actually effectively use coding agents because if you don't have a good engineering rigor, good CI/CD, good test harness such that you can get fast feedback, all your coding throughput is just as useless as it can
be, right? So, you really need that good engineering ecosystem around that because without that, you are not going to get the ROI and the speed that you're really looking for because you're only optimizing for uh one part, right? And needless to say that AI will amplify the good parts and the bad parts in the same measure, right? It will not bother whether the AI is, you know,
it is it good or bad? It's not going to make things good automatically. If your practices, your system are broken, you know, your processes are not optimal, your code is not clean, you do not have a good test harness, AI is not going to magically change that equation. It's going to actually amplify that bad to worse. Right? So, AI will amplify indiscriminately, right? So, if you're not
thinking about engineering practices now with more intent and just giving into the AI generated slop, all the best is what I can say. This is, by the way, from uh DORA report 2025 and it says that the teams, uh based on the cohort, that the teams who actually invested in value stream mapping saw a medium increase and a small increase in how they are rationalizing the speed
of AI or the usage of AI across their team, right? So, this is an indicator that even the industry is noticing the the difference in the posture that the teams are adopting, right? When it comes to value stream mapping. The second one is scaling AI across the org with proven workflows and agentic developer platform. Now, this is the most interesting of all the things that we have
discussed so far, right? You I spoke about the tool fragmentation as the first challenge, right? I spoke about the missing workflows. I spoke about how different roles in different teams are getting local speed, but there is no collaboration or the no no relationships or no team relationship that are getting enhanced because of AI. This, in my opinion, is the missing link, right? The workflow. How do you
go from researching about what to build from a product perspective to going and doing anomaly detection in production, right? Like the entire workflow, the entire value chain. And that's where I had like a very phenomenon happened when we're talking about it internally. I'm sure all of us you know, the importance of doing observability in the production system, right? We look at the cost, we look at the
security, we look at the infrastructure, we observe all of these things. And that was a very interesting moment when we said that now we need to do the same set of things for our development environment. Does it make sense? Why? I spoke about how there are new attack vectors that are associated with your coding agents. Now, if you are not monitoring and observing how your environment is
shifting, what tools are being used, and how they are being used from a cost perspective, from security perspective, from performance perspective, it's going to be a challenge, right? So, there is sort of a shift left from just monitoring the production environment to actually monitoring your development ecosystem as well, right? And that's that's a mind shift that you have to develop. I think it's very subtle here that
instead of just monitoring your production systems, you now need to monitor the entire software creation process itself, right? Which is your development ecosystem. Now, we spoke about the tools sprawl, right? There are so many tools. But what about the workflows? What about the things that connect these tools? The connective tissue between these tools? The tools fragmentation, which is a real problem. The recommendation is to apply platform
thinking. Now, I'm going to, you know, break this down even further for you to understand. What do I mean by applying platform thinking for doing all of these or stitching all of these tools together? Now, this is a logical view of an agentic delivery platform, right? Which will be used by cross-functional teams to do software delivery. Now, the first part is the channels. What are the channels
typically software teams engage with? It's the terminal. It's your integrated development environment. Probably it's a dev portal, right? Now, these are the various where developers and the team are spending their time, right? So, the agentic delivery platform needs to ensure that it provides its capabilities where people spend time. Right? It cannot be outside. It has to be at the place where developers and product teams are spending
their time, right? So, that it has to be available. The capabilities have to be available in the terminal, in the IDE, it accessible through a dev portal, right? That's a channel part. Then comes the agentic products. Now, what are some of the agentic capabilities that are required by our teams at large? We need reverse engineering, right? The ability to understand legacy code. We need product thinking, the
ability to build quick prototypes, do quick research, competition analysis, market research, right? So, we need product thinking and we need forward engineering, right? Given that there is a specification or there is something that we have already carved out on how to do, we need forward engineering capabilities. So, this is the layer of agentic products or agentic capabilities that that span across multiple tools, probably, right? It's not
just one single tool, but probably it's spanning across multiple tools. Then comes the static context. That includes your coding guidelines, right? What are some of the guidelines and architectural principles that are important for your organization? What are some of the reference architectures that your organization needs refer, right? What are some of the ways in which data models need to be created? That's the static context. Think about
all the good ways to distribute this context as agent skills, as MCP servers, as plugins, right? There are plenty of ways in which all of this reusable information can be made available. Then comes the dynamic context, right? Where we're talking about domain knowledge packs, right? How do you do a pricing service in retail in general, right? Like that's something that probably your organization has already solved. Can
that be made available as part of Dynamics context or specifications so that many teams can use it consistently without trying to reinvent the wheel? Sensible defaults from a tech perspective, data products, MCP tools, right? All of these are dynamic context. And last but not the least is the control plane. You remember we spoke about the governance uh aspects, we spoke about the observability for the token spend.
These are the elements from a control plane perspective where we are observing the token spend, where probably there is an AI gateway through which all your traffic, enterprise traffic, is being routed such that you know which team is incurring what type of spend within the organization such that you can put rate limits, such that you can put budgets for different teams so that they're not just burning
money at unprecedented pace, right? So, these are the elements of of the control plane where there are capabilities that take care of evaluating your code, right? I'm sure you have folks have heard about this interesting term called hardness engineering. I'm going to talk about this tomorrow as well in my other talk, but the whole premise of giving feedback to your coding agent based on deterministic and non-deterministic
aspects such that it can course correct itself, right? All of these things can be embedded as part of their of your agentic delivery platform that's powering your software teams at scale at large where they're not reinventing a lot of these things, right? And that's where, you know, if you just put it all together, think about the various capabilities that we're talking about, the guardrails, right? Whenever you're
talking to any AI service or AI agent, there are guardrails that are already embedded through your control plane which makes sure that any sensitive information is redacted before it hits the external servers, right? There are there is a uh MCP marketplace where all your vetted MCP servers are available so that you're not just downloading any random MCP server from the internet which can create the attack vectors
that I spoke about earlier, right? There is infrastructure on what are the set of models that the teams can use with confidence, right? So, all of them are available in the model store, right? So, think about this as a way to streamline software development for your product teams such that they are not doing anything rogue, such that they are not using tools that they are not supposed
to, and all of these capabilities are available to them in the places where they spend their time, right? I think that's the whole premise of applying platform thinking to your agentic software delivery world, and without this I can guarantee that you can keep on using tools, but you would never be able to create that economy of scale with which which with which you are trying to actually
make those investments, right? I'm sure in the past we have proven this fact that platform thinking is the only way to actually kick in those economies of scale because you are generalizing, standardizing, tooling practices underneath capabilities available to your teams to use confidently and reliably, right? Think about uh agents, right? I have a product research agent which has been built and available in my agent marketplace which
can take care of doing the competition competitive research, can take care of looking at what are the customer tickets that, you know, my customers are logging in my Zendesk and create a good product requirements document based on all of this, right? There is a workflow that's getting created here, right? It's not that I'm doing individual research or I'm looking at my customer tickets. It's a agent that's
taking care of all of these steps in a workflow that makes sense for your organization. These agents that you see on the right right hand side are just exemplar, right? Each organization will have their own typical workflows that they would want to, you know, automate into a workflow for for your use case, right? For example, I love to take example from Google, right? Now, all of us
understand that Google has multi-million lines of code across their various products and repositories. Now, they migrated from int 32 to int 64. Now, could you imagine how much time would have would it have taken if they would not have used AI? It would have taken them 3 to 5 years to migrate from int 32 to int 64 across their multi-million lines of code and multiple multiple repositories
just in one part of their product line, right? But with AI, they built a custom workflow that allowed them to do with a human in the loop to go much faster and complete that work in say 8 months, right? So, from multiple years to less than 8 months, but with a workflow that was required for their organization. And I I think these agents are exemplary. with make
is that within your agentic delivery platform, you need to be able to identify the workflows that are needed for your organization, agents that automate the steps consistently in a streamlined way such that different teams don't have to do all the heavy lifting, right? So, that's the whole premise of an agentic delivery platform. And this is again from the DORA uh metrics report. And it says that teams
who have access to quality internal platforms show a very large increase in in adopting and scaling AI across the organization, right? So, platform plays a huge part in the scaling journey for enterprises. The last one, uh the the the mitigation is getting the crowd or the people involved in this change, right? I mean, I'm sure all of us understand that technology is never the problem, right? It's
always the people aspect. It's always all of us who have to really believe that this change is useful, right? So, getting the people involved is the critical element, and that can happen with these four levers, right? There needs to be basic AI literacy across the organization, right? There needs to be a way to foster champions and communities within the organization so there is a lot of cross-pollination
that's happening, right? What's working well for one team needs to be shared to the other team and, you know, create that network effect, right? So, champions and and communities of practices would play a very, very important role in making sure that we are creating a feedback flywheel for the entire org, right? Comes use case-driven adoption. I'm sure you will recall when I said that instead of just
running behind the new shiny tools, you are evaluating your value stream, your software value stream, and you are identifying, "Hey, what are those use cases where AI intervention is going to be useful for me?" And that's where you do a use case driven adoption instead of a tool driven adoption, right? So, that's use case uh first approach, and then an AI playbook. Right? A record of across
your organization, which is maintained by the community, which says what's working, what's not working, so that we can create that, you know, learning organization ecosystem. And And any organization that's not thinking about identifying these levers of change is just pumping money. I can assure you of that, right? There is no good way to get ROI if you are not getting the people in the ecosystem excited about
it and creating all of these levers for them to understand how to go about it, right? So, getting the people People centricity is going to be very, very critical. Last but not the least is the metrics that matter, right? I'm sure a lot of you who are engineering managers get asked, "Tell me those four metrics that you are tracking for knowing how productive your teams are." What
What metrics are you using? Just for my reference. What metrics are you tracking? Don't be shy. Tell me. Acceptance rate. What else? What are the other metrics you folks track? Okay. Good. Good to hear that, right? I think the favorite answer that I usually get is, "We are looking at the team velocity. It's going up." That's a fantastic metric to track. Really? Number of lines of code
generated by the AI. Really? Now my enterprise code base is 50% AI generated. Really? All of these things are shallow metrics, right? They are not giving you real indication on what's the impact of the overall direction that you have taken, right? Now, I I think this is the number one question that I get from all our customers, right? How do you measure what's the impact of AI
on on productivity? And that's where I have I decided to just break it down. There are three elements in my opinion. The first one is the human attitude and opinions, right? What this really means is how are the people on the software team feeling about this change, right? Is it Is the Is the process of accessing documentation easy? Is the developer well-being good? Is there ease of
debugging in in your ecosystem? Is the developer satisfaction higher, right? So, these are human opinions and attitudes. Then comes the system and process behavior. What about your build time? How much time does it take for your build process to complete? How much time does it take for your test suite to run, right? How much time does it take to identify the root cause? All of these are
process-related metrics. And then comes activity-based metrics like number of lines of code written, story points shipped, what's the velocity, code test coverage, and all, right? Now, my hypothesis is that there is no one single metric that can give you an answer whether you are being productive with AI or not, right? It has to be a combination of across these three aspects that I just You have to
think about human attitude and opinions. You have to think about system process and behavior, and you have to combine them effectively with some activity-based metrics to paint a holistic picture on whether we are trending in the right direction, right? You cannot have deterministic goals. You need to have keep on asking this question, keep on looking at these metrics, and say, "Am I making the right moves? Am
I going in the right direction?" Do your experimentation, add your new tools that solve a certain purpose and again ask this question, right? Are we trending in the right direction with the help of all of these three buckets of metrics that I'm talking about. And as all of you said, the Dora metrics, I think they are the golden source. I really like in the 2025 report the
rework rate, right? Now with AI coming in, I think it's going to be even more important responsible for all the product teams to understand how much rework are they doing, right? Now rework is going to be a good signal of the software instability in your software value stream, right? In the In the latest edition of the technology radar at ThoughtWorks, we put Dora metrics again in adopt
because we believe that these metrics are a golden way to look at how you are trending in adopting and scaling AI effectively in your teams and most specifically this instability part that it's getting challenged with AI more and more, right? And this is my last slide, I promise. The journey from A to B, right? Again, these two screenshots are from the Dora state 2025 on A is
a team that's struggling. There are lots of friction points in their value stream. Uh there is burn, there is friction. They are doing less valuable work. Their Their software delivery life cycle is instable. On on the right-hand side, there is a harmonious high achiever where we can see that their team performance product performance is quite high. Their friction and burnout is quite low, right? But it's going
to be a journey, right? You cannot just leapfrog from A to B. It has to go via the learning loops and those loops can only be possible if you're monitoring your metrics and trending in the right direction and taking the right steps and and sort of combine all the things that we just spoke about. And with that, thank you so much. I hope you could make something
out of this. >> [music]
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59