ISTA Conference 2025

The AI Advantage: Measuring Engineering Productivity Uplift

27:08 · 16 Oct 2025 · YouTube

About this talk

This talk addresses the integration and measurement of artificial intelligence (AI) in business operations, highlighting the challenges of quantifying AI's contribution to productivity and efficiency. The speaker discusses the evolution of AI from experimental labs to core operational functions, presenting case studies from companies like Taiwan Semiconductor and Salesforce to emphasize the tangible benefits of AI investment. A roadmap for measuring AI's return on investment is introduced, which includes stages of data collection, monitoring, analysis, and reporting. The speaker stresses the importance of a standard approach to evaluate productivity metrics across different teams and suggests incorporating AI into various engineering processes. In conclusion, measurable AI integration across the software development lifecycle is presented as essential for businesses to derive true value and enhance operational outcomes.

Full transcript

Hi everyone. Okay, it's uh it's past three, so I'm going to ask you to do something. Everybody, please stand up if you're able, if you're willing. If you're holding something, please leave it. Okay. Okay. When you're ready, let's do a big stretch. towards the ceiling. Okay. And now a deep inhale and a slow exhale. Okay. Is everybody awake? Thank you so much. Thank you. Okay. Now we

can talk about AI. All right. A quick post check around the room. Please raise your hand if you're using AI in your day-to-day job. Okay, not surprising with this crowd. What about raising your hand if you know how much exactly does AI help you every day? Can you put a number to it? >> Okay. >> No, no, no. A number to to calculate how much AI helps.

Okay, there's a few people. Okay, that's good. Honestly, I'm I'm not surprised. This is this is very interesting because all of us nowadays are using AI but it's very difficult to put a number to say how much exactly does AI help us every day and we have really seen a huge shift over the past couple of years because AI was really tucked away in some test labs

or some specific teams. They were looking to do a test use case or something to see if AI works. And nowadays AI is really moving into our business and operations and it helps us drive growth at scale. However, it's very difficult to say how much does AI help us. We know that the payoff is undeniable. We're seeing a lot of positive impact from AI. We're seeing that

AI helps us to create more productive teams. We know that AI helps to remove some mundane tasks. We know that people are um spending their time into more value adding activities. We know that it helps us with cost efficiencies. It's helping us to reduce waste from our processes. It's helping us to create more streamlined operations. We see faster times for delivery for our clients and we see

our quality going up and the results overall are very promising. We're seeing stronger teams. We're seeing better products. We're seeing clients who um see value faster. But if AI really helps us to drive all of this, how do we know that um how do we know exactly how much does AI help us? It no longer works to say our teams are faster. It no longer works to

say our code is cleaner. Um because we are making investments in AI. We need to see how much the payoff actually is. Businesses need to see tangible results and managers and engineers alike need to know if AI really is being put to use um and just optimal use. Now there isn't really um a one-sizefitsall formula for calculating the return on investment from AI. However, there is some

um testing happening. There are some experimenting. There is some experimenting happening. We're seeing the tech leaders um finding a lot of information about AI and that really gives us some valuable insights. For example, the Taiwan semiconductor company, you know them, they're very popular. They're um manufacturing chips, microchips for the likes of Nvidia, um Amazon, uh Apple, uh Qualcomm and and so forth. their CEO um CCY recently

shared that by integrating AI into their core engineering functions, they have um realized a tangible productivity benefit of around$1 billion um Taiwan dollars which roughly translates to 30 million US. Now this really um showcases how embedding AI into the core engineering processes um this is where you start to see the benefit in the return on investment from AI. Then there's Salesforce. Mark Pyth CEO recently shared that

AI now drives between 30 and 50% of some of their um processes specifically with engineering and with customer with customer support. This doesn't really give us a precise ROI formula, but it does share something valuable. AI is no longer just a helper tool. It is becoming a coworker and it really shifts how and where human effort is spent. Some more interesting numbers. Um this year's Stanford uh

AI report shows that there is a significant difference um in terms of the productivity boost we see between corporations and startups. Um startups are seeing up to 40% productivity boost compared to corporations where the productivity uplift is only up to about 20%. I guess it's not a surprise given that in corporations usually AI adoption happens slower. It's more controlled whereas startups can be a little bit more

flexible. Um similar differences are observed with uh between junior developers and senior developers as well. um the lowability workers are seeing much higher boost of their productivity when they start using AI compared to the high ability workers who already are very um productive. However, 15% is still um um that's still um an impressive result. But when we look at those numbers, we can't help but wonder and

ask ourselves, how do we know how these are calculated? How do we know if these are consistently measured? How do we know that whether it's Stanford or anybody else that's, you know, going around running surveys and reporting these results? How do we know that they're truly comparing apples to apples um when they go through those different companies or even through different teams in the same company? How

do we know how they're measuring those that productivity uplift? Now, my job allows me to work with all sorts of teams. I am able to work on solving all sorts of business problems. Um, and I am also working on all sorts of business strategy. Because of this, my job recently with engineering teams has allowed me to see firsthand the need of such standard approach with which we

can measure those numbers because we know that this is now the big question. Everybody's looking for the standard approach in which we can measure those productivity numbers and this is exactly what I'm going to share with you today. So I'm going to walk you through a very simple road map. This is uh something which we've come up with based on findings from the biggest and baddest in

the tech world and how they are measuring their productivity impact. Um it's a universal approach. It can be applied across corporations. It can be applied across startups. It doesn't even matter what type of AI tools you're using. You can be using the GitHub co-pilot. You can be using Corser. You could be using Gemini code assist. Whatever. It applies to all. It applies also to all uh types

of tasks, not only code generation. Okay. So, I'm going to walk you through this and you can actually take this away and start implementing this into your business tomorrow. So from the perspective of you being the business or the CEO, you can start by doing the following. Stage one, we need to understand where your teams are today. You need to collect baseline. The first thing you can

do is run developer experience surveys. Understand how your people feel today. Understand how much time your people say they're investing in non-value additing activities versus how much time your people are investing in value adding Understand whether your people feel that they're equipped with the right tools to do their job. Understand whether your people are feeling that you as a business are providing grounds for innovation for automation.

Okay. So understand that human side of things. Then pull all of your data from um Azure, Jira or whatever you may be using and complement this to measure your core engineering metrics. Um what's your engineering output per day? What's the code quality? What are the cycle times? What sort of AI tools have you made available in your business? How many people have access to them? How many

people that have access to them are actually using them per day? So get all these sites and pull them together. Stage two, you need to develop and consistently use a monitoring system. So this data flows over time. you can pull this data from your different sources whether it's your engineers and that survey whether if it's your um systems and applications you're using insight whether it's the the

AI tools accessibility um pull that data consistently track your AI tools adoption see how it's going start introducing new AI tools in your business to understand how um adoption is changing as well as you're doing this continuously run surveys with your people. Understand how their experiences are changing as you're introducing more AI capability. Continuously collect telemetry from the systems to also understand how your core engineering metrics

are changing with this whole thing. Now, the first two stages might sound like something you already have in your business. Um, a lot of businesses already are um tracking a lot of those metrics. You might even be using a smart software um to help you track this. There plenty of softwares out there. Um um there's Blue Optima, there's Dorometrics, there DX, so many of them. uh if

you don't have this in your business, now is a good opportunity to establish this as a system and even start testing and deploying specific AI tools within selected teams so you can start pulling those comparisons early Now, be careful with the data quality. This is one of the biggest pitfalls here. Um the whole AI hype over the past couple of years really unlocked a data quality hype

and those two topics now go hand in hand. Remember the garbage in garbage out rule. So your data needs to be consistent. It needs to be able to tell you something about your business and your performance. If it doesn't tell you anything, you don't need it. So don't collect data for the sake of data collection. Okay. Then we move to stage three. Time to run some data

analysis. Um, usually this is the time when I sort of confuse half my audience when I'm talking to senior leadership. But um, data nerds and statistical analysis nerds like me really get excited. Um, but this is where uh, we can run some statistical analysis. When you have all of your data, this is the time to unlock it and understand what it's really telling you. You can run

some correlation analysis and understand what is the whether there are any significant statistical relationships between the use of AI and in example um your daily output does that grow? Your code quality does that get better with the use of AI? your cycle times. Do these get shorter with Now, be careful when you're running this analysis. There might be multiple factors at play. So, make sure to run

multiffactorial analysis and understand what else is there at play that may be skewing your data. In example, engineer seniority that can give you some, you know, skewed results. Complexity of code also can give you skewed results. Sometimes even the physical location of the engineers can give you skewed results. So make sure you find what those other factors at play are and exclude them or take them into

account when you're reading your analysis. This is the moment where you can start grouping your AI users into an example heavy frequent occasional non-users even and just run those comparisons see and have data confirm what you're seeing. You can take this to the next level and compare the performance of a team using the GitHub copilot versus a team using Corser versus a team not using anything at

all in example and understand where your optimal performance is. Gathering all of this information really is going to become the blueprint for your continuous improvement opportunities. And having all of this data available is going to really help you improve your or onboarding and training programs. Now I'm going to show you how such analysis can look like in a minute. We're going to go through an example. Um

but in terms of the road map, don't stop here. Okay? Don't collect data once to prove a point for your AI investment. You need to do this continuously. And I think that stage four of this road map is really the most important stage. You need to create momentum with regular reporting. Okay. You can in example do monthly um reporting with your engineering leadership teams. Report all of

your findings and that analysis to them so they can take actionable um items out of this. You can do quarterly deep dives and you can have really people partnering all sorts of roles partnering and sitting around the table with you um to to look at this holistically. You can have managers, engineers, consultants becoming a partner in this process. Adapt your um AI operations. The world out there

is dynamic. As you're organization, there may be new effects over your engineering metrics. Okay? So, make sure that you can adapt your measurement system to capture those new effects so you can continuously pull that meaningful data that we talked about earlier. And once again, as as I'm talking to you about this road map, don't think only about code generation. Contrary to common assumptions, the biggest impact from

AI comes in not only in code generations but in all sorts of of use cases. So think about stack trace analysis and debugging. Think about code refactoring and cleanup. Think about test generation and documentation in even learning new frameworks and languages. Okay, is everybody asleep? We okay? All right, let's see how this applies in reality. This is an example of such Um, I am going to illustrate

some of that. For the purposes of me giving you this example today, I've had to clean up this data. I'm not allowed to say which company's trend analysis this is. I'm not even allowed to say where the engineering teams um are located. I can confirm it's a very popular offshoring destination, the large pool of tech talent. Okay, this analysis showcases the use of AI um actually the

use of the GitHub co-pilot within specific engineering teams. It's pulled from Blue Optima because this is the company's um software of choice to track their telemetry. It's um using data collected between June 2023 to July 2025. Um so it's two years trend analysis. Now I'm not saying that you need to wait two years before you can run your trend analysis, but be careful because this is another

sort of pitfall. If you run a trend analysis before your AI adoption is mature enough, that can actually give you skewed results or even discourage you. Okay? So, make sure you're starting to pull that analysis in a moment where you have enough AI adoption to tell you something meaningful. So, what you see on the screens, let me not get in the way. Okay. The first um graph

shows the daily output per engineer because the data is pulled from blue optima. As I said the uh the scale here shows the BCE per day or the billable coding effort as blue optimum understand it. Um it's uh it includes all types of engineering tasks. It's not just about generating code. They define it as the quantification of the intellectual effort of developers per day. If you're interested

to understand more about it, please Google BCE. I'm not here to talk about Blue Optima. Not here to sell it to you. I'm not sponsored by them. Okay, that's just a business example. The important thing here is that we're seeing a strong positive relationship between the engineers who were using the GitHub copilot and their daily output per day. Over time, it grows. the graph in the center.

Do we have anybody working in finance here? We're dealing with finance data. One person. Yay. So that's for you, right? The the middle graph shows the daily cost per unit of output or in other words, how much output you're getting from your engineers for the the investment you're making in them per day. This shows that over time your cost decreases or in other words your daily output

becomes larger which sort of relates to the first graph makes sense. The third one shows the quality of code. Um blue optima calls this abery percentage or the percentage of errors you have in your code. It shows that the engineers using the GitHub copilot over time their their percentage of errors in the code over time decreases. the underlying data behind those three graphs, everything that sits behind,

there's so much data collected that that company is able to pull an exact dollar amount to say what's the return on their investment in the GitHub copilot licenses. Now, remember what I talked to you about, you know, correlation does not mean causation. understand what other factors are at play. We've done a lot of multiffactorial analysis here. Uh and there were some factors at play. We had to

clean them sort of um from our data. In example um engineer seniority of course senior engineers um code of quality is going to be higher. That's going to skew your statistic. So understand what are those other factors at play. When you know what these are, when you know what's causing them, just remove them from your raw data. You need to make sure that any analysis you run

is representative for your large largest pool of AI users. So don't just look at the top performance. Don't look at like teams that have low adoption. Try to keep it at the largest pool of AI all of this may sound complex but truly what it is is gathering the right data and then knowing what to do with it so you can pull the right enabling data for

you that enables you to drive business decisions. Remember what I said about being about adaptability. We need to be able to adapt our measurement system, especially with the world we live in. Um, Gen AI world continuously evolves their AI tools that are emerging every day. So, we need to be able to um adapt our measurement system as well as we're introducing those new AI tools and capabilities.

So, with this in mind, let's quickly go through what the future what the future holds for AI, right? So how much of the AI generated code goes into production? It's one thing to AI for AI to draft code. It's another for this code to be tested, approved and then deployed into production. Currently there is a lot of research and tech leaders are looking how to measure this.

So this these are some examples of how this can be measured and think of this and how to integrate it in your measurement systems within your business. code commits versus production merges, pull request acceptance rates and feature delivery times. What else? How does AI help augment teams? Now, when we think about the engineering teams, the most recognizable engineering structure today is a scrum team. So when you

think about the scrum team usually, you know, we picture the daily standups, we picture um the the cod and quality reviews, the sprint planning, the backlog rooming, the retrospectives. But think about all of these rituals with AI woven in them. In example for sprint planning, we can have um forecast how much work we can actually deliver with the set criteria within the set time frames. For daily

standups, AI can draft statuses based on commits. For code and quality reviews, AI can flag risks. It can propose fixes. It can even help with consistency across the codebase. And for retrospectives, AI can actually run trend analysis based on past um deliveries so that scrum masters actually can have more data to enable them to run continuous improvement in their teams. Here's the most important part. AI does

not replace the team. It augments the team. It lifts the burden of admin and you having to spend all the time in these admin tasks. So you can actually invest your time and effort where it matters most. What else? MCP. By show of hands, how many of you have heard of MCP? Okay. A kag. There's some people here who hasn't. Okay. Um, right. MCP or model context

protocol. This is a universal connector. It really sets a standard for AI and how it interacts with external systems, databases, and other sources. Think of this as the uh USB C type port. We have it on our phones, our laptops, our headsets. That's the standard today, right? So that's the USBC type port is the standard for connecting variety of devices. Same with MCP. It um really enables

AI models to perform various tasks such as calling other databases, interacting with APIs and other systems. This is a huge area of research today um because this can definitely move um the whole productivity game of AI to the next level. So keep an open mind for MCP and think about integrating this into your measurement systems in the future. What else? End to end automation with So far

we've talked about AI supporting engineers, AI augmenting scrum teams, AI boosting productivity. But what about AI not just being a helper tool but orchestrating orchestrating an entire end to end? Imagine AI generating and testing code irrespective of its complexity. Then it changes your automation so it can integrate it and safely um release it It monitors product performance um and it acts to selfheal all the while. This

isn't fiction honestly because companies like Amazon, Tesla, and Netflix are already testing this out in some of their operations and they're reporting some very impressive results. So in conclusion, productivity coming out of AI is there. We know it's we know it is, but we can measure it. We just need to make sure that we create a consistent model for this and this model needs to be there

to enable our people to learn that it is sole purpose so we can make adequate decisions going forward. In order for this consistent model to exist, we need to collect data, we need to analyze and expand our uh AI capabilities. And we also need to think end to end. How do we integrate AI across the entire software development life cycle? Okay, measure it, optimize it, scale it.

This is how you're going to move AI from just a promise to driving tangible impact in your business. My name is Pete. I'll be very happy to connect outside or just message me on LinkedIn to share your impressions of today. Thank you. >> Thank you, Petia. What do we want here?