jPrime 2026

Agents With Seatbelts: Practical Ways to Keep AI Code Gen Under Control, Jonathan Vila López

41:46 · 03 Jun 2026 – 04 Jun 2026 · YouTube

About this talk

This talk discusses the challenges and best practices associated with integrating AI into software development workflows. The speaker emphasizes that while AI can offer efficiency gains, it often slows teams down due to trust issues, context bloat, and the need for extensive code review. He introduces concepts like using agent-based systems with context isolation to improve outcomes. Best practices highlighted include creating software design and development documents (SDDDs), using multiple specialized agents rather than a single monolithic agent, and ensuring security in AI calls. The benefits of doing AI integration correctly can lead to a 55% increase in coding speed and improved metrics in terms of production code and build success rates. The talk ultimately stresses the importance of careful implementation and holistic measurement of AI's impact on software development processes.

Full transcript

Okay. So, um, who of you speak Spanish? How how many of you speak Catalan? Huh? I know this guy. Yes. He's holding one of the best conferences ever, Spring.io. Huge applause to Sergey, by the way. Thanks for coming, man. Yes. And I just learned that the right pronunciation of Barcelona is actually Barcelona and not Barcelona. [laughter] So, don't say it's the right word. So, I I'm I'm

really happy. I am really happy that our great great great friend Jonathan made his way from this amazing city of Barcelona that we all love. Many of us have been there and uh he will say not how to actually use maybe agents or what is the not to expedite use of agents but how to a little bit push on the brakes with agent agents. So, uh, I

guess that's the talk we need and, uh, huge applause to Doberen. Let's start with uh, my presentation about Asians with seat belts. And I want to warn you from the very beginning, I not I am not a trend follower. You will not see me on the bright side of AI more on the half of the spectrum and sometimes on the other side. But let's start with this.

I want to impact you with few numbers. Some reports are saying that in practice AI is slowing us down 19% of the time. So teams using AI with this report are going 19% slower. There is zero increase on DORA metrics and there's only a 44% trust from developers to AI. Now let's start with raising hands. Who agrees with these numbers? Come on, who doesn't agree with these

numbers? Okay, and the rest there was not a third option. So either you agree or you don't agree. Okay, so uh but this only happens when you do it wrong. It's very common to do it it wrong. Let's start by introducing myself. Uh just to give you a a bit more of context about me. I'm Java Champion, Oracle Ace and IBM Champion and also one of the

members of the Barcelona Java community and founder of three conferences JBCn Debbcn and also AI for Devs. Uh you can find my content in several places dev.2 to Fugj. Who follows Fugj here? Okay, if you are a Java developer, you should follow Fug.j. It's an amazing um free uh content gatherer from Java. Also, you can find my articles in medium and uh Java Pro. But I've been

developer for more than 35 years starting with Pascal and C and then using Deli, the one that has my heart and also Java. Of course, maybe you know me from my past position uh that I was senior staff developer advocate, but you know layoffs happen everywhere. So now I'm principal AI development enabler at GK software uh German company. If you want to know more, you can just

uh scan the the QR code. At the end of the presentation, I will give you the the QR uh that you can use first to give me feedback and then if you do it, you will have the link for the slides. and I'm super proud of being here for this is the third time that I've been here at uh at the conference. So I'm very happy and

uh who has uh followed my past presentations at J Prime. Well, thank you. Well, I come from Barcelona, this small city in that corner, Catalans, we are famous for uh building human castles. uh also for having an uncompleted that now is the tallest church in the world now because we have put the cross on top of it. This is the perfect year to go to Barcelona because

this is the Gaudi year. Who hasn't been in Barcelona yet? You have two perfect excuses. my conference and uh the city and the beach, cheap sangria, you know, all those things. Also, we are famous for our beaches, our mobile world congress, but I would say we are the uh we are super proud of being known by another thing. This is a picture from my yesterday's flight and

I was very touched uh by Let's go with the presentation. Uh why AI is slowing teams down? Well, first because people is not having a lot of trust on on on AI. So because we think with very um light questions, AI will understand all the implicit context that we have in our mind. that doesn't happen uh especially for me at home but also with AI because uh

people with having uh well a lot of experience in in in the developing uh are creating prompts not giving proper uh context to AI. For this we need to uh use SDDDS uh documents. I will explain it later also because AI is producing a lot of lines of code. We will see that this is a very huge or uh yeah it's a very huge uh impact on

our because uh in the end estimation perception versus reality is quite different. So it is slowing down because we need to validate the output because we all review the code generated by AI. Okay. Let's put it that way. Who has a friend that has a friend that is not reviewing the AI generated code? Okay. Okay, that's fine. That's what I thought. I think we share the same

friend. Also, there's a lot of context bloat. So, we need context isolation uh with our AI tooling. And also, it's important to understand that AI is not perfect. In reality, it's far from being perfect. It's vulnerable. Here you can see several companies big names that suffer from uh vulnerabilities by using bipe coding. Who is a fan of bipe coding? And the rest are with me. Fine. Even

Amazon suffer from uh a major vulnerability. So I assume that now you are kind of surprised. Maybe this was not the talk that you were expecting. You thought, I'm coming here just to hear uh how can I use AI the right way or how great and amazing is AI. Sorry to deceive you, but in in reality if you do it right, yeah, there's a lot of positive

uh result on this. You are going to code Hear me well code 55% faster production code is going to increase in a 30% and successful build rates is going to increase in 85%. But this is doing it right. So I will share with you several practices. First the composing not using like a huge uh agent doing everything choosing the right model and being precise on what you

ask the model. We will see differences between models creating a specifications. There are seats in the row in the first row if you want to just sit for the whole uh talk. Um, also context isolation between your uh, agents and AI tooling, securing your MCP calls. So this morning, William was talking about uh, securing the calls to the MCPS. But I will give you a bit uh

a bit of information on a side part of uh using MCP calls guiding the response of your agent. It's this is also super important in order to have the right response having a human in the loop. Yeah, I know that you can have an agentic system with a person that well a person that creates SDDDs, then you have an agent that generates the code, another agent that

creates the tests from the code, and another agent that reviews that code, and another one that deploys. What What can go wrong? And finally, and this is super important, at least for me, is measuring reality in a holistic uh way. From uh in my research to create this presentation, I've seen a lot of messages saying we are going super fast using AI, but they are only measuring

the coding phase. And for sure you agree with me that if you are a developer, coding is only a tiny part of your day. You read code, you understand the code, you you have domain knowledge. Sometimes your nos, the times that you say no are the most valuable. And come on, AI is always saying yes. So this is super important measuring reality. Let's start with the the

composing. So instead of using I don't know wind surf who uses windsurf claude the famous guy on the blog now. Um I don't know codeex copilot chat GPT okay interesting um instead of using only one agent doing everything decompose your tasks having soup agents doing specialized things you will have better results uh and also you will isolate contexts between them. It's very easy to create a sub

agent. These are examples for clo but it doesn't matter. It's the same concept. So it's easy to create a sub agent. It's an MD file as everything and you you tell the agent even which is the model that is going to use which kind of uh isolation which are the disallowed tools that it can use all of several things to the agent and also you will give

instructions to the agent. This is how this is the way you have to behave and then the context is not shared. So you have s sub agents and they will work. And what do I what does I mean uh what do I mean? What do I mean by isolated uh contexts? Well, if one agent is doing something not the best way possible, it is not polluting the

others. So this is a good way of doing uh this second. Well, you could ask why the first best practice starts by zero because we are Java developers and arise starts with zero. Okay. So, choose the with uh Sonar. My former company creates an a super super interesting LLM leaderboard. They have an experiment with 4,400 Java exercises. and they ask those models in the list. Now the

list is has around 60 models and then they analyze the result with sonar cube and these are the results. They check for uh pass rate so accuracy they check for bugs. They check for vulnerabilities and they check for complexity. So if we take a look, we see that Gemini 3.1 Pro Hike is the best one in terms of pass rate. do any of you finds a 100%

accuracy rate? There is no such a thing. The highest is an 84%. But if you take a look to the uh next column, you can see the issue density. So we have we find almost 20 issues every million lines of code and the lines of code we will check that if you compare Gemini 3.1 Pro 300,000 lines of code. If you go down GPT 5.4 uh hike

more than one million lines of code. So just changing the model we have multiplied by four the number of lines that you are going to review. Well we know your friend isn't is not going to review those lines but you I hope you are reviewing those lines. In this list there are other models in green. We have GLM5 and we have Kim K2. These are open models.

The world is bigger than the SAS models. Who is using a hosted model for their AI thingy in their company with an open model? One. Two. three, four, that's nice. The rest you should talk with them and see why and how they do it because definitely it's a good option. I had a conversation with G one guy in Zurich and he works for a bank company and

they are not allowed to send code to the SAS models. So they are uh using um Quen 30B and they are super happy with the result and you don't need to use SAS models. You can control your AI. uh in fact my my recommendation would be start slow with local models small ones and then you can grow and maybe one day you will realize you need or

not if you are using AI for code completion and to create uh methods instead of big com big big applications from scratch maybe you don't need super big models but you can see vulnerabilities is they are introducing a lot of vulnerabilities too and in terms of bugs the best one is GLM5 it's an open model so lines of code as we saw before depending on the model

you are going to invest way more time reviewing than with other models so this is also a criteria that you should you should When you use those models, you need to be precise on the question that you're asking. And usually giving the agent negative constraints, it's going to help a lot. Don't do this, don't use that. And also, of course, check security, check this, and check that.

But negative constraints are going to be very powerful. creating a specs before coding. Well, you need to um give um good specifications to the agents in order to have a bet uh a best um result. Usually what happens is the agents do not throw exceptions. They are just simply going to give you answers that are not good enough. So you need to check um or you need

to give good prompts or specifications in order to have those agents especially if you are having um parallel agents. In this case a spec is a very simple thing. So you are just simply creating a use case. To be honest the older ones in the room they will connect with this idea. a big document of a specifications where you were putting why, how, entities, examples, everything. I

was doing this 30 years ago. Now, apparently it's cool and we call it SDDD. So it's about putting everything that you think first that any element let's say a human or an agent will know how to create that code uh to fulfill the specification but also in order to guide the agent definitely. What are you going to include in the specification? Well, everything, even JSON examples, um

the entities, the fields, everything, even plant UML models to help the agent. There are several uh frameworks. These are only three. Maybe you know more. Uh but AI unified process is from a friend um called Simon Martinelli. uh uh it's a very nice one and he's uh creating the the framework while he's using it for his work. He's a consultant and he's freelancer and he's using this

in order to uh work for customers. Specit from GitHub or openspec that it's also a a framework to to create these specs. uh you can you don't need a framework you can do it manually but if you want to follow a a framework you're to have the same uh process uh across your teams that's a good one the specs are very easy so I mean using open

specs you are going to use commands in order to create an uh an spec and also in order to build the the files that are going to create a plan uh as you can see well it's the one that it is at the back. Uh it will create a plan with all the steps that are needed to be done in order to implement this these features and

it will check every step that the agent is doing. And why this is important and we can connect with u my previous message about having an agentic system where there is only a human creating SDDDs. What happens if you create tests after your code? Who is using AI to create tests after you code and the rest are shy, right? Okay. Okay. What can happen? Okay. You have

five seconds to find an issue in this code. it's a country that will give you a deduction or your income tax. You have the answer. Just raise your hand. You have come to my talks. You know already. Nobody. Come on. It's very easy. It's not Java related nor anything. Yeah. Sorry, it's rounding. Okay, that's one issue. Any other? Okay, the problem here is that the deduction can

be higher than the tax that you need to pay. At least in Spain, you will never get back more money than the money that you have paid to the government. in my country. In this case, what happens is that if you have to pay 500 euros as tax and the um deduction is 1,000, your your government will give you 500 in return. The problem is that if

I use AI to create the tests based on the code, it creates that test. So it checks that for a low value there's a negative value in return. That shouldn't happen. It is not detecting the problem just because source code is not a reliable source. AI doesn't understand the the intent unless you specify it. So the way to use AI in this case is connect it to

this codebase. Yes, but connect the agent to your specifications, your Jira tickets, your documentation everywhere just to find that this method is wrong. There's another way that you can use TDD. So create your test and then ask AI to generate the code. context isolation. Um well, sub agents are not going to share the context. So they are not going to impact uh one to each other. And

you can have agents or sub agents that are more focused on specific things. And then for uh repeatable tasks, you can have skills. I need to create a screen. I need to create a table on the database or to connect to the database something. You have skills that have all the information that you want for the agent to uh execute that task. So you can have all

of them and then redistribute them uh among your developers. I put here this example that it is um a skill from Quest DB. It's a review uh skill that will be executed every time there's a a merge request and it's going to check a lot of things. In this case, it has eight parallel sub agents checking different things and each agent is checking for a particular part

of the code. So this is also uh a a nice example of how you can create subbations and use them. But now you can use uh APMS that it is simply a way of packaging everything skills prompts instructions even MCP servers or plugins in uh let's say in in a package in a folder that it's easy to distribute among your uh securing your AI calls. This is

not only about how do you um authorize or authenticate um users to use the MCP servers. It's about something else. It's very easy to install MCP servers. Now you go let's say Windsorf you go to the marketplace and you click install install install install install install. You have 200 MCP servers installed. they are running locally uh Potman or Docker and that's it. The problem is do you

know what the MCP really is doing? Do you know the security of the code of the MCP server? Who uses more than one MCP server? And who has who has checked the internals of those MCP servers? um that's that's very common. It's very And you can leak credentials because you never know which is the content that the agent is going to pass to the MCP server. It

will decide but you cannot control which information. So it could be that it is sending I don't know credentials because you have them in a file that it is not committed and in the ignore file but the agent has access to those files. Uh it has access to your yeah your uh file system or a lot of So it can uh it can leak secrets that you

put in your prom. you're just simply putting I don't know uh passwords uh users passwords in your prompt and you think it is everything uh isolated but it isn't because you haven't checked what the the agent uh the MCP server is doing or it could be a double agent it could be that you are downloading the GitHub MCP server by the way if you go to one

some of the marketplaces for MC MCP servers. There are dozens of MCP servers called GitHub with the same uh paragraph for the description and maybe you click on the wrong one and they are doing what you expect but a bit It could be that they haven't checked for the vulnerabilities and then they are suffering from lock for shell maybe. So you are sending information to someone that

you don't know if they are secure or not. context uh pollution and po and poisoning because if you have a MCPS per soup agent that's one thing but if you have a lot of MCPS for the big agent you are going to uh pollute all the context this is a very interesting thing I was playing with the sonar cube MCP server also with the GitHub MCP server

and I asked my a my agent give me the most critical issues for me this project. The thing is both MCP servers are let's say issue trackers. So the agent sometimes is going to use the GitHub MCP server or sometimes it's going to use the Sonar MCP server. So what what it can happen is that if you have 20 MCP servers, it doesn't know really which MCP

to call each time. I would recommend uh stop using MCP servers and use uh skills. It's uh way more productive and controlled. Okay, this is the example about the GitHub MCP server. You can see several of them are having the same description and the same name. I can click install in one of them expecting that this is the official one, but none of them are the official

one. Isolation. How can you do isolation? Because what happens if I use the wrong GitHub MCP server and I want to be sure it is not going to send information to a third party destination. Well, you can use sandboxes from Docker and it will isolate which are the network destinations that you can use. Tool hype is a tool that you can use also to uh the address

list or if you are using VS code also you can use the sandbox and specify which are the allowed destinations for the MCP server. This way you know that the MCP server cannot go to places that you don't want to guide the agent. Again, you need to specify in the SDS or in the guard uh guidelines document which are the output contrast the way that you expect

the information to be generated which are the guard rails. Okay, I don't know never use a spring and use squarus that would be a one um or uh never include card numbers whatever. So you need to guide which is the output for the agent and verification check everything and maybe check everything with sonar cube MCP server and see that the output from the agent is correct for

different agents you have different uh documents or files this is something uh I can understand but it's the way the market goes so you are going to create a document an MD file with all your instructions for the agent. Easy. Or you can use hooks. Agents have hooks. Uh depending on the agent, you will have more or less hooks. And you can connect these hooks to scripts.

So for instance, user prompt submit. And then I can have a script that checks uh and deletes uh sensible information. Or I can have a post tool use and I can send this information this output to the sonar cube MCP server and then you know that the output from the or whenever the agent has uh finished send the message somewhere. It's uh depending on the agent you

will have more or less hooks. In this case cloth has 23 compiler 12. It all depends and it's very easy to create a hook. You just define which is the trigger, the matcher uh when it has to be uh triggered and which is the command to be executed. Very easy. Human in the loop and deterministic tools. This is a very important thing. AI is not deterministic. If

you run 10 times the same, maybe you will get uh several responses. So what I suggest is that AI multiplies you but not replaces you and humans give confidence to the system. Uh using good practices, TDD, per programming, uh or review whatever. So all the best uh that you would use no matter if you're using AI or not team knowledge sharing knowledge uh domain knowledge specs written

everywhere and deterministic tools you can use sonar cube you can use PMD fine box sneak lot of things and those are going to give you deterministic results because every time that you run them you have the same what I think is that AI will give you time. Yes, that's fine. But it will it won't give you confidence. It's up to you to have confidence. uh it's you

who is the element that it is providing confidence to the whole system using as I said deterministic tooling like well in this case is sonar cube but you can see using the sonar cube or any other llinter in the ID is super important to use llinters they are free and they are give you the Las Vegas illusion Everything that happens in your IDE stays on your IDE.

So you are going to have someone that it is warning you about all the issues before you send them to the repository before you spread those issues among your team. So definitely use llinters and they are free and learn from them. In this case, sonar cube, I think it's one of the best uh for me, but it's giving you why you have this issue and ways of

solving it. You learn from those tools. That's important. And also use tools that are going to give determinism to your CI/CD pipeline, not allowing you to merge to the uh to the repository. We are arriving to the end. But first, I need a picture because I think you are a bit sleepy. So, I will do a video. Okay. One um I I've been I've been said that

Bulgarian people are kind of passionate. I don't know. I don't have any proof. Maybe this will be the first proof. One, two, three. Awesome. Amazing. So measure holistically. So AI is making us going faster. Epics completed plus plus 66% Task throughput per developer 33 full request merge rate per developer 16%. But this is not the whole story deployment we are deploying less. H interesting. So it is

giving us an increased throughput but it's giving us a lower stability. because pull request size has increased in a 51%. So we are going to invest way more time reviewing that code because we do review that code not our friend but we do and files edited by per pull request it's again increasing so we have more output 200 so double the incidence code turned means code being

touched 800%. We have 400% more time in review 25% more comments but this is the best one for me. This is the best one. 31% more pull requests without being reviewed. Come on, man. Our friend is doing a lot of things. So, because a lot a lot of um messages that you see out there are only focusing on code and testing and there are reports saying that

we only invest uh 10% of our time writing code. The rest is reading code. So we are increasing a 55% on coding. Okay, that's fine. But we need something that measures the whole story. We need Dora space debx or well you can use DX core for that it's a framework that will measure or put everything together because we need to find also we deploy but how many

bugs do we need to fix coming from AI uh and which is the time that we are going to spend on finding and fixing those bugs. So if you do it right with the proper guard rails and proper best practices, you can increase uh the test that pass 53% uh you reduce the um the readability uh space or yeah I don't I don't find the the right

word and you are going to increase also the delivery. So doing the right thing can give you good results but you need to do the right thing. So we are reaching the end of our presentation and I will give you some references first. I will give you the this. So you can click not now but if you have the the slides you can click on the article

and you will see more detailed information about all of the things that I have uh shown here uh in my article about the best practices. Here are more uh links that you can go and and check uh if you want to know more about all this process. And basically that's it uh blogaria and hope you enjoyed the the talk.

From event

jPrime 2026

03 Jun 2026 – 04 Jun 2026

All event videos
Back to Watch