Volodymyr Malyk - AI Org Change Management: Team Transition to the Agentic SDLC and Ways of Working
About this talk
In this talk, Volodymyr, an engineering manager at booking.com, discusses the transformative impact of artificial intelligence on software development processes. He explores the concept of generative AI (GenAI) and highlights the importance of effective context management in enabling AI to produce high-quality outcomes. The speaker shares insights from a case study on transitioning teams to a generative software development lifecycle (SDLC), emphasizing the necessity of generating clear specifications and documentation. By leveraging large language models (LLMs) to create architectural diagrams and detailed technical specifications, he illustrates how teams can automate much of the software development processes while maintaining quality control. Volodymyr concludes that successful GenAI adoption requires a coordinated effort across all teams, with the understanding that AI functions best when incorporated collaboratively into the existing workflows.
Full transcript
Okay, guys. Can you hear me well? Awesome. Okay. Thank you for joining. I know it's a second day. It's 4:00 p.m. I guess everyone is quite tired. So, thank you for coming. I really appreciate it. Uh my name is Volodymyr. I'm engineering manager at booking.com. And I'll tell you about Oh. Yes. And I'll tell you about like I'll tell you a story about AI or change management,
like real case study. Uh how we did uh teams transition to a genetic SDLC and a genetic ways of working, right? And by end of this talk, we will know for sure why GNAI is a team sport. couple of words about myself. I'm in industry more than 18 years. Uh I started in nuclear industry. Then I was in entertainment. Now I'm in fintech. I've been in different
areas. I run in a few communities, like Amsterdam Scala and Ukrainian Software Architects Club. So, that's my LinkedIn profile. Feel free to connect. agenda for today uh three parts. Part number one is the biggest question of 2025. Next one is org GNAI transformation and some lessons learned. Let's start with part one. And what is the biggest question of 2025? And the question is the following. Is AI
really capable? Can it do what is stated it can do, And basically, if you ask me about 2025, it was like the biggest promise, the promise of artificial general intelligence. Early 2025, and it came into two things. Into fear, it can replace people, ruin businesses, disrupt everything. And a huge hype that everything is going to change. It's a beautiful technology, and it's going to change our life.
But what happened next, right? Uh some coding momentum appeared, and we all met what is called wipe coding, And wipe coding, if you ask me again, it a disappointing moment, right? And low quality of wipe code results came up with quite significant skepticism in the industry. Where is AGI? How it can work? Why 2025 so much mess, right? That was 2025. And if you ask like, what
was the biggest challenge in 2025, you know this noise and hype is to is basically to filter right voices. Because too many voices, too many noise, too many opportunities here and there. But how to understand what direction is, And for me, it's quite simple, right? Uh this guy is in the car party. Uh he was one of the founders of OpenAI. He's quite a visionary, right? And
basically, the guy famous by inviting inventing the term itself, wipe coding, right? So, I would suggest to use him as a lens, as a filter, to to look back into around middle of the year, concept of context engineering emerged, right? People started talk that white coding is not as good. Prompting is not as good. What about context? If you have good context, you can build on this
context really good software. Right? And we basically tried it. And what we kind of discovered, identified is very simple but very powerful concept. First thing we did, we created solution architecture, application architecture in a simple markdown with AI. No coding, no white coding, just architecture, right? Simple C4 model. And it was surprisingly good. And the discovery was that once LLM can produce a lot of architectural artifacts,
you don't need it in text. You want to have it in diagrams. And the reason why? Because as a human being, you can easily validate diagrams. You have this built-in accelerator in your mind like your your eyes recognition, right? And you can really fast validate But for machine, under the hood, for LLM, it is still just a simple text. And for LLM, it doesn't matter that it
writes you simple English or in text defined diagrams. And guys, it worked as a charm and it worked as a charm right now. Text defined visual for your architecture artifacts is mind-blowing. Okay? Uh there was something and basically it's aligned with what Karpathy also tell like the biggest uh challenge is this contract between AI able to generate a lot of content and our ability to validate a
lot of content. And text defined visual is a key enabler here. But the next question was what if we have very detailed solution with a lot of diagrams with every flow described in markdown with TechDefine visuals, can AI generate for us something? Can we use this architecture as a context to make it happen, right? And we had like a POC for one single component in mind, but
what really happened blow our mind. We took this back kit spec driving framework tool and the whole service was implemented. The whole thing in not a single class, not a single component, everything. Whole application, whole test, everything. Okay? You can tell me like it's like right? Not high-quality code. I've checked it myself. My team checked it themselves. Our principal engineer checked it, right? Functional requirements are in
place. Non-functional requirements are in Test, business logic, everything, right? for me it was like no way back moment that AI can really do the things if you ship to it really good specifications up front like architecture, right? And when in December I saw this tweet from Boris Cherny who created code that AI can do this, that not a single line of code was created by him, it
really works. I knew it back then, right? And when Karpathy was telling that AI is capable, it's a skill issue. We just need to find out how to work with this, it was crystal clear to us as well. It All the bits and pieces are there. The question is how you can create this context up front so LLM can deliver all the stuff end to end. 2025
grand finale AI is capable, right? Intelligence is a commodity. It's happened already. But so what? What does it mean for your own organization? What needs to be changed, right? And for this, welcome to year 2026, and it's the year of org gen AI transformations. And you remember about all the harm which hype could indeed to the industry last year. A lot of skepticism, a lot of ruined
promises, right? And according to Gartner hype cycle hype cycle, we like coming down. That what was in early January, basically. Huge skepticism. That's it. And at the same time, with recent models, with right approach, some companies and teams are building dark software factories, fully automated software development. And it's happening in parallel, two different worlds. Huge skepticism and tiny fraction people who understand what really happened last year,
at the end of the year, right? yeah. we build a POC, like I started to scale it. And the vision is the following that AI is fully capable, but people and processes are not ready yet. And the first pillar of everything, guys, start with the context, with specifications. And in broad sense, it's like any digital trace in your team. Any piece of documentation is for LLMs, not
for humans, for LLMs. Any Zoom recording transcriptions, any meeting notes, any document, there is a low probability that human will read it. But the AI will read it for sure because what people do right now ask in Gemini, please summarize, please give me an answer, right? All digital trace is accessed through the agents, right? And what does it mean basically? It means that you need to know
your raw knowledge assets, right? Zoom transcripts, Google Docs, Confluence Slack, everything is a raw knowledge assets which needs to be eventually refined it in a source of truth for your agents into product specifications, into engineering specifications, into testing plans, right? Into documentations, into knowledge base so agent can trust your source of truth and it can really deliver. That is the future, right? And if you ask me
how we can do it, should you write it by hands? The answer is no. You should use agent for that. We have these agents already. Cloud code, Codex, Antigravity open flow. You don't need to do it by yourself. You should ask your agent to do it on your behalf, to access raw assets, extract information, verify, and persist it as a source of truth for agent downstream in
the delivery pipeline, right? And really crucial piece here is agentic skills, guys. For those who don't know what agentic skill is, basically a standard from company Anthropic. And it's just a markdown simple document with some tools and scripts. It's like a piece of documentation for agent how to do things, how to perform certain tasks. It's kind like actionable documentation for agent how to act on your behalf,
But if you are in an organization, it's more because in organic from organizational perspective is organizational guidelines for the agents how to behave in the company, how work needs to be done, what are the standards, right? So, for example, if you have PRD product requirement document creation skill it's not just a skill. It's a standard. It's template. It's everything. So, when PM creates PRD with this skill
agent will ask who are stakeholders? What is the scope? What is the security requirements? Like everything. And on your behalf, all the missing information can be fetched from the your knowledge assets to fulfill PRD on your behalf. You just need to control. You need to steer to tell it what is good, what is wrong, right? Use text defined visual to speed up a validation of the result,
right? But in the end skill is organizational guideline for your agent, right? And with this in mind how it really work it in practice in company? It always starting with set of personal skills because it's easy to create a skill for your agent. It's like simple English. Really easy, really fast. So developers can do it. EMs can do it. PMs can do it. And PM do it
much more faster and more than even engineers because they know what they need in simple English interface. Please create skill it does, right? So, it started with personal skills for everyone. Next step is a team because team has coherent scope and some common stuff, some common knowledge, skills, whatever. And it comes to team skill repo. Okay? On top of it, project specific stuff. How this project needs
to be deployed, how this project is maintained, how on call needs to be done. And it's built in as a skills for agent. You don't maintain it by yourself, but your agent do on your behalf with the skills you created, right? And when you have it on a team level, team does the automation. Next level to the top of the like like yes, agent do you whatever
you need. But next step is scale it up to the org level. Because you have your standards area, business unit, the full organization. If there is a quality guideline how PRD needs to be created on level of business unit, it needs to be a skill. It should not be a conference page. It should not be Google Doc. It should not be instruction. It should not be a
workshop. It needs to be a specific skill for which PM can use. And the whole thing is organizational hierarchy of agentic skills, right? And your agent can have skills from each and every level to do this specific work well. You as individual contributor, you as a PM in specific area following company standards and guidelines, right? And basically org structure convey low. Nothing you hear. that's a big
thing here, okay? So, you can imagine teamwork like that. The whole team, PM, AM, developers, right? Everyone has an agent. Like everyone has an agent installed, right? There is a team digital space, meeting notes, second brain, whatever it is. PRDs, solution design, everything in LLM readable form, right? And then you interact. You never edit documentation by yourself. You don't need to do it. Agent does it for
you. Summarization, agent does it for you. Stakeholder report, agent does it for you based on digital trace you provide, right? And the standards for it agentic skills. So, all the digital trade is standardized through the agentic skills of the whole And that's why GenAI is a team sport. You cannot have single AI champion here It's a whole team. It's a whole or it's a whole digital company
trade transformation. And now we come to implementation phase. And what does it mean? Basically implementation needs to be delegated to agents. Period, right? Your SDLC artifacts like solution design, testing plan needs to be LLM ready, LLM readable. So, agent can understand it end to end, right? And how it starts? PM with an agent extracting necessary knowledge from raw knowledge assets like meeting notes, ideas, whatever into clear
product design. And product design in green box because it's standardized through the agentic skill according to the company standards, right? Next step, on top of product design solution design is created. Standard template to agentic skill with flow diagrams, with ERD diagrams, with everything, right? But important thing here is information about infra, right? What it is. Agent needs to know not only information about your product, but about
your organizations. What services are available, right? Like databases, authentication, whatever it is. And this context also needs to be LLM ready, So, having these two inputs, your agent can create the fully engineering design giving all the bits and pieces to connect, right? Works well. Same thing in testing plan, Right? You can create testing plan is Gherkin notation with all the bits and pieces in place. And by
the way, it's really important. Engineering design and testing plan needs to go independently in parallel. I show you later why. But everything is green, everything is for agent, not for humans, for agents. Human just validating, right? Once you have it, next step is like low You put it into you extracting from like functional requirements in solution Basically, application level architecture like which language it is, which framework
we should use, which specific database, SQL, not SQL, whatever technicalities, right? Deep technicalities, how it needs to be implemented, right? And at the same time, from testing plan, test cases are generated, right? We prefer behavioral driven ones. And very important that these test cases are not biased by solution design at all. It's fully based on product requirements, right? And next step is code gen, right? You have
detailed technical specification. connect your coding standards, you can implement everything. Many ways like Cobra super powers, go with it. Spec it, go with it. Open spec, go with Gas factory, do it. It doesn't really matter. What really matter is digital trace prepared up front according to standards, coherent with all the necessary information in When it's done, we're doing the evaluation because any code generation is a black
box. People can review it, but eventually, they can miss something. That's why we have these two parallel streams. Testing requirements in parallel to the engineering requirements, but at the end, we do the evaluation. We validate the implementation as a black box if it really does what we need, right? You can Google it. It's like it can be called like context engineering, it can be called like spec
driven development, harness engineering, all the same thing. Yeah. You need to create specifications up front, guidelines, and let model to implement the whole thing you end to end and validate the results. Just that simple. Okay. And the final step, of course, is production. Preferably with feature toggle, right? You know. what is really interesting here, guys, is that that blue boxes it's what people call like disposable software,
right? Because code generation is so cheap. Of course, you pay for tokens, but trust me, tokens is going to be cheaper. Only cheaper. Inference going to be cheaper and cheaper and cheaper in the future. It's a matter of time it's going to be free. And the coding part is going to be done like in one night, in one day, right? You can implement in for AWS, tomorrow
for Google Cloud, everything, whole thing, you right? And what really matters is your specifications, validated by humans. But implementation is disposable. Regenerate next day, next year, whatever You don't need to maintain all business logic anymore. It's going to be cheaper to re-implement it from scratch. That what is happening, right? And how is different from last year, like year ago? This red box, All the company spent 80%
of effort in the red box with developers, with Scrum, creating Jira tickets for implementation, guys. Imagine. You don't need it at all. Not a single Jira ticket to implement change, to do the bug fix, to implement component. It's not needed anymore. Period. What you really need, you need SDLC around product design. Quality check that product design is good enough. It's not a high slope. Good scores. It's
not with a meter, right? And say enough of PM that it's okay. Next one, engineering requirements. Same thing. It's not a slope. Sign off. It's okay, right? That's what you need Jira for. And at the end, code generation is semi-automatic. Okay, and that's the whole thing here in one big piece. And that's why it's called GenAI transformation because you cannot reach this with AI champions here and
there. You need change the whole pipeline, the whole digital trace, the whole way how teams are working. The whole SDLC, the whole ways of Um yeah. In part three, guys, lessons learned. One simple slide. what works really well is to start GenAI transformation not with engineers at all. With PMs. you interact with agent in simple Agent hides from you all the technical complexities on your behalf, right?
But PM can fast track use like PRDs, explore knowledge bases to build second brain. And once team can see that PM is doing that, agentic stuff here and now, it is not a slope, it's real. The team like starting to ramp up again. There is no skepticism with engineers. Is it work? Because PRD is great. Jira ticket is awesome. Good description. Everything in it. It works, right?
Moreover, in general for some reason, I genetic tools works much better with non-technical people. Right? For example, someone who is like procurement manager who works with suppliers. And they build knowledge base, right? About meeting notes, about supplies, about agreements into knowledge base with an And it's really speed things up, right? And if you ask me who who spends the more tokens in the company, like not company,
but in my like area, it's not developers. It's PMs. PMs who create personal knowledge bases, who like doing like investigation around product stuff. They top token spenders, not engineers. That's the reality. Next one is bottom-up works better. What does it mean? Uh Many people can expect that company will tell them what needs to be done. Top to bottom. It doesn't work because each team is unique, each
scope is unique, every situation is unique. It works only bottom-up. When you gave like agent, structure, tools, guideline, and each trajectory of each and every team is different. Some of them use like for brownfield projects over superpowers, for example. Some is ready to build gas town. It's a different level of culture. And it works bottom-up. Also important to keep an eye on OPEX, We never expected that
PM can spend so much tokens. Just keep an eye on it, keep it under to not spend too much. And the final thing, yeah, the common ground for everyone. There is no way back. No matter who, developer, PM, engineering manager, senior one day know how to do it they will never go back. They tell me, "I need it. I want it. I use it." forever. All you
need is an agent with skills, and if you have any questions I'm happy to answer. Any questions, guys? Any skeptical voices? Okay. Uh the question is is this way is the whole company or just a team? Uh Booking.com Like really big. Many teams, many Uh what we did, we identified these opportunities. We started like POC transformation in the team. But very fast we got traction from other
And now it's growing. Growing fast. Growing fast. Different results, but what I showed is a common ground. Everyone agreed that digital trace for LLMs we need agentic skills, and we need agents. The rest is varies. But the whole idea is the same. Yeah, please. >> Uh when you're talking about um documents like PRDs that are LLM ready, are you writing them in markdown? >> Yes. Under the
hood it always markdown, right? But you don't write it by hand. You have conversation with your agent, quad quad, cursor, whatever. I want to create a PRD, and that's source knowledge assets like meeting notes, stuff like that, right? And when agent start, it speak up PRD creation And it know that prompt is to analyze this piece of information. Skill tells me to create a PRD. And agent
creates you a flag a draft. A draft of your PRD in markdown, but it's for you it's like rendered. For you rendered markdown, not like Let me show it to you. It looks like like this. For you it looks like this. For machine it looks like that. But for you it looks like this. And you have your conversation here. You are telling it what needs to be
You edit it, you adjust it, and you check the result. And once it's ready, you publish it. Whenever you want to any system which supports markdown with text defined visuals built in. you can use VS code, you can use whatever you want. Any agent, any local agent can support this. Good question. Thank you. Thank you very much for this talk. Uh I'm interested how how do you
actually organize sharing these skills across the organization and across the different tools that you can interact with LLMs? It is hard. Because you can do a workshop, you can tell people never work. So, we did couple things. First thing is we create a really good guideline how to onboard for non-technical people. Right? And part of guideline is interactive live code, that I would say. But it gives
you perspective like what a genetic skill is, how to install everything, how to start work. And at the end there is a demo 4 minutes how to create a PRD using a genetic skill. 4 minutes, pretty simple from PM 4 PMs, Once you have it, second thingy, you need to give people skills, right? And now we are talking about this organizational hierarchy of skills. People just connect
to them and enable my role like I'm an APM and there is a plugin with skills for PMs. You enable it, you have everything. You start working. Because all the guidelines is a skill, you have agent with all the MCPs and start is just let's create a PRD. And it takes for PM just one day to feel the difference. How it really works, how to publish it.
How to create Jira tickets not by hands, but through the agent. And Jira tickets become to be much better. PRDs much bigger, right? Much more detailed, That's how it works. Not through the workshops, not through the guidelines, through the set of agentic skills provided as an organizational hierarchy. There is a question if you can take it. Oh, behind you. Hand you. >> All right. Yeah. Thank you
for the talk. Uh really informational. Uh we are also working on something similar on this, but uh we are >> Can you be a little louder? >> Yeah. Okay. Can you hear me now better? >> Yeah, much better. Thank you. >> So, the problem we are facing is the cross-functional team agents communicate to each other. Because right now Cloud doesn't let you do that, right? You had
a slide where you do the PMs and EMs and all those things. So, if they have their own agents, how do they talk to each other? >> Uh the question is how agents communicates to each other in such an DLC. And answer they don't. They basically don't. Because in step one, when you create product design, You sign it off and you saying that this PRD is ready
to >> And this PRD physically is an artifact. Is a markdown somewhere, right? And it's ready to go. And then next step is people just giving an agent link to this artifact and tell him like, on top of this PRD, on top of our infrastructure, let's create solution design and let's validate it. That's how it work. Uh can you be a little uh >> Uh how can
we manage the duplicacy of agents in this case? So, there is like similar kind of skills >> Okay. >> uh which they which are being created as agents by different teams. >> which one is better and which can be used by other people as well? >> manage that? >> It's an excellent question. Let me reiterate on this if you have time. Because it's exactly what's really happening,
right? Once you doing that, each and everyone in your org because to create skill is just telling, dear Claude, please pack our conversation into a skill. That's it. That the whole thing and you have a But when then each and everyone started create their own skills for the same thing, you can and like for the same tool 20 gigantic skills. And people like these numbers. We created
thousand skills. You don't need thousand skills, right? But what you really need we need this organizational infrastructure. Right? And when it's on individual level you can use whatever you want. It's your personal style. What works for you works for you, okay? For the team the same. 7 to 10 people they can agree. But above that subject matter experts the people who really responsible for to make sure
that racing metrics is created according to the standards in the PRD for example, right? And this shipped as a part of organizational structure to the team downstream. That's how it works. So it's like it's wild west on level individual and team above that you still can have wild west. But it's better to organize it through the subject matter experts. That's the way. Oh, yeah. There is a
question. >> When you are producing PRDs and technical designs this way, they tend to become quite big. >> How do you manage the the cost of the in context for the agents to translate that into code mention? >> Excellent question. Excellent question. And that's why you need this step in here. Because engineering design it's big, but it's not as big. It's like still in like model is
still do not degrade this reasoning for one PRD, right? But that's why you need to decompose it into technical specs. The derivative for engineering design is tense of specification with database schema with architecture with different step is implementation plan. Right, many files are generated out of this engineering design. And [clears throat] once you have the technical specs, there are frameworks for this like open specs, spec kit,
that they give you a standard how to decompose things, right? And once it's there, we come into code gen. This piece. And it's done by multi-agents. It's also built into this frameworks. Basically, for each small task in the implementation plan, new conversation with code starts under the hood. Task is implemented. It's check marked in the markdown and next task is taken in a new in a new
session. You can use like guest down, you can use different ways. This problem is solved, basically, like to some extent. Yes, but the execution plan is created up front from engineering design. And the execution plan is quite small by itself. It's manageable, but then it's like decomposed. we don't have One more question, guys. >> Okay, thank you. Uh have you experienced people using the skills and the
agents to do cumbersome administrative stuff besides of programming? >> Oh, yeah. Uh it's really gets traction not only through the APMs, but also through the engineering managers. When you need to collect all the information from project basically like like a small company brain, team brain, area brain, whatever it is. So, it's basically piece of knowledge maintain it up to date as the source of with all of
the inputs, right? And on top of it, you can create like if you use for different stakeholders, if director ask you, what needs to be done? LLM can dive deep into this knowledge up-to-date and give him a report. If I'm an APM, I can ask If I'm a PM, I can ask, right? The magic of agent with skills is that once you have standardized knowledge base, once
you have agentic skills for that, you can have agentic skill for executive level. An agent following the skill, following this guideline, can extract information for him or for her. Indicate a dashboard, one-time dashboard with visualizations, rich stuff, but it No one need to read the knowledge itself. You always work with information through the agent. It's not the same everywhere, but people who started doing that, there It's
just no way back. Uh thank you, guys. Thank you for having me and enjoy the rest of the day. Bye.