About this talk
This talk presents the work of the ATB Institute, a private research organization in Germany focused on AI and robotics. The speaker discusses their project on AI-augmented software engineering, aiming to enhance software practices for intelligent robots throughout their entire life cycle, from requirements gathering to implementation and maintenance. A significant concern addressed is the trustworthiness of AI systems, particularly in safety-critical environments where robotic systems operate. The speaker highlights the necessity of using open-source tools and methodologies to ensure transparency and reproducibility in AI applications. Collaboration with various partners is emphasized, along with the development of tools for generative AI, code generation, and dependability assurance within industrial use cases, such as those involving manufacturing and service robotics.
Full transcript
My name is from ATB um maybe very quickly ATB Institute. So typical German name for an institute. Therefore from the beginning our founders decided ATB might be easier for everyone else to remember. So we are located in Braymond. It's a private research institute and um yeah we are also active in software engineering projects in AI domain h maybe um AI augmented software engineering for intelligent robots. So
from the name you already see our focus is on robotics. So um yeah it's ASI is about how to improve software for robots. Um so the whole engineering process and we focus on the entire life cycle really from requirements to the design uh to design to implementation testing and also maintenance. Um and our key idea is to use AI for this. Um but yeah a very important
point is also trustworthiness of AIS. Therefore we have also one pillar you will see on one of the other slides um how we can or you can ensure trustworthiness of the LLMs that we are using or you are using. So project partners we are nine partners um four research organizations ATBC University of Warwick and Cyprus University of Technology we have uh three pilot partners unparalleled KUKA and
PAL and then two additional technology providers it's the open group uh and Aragon um we started December 24 so one month earlier than Um and yeah we are nearly in the middle of the project. So we are about to finalize our early prototypes now and then uh finalizing first period and go into the second period uh where we also want to uh create more public demonstrators and
use this in public. So what is the challenge? When writing the proposal, we looked at um how software engineering and robotics works and we saw that in traditional robotics, some parts, especially with drones are very innovative. Um but in traditional robotics software practices, they remain largely outdated. So uh I first thought to put it on the slides, but I think it's better just to to say it.
they are still in medieval age and the challenge is not to how to bring them in future but how to bring them to recent days it's it's really really amazing uh what we saw there so what are the challenges besides what I just said um we have very complex horigene um robotic systems often these robotic systems are um one piece so it's built per customer um we
have very safety critical environments. So um having there a problem is uh is not possible. So very critical environments and then on top AI often introduces uncertainty. So that's really really a challenge. Then we s thought um yeah why open source can uh help us here already from the beginning and yeah transparency you need this to understand what an I AI is uh suggesting um yeah then
uh you need reproducibility trust in the results that's also not always given but we have to achieve is and then the last point interoperability um we need to have systems working together especially we also um are doing a multi- aent approach like you are doing so therefore interoperability is really a topic um yeah overview of AIA um our idea is of course the the three pilots and
then we have three pillars trust worthy generative AI for software engineering. Um so mainly two parts. One is LLM testing. So to test the LLM that we want to use to create software. Um and the other part is an idea to integrate knowledge graphs into um LLMs to reduce the effect of hallucinating. And then the second pillar is um short from concept to code. So um taking
and be able from the requirements using AI tools to generate an code and also uh test this code. And then last but not least is about the dependability assurance. So monitoring using generative AI. Um so it's more during runtime and then for for maintaining the system. Um industrial use cases I already said KUKA we have PAL and unparallel KUKA classical robotics manufacturing use case um so we
are working together with Kuker located in Brman and they are doing um manufacturing cells they are building they're building not the robots but the manufacturing cells and um as I said they are always built per customer and um when you look how they are doing they are coming from the hardware and when you see how they do um for example the the requirements engineering process it's all
Axel fantastic tool sets on Axel but yeah really um par robotics is um about service ah and here for example first one I forgot to say Um it's really a safety critical area when such an robotic arm hits you. That was the last time. Um second one is also quite critical um because it's about um service robotics. So for example helping elderly people with a robot but
completely different compared to this one. Um their main topic is how can they you evolved their source code from ROSS two to Ross uh ROSS one to Ross 2. Um and this by using AI tools. And then um the third one is unparallel with drones flying over wine yards to identify um sickness or illness of the the the the plants. So I already said this a little
bit um we thought about open source um already from the beginning and by design and our idea is to have this in three pillars. one is reusable components then benchmarks and data sets and open source exploitation. So reusable components especially um when you go look back on the the the slide before with the three pilots they are very different. So you need different agents that you can
plug into um your architecture um to execute the the task that is required. Therefore, reusable components. I think that's really something where um open source by design can help us. Um then the second part is benchmarks and data sets. I think it's really important when you have several agents that you can benchmark them and also that you have um data sets that you always can reuse. um
because we have gazillions of AI tools um but you do not know um how good this works in your case. Therefore, we thought uh from the beginning about creating such benchmarks and data sets that are open source to be reused by others also. And then last but not least is of course this opensource exploitation. That's more or less the classical approach. Um what AIA contributes to open
source various tools. So we thought mainly about these eight different tools. Uh first one is what I already said knowledge graph integration into LLMs tools for testing of LLMs. Then from concept to code requirements engineering tools, um system design tools, code generation tools and then for dependability assurance testing tools, security tools and tools for runtime monitoring and each of these tool set will be made um open
source during the project lifetime. Now this is a little bit typical H horizon Europe slide targeted impacts. Of course we have um uh thought about these impacts. We have three impacts scientific, economic and so societal. Um I think the most interesting one is maybe when you go for um an open-source approach from the beginning, you really have the basis for several potentially new tool sets or products.
So it could be the the workbench that we want to um develop or we are developing um all the tools by itself and these benchmark and and we hope that um the acceptance in using AI will be increased and with acceptance I do not mean this ask chat GPT how to write my email but how such a tool can help you really developing your Okay. Then about
a little bit about collaboration opportunities. Um because this workshop was organized by by uh three projects um we see these four um collaboration opportunities. It could be shared infrastructure, could be joint benchmarks, common data sets and standards alignment. So standards alignment okay is one thing. shared might be difficult but I really see potential in these joint benchmarks and common data sets. Yeah. And then AIA moine and
also other initiatives of course could align on such uh shared benchmarks or data sets. I think that would be really interesting to try something out there. But of course this also has challenges. we have to check the int intellectual property versus openness. I mean for universities and research organizations it's maybe not that critical but when you have companies like uh KUKA involved definitely they are looking at
these topics then um yeah open source is very nice but it also comes with this problem of maintenance of open tools especially when the project is over is it still maintained so we have to be aware of this and then industrial data constraints what can you share from industrial partners and what not and then maybe to conclude I put there as title open source is collaboration infrastructure
so open source um you see my gray hairs I'm already a little bit older but open source is not just about sharing code so many of us still think about source code but I think it's not only source code it's about creating creating an infrastructure for collaboration and you have to enable trust reuse and scaling of innovation because I think that's quite important and of course you
can also look in our website and if you have any questions feel free. >> Thank you. I don't know if we have any questions. okay if there is someone else otherwise I will Yes, there is another one. I don't know. We can debate the first or we can ask some agent. No, it's okay. >> It's an urgent one maybe. Um you see you're talking about infrastructure and
I think that's an interesting point. Um basically in almost uh everywhere when we talk about infrastructure we talk about something that is funded one way or another to guarantee a service like roads, railways, uh electrical networks, communications, telecommunications and so on. So in your view if you talk about open source as an infrastructure as a collaboration um how should we look at that aspect you talk about
maintenance also of open source and I think it's related yeah so one point is for example especially when you are dealing with um LLMs one point could be where to execute these LLMs This has to be uh defined in such an infrastructure uh because otherwise everyone uses its own and and that's maybe what we could or should avoid and then of course also you need to maintain
this infrastructure and not just put it somewhere. I mean um for example, KUKA will definitely run this on its own infrastructure but we as researchers when we want to to work together and collaborate on this then it could be maybe an option to have a common shared infrastructure for this >> with infrastructure you can mean foundation >> I don't know I was just But like because we
are Eclipse Foundation and of course I need to say that but yes we are a kind of hosting for this infrastructure open collaboration and so that could be maybe an option and we are we are really happy if you want to join. >> We already joined. >> Okay. I wanted to ask you I saw in your slide with the benchmarks uh that you also said that you
measure trust. So how do you measure trust? Is it you ask people or how does it work and on the next or maybe previous? Oh no here >> the matrix. Exactly. Yeah. So >> uh yes exactly um the idea is to to ask the developers. Okay. So this will be the trust of the developer in the code that was AI assisted some. Okay, clear. >> Do you
have some results out of this or it is just you're still in planning at the moment? >> No, we are starting next month is the plan to start the testing or evol first round of evaluations in the pilots. >> Okay. Thank I think it's a it's a great idea. I mean it's important to do both because I suppose the others are some are automated and this >>
I think the combination is very good >> and the especially if you have in each bigger company you have different type of developers or people involved in this and they look on these tool sets completely different way. >> Um yeah I have a question too. um what is specific of robotics in your research contributions in the in the four pillar because I they as you presented them
they sounded much more g I mean that something that could be applied to a lot of other domains. >> Yeah that's what we are hoping that we can apply this also in other domains. Um but what is very specific to the robotics domain is that you uh these the safety critical um aspects because of these um yeah moving things and of the robots and and for example
>> and of course um the code that is generated is often for example in Kuker's is it's their own uh robotics language which not many other tools know or it's PLC code which is also not that easy to generate because there are not so many people that's maybe also the second aspect okay and I have also another question bit more technical so um you build agents do
you use some of the shelf frameworks to build agents that are in the market or you everything built from you from scratch. >> No, no, no, no. We we for example for for some parts in the beginning we use these papurus for Eclipse and light LLM and and so also offtheshelf stuff. Are you developing everything from >> No, but I mean you no absolutely absolutely not. But
I my question if you use another agentic framework I mean you know you use lightn to connect to the to the to the lms or you use um papers to design the system I understand that but you don't use an an aentic framework so the agentic part you are developing it I >> okay there are other questions u otherwise I think that we And um I mean
you are sister projects. So you in theory are working on the same topic. Okay. In different perspective. Uh in this case is robotic. Okay. And I heard you say Ross. >> Mhm. >> Right. Okay. Um I I know Ross. So that's why I remember that. And Ross too because you were talking about integrating that. Uh my question since we are talking about collaboration and you know open
collaboration and so on uh it's early and you are in this in this way you are not competitors so can you see a kind of uh collaboration between the two project I mean except the technology that maybe is different but one can reuse the because Ross is a a platform that you know is very famous using for robots and you are using agents and you create a
platform for taking decision. So I see the one on top of the other maybe I'm wrong. >> Yes. Uh yeah, I mean this is the slide where he was uh listing let's say I I I would of course from our side we are very interested in shared infrastructure. Okay, let's be let's be clear. Um for instance for benchmarking uh we we provide a a benchmarking infrastructure for
continuous benchmarking of the agents, right? And um so it's true that here in this slide the shared infrastructure is not in bold um because it it's a cost. It's a it's a big cost. >> Um but the cost does not make sense if there is not a gain uh that is comparable on the other side. And um so the question is u uh for instance if if
in a platform for um continuous benchmarking you find a gain that could justify the integration or things like uh then I don't think that in your case you have a lot of agents right you you're developing like one agent per phase for instance if I understand correctly one for requirements one for uh yeah for design etc so you're maybe not that interested in the large scale collaboration
that we are studying um but already the fact to to to evaluate against metrics studying the the KPIs I mean this is a big part of our job at least we provide infrastructure to do it. >> Yes. Because besides that the a curious question that we cannot solve now will be if uh one solution can be applied to the other use cases and vice versa. Of course,
>> that's a good point >> because usually we say yes when we develop something this is these are the main use case in which we applied our solution and then can be applied elsewhere. But I I never believed so much that because sometimes like you said uh in some specific scenario like very specific industries is not really so easy to apply everything. So that will be a
question maybe for the future or for our next meeting that will be in one year here and we can discuss if we can really uh do that because uh since you are sister project always my when I look at sister project I also have a project that has sister project and I also wondering why I I also wondering if it's really possible to apply the same solution
Even if it's the same, you know, the same topic, the same a not the same path because the two of you follow a different path of course because it's a different approach but the goal is the same and you provide the you know the results on different uh solution different not solution but use cases and I I want you to think about that Maybe we can discuss
about that. We can have a paper about that probably because it will be interesting. See uh also theorically if we can do that. Maybe I'm just bubbling around. >> No, no. And that's a little bit also the reason why I put these two in the middle joint benchmarks and common data sets in in bold because um I think that could be something what you can really extract
from the projects and then reuse in the other project or use or easier at least and because naturally you build your project around um the pilots you have. and and therefore it's not not always that easy really to one one reuse the one idea >> to another >> to another. Yeah. >> Yes. The data for sure but then we have a question for data for sure but
then we have a question about the standard because interoperability it cames also for data. So we need to have the same data structure otherwise we need to you know adapt uh the data that is really easy you know I mean you just have to translate the data but it's still something that is not very fast and I was thinking more a solution like okay let's try this
or let's try this and that's not really easy. Yeah, indeed. Um, unfortunately say I don't see clear overlapping among our um um use cases. Um I'm I'm going back just to infrastructure to say to say one thing. So one thing we really cared about because collab collaboration is really a a core aspect of our project. Uh so we really care about the fact that we are based
on protocols and that these protocols are not not come really from us but they are more or less standards in industry. So um if your agent already supports 82A and MCP that are let's say the two main protocols around there today to ACP but this is another issue. um then already they can be integrated at at some level okay you will not have u I mean some
features that we provide only with the extension but already with the benchmarking part for instance would would already work. Um so um what I wanted to say is that typically in these projects it's clear that you will never use as an agentic framework our agent framework because it's a it's in progress. So you already have a lot of sorry a lot of bugs coming from your own
code. You don't want to have also the bugs coming from our code. >> Yeah. >> Uh so you want to have a the the the components that are not developed by you. You want to have them of a high industrial quality let's say. So I would expect project leaders to use um I don't know Microsoft framework or Google framework things like that um also in MIT license
etc but of of much higher quality uh today um but that's why I think protocols help here because uh uh even so if you use any framework that supports A2A okay you can have for now your demonstrator and then at some point you still have 8 to8 components that could be uh compatible with our framework frameworks and this can make really concrete collaboration. Sorry, I'm I'm one
of the guys that thinks that open source is still source code first. >> yeah, sorry. >> Sorry. It's just because source is in the name. So uh so I think to that uh first >> no that's fine. I mean yes I was suggesting that um yes assuming that we have a stable components in open source uh that will be more easy and of course nothing is uh
100% secure nothing in 100% without bugs uh I am a computer scientist so what I say about the software is that there is bug I just didn't find yet >> so that's So that's um that's something that we have to face anyway. But it was a it was curious on my side this aspect. I think that we can uh yes we have still it's not an interrogation
for you. You find yourself standing here but it's just an open discussion with the other project. So I also have a question when you talk about trust right you said that you the trust and the the reputation and everything comes from the evaluation that the developers have to do when they are receiving a solution right the feedback. Um so and we were talking about how the two
projects AI and Moso can collaborate. So do do you think that uh the Moza platform could be used for you to have another layer of trust that comes from the the whole uh process of the agents you know with the supervision with the the need to be uh to have a consensus? Do do you think that adds additional trust that you don't get only by having the
developers in the loop to have a more refined solution before it even goes to the developers? I would answer like this. In principle, yes. But uh then we should organize a meeting so that um we go into a little bit more detail how this could work because uh sorry but on on the slide it looks nice but if you have to integrate this into your stuff then
yeah I'm happy to that we organize something >> Sure. I mean I think that makes total sense and that's the well one of the reasons uh um it can collaborate with open source because everything that we are doing uh at least in mosaic is is open you can download it you can install it you can do whatever and u then even before we have the discussion you
can already look at what we are doing right and um definitely I think organizing the meeting I mean I cannot speak for Maximo. But I think that makes sense. >> Thanks,