Trustable Software Framework for assuring integrity with open source based systems
About this talk
This talk presents the Trustable Software Framework, aimed at ensuring integrity in open-source systems. The speakers discuss the challenges of software reliability and the need for tools that help manage software behavior in complex environments. They explain how version control, particularly with Git, is essential in maintaining integrity and tracking changes over time. The framework incorporates a model that represents relationships and objectives, facilitating the analysis of software behavior and compliance with safety standards. The speakers also provide real-world examples of the framework in use, particularly in projects involving safety-oriented software development, highlighting its applicability across various regulatory contexts.
Full transcript
[music] >> My name is Lucas and um Yeah. Uh my name is Rowan and we'll be talking about the Trustable Software Framework um for ensuring integrity with open-source based systems. Uh so Lucas will uh run through the rationale and a bit of an introduction um and then I will talk about um a bit more a deep dive into some examples and some like use cases that we've
used this um and then Lucas will finish off looking at it for an application. Yeah. And before we start uh what do you think when you say we are checking the integrity uh of a product or even with uh your work. Do you think that uh do we need something to help us to define how software complex software should behave or can we guarantee some multi-core hardware
environments the behavior when they we are looking then can you be sure for example, that you know the behavior of your product all the time. It's hard uh how I say it the correct behavior, but uh you can have a guess and uh try guarantee that it's always going to be the same. And uh at the same time, we need to think about the safety standards that
we are trying to be compliant with our product. Sometimes, they are only focusing in certain aspects of it. They just think about the hardware fault and they ignore the random behaviors that we can uh we the environment can can introduce or even some piece of software together can make them uh coexist. And uh impact the your whole idea about how the system behaves. And of course uh
if you think about the complexity, sometimes is not really easy to check. So, as a result, sometimes software are not safe, secure, and of I don't know a software that is bug-free at all. I mean, I don't believe softwares can be bug-free. Okay. And uh the idea is achieving same 100% confidence in most of the software is not really realistic. So, how can we do better about
it? First thing, if you we have requirements, ideas, uh what we software should do, we should control version from since day one. So, we can know if something changed when we are over time or like our initial requirement wasn't correct, so we know when it changed. We can maintain integrity from the single source of there's a proof to reduce ambiguity. And the best part when you are
have a different teams working at the same time, you can always keep auditing and checking the change in a way that we know when that can happens and how and when and in efficient manner without having to ask every single group or every single person. And if you have documentation around it, we do we can always keep them up to date with checks or anything. also, we
can from time to time review them and adapt to match any requirements. And we can approve that whatever we are doing, we can connect them from links to our require from our requirements to implementation and also verify if they are true and they are correctly. I also show how the progress from the concept of our idea to the delivery of our product or our package is connected
to every single piece of the software or product. And then, let's think about it. If we needed to to organize this whole thing, how we were we could do this. The first thing that we think as a software engineer as our close our users related to something to version control nowadays is Git. Git is quite popular, and is the best and I think if you don't like
Git, well, we are in the wrong talk, sorry. Why we love Git for this kind of work? It's familiar to almost for every single engineer nowadays. I think for the 15 around 15 years, Git is the most common tool for this kind of work. I remember when it was before we had version one, version two around FTP servers, and so and so forth, and a lot well,
different tools for this purpose wasn't reliable at all. Like uh I don't want to get into this, anyway. Uh Git helps uh as a transparent tool, and it's also easy to engage because when we talk about Git, we are talking easily about GitHub at a minimum or GitLab or anything similar. So, they are just Git forks with nice a nice interface to open change, which is quite
uh nice and easier for reviewers to engage with. It's easy to transfer knowledge, so when you are doing review, you have a fast cycle for feedbacks. And you can um share the responsibility about something for approving, reviewing, so and so forth. one of the most interesting the features is that you can uh always add some automation around your pipeline. Like you have concepts like CI/CD where you
can add some verification like when you are adding one change to your product is this change breaking everything is what I hold as truth. Still true still true at the end. the last thing about Git is nothing more beautiful than everything inside the version control where if you need to double-check when it was the change one change has happened to a file you can always trace back
to the beginning of it. So, you know every single change and why that they happened to it. And this is the most most exciting thing for me at least when we are talk about Git. Okay. But do you think Git is enough to help you to build your product in a sane manner for documentation checking if everything is working as you expect to my requirements are right?
Yes. But you you'd have a lot of discipline. And what if sometimes discipline is not enough? And if you want to have something to help you to be in compliance with different standards like ISOs, IECs, or even different kind of standards like CRAs, ITU COBIT or whatever. you don't want to be locking into a specific vendor like GitHub or GitLab. You just need like I need I
need a Git repo and something to execute my CI/CD. And if that thing could be open source. And if that thing could be even extensible I don't have my maybe piece to do what I want, but I do have some I know something that maybe can help me, maybe by someone else that I can adopt to my use case. Something that I can be grateful your community
around it the the this tool. And you can even reuse as it is if you need to. And then we we are we introduce you guys Trustable and the Trustable Software Framework. With Trustable, we want to achieve something that help us to analyze and soft from software and specific use cases. We want to analyze what could go wrong for those use cases as well. We we want
to plan something to fix monitor and even mitigate this the problems that we find it around it. We want to demonstrate that what we think about our product our software is defining a certain set of behaviors where we called okay, we know what it does and I can prove that they are always going to be true all the time. And if we find failures for our goals,
how do we mitigate them? We want to demonstrate the mitigation in place and how do we monitor that mitigation still works all the time. And also we need to show how do we believe that our mitigation works how confident we are on everything that I have to just said. Just think that this is not one process thing. Like it's not going to I'm just going to do
once and I'm going to be safe all the time. It's something that you have to practice you have to add to your pipeline. You have to repeat all the time. So it's something that it should be practical, easy to use and repeatable. can we we apply this other how can we apply this in a manner that it doesn't impact the workflow from day to day for software
engineers or anyone that I use Git as uh guide your product who's there your content. So, as part of framework we are just going to focus on on three aspects of it. We are going to talk about the model where we try to how to express complex ideas or focusing in certain elements or relationships. We have the objective objectives that defines what is important for our goal
with this product. And the two and Rowan is going to talk around the tooling and the uh the current implementation to help us to achieve this goal. Let's talk initially what is a trustable model. Trustable model simply is a graph information that we connect and we describe uh their relationships. Basically, a trustable model is a set of statements which are connected by links as you can see
the image on the right side. We have statement X connecting to a statement Y through via link A. Links are just connections logical connections indicating uh statement the X implies the other or the other. The set of statements also can have a product and their links form create a bag, a directed graph, which means that easily you should they shouldn't be cycles around the graph. To remove
any kind of circular Okay. Sorry. I had Statements can be classified into overlapping categories, requests and claims. As you can see in uh I have a request X which can be also interpreted as expectation. we can and this expectation can be connected to another request or a new claim which can also call assertion. And we have at the end connecting to a claim. At the end we
also going to be a premise. can produce a combination of a trust graph which is a combination of statements and links where uh you generate the model. And here is our way to describe your statement the good practice for a trustable statement. Try to use indicative mood third person perspective all the time. Uh present tense. Make them always affirmative. Stick to single sentence when possible. Don't try
to create complex things. Like for example look at the good examples. We have this trustable product provides a set of implementation Python. It tries to be expressly if it's true what it does and how what it offers. At the same time a a bad statement is when you maybe a the trustable product should use Python should you use a different language. Uh what about Python? What kind
of tooling do I need to specify the documentation? No, it should be clear. We use X to for Y reason. So I'll start talking about the actual trustable software framework and I'll ground some of the stuff that Lucas has introduced with some real examples. Um I'll first talk about the objectives of the Trustworthy Software Framework, the kind of objectives that we believe applicable to all projects, and
then I'll ground these and use some actual examples for some real projects that I've worked on. Um so the Trustworthy Software Framework is is lightweight. It's meant to be implemented in your workflows already. I don't want you to be changing your workflows massively. We want you to be just using what we're doing to augment your workflows. so let me introduce some of the tooling. So here you
can see TSF provides you the tooling to do this. So provide you the tooling to manage these graphs that Lucas talked about earlier. And we'll see some more examples of them soon. And provides you the tooling to produce these reports based on the graphs. So as someone who wants to know how my project is going towards its objectives and its claims that it's making towards safety or
compliance with a standard, I want to be able to produce a report. Here is an example of part of that report. I'll get into some more examples in a bit. Um so the objectives, these are the objectives that we the the TSF maintainers have come up with. We think these are applicable to general software um projects. And they are they serve as a starting point. So individual
projects may care about different things, but this serves as a starting point for you to base um your objectives on top of. So we split these down into two categories, tenets and assertions. So let without further ado, let me show you what they look like. So here you can see that graph structure that Lucas talked about earlier really starting to come into effect. We have the top-level
claim, which is this project, so XYZ is trustable. Um and then we break that down into what we think makes a project trustable. So, these tenets here, um proven providence, construction, changes, expectations, results, uh and confidence, break that down further. And then below those, we then have assertions. So, breaking those tenets down even further and then allows you to really build out from that. Um included in
the TSF documentation um is also examples and advice on how to evidence towards those claims and how to map safety standards um on top of those claims. So, this gives you a starting point, but obviously it's your graph for your project. You can change it as your project develops. So, let's zoom in and talk about some real things. As a developer, you know, I want to talk
about some real things. So, uh let's take one of the assertions. Um let's take TA validation. And TA validation basically mandates that tests should have both normal and edge cases. Seems pretty reasonable. I'm sure that everyone agrees that we should both be testing our edge cases and our kind of normal expected paths. And so, as someone evidencing that, as someone wanting to prove that that is true,
I would then break that down further into I'm going to write an edge case test and I'm going to write a normal case test. And then underneath that, I may have the actual results of real edge case tests passing and the actual results of real normal case tests passing. So, those are the red boxes behind me. then obviously if I'm writing a test, I'm not just evidencing
that part of the trustable graph, but I'm also evidencing some other parts of the trustable graph. So, for example, TA behavior says uh mandates that required behaviors are identified and validated. So, my tests my tests that a my normal case is passing also shows that I'm validating one of my behaviors, my normal case. If I'm adding a product option and then I'm validating the normal case for
that product option, then I'm validating that I've identified some behavior and then I'm validating it. So, you can see here how from a test at the bottom, you're evidencing multiple sections of the graph um because that's quite important. So, I if you change your test, you want to make sure that you've updated your behaviors to account for that as well. posing all this together, we end up
with a kind of feedback loop. So, the idea is you have your you start with your TSF graph with your statements, objectives, all of these things. Um this then links to evidence. Your evidence could be generated by CI um and then you do the validation of your graph in CI and then that tells you where you're missing or what evidence you're missing and then you go round
the circle again, generate new evidence. That evidence could be generated by CI. You then put that into the trustable graph and it identifies new areas that you're lacking in. So, the idea that you're constantly feeding back and then updating your trustable graph and your overall score. um ground that in a real example for a real project. So, let me introduce the So, we use it we use
TSF on a platform project. So, a platform project builds an image for a piece of hardware. Um and we want to make claims about that platform project. Platform projects are massive projects, so they integrate services, drivers, applications. They're huge things. For example, an entire Linux distribution will be a platform We have a platform project internally. We call this control. Um and this is a Linux-based operating system
which aims to produce a safety-orientated Linux-based operating system for the automotive industry. control is a as you can imagine a massive project. It has many components which need regularly updating. It has many engineers working on it um and therefore we need to track our progress towards these claims that we're making about it and make sure that we're understanding the impact of any changes we're making to this
project on those So. Uh these are some real statements that I've pulled out from the control graph just to give you some examples. So, under TA releases, I have a I have three. So, you can see that we number from this you can probably guess that there are other changes one two at least. Um but changes three says changes are not made by shortcutting or subverting the
normal process controls which I'm sure we can all agree is a good thing and we should not be subverting the normal process controls when making software changes. And then I may have a evidence item. So, changes evidence 11 um supports changes three. And changes evidence 11 says the GPC configuration for control ensures that a merge template is set. So, that is helping us to not subvert the
normal process control mechanism because I'm saying that any changes to this project have to go through this merge request template that I've set. another example, if we go back to the kind of test example that we had earlier. So, we have an actual statement. So, validation pre-merge. It says, "Pre-merge test scenarios must include exception handling test cases." Um and then underneath that we have specific actual fault
induction tests uh that are linking and supporting that statement saying, "Hey look, I'm running these as pre-merge. Here are the results on pre-merge. Whether they pass or fail, I'm running them as pre-merge." So, you can see how for a project you can start to build out your graph for your specific use case. So, I've included specific technologies that we are including there. So, GPC, uh which is
GitLab project configuration configurat- configuration configurator or configurator. Configurator. Yeah. Anyway, configurator. So, where I'm including specific technologies that my use case is based on. Um building that graph out specifically for this project. once it's up and working, these are actual reports produced by Control. Um and you can see here obviously the report is humongous, the graph is massive. Um has thousands of ele- thousands of nodes in
it. Um and so, some of here at the top level though, I can track my progress towards certain goals. the score from these bottom items, if they're evidenced or not, bubbles up to the very top and I can see an overall progress towards certain goals. And you can see here that we have both the compliance for Trustable, so these are the tenets from earlier, and we can
see that I'm also tracking the compliance for Control. So, these are statements towards a safety argument and my progress towards them. So, as a developer day-to-day, what does this look like as an engineer on the ground using it? Someone that's going to be making changes to this project that's using this. So, we'll go through two scenarios. We'll go through making a change to the project. So, when
I make a change, what do I have to do to continue that compliance and make sure that nothing I haven't broken anything by making my change. Um and then we're going to talk about adding a new node to this is actual screenshots I've taken. So, say I'm going to go through this example here. So, you can see here that in my terminal I've changed the CI system
to CI CD system from using Podman to using Docker. I know I you know, a change. I then want to understand the impact of that change and how that affects my project as a whole. So, I run Tru Graph, which is the which is the name for the tool, the TSF. So, I run this lint command to check if anything has been affected by the You can
see here that I get this output. So, I've I can see here that this specific item has been affected by my change. And so, I've opened up that item on the side here. uh references and a score and this the item statement here. My head's kind of in the way of the items statements, but basically the item statement says that's control pipeline runs on the smallest runners
with no internet access. So, this you can see is unrelated to what container I'm running my pipeline Um but the way that this change detected that this was relevant is that I've referenced this file that I'm changing here. So, any change to this file here would mean I would have to check all the statements related to that file to check that they are true. So, I fully
understand the impact and possible impact of making that change. Um nowadays, we have slightly more advanced things, so you can reference specific parts of files um and things like that, and that's an extensible interface, but the main point is that I am I've made a change this file, I need to understand any statements or claims that I'm making that rely on that file and check that they're
still true. As the person that reviewed this item and said that this is true at the moment, I can then set that item to true um using the true dag manage set item. I can add it, and I can commit it to my uh project. So, for adding a new evidence new evidence to the graph, so I'm going to add a test to the graph. So, here
you can see the new test being added. and so, I've added a fairly basic test. It I'm checking for duplicate services have not started in error. Then, I'm going to add a new evidence node to the graph that references that test. So, this evidence node um I'm going to reference the test that I've written, so you can see that at the bottom here. Um and so, any
change to means that the person changing that test has to evaluate if my statement here is I'm going to use the evidence item, so instead of having the manual score that we saw earlier, I'm going to automatically validate this. So, when you run the test, this this validator here is going to check the output of the test that it's true. And if it's not true, it sets
the score to zero, and if the test did pass, it sets the score to one. And so, here are the terminal commands that I ran. Not a massive amount extra, so I was going to create that test anyway, and now I've just created an evidence item uh to reference that test and then I've linked it up at the bottom to the rest of the graph. Now, obviously
it is not a mint like, you know, it's not a no workflow like it adds some time to your workflow. So, let's kind of evaluate this compared to some of the more traditional workflows when looking for compliance. TSF workflow is to make a change. It required running a few extra commands. We had to view some files for adding a new node. We had to create an extra
file that we wouldn't have otherwise and run two or three extra commands. So, the overhead is minimal and it's spread across the entirety of the project's development. So, I'm doing this as part of my normal workflow. I'm not adding a step at the end of my project development where I have to scramble and gather like, you know, thousands and thousands of pieces of evidence to go for
compliance for a safety standard. I'm collecting that evidence as I go along. I'm tracking my progress towards that evidence as I go Um and I'm I can understand better where my current project is at and how far away from perhaps going for a certification I am because I'm tracking progress towards if you're in a more traditional workflow, if you are trying to do this as you go
along, what this would actually involve would be chasing safety experts. Like, you would if I made a change, I would have to chase safety expert to assess the impact of my change and assess if I needed to collect new evidence. What this does is with managing it as you go along, the experts have to be can be less involved and they can perhaps if you need to
add a new evidence node, they can review that new evidence node if it's critical, but if it's already been mapped out, then it's completely fine. So, we've kind of seen how TSF works on like a platform project on a big Um Lucas is going to talk about how it might work for an individual component project, which is a bit different. Yeah, and uh also uh before getting
into a component level, we should talk about the community around TSF. Currently, we have uh some work uh by the guys from from Define. Do we have someone from Define here? No, okay. Sad. Uh we have also some more work from Code Defending a open source called Safety Monitor, which uh we applied TSF uh on this product which is a an older version, but uh it's still
I I would be able to wait to see this. And then we also have some more work TSF with Deno uh is a part of the Eclipse Foundation member. I think. The one that uh we are going to talk as part of a TSF in a component, we are going to analyze briefly uh the work done by the Define guys in the new Nolom Nolom JSON library.
I am I'm not sure if you know this library. It's a well-known library C++ library. Which is a small uh library uh doesn't a lot of dependencies external dependencies. Just a good set is a they have a good test coverage. It is really extensive and well documented. bear in mind that this library wasn't built with the safety in mind. It's just JSON library to parse the files,
to create a JSON files from C++ context. So, the first thing that uh they did it was like uh they added TSF to it. So, they added that they added the the assertions from from Trustable Upstream to this product. Before applying TSF to it, they They had to understand what are uh what were their goals related to what do I want to show? What is my evidence?
How do What is a good test? Or what is What is the result that I want to add as part of of my or my TA my TA and the TAs and how do should I collect this kind of evidence? And they learned the framework intrinsics and uh also as Rowan said, a historical introduced a euro reference and validators. The PSF allows you to create your own
reference plugging or validator to fit your own needs. For example, the repo. If you want, you can access the QR code to their this specific version of the report. As you can see, you can see the whole trust with graph. You can see the compliance side for certain aspects of it or AU's and so forth, which is defined by the product. And at the end, you can
see some of the trustable work The URL is at the top, so And this is one of the items. I picked at this one because I I was quite interested when I saw the the specific part of the JSON library we're analyzing and how they evaluate the upstream uh product, which is not how they see if they some change that aware or not. So, it's it's it's
telling the whole process from the TA updates to around here and how they evaluate this thing. And as you can see at the at the end of the image, you you see a certain reference to a website and explaining how they do uh how they evaluate some aspects with a one single sentence. if we deep dive how they do this, it's a simple matter. They created one
reference type for these specific needs like a I have I don't have this type for my product. This app doesn't provide something. Can I create my own type? Yes. And they created this specific type called product website, which is just parses a a given URL. as evidence, they are collecting some response from this URL. And you have a certain aspects like a two seconds to uh for
a specific query. So, they just have all of this process and if it's true, this evidence have a good score. And at the end, you can also note that uh there is some explanation why how it's evaluated. In this case, they are they are not that there is no purchase uh from this library upstream. So, it it's fine. It's well known as acceptable. So, the branch is
the same as they are using. And now we got to the part where you can if you want to see the details of how is it the whole report or how the file is, you can go to this URL or this QR code to see the same file that I was presenting the last image. Now, I have a question for you. Do you have a product where
you think is suitable for Trustable to help you to engage and show that uh what I'm doing I believe is enough or can I show that uh my product is work as it is for any kind of component? Think about this. I have a component that uh is not built for safety, but uh I think my component is stable enough or it it's good enough to be
used in own product uh part of a big product like a road safety for a platform or anything. First of all, to help you to guide into this direction, do you have anything in mind? Hm? Wait. Really, what is this? What what kind of product This this one I don't know, to be honest. Yeah. So, oops. Yeah. Yeah, Eclipse rich client platform. Mhm. Um yeah, it's a
Java-based application development platform. And the final application actually would need something like this Mhm. to show it it applies to or or it fulfills whatever needs to CRA things like that. Mhm. and to start with >> [clears throat] >> uh I think it should make sense to have Eclipse rich client platform as such to show it it it it complies to to these yeah, requirements. Do you
have any other? I mean, it bear in mind that it's not only for safety or even for a specific ISOs or IECs. You can also think of this about a product for for a CRA or even for your own organization as well. I have no project, but uh actually a generic answer which may lead to a a question to you, okay? So, any project where we have
regulations such as actually mentioned CRA, safety, AI Act, Data Act, where you need to to make available evidence that might be checked. Would have this in mind, okay? So, my question to you is probably this is the intent of your your project, whether we can also categorize them. So, now this guy is focusing on cybersecurity, but the other guy is focusing on AI Act. So, they add
the additional things to your system so that they produce the evidence and the evidence are dispatched according to the Yeah. Is that the intent of your system? >> Yes. Yeah. So, something we are working on at the moment is publishing reusable argumentation. So, if you are trying to be compliant to something like the CRA, we would have a set of default argumentation that builds on top of
the trustable foundation that we talked about, and then you as someone trying to be compliant to that would consume that as your base, and then it would help you organize your own argumentation underneath that. So, collecting your evidence specific for Um So, so we're interested to check and how this can work smoothly, yeah. It is really a need by the way, so we hope that this can
be a success in that direction, yeah. Awesome. Yeah, thank you, very interesting. So, following up on that question, in the specific use cases of safety, which are the safety standard have you tried to apply this to? I think your abstract suggests A61850, but have you applied something else into railway or automotive? And following up on that, do you see the opportunity to working with some of the
standardization body to provide templates that you can use for each of the safety standards so that adopters can just have a a building block to start from. Um so, we have a mapping of um the the TSF to um IEC 61508. and we've been told that that mapping is then sufficient to cover ISO 26 262. Um so, we have applied TSF towards that and we have a
safety body that has agreed that that is a valid and reasonable approach um and we're looking to apply this to other ones as well. Is that mapping available as part of the open-source project or or not? Do we have to give the email to? >> Sorry. Okay. The We did explain our long-term mapping efforts. So, the long-term mapping effort 61508 first we just did yesterday. We talked
about CRP, which is the sorry EDP, the Eclipse Development Process itself. We did another mapping. Actually, Daniel from ETAS did another mapping, which is now being made public. Um and we have requests and plans for DO-178, IEC 31, whatever the um consumer good devices. So, we we're showing you how to do mappings. And the purpose of doing it is if you then aligned as as the gentleman
we're talking about is Rowan and um Lucas were talking about, if you actually assess to TSF what you would do is you say, I want to assess to TSF. The first question from the assessor is, what mapping? You would then adopt the mapping 61508, 26262, whatever. And now you just apply your TSF. What you get is the actual cert from So, you get an IEC 61508 cert
or an ISO 26 So, you get the cert you want, but the manner in which you're doing it is instead of following proprietary confidential, contract covered spreadsheets from different assessors, you follow an open process and get the actual cert on the outside. Yeah. And all that is open. And our long-term goal is to put it in the project fully open. But right now we haven't got our
own cert, so we're holding it, but it is an early preview and you are free to see it, comment on it, use it, etc. Yeah, I I I understood that this is uh coming from 26262 IEC uh 61508 exactly. But the problem is very generic, okay? So, if we are able to express what you express in a very generic way, and we can standardize this, okay, as
a framework, then you contextualize according to the needs and the verticals and actually the the characteristic. Safety, cybersecurity, interoperability, trustworthiness, etc. So, we need to have this kind of framework. And this tool could be the one that shows that you can implement it, yeah? So, we need to go into that direction. Um I completely agree with the gentleman. So, also I'm from Code Think here and I
so um I was helping facilitate a workshop yesterday that was just precisely on these points. So, Code Think has successfully through client engagements already um been using the TSF and the associated methodologies with 9001, with base by assessments, with ISO pass 8926 that illustrates that ultimately it is the um the the the the the general nature of this methodology can be applied to many many different contexts
to the point that the gentleman has made. So, since we're like that, we can try to come up with some ideas where we can promote this framework and standardize it. Maybe SC7 software engineering, maybe elsewhere, I don't know, but let's discuss it, okay? Mhm. >> [laughter] >> Thanks. [snorts] We have the next presentation starting, so we do have to wrap up. Awesome, okay. >> Thank you.