Open Community Experience (OCX)

Benchmarking trust: The launch of Eclipse PanEval

28:34 · 21 Apr 2026 – 23 Apr 2026 · YouTube

About this talk

This talk focuses on the launch and objectives of the Eclipse Pan Evil project, which aims to address benchmarking in AI by providing an open-source evaluation framework for foundation models. The speaker discusses the challenges in AI benchmarking, including the trust gap due to rapid advancements in AI capabilities compared to the pace of reporting standards. They highlight the implications of the EU AI Act that mandates compliance and assessment of AI models, outlining the project's community-led governance model. The speaker emphasizes the need for transparency and collaboration among stakeholders to successfully implement AI evaluations, aiming to ensure responsible AI usage and aligning with legal requirements.

Full transcript

[music] >> Uh welcome. Uh thank you for joining me this uh afternoon. Um I've recently joined as head of AI of the Eclipse Foundation. So, I've been with them um about 2 and 1/2 months ago. Uh in that time uh one of the things that we uh launched was um together with uh Pan Evl, the Eclipse Pan Evl project, which is all about uh benchmarking. And I

know we are hitting the afternoon and you guys want another coffee, but I'm still grateful that you're here and and listening to me. So, um I will try to to walk you through the the reasoning behind it and and some of the methodology and then you can decide for yourself whether you find some angles to contribute or um to to to to work on this. if you

are close to the AI and uh use cases uh in that space and and the news, you're probably familiar with the annual Stanford um AI index report. So, just um as they are very very detailed with many graphs, I decided to to pick this one out. Um and the reality on the ground is when when it comes to the benchmarking that it's either um extensive, expensive for

you to do, um but it's also clear that the capability is running ahead of the reporting on that side. So, when you wanted to um reporting on um responsible AI cases, um as you can see here, the the capability has scaled by factor uh 3.3, uh while the reporting capabilities have pretty much uh over time um ch- um stayed the same. So, what you have therefore is

a a trust gap, and that is structural based on the the speed of innovation, the speed of at which uh models are getting published, less and less disclosures. I mean, from an open source perspective, of course, when you're familiar with the open source AI definition, you would ideally like to have a scenario where people don't just publish the model, but publish the weights, the data, the rules,

and everything. Um but that's is not actually happening in the market at all. The competition is so fierce that the disclosures and what went in there in terms of secret source is declining somewhat. And that means that you see some saturation, you see some benchmarks that are not relevant anymore, and new benchmarks being taken on. And it becomes more and more difficult to compare performance over time.

And therefore, you kind of have a case for independent testing, particularly in mind when when it comes to the EU AI Act. Sorry, now we jumped. The other thing is also that some of the the trust in the developments in that area, when you look to the right in terms of the 31% of the US public trust that the government can regulate AI effectively, and there's been

some changes. In Europe, it's is still somewhat higher, but you can see a trend to disclose less and less. But at the same time, companies and individuals would like to know more in terms of having more insights in terms of what's going on, because it's very very difficult to to keep up from from that front. So, what I would like to cover today, I've split it up

in four different parts, is to kind of talk briefly about the the regulatory space, then introduce what we're doing um together with our partners on on Pan Evil and then um go through the the setup a bit with the stewardship and to finish off with how to to join this and and and how to contribute. So, let's talk about the regulatory function. I've called it forcing. You

can call it different means, but it's the AI Act that really turns the benchmarking from just something you do because you want to have a view to your models and kind of some drift in the data and the performance to a a legal obligation. Most of you know this. We've seen kind of the first parts uh come to to force in August last year where the obligations

begin, um but where it gets really interesting is then August this year where we have the full enforcement for the most providers and then of course you've seen the the work of some of my colleagues on the CRA um side in terms of the Cyber Resilience Act where I feel there's um some some overlap in in a way between those things. But, what you've also seen is

that the risk or the the use cases are getting categorized in terms of different risk um categories from the prohibited or unacceptable risk to high risk and then a limited risk and minimal risk. So, you could say for the minimal ones you don't have to do something, but there's voluntary codes, best practice, and then necessity um to to still do something. So, even though you don't have

to do something, a second opinion um could still give you some insights that um you find meaningful or something that gives you uh insight in terms of training and and so on. So, what the act asks in terms of the providers is is not only the the technical documentation, but also the model evaluation and that could be different adversarial testing and and systemic risks. But then also

goes over to the the safety of the model, copyright and provenance. In the earlier days of kind of the coming of of responsible AI in in in my old life, we provided a test kit where you could look for biases in your model, you could look for how sustainable it is, how it whether you can kind of trick it, whether it shifts with different data and in

terms of doing some tests on the robustness and so on. Some of this has moved on. But when you look here for the the copyright and the provenance, while this is important, it's important for you also to realize that some of the legal battles that are being fought on the copyright front, they're fully decided yet. Yeah, so it might still be a year and more out until

let's say at least on the on the US side some of these are decided. For the EU, of course, copyright law is is is a different case, but there are still cases being decided even in in right now. So, I've I've done some research. Maybe you can have different views on this, I don't think there is a open framework that really covers all these three things that

wants in terms of aligning with the EU conformity, but then also doing something on the cyber and and safety front. And then of course being in the Eclipse Foundation, something that is vendor neutral and community governed so rather from from one provider, something that different stakeholders can take interest in and contribute. So, from from that Eclipse Eva is trying to to close that gap, to address all

of those, but maybe in the first instance it will not fully address them, but it's trying to play to those three main themes, you could say. And then, of course, why Europe? Why why something now? We we've seen it in the slide earlier, the the trust in terms of the the regulated to manage the AI in the US is around 31%. We've got China with 41 and

then EU with 52. Um Our headquarters is not uh incidentally uh near the European Commission, but it's there because the European institution need a credible neutral counterparty to um rely on when it comes to uh looking into these systems and you could say that open uh strategic autonomy is is really the the bridge to and and open source is how it will be done in in in

that way. So, let's talk more about what is it exactly we want to do. Now, we talked about the the why and why now. So, now let's dive a little bit deeper on uh what is kind of under the hood, um what is being evaluated and and why does it differ. the idea um and and and this was uh started um by the original developers in terms

of BAAI um in in the last uh couple of years. Uh it's an open source framework and a platform to evaluate foundation models in terms of the performance, the safety, robustness, and bias. So, you could say the responsible part is kind of fully built, uh, in, right? And when it comes to the the different components, and we will see this over the, uh, coming slides, you you

you don't just cover different language models, but also, of course, the multi-map model aspects, the the vision, and, um, different speech and and, uh, language models. So, let's kind of do it in the in the first layer, given that we have got three. So, we're looking at about 40 or so, um, capabilities and different task types. So, we want to see the the reasoning capability, the the

knowledge application, coding, safety, and task, and and a bunch of other things. So, that's kind of like the, uh, capabilities in then. But then we in in the middle part, um, it's it's all about in terms of what goes in from from the vision side, um, the perception, speech, audio, spoken languages, um, speech generation, and and all of this. So, this sits in the middle in terms

of trying to understand, you could do, uh, intent recognition, and so on. Trying to understand what it's, uh, um, an analyzing here. And then, as the third component, is the, um, multimodal, uh, models, where you want to be able to compare it between, different ones. So, that's why that cross-model understanding comes in, the uh, and and the image text understanding, video text understanding. So, you're not just

looking at one output and how good this is. You're looking at when you convert it or you you change the input and output modes in terms of how good it maintains to be. And um you could ask why are we doing this together with the the team from the Beijing Academy. It's something to do with the kind of the history of of of this case and Most

of you kind of maybe remember some of this story here in terms of chat GPT being released in November 2022. And you could say this world of AI that I've been part of now for some 13 years or so has been severely disrupted and overhyped since that point in November and that was the case for some of the development teams that I used to lead in We

saw the emergence of GPT-3. We did use it for some of the cases, but for example, my time in Digital Reasoning um once upon a time Digital Reasoning had models that were had more parameters than anything that Google had produced, how can I say the community was not interested in? It was maybe presented at some obscure data scientist conferences in in Nice in Nice in in France

to a audience of other data scientists. And it was maybe a news item for a week and then long forgotten. And this is why I want you to kind of keep this history here in mind, but then also look over towards China and look at this um Wu Dao model. The Wu Dao model in 2021 um was released after GPT-3, yes. But it had 10 times more

parameters. So that was an early release by the same team that has contributed to the Pan Evil originally repositories. It was an early release, hardly neglected you could say by by the West. but in terms of if you were just looking at the most parameters or the most powerful model, it should have registered on on every radar, but it never maybe got the marketing or wasn't widely

used in in the West. And and that team has continued on to to build a number of other things. And it's been been interesting to discuss with them the the different aspects in terms of what this could be there. So some of the experience that they had in terms of the model building, but also the evaluation and leaderboards went into this. You can see the different aspects

from um IEEE, but also the multimodal framework that went into this to kind of really start on a global collaboration with a number of different organizations. So my colleagues have been talking to them for for some time. And it's been great to have that launched in over there in about a a month or so ago, and we're now looking to to build this further out as as

we talk. But before we go into the the call to participate, I I wanted to go a bit deeper in terms of what are the different components of this. You could talk about it as a three-dimensional system in you look at the capacity in terms of the scope of the model capabilities, the actual task that you are giving it in terms of how to evaluate, and then

you look at the matrix of the KPIs. Okay, now that you've decided the scope and the the task, how do you measure this? And and then you are talking about a three-dimensional room where you kind of trying to compare things like for like but on a complex matrix that is not so easy to manipulate or to to work towards. And and maybe this is a theme that

you've seen in the competition in the market where where some of the companies you sometimes have the the feeling that the models were directly developed to to beat existing benchmarks. And while they beat the benchmarks, the the actual performance that you might see for your specific use cases might not have changed or in some cases have worsened or the the language or persona has worsened. So when

we we have a more complex system it might be an interesting one to to to look into. They started with about 40 evaluation tasks at launch and they are getting This is probably hard to read in terms of the diagram at the top but let's just go through the the pipeline that is fully reproducible and and end-to-end. So we we talked about defining the the start and

then really running this in terms of the the batches, measuring it on an ongoing basis, reviewing it with human in the loop so it's not fully automated to the machines yet. Um One component of contributing to this could be either sharing data or effectively reviewing and then publishing it openly. >> [snorts] >> Now let's let's talk in the third part let's talk a bit more about the

the the setup in terms of given that we want to launch it as a community led project. this is um, the the history that I briefly explained earlier where I showed you the the background on on the developers and and the WDA and and so on. Um, so what they've done over the the last couple of years was to Sorry, let's go back um, to have a

comprehensive evaluation structure with a lot of different models uh, and um, a lot of standard work. But, it was decided sometime um, towards the end of last year to um, contribute it and that means that they decided they want to make it a vendor neutral to look for other contributors to um, really try to um, give a different perspective in terms of the EU AI Act and

and that created some additional demand and request from people um, to to contribute. And so, they submitted the um, membership proposal and the project proposal in um, Q1 uh, and now it's it's the case of um, the the trademark being filed and then uh, working together in a Oneiro style dual sync model that I will talk a bit more about this and for this uh, last stage

to happen, it's the um, community really to have um, other committees that are getting uh, elected in this process to have um, contributions from um, other global actors and to have like a uh, regional autonomy with with some uh, synchronization uh, globally from from this one. So, yeah, you can see before this transition, it was really a a single organization technical direction that decided what was going

to be in terms of the commits and the road map and and so on. Um once this was kind of transformed into a open source project on on Eclipse. Um the committers the the will be elected by their community by merit of course and then you earn the the rights to commit and um shape the the road map going forward and the IP in terms of the

due diligence and the providence tracking. And providence in terms of figuring out where model building blocks come from in terms the data and and the stage will probably play even more important role when you look into the future in terms of AI LM evaluation and and that. I spoke about the decoupling. So the way that this is being out is not that um kind of copied over

but in a dual um models. That means the the core is being shared as being independently evaluated and then the releases will be synced. So you could say in in terms of BDAI track that has the original code base but it's now been migrated into something that is relevant for here. So then for for Europe this is the the Pan Evil track the governed version in terms

of the EU compliant data sets and kind of playing to the AI Act and the CRA alignment. Of course then when you think about some of the origins of that model, that also means a different type of language data set, a different coverage, and and and so on, to have a a chance to kind of achieve some some interesting insights as we progress with this. when when

we look at the the decoupling and and why this kind of matters, it's a regional de-risking. So, each side operates it under their own governance. So, then there shouldn't be too many geopolitical shocks or anything that affect each other. Preserves some of the the the learning from each other, but it also has this compliance by design built in in terms of being fully aligned with what's happening

in here in Europe, and coming from a audit firm background in my last role, I think this is what the procurement teams and the auditors will ask more and more for going going forward. you're probably familiar with some of these things, but from from the ecosystem then it it's certainly public sector, the researchers, the regulators, but kind of really leveraging the whole ecosystem that we have here

to to really try to get something some valuable insights from from the global community in terms of AI and and and and such. And then, as the last piece in terms of the the fourth part, um how to join and and what comes next. So, as we've just launched it a short time ago, what are the things that could be relevant for this to to really take

off. What are the things that they're looking for as a contribution when you've dealt with Eclipse and the different things that we already have in the space of AI, you've probably seen some of the presentations from the colleagues from Thea AI over the last few days. You might have heard from Elmos or Edge. And probably you've seen something from Open VSX and and the direction that this

might take given the vibe coding and and the the hype around this. So, where where the Paneva will really fit into the existing ecosystem of applications is to have a neutral open evaluation framework for foundation models that are then leveraged, let's say, in Thea and in Elmos and and others to either drive agents, agentic AI, but something that is aligned with the EU AI Act or at

least gives a second opinion based independent of of public benchmarks and allowing people to assess that for for themselves. High-level roadmap, but one that you can contribute to. Um Like we we we we looked at this in terms of the the project has just been created and chaired. They have just become members. But there will be some the incubation time in in this quarter, if you want.

Some some of the committees are getting confirmed. The the GitHub goes live and and so on. And then um, in the year, maybe some some of the further releases where the the language coverage will be improved, some additional safety mechanisms and and so Probably more for for next year is then the, uh, conformance pack where, in terms of article 95, the compliance to kind of make sure

that we are fully on track on this. kind of the the domain templates and so And while this is not, uh, alone my decision, but, um, as far as the Eclipse Foundation goes, I we have a number of tools in the AI space. I think we will get, uh, a few more, um, but we are, in terms of anything from, uh, data science workflow or value creation,

um, there's a few more white spaces that, uh, we we could cover with open source tools, um, but we're not a million miles away to to launch this, but I I wouldn't put a month to it just yet. I would probably say sometime next year, there should be either a collaboration between existing working groups or a a separate one that just depends on who's active and and

and so on to to see what the the the future holds from from that point of view. And then, um, as the last slide, um, I mean, when it comes to responsibly AI or trust in AI, I feel there's really no better place than to do it in a community. Um, I spent the last 13 day 13 years, not days, um, educating clients about AI, the the

upside, the downsides, the risks. Um, I spent a lot of time, um, doing knowledge sharing or publications and and so on, and I feel I learned a lot from community and meetups and exchanges in my many years in London. Um I'm excited to be part of a open source movement and where we can share data and insights openly and where we probably learn a lot from each

other. And I'm just going to hopefully encourage you to if you haven't done already to look at some of our other offerings in in the AI space um or um in in ways to shape and and form some of the new things that are to um LLM benchmarking with Pan Evil or some of the the other initiatives that we have. So I kept this um short on

on purpose to give you a chance to ask some questions. I know it's only a few here after after lunch and and coffees but by all means you've got my full attention. So um if you like please ask them some questions. I've seen some of you taking pictures so maybe there's also some questions you would like to to to ask um so far. Have we got any

any questions so far? Have you got any reflections on our work there? Okay. Good. Thank you very much for your attention. >> [applause]