SEDIMARK - An open source data space framework for a decentralised data and AI marketplace
About this talk
In this session, Tar Alsale from the University of Surrey presents CityArk, an EU-funded Horizon Europe project aimed at establishing a decentralized data and AI marketplace. The speaker discusses the issues with traditional centralized data marketplaces, such as data sovereignty loss, privacy concerns, and energy inefficiency. He elaborates on the advantages of a federated data ecosystem, emphasizing how their open-source data space framework allows data to remain at its source while facilitating secure exchanges. The presentation also covers the system design, which incorporates edge computing and decentralized trust frameworks, as well as the methodologies for data management and AI model sharing. Overall, the session champions the need for a collaborative environment where researchers can share and utilize data without compromising privacy or control.
Full transcript
[music] Great. So, uh thank you very much everyone for attending uh this this session. So, my name is Tar Alsale. I'm from uh the University of Surrey. U I work uh as a um research software engineer. So, that's that's basically um my my role at the university for quite some time. Uh I've been involved in quite a number of EU projects. Um and uh this this is
basically the latest one that I've been involved in uh uh called CityArk. Uh it's a uh Horizon Europe uh innovation action project uh which targets kind of like the higher TRLs uh and basically uh involves a consortium of of partners from industry and and research and I'll be going to them a bit later. Um so what is basically the kind of solution I want to present um
is essentially an open source uh data space framework for a decentralized uh data and AI uh marketplace. Um what I'll be going through in this session here is not even though this this project has has been completed last year um I will not go into too much about lessons learned but rather uh probably just go into how our methodology and our approach to to to to the
project and how we uh you know came up or how we did de design and developed uh the the uh data space framework solution. So essentially uh what I'll be going through is is just just essentially the the kind of why uh you know why we're doing this uh and you know uh from since we started back in 2022 what was kind of the landscape with regards
to data marketplaces uh and and what basically dataces uh can can provide. um then basically present what what our solution is uh essentially and then I'll be going to kind of like the how and how we the approach we we took to to uh design develop uh the So essentially when when you're talking about data marketplaces um usually or traditionally um you you're dealing with with a
kind of centralized uh uh uh entity which basically where you have data providers pushing their data to this to the central entity. Um obviously issues that come with that uh especially when um it's it's a kind of third party party entity uh is that obviously what you know storing all that data into one one obviously centralized node can be a possible uh uh single point of failure
um and also primary target for for cyber attacks. Um also there's the issue of loss of sovereignty. So basically once your data is uploaded to this traditional marketplace uh usually the provider loses control of how uh the data is used or how it's consumed. Um there's also issues with with uh uh p privacy as well. So um you know sharing sensitive data often requires uh more data
to to to the consumer which then obviously risks leaks and violating uh policies such as GDPR uh and possibly also you know uh exploit uh exposing trade trade secret uh protections. Um so obviously you know especially when you're dealing with with with with you know uh well-known vendors obviously you know that that kind of debate come comes up quite a lot in terms of you know you
know should we really be pushing our data to these to these third party entities um also with regards to uh um terms of actual energy energy use. So the other thing is basically obviously these centralized data centers are quite well known for consuming a lot of energy um especially when it comes to transferring mass amounts of data. Obviously that data needs to be transferred and stored and
all that obviously takes quite quite a lot of energy uh um and and and hence obviously you know from what we're seeing today with uh you know with data centers obviously it's becoming a or has become a a uh an issue of concern. So um and then also the last lastly is basically with regards to vendor lockin. Obviously um relying on proprietary APIs and closed standards can
make it difficult for researchers andmemes uh to switch platforms or even collaborate across different ecosystems. Um so yeah essentially obviously when when you you know you're a research group you're you're kind of working on a particular solution you kind of you're already you know you already kind of get used to a certain uh uh uh setup uh and then the issue comes where you if you have
you know you you you need to migrate for whatever reason you know having this kind of locking can can provide uh uh problems uh for um evolving your your solutions. So what's been happening basically in the past few years obviously there's been this kind of shift to to what are called federated data ecosystems uh where basically uh it's essentially decentralized by design uh where you basically got
data that remains uh at the source uh uh with the owner and is only exchanged when a specific you know agreement is reached. Um so and also you got uh sovereignty uh via sticky uh policies uh where uh you are able to enforce uh uh or uh um uh usage rules and how how the data should be used by by the consumer. Um there's also uh the
aspect of obviously kind of a shared trust framework where common set of rules, identities and security protocols uh that allow participants to interact uh without a central authority. Uh and then also quite important which is basically uh interpability where essentially um you want to be kind of relying on open standards uh which are easily adoptable. like for example ids or orex which have been kind of at
the forefront and defining kind of like the data space uh specification and protocols uh so that different platforms can talk uh to to each other seamlessly. So uh what's the solution? Um Simark. So what Sedark uh is or what it's set out to to do basically is to kind of provide uh uh uh an open source solution uh which is communitydriven uh uh which is fully transparent
uh and based on the uh EPL EUPPL license uh as a framework which is designed for research adoption and and regional replication. Um it adopts a first uh edge first uh approach. So essentially rather than the computation being at you know the cloud it's actually basically uh the comput computation stays closer to to where the data is being generated or being uh uh uh uh um sourced.
Um with regards to uh uh privacy preservation, obviously you know keeping that data at at the edge uh allows uh u you know full control of your data uh and basically we enable this enable this by by using techniques techniques such as uh data minimization uh and employing uh secure uh database connectors uh which ensure that that the these raw sensitive data never leaves its origin origin.
original domain unnecessarily. Uh we also uh set out to to employ a DT to to uh enforce this kind of uh uh trust framework um to to basically uh provide uh transparent immutable uh logging for transactions uh and smart contracts uh without uh a central broker. So the DT obviously itself is is uh uh decentralized as well. Um and then when it comes to obviously semantic intelligence
in terms of the uh interpret interpretation of the data and inter interoperability of the data uh also employ uh AIdriven metadata and common ontologies to make the heterogeneous data uh fair uh findable accessible interoperable and reusable. So um again so the really the aim here you know for the project was to kind of empower the research ecosystem. So essentially uh develop a modular toolbox where uh you
have this kind of readyto use open source components whether as a whole or as kind of like your um as a separate kind of tool for a particular if you you know if required for a particular purpose or to be adopted for for another uh uh project uh which is all available on on GitHub. um obviously aiming to to lower barriers in terms of reducing technical debt
uh for researchers andmemes wanting to pro participate in the data economy uh without building uh infrastructure from uh from scratch. Um also wanted to kind of foster this kind of uh sense of collaboration uh like you know cross- sector manner where basically uh different uh use cases or application um or data sources can can you know come together you know in this kind of marketplace environment to
provide a shared neutral ground for uh data exchange. And then finally also to to basically call to action to to to promote obviously Senark has initiative to operate uh to to operate uh uh um your your own uh regional marketplace uh to contribute to evolving uh codebase um the open source code base. So just to go a bit uh deeper about uh the kind of like the
um design of of of of uh Smark in terms of the system design. Um essentially you you have uh um um here what we're trying to do obviously is to transition from from a centralized kind of approach to an edgebased part uh uh approach uh where essentially you have nodes in a in a decentralized manner where each node has what is called a toolbox uh and with
this toolbox it basically allows uh allows you to uh act as a as a for example as a data provider where you can uh process your raw data uh to generate uh a kind of what we call data assets. Um also play the role of of consumers that allows you to uh discover uh uh find kind of data assets uh and then purchase them. Um and then
also um most importantly uh uh the role of marketplace operators where where essentially uh there needs to be one node that that at least one node to to to host the what was what we call the baseline infrastructure uh that will enable the the the marketplace uh uh operation. Um also uh alignment of with standards. So obviously we were we're we're essentially following the blueprints from IDSA
and GX to ensure uh European cross- sectoral uh interoperability. in terms of sovereign connectivity, obviously we we have every node utilizes uh uh a an Eclipse database uh connector for um aspects such as offering management uh or what we call asset offering management uh and also peer-to-peer uh asset exchange. Uh so essentially the core assets that we're dealing with in the marketplace. So obviously what what what
are the products here? Uh essentially you a data data set uh a data stream uh or AI models you know kind of binary AI models made available. So uh in terms of the approach um when it comes to curating uh an opal tool to toolbox for um research and adoption obviously you know [clears throat] at the beginning as everyone would do is basically uh uh uh start
a kind of state of state-of-the-art analysis of open source solutions. Um obviously you know we've identified the key components for for uh a data space or essentially also a a a marketplace for uh in the sense of how we want to design sedimentar uh is that we obviously first we need we need a DT registry uh which provides a a kind of like fearless uh means of
of transactions and priv privacy preserving uh immutable streams um orchestration so uh being able to uh orchestrate pipelines for data uh data processing to generate your data assets and also uh your AI model we call AI model assets um as well um providing a uh a framework for federated learning. So obviously u because of the fact that we we we are involving this kind of decentralized approach
to uh uh creating assets as opposed to centralized approach where uh you know you have um most of your comput comp your comput computational resources uh hosted. Obviously in this case you would rely on on a kind of federated approach to uh to uh especially in in uh um tasks just such as uh uh um training AI models uh uh across uh different data sets. Um there's
also the aspects of of of knowledge discovery. So obviously we want to host uh kind of like what is called a catalog. So it which will enable basically uh consumers to discover uh uh assets but also what we want to do you know not just not just as a simple means of discovery but but but to be also to to link that information to other uh sources
of information that that can provide more context about about the asset. Um and then lastly there's also uh the aspect of context brokering which is essentially uh what we'll use to to basically host uh the data sets and data streams uh that'll be made to be available for for exchange uh between consumer uh and producer between producer and consumer. So um that's in terms of like kind
of like the shopping list. So uh in terms of obviously uh the expertise uh obviously we're as I said we're we're consortium of of uh um uh industry and and uh um research partners. So each each of one basically having a you know certain uh uh specialtity uh although we do kind of overlap uh but but just to kind of kind of highlight you know you know
the kind of partners. So obviously leading the project we had Evan Atos um basically who who are you know quite experienced in data marketplaces and federated uh learning uh frameworks um Seammens as well who kind of focused on uh orchestration pipelines for for data and AI. uh Wings ICT uh who uh is a um theme that focuses on um machine learning prediction uh and also uh has
been kind of uh the backbone for the kind of software engineering uh uh coordination for for for the open source obviously activity. Um EGM uh the French theme as well um focused on data brokerage and and federation. So they they uh uh basically a member of the fireware foundation uh and they also provide one of these kind of context brokers which are essentially kind of foundational building
block building block for um for data exchange uh inter interoperable data exchange. Um we've got also in terms of uh research specializations so University of Cantabria who who are experienced with data spaces and information modeling uh University of Surrey focused on edge based AI optimization and knowledge graphs. Um the University College Dublin um who specializ in also federated learning but also recommener systems. Uh we also had
Enria uh Paris um who specialize in AutoML and interability uh links foundation uh who focused on uh security crypto cryptography and and DT infrastructure. Uh in terms of um obviously use cases obviously which we require for solution validation. Uh we've had sand municipality uh where we had the use case focus focusing on uh future cycling route planning uh for virals focused on uh efficient municipal uh or
efficient mobility management for municipal services such as uh slow snow clearing. Uh Mittellaneos uh who basically are uh um focus on domestic uh energy valorization. Um so basically uh uh studying uh data that that comes from um from from domestic homes in terms of energy consumption and customer turn. Um and EGM uh as well as a technology uh specialist also providing a use case on water quality
monitoring in the south of France. So um so I kind of showed earlier a diagram um if I just go back This basically which is essentially kind of like the system view where you essentially have a toolbox on the provider side and a toolbox on the consumer side and then also you have your baseline infrastructure which essentially uh enables the the marketplace environment. So our kind of
first approach here uh in terms of develop development um we we took a kind of like a uh proof of concept uh approach uh where essentially we focus on certain themes uh for the system. So for example, we focus on aspects of onboarding which basically involves uh the uh process of of um uh onboarding uh a participant into into the marketplace. um um essentially you know providing
them a a uh um uh you know identity and and and also the the means of also uh uh starting up their um uh their toolbox. Um in terms of other scenarios that uh that we that we kind of split up we focused on aspects of data quality improvement. So uh uh that involves like uh development of um kind of orchestration of of pipelines for data processing.
Um we also [clears throat] had um proof of concepts for for what what is called the offering life cycle. So these are essentially uh the involve the kind of descriptions that um basically describe the assets that you that that uh uh provider has. basically how that starts from creation to to to registration then to population into the into the catalog. Um and then also basically uh another
proof of concept that that obviously we've kind of uh contained is is the asset uh exchange process. So once the the um uh purchases the the asset um and then basically you know um what are the processes involved in in in the exchange of of the asset itself and then we moved into a second iteration uh which is what we call the MVP approach. Uh this involves
um uh two aspects. So one aspect is is the uh minimal viable uh intelligence which focuses on data and AI model orchestration. Um so essentially here basically we we had these all these kind of like uh seven uh proof of concepts which we then um evolved into just just two two main sub subp parts of of of what we call the MVP. Um so the first one
being the MV MVI which focus on the data and AI model orchestration um which is on the green side here and then uh the MVM uh which focuses on the uh minimal viable marketplace. So how do we enable uh the kind of commercialization of of the assets in the marketplace? Um and then finally obviously uh once obviously that stage is done uh releasing the obviously the the
uh the the uh the platform onto uh the Sedar GitHub uh repository. Um so I guess one of the aspects here you know especially with this event is is uh you know open source adoption. Um uh you know obviously we don't want to kind of start from scratch and we obviously you know wanted to kind of see what what um tools we could we could use as
a techn technological foundation to to build our uh uh marketplace system. Um so essentially again so the approach is basically to to involve a kind of like state-of-the-art uh baseline where involve extens extensive analysis of existing open source uh options for data space system entities terms of selection criteria uh we look at maturity and community support uh uh architecture fit so in the kind of sense that
we want it in a decentralized uh kind of uh framework work um and also license compliance as well obviously. So um terms of you know preferably obviously wanting them to be EU aligned uh or or at least they they don't uh contradict uh EU aligned licenses. So just to give obviously obviously I mean we we we have made use of quite a number of tools but just
to kind of focus on the main main ones. Uh so for example for kind of orchestration uh we selected a tool called ma mage AI um so even though there are kind of well-known orchestration tools like Apache airflow uh we kind of saw it as as a better fit uh in terms of modern kind of data and AI integration um and also um the fact that it
it also provides APIs where we can actually plug in a UI which we we developed uh to to kind of uh um ease the process of uh uh creating pipelines uh for uh data processing and AI uh model generation. We've also uh in terms of the uh DT uh we selected the uh IOTA uh framework uh which is a well-known uh European framework for for for blockchain
uh which essentially involves uh feless um and employs what is called the Stardust protocol for encrypted granular uh access control. Uh as for the data space connector we used obviously um uh database uh connector um obviously the fact that you know it's an industry standard for sovereignty uh high sensibility for citymarks uh uh note concept so really fit the bill there. Um in terms of uh AI
uh collab collaborative AI or federated AI we we chose the the flower uh which is flower um federation framework which is a well-known framework for federated Um and for metadata catalog uh um uh management and uh uh hosting we we we used the Apache Jenna uh Fuseki uh uh tool uh which provides uh essentially robust RDF uh uh spark and sparkle support uh required for knowledge graph
Okay. in terms of obviously balancing decentralization with privacy um the IoT ledger uh and essentially the the tangle which is the the DT network allow Smark to create uh a mutable encrypted uh data channels uh on a permissionless ledger ensuring only authorized consumers can decrypt specific transactions uh info uh such as personal info which is obviously something that's um obviously one of the things that we focused
on is is not to expose any personal info even with providers and consumers um um and is only done with their uh explicit obviously consent on a on a one-to-one basis. Um in terms of uh the commuting to data shift obviously using flowered learning uh the framework avoids uh privacy compromise. So basically any data that's hosted on uh uh any participant stays there and basically uh uh
the federated learning will just uh apply um machine learning and training on onto the data uh without uh compromising uh any aspects of privacy which could be uh in place. uh and also in terms of resource efficiency using the maji for orchestrating uh the use of of obviously local participant compute resources uh fundamentally uh provides a more sustainable uh approach than than consuming cloud consumption of traditional
model of marketplaces. Um so again here just to obviously in terms of our expertise obviously we we kind of did a preview selection um each partner focused on a certain aspect. So for example, sorry we focused on uh uh uh interoperability and AJI uh links foundation experts in uh in DT and cryptograph cryptography. Uh and then also when it comes to validation uh to the use cases
obviously we've kind of validated on on the um uh kind of real deployment of the solution with regards to u data sources from uh basically uh water quality uh sources um and also refining the federated learning stack uh um for recommendation systems. So when it comes to integration release um essentially yeah so going back to the obviously the the proof of concept issue um uh approach uh
we had seven uh proof of concept pillars uh on boarding uh data quality offering life cycle uh a exchange uh AI uh uh model management uh um graphical graphical user interfaces uh and uh open open data enablement. So other than just uh having explicit providers joining uh sedimeark, we also want to make open data uh uh that's already out there available also in into Um so essentially
when it comes to consolidating the scenarios into the MVM MVI so basically transitioning from an independent proof of concepts to three parallel integrated streams essentially the MVM which you mentioned earlier uh which integrates onboarding offering registration and secure peer-to-peer exchange in a co cohesive life cycle. Um also uh with the MVI focus on AIdriven capabilities including federated learning, data curation uh and also emphasizing uh the aspect
of privacy and uh of of data uh within the marketplace. Uh and then essentially we kind of then ended up with bringing them together to uh um create the the the sedimear platform. uh and then also uh providing obviously validation tools um as part of uh the solution validation with with with the use cases. So in terms of the ecosystem essentially just to kind of simplify the
kind of process with regards to uh uh how the marketplace work or how the cityark uh uh ecosystem works essentially you kind of what you'll do basically is that you you'll target your original data asset uh you'll define and select a pipeline and then what you'll get is basically at the end of that pipeline is is the actual asset whether that's a data asset AI or AI
asset uh and you also get a a what is called description which provides just the kind of uh uh basic metadata about the the asset when it comes to then obviously that that data asset is then stored obviously within the participants domain. Um so when it comes to actually uh pushing it to the marketplace or advertising to the marketplace um essentially what a a participant will do
is uh select the assets that they want to package within an offering uh and then choose the method for generating the offering descriptions. Obviously there's there's the way of doing it manually. So you'd have to create your kind of RDFbased descriptions uh which is not the best. uh there's the UI option also uh uh where you kind of fill in a form but then there's also uh
an LLM based option that we also uh developed at at University of Surrey uh which basically uh the participant provides a textual description of of what offering they want uh and what assets they want to target and basically it what it will do is basically it will fetch the asset descriptions uh and also So um being kind of schema aware in terms of the the information model
that that we use uh in in Sedark it will create the uh kind of the the basically the RDF equivalent uh uh description which is then uh essentially um submitted to what is called an offering manager. Uh and this is basically the offering manager also contains the the uh the the the um the EDC connector. Uh but what it does basically is that it validates first of
all it validates the description is is is is um in compliance with the uh information model or or the information that's required uh uh at that stage. Uh and then is then populated into a what is called a self-listing. Uh and then once that's done basically the offering is then registered uh at the uh DT. So registered means uh that the description doesn't go to the DT
but what goes to the DT is basically uh basically just a a claim that this I have an offering that I have available and what is also passed on is is a is a hash of that um description and the reason why we're using the hash is basically for for for for trust reasons uh and and verification. So in the sense where basically um if uh if
for example an asset uh description has changed uh in some capacity uh that would need to be sent to the registry uh which is then you know registered as as a new new asset. um because obviously if it doesn't do that um what it will do is that when um for example the the catalog so what once the registration is done uh the catalog is notified in
in the sense that basically there are new offerings available it'll it'll fetch the um the offerings from the um from the from the participant through the um uh connector um and then it will obviously will compare the the hashes. So if if there's if there's the if the if the hashes are inconsistent then basically the offering description has been tampered with. So it won't accept it. Um
so essentially that that's basically the the the process in terms of uh going from um inventory to to advertisement onto the marketplace. Um just go through quickly to here. So on the MVI side again we kind of employed the the mage AI um orchestrator basically involves uh processes that deal with data curation uh data storage data formatting model training uh and also um obviously we employ uh
uh MLOps tools uh to support the uh AI processing such as MLflow and Minio um and then also the federated frameworks are also part of the uh the intelligence which kind of mentioned earlier which which involve the kind of federated learning um processes. So again so in terms of uh reducing the barrier to entry for providers uh and consu providers essentially we provide this kind of low
code interface where uh you have this UI where you can uh uh create your own pipelines visually. Um and then obviously when it comes to essentially enabling this kind of framework of trust within the marketplace. uh we employ a number of u uh of core components here. So we have something called the DT booth which is essentially uh the component that that uh talks to the DT
uh um between obviously the part the the participant and and and the DT for especially for cases for on boarding and registration. Um also in terms of identity and revocation. So um the DT obviously makes use of the uh uh smart contracts or employs the the smart contract L2 uh type um to manage credentials and enable revocation capabilities. Um also uh employ this kind of tokenization of
assets. So essentially you have um u the uh regist the offering registrations represented in an NFT form uh and also ERC20 tokens to represent uh signed agreements uh or ownership of of the contract for for the offering. Um and then in terms of obviously negotiations. So once a consumer finds the asset on the onto the marketplace. Um all this is done through the sedime connector which is
obviously based on on the uh Eclipse uh EDC uh which essentially automates the negotiation process and also records uh agreements immutably on on the DT ledger uh ledger. So uh when it comes to establishing trust in the um decentralized u what we do is that basically provide the simplified user experience. So essentially you have a you'll have a marketplace front end which provides a uh userfriendly uh
inter interface for uh masking complexity uh for underlying the the decentralized system. Uh we employ obviously uh self-s serving identity uh so that new basically new users uh new users uh follow automated steps uh to become uh foreign participants by generating necessary cryptographic objects and credentials. Uh also employ uh local cryptographic management. So again mentioning the DLD booth mentioned earlier and also secure storage of private keys
and and ver verifi ver verifiable credentials uh uh which are also stored um within the the the the participants domain um and also within within the within the DT um and also uh providing automated cred credentiing uh where the system facilitates direct interaction with the issuer to obtain the valid uh citymark verifi verifiable credentials. Uh and then last but not not least the um the terms of
trust and accountability. Uh we have basically provide uh the verifiable data registry which handles public storage and management for um decentralized ID um ident identity documents and uh identity sorry smart contract uh which which resides on the smart contracts platform nearly there. Yeah. Okay. So uh yeah. Okay great. So um just just to kind of really go quickly here. So in terms of commercializing assets with trust
and transparency um essentially you have your offering description here uh that you create I mentioned earlier and that goes through a process of registration and population onto the catalog. Uh obviously this is done through um validation checks and so on to make sure that it it's compliant. Um and then basically uh the in terms of the catalog uh manifestation obviously it can be done in a centralized
manner but obviously because we're promoting decentralization uh we also have the the concept of uh catalog coordination where you can actually uh um decentralize the catalog onto the nodes that are part of the the marketplace uh uh at a particular time. So if you had 10 participants, you'd have your basically catalog distributed among the 10 participants and so on. Um yeah, I think that's [snorts] pretty much
it there. Uh in terms of managing assets and offerings, obviously we provide the user interface where essentially uh you'll have access to your your kind of offerings uh and also any kind of smart contracts that are uh that have been uh generated from uh a consumer transaction. Uh and also when it comes to browsing the catalog, we also provide a kind of a uh user interface where
you can basically search u which is done in the back end through to um sparkle which is essentially the language that we use to interact with RDF based uh um data stores. Um and then basically you have your kind of uh once you kind of find an offering for example give you kind of details about uh when it was published uh and then who who the provider
is terms of stakeholders in the marketplace essentially who who they are basically you got your marketplace operator um which is responsible managing the marketplace infrastructure marketbase developers which uh are are are basically um uh they extend the customize the platform to meet domain specific needs. Providers essentially organizations which registered and share the data via services. Uh and consumers which are essentially public and private organizations that access
data sets AI models and analy analytics services. Uh in terms of solution delivery basically we have uh it's hosted on on GitHub. Um again we we it's all been dockerized obviously container containerized and hosted on on GitHub and uh basically yeah we're just obviously invi inviting the open community to uh check it out and and and um you know see whether they want to host a marketplace
or at least um explore the the tools that we have uh made available uh for particular aspects uh dealing with uh data data and AI uh management. Uh we also have our information model which is quite uh docu well documented that we have hosted on uh on on GitHub as well. Uh just to acknowledge obviously the project is funded by by uh the EU um under Horizon
Europe framework and also uh through the UK research and innovation uh council. And just uh bas basically big up the the whole kind of uh development team. So obviously uh these are all contributors to to the uh um to the city mark And that's that. Thank you very much.