Beyond the cloud monolith: Building transparent and distributed AI with open source
About this talk
In this talk, Rafael Cime discusses the relationship between open source and centralized AI systems through the metaphor of a llama seeking freedom from a monolithic cloud. He provides insight into the current landscape of AI, emphasizing the dominance of a few major players that lead to concerns like economic oligopoly and ecological impact. Cime argues for the potential of open source to foster a more sustainable AI ecosystem by promoting control, reproducibility, and distribution. The presentation is structured in three parts: the meaning of openness in AI, an overview of the AI tech stack for local inference, and the implications of distributed and decentralized AI. He concludes by addressing future directions in AI technology, particularly the importance of localized and autonomous AI solutions while highlighting the possibilities of decentralized computing and data ownership through innovative projects and protocols.
Full transcript
[music] All right, let's start. Can you all hear me? Maybe. Yes. Right. Okay, let's start. All right. Welcome. Thank you for being here to today. Uh my talk is about the story about a monolith and a llama um an animal and it could be a dance between uh open source and uh closed and centralized AI system right actually I could have called the call the talk uh
the planet of the llamas but maybe you won't have come so let's stick with this behind the mallet so yes it's about a llama who is trying to break free about from the cloud monolith and get rid of the the strengths that he has. Uh and you'll see he has a superpower that will prove very useful during the the talk. But uh let's uh introduce myself. I'm
Rafael Cime. My name is pronounced. I'm from France in Paris. I head of tech advocacy at WLine. Who ever heard about Wline? Nobody. We're in the payment business. That's enough to say, right? I'm a senior architect. that is uh something important also when we analyze a complex system that's what I do I try to break down things and organize them to understand the complexity uh I'm also
a certified yoga teacher and I'm a geek hence my nickname yogic Rafael yogic right so namaste uh and I'm a geek because I'm in open source for years and years I design games and more recently I started to publish uh novels sci-fi novel uh with the the assistance of AI but we'll speak about it at the end of the talk you'll see why right opinions are my
own so I don't they don't represent the the positioning or the the opinion of my employer right and I say that because I think personally that today regarding AI we are in a intelligence oligopoly right where you have a few uh companies a few countries who centralize the power and the the the yes the power of AI in their hands and it has a lot of impact
on the the planet and on the many many different aspect that I won't detail there you heard about it you thought about it you you know about it uh but it's economic issues that are now transforming into geopolitical issues with you know predation reflexes that we observe on how to to take the resources from other countries to build the AI and create nuclear plantants and stuff like
that. There's there are obviously technical issues uh of dependencies for instance we speak about a lot about here in Europe in about resiliency sovereignty that kind of stuff right so on who do we depends and of course all of that has a huge impact on ecology and on the planet uh speaking about resources like energy but also water and all that kind of stuff right so centralized
AI seems to me not to be the only way because it doesn't seem sustainable uh on the long run, right? So, how can open source help to build an AI that we can maybe distribute a little more and uh have more control upon, right? So, let's see if uh uh can open source save lama because I I told you he has a superpower. uh and to dive
into that uh I'll my presentation will be break down in broke down in three different parts. The first one is just a refresher about what open means for a model for AI. That that was a talk I gave two years ago in this very same conference. Then we'll have a look about uh the AI tech stack. When we speak about local AI inference test tech stack, what
are the different component that we need to have a a model running on on this PC or in a in a container on my own server, right? And then uh we'll move into the world of distributed and even decentralized AI where the craziness happened, right? So bear with me and let's start with uh what openness uh is in AI. So what's going on? Yes. Ah that's the
superpower of super lama. He has a time travel power, right? So go back the last last edition of OCX. I won't do the same talk because there's no point of that. But maybe uh if you didn't uh see it, you can scan this uh QR code and then you'll have access to it. Basically uh I'll give you a very very 101 on that. Uh so I I
I made this talk and I created this talk like before the open source AI definition from the OSI but basically that's the same idea there. So that's stack number one, right? To understand and to map the the landscape of AI and open AI, right? So I spoke with uh AI scientists in my company and I told them what should I look at when I uh I want
to evaluate how open a model is, right? So they they told me okay look at how it was built and do you have access to all the component that you need to reproduce this? If you can reproduce it from zero, it means that it's completely open. Scientifically, you can reproduce it. Right? [snorts] So here is a very simplified schema on how models are built. Very very simplified,
right? But here you [snorts] have like the big model foundational models that are uh trained in a selfs supervised pre-training phase and they're trained on large data set. Basically they're very hungry. They try to eat everything that we produce and now they start eating what they produce themselves. So we'll see the matt car problem will will become a reality then sorry then you have um here uh
fine-tune models that you can adapt you can fine-tune you can derive all right so it looks like open source you reuse the open source product and then you derive it you adapt it to your own need uh and then for that you use like another data set with human instructions and responses for instance to fine-tune it and then you can also align it or reinforce it whether
it's based with human feedback or AI feedback now uh and you also have another data set here and I told them okay so based on that when I want to assess the openness of a model what should I I look at they told me first thing the model and the model is not the the code to design the neuron network everybody every AI scientist know how to
create a transformer stuff and that's not the the the point the what we want to to to know if if we if we have access to the weights. The weight are the result of the training. So basically the configuration of the neural network after the training, right? So knowing the architecture, having the the weights, you can reproduce the model, right? So that's the first thing that you
you look at, right? And basically when you have that and only that, that's what we call open weight, right? Uh then of course do you have access to the data, right? Whether it's the big data set or the the smaller ones, right? And then of course most of the time when it's reinforced there is an intermediate model that is called the reward model here. A lot of
project don't really share or speak about it. So if it's available it's a good sign that it's open and that people really want to share things share the internal recipe for you to reproduce the experiment. Right? And of course yes some code but again not the code to implement the the the model itself the neural network but the code to do all the process all the heavy
lifting to prepare the data prepare the training uh etc etc. So based on that I made a an analysis of the of the landscape and uh each time I give this particular talk I update it and stuff but again this is not the the main subject of this talk. So where do we stand here in in 2026? Basically we can say that we are in the era
of open weight right so open weight is the new black uh with two different versions the first one is restricted open weight basically it means that you have the weight you can reproduce it you can even adapt so you can modify the the model fine-tune it right but there are some restriction and it's uh framed by the license so the license is everything here so models of
this type are llama schwen sen version of trend that from Alibaba uh bloom for in instance which is a French well European uh old one model uh but they have some uh restriction like responsible AI or basically it is don't do bad things with it but what is bad is defined in the license so we have to read the license before because you might discover that the
the definition of of bad especially today uh we don't know what bad is anymore right uh the new trends sorry Sorry. Again, the new trend is permissive open weight. So remove this restriction. So you can do anything you want with it, but you won't have access to the data. The data are still private, right? So it's like remember permissive copy left open source. It's a little bit
like like the same, right? So here you have models like Gemma, GPTOSS from OpenAI, a certain version of Shen, etc. the the actors that you find here are are basically the challengers that started this positioning from the early beginning like Mistra from France, right? Uh or Deepick. Uh and the people I call the shovel seller, you know, in the gold rush, the people who make money is
people who sell you the the shovel, right? So why they do that? those those big tech giant because they want you to uh to specialize Jamaa or even GPT uh OSS right because at the end of the day you'll need a resources to to make it run right so they will give you all the information to to to to specify to derive it and to to run
it but on their infrastructure possibly right so that's a positioning but of course at the two extreme you still have the flagship model from the same ones for instance that are closed the more capable models Right. So it's black boxes behind APIs. So here you find Gemini, GPT, 3 etc. Five. Now uh Muse Muse is interesting. It's a brand new family of model from Mita. So Mita
who started the open weight model movement with Yama now stopped with the new family. It's completely closed. No more open way, no more closed proprietary, right? And on the other side you still have some project who are trying to be the most open possible. Most of the time it's in research uh projects but not only where they really try to propose the full reproductibility of their model
right so it's possible but it's still rare so why open why open source I say open because you understood open weight open source there's a lot of open source washing uh greenwashing today but anyway it's interesting we're in an open source conference I don't need to go through resilience because you know you have a no not a single point of of a failure of the strategic dependency
towards a unique actor innovation because you know building in collaboration and within a community is a is also innovative let's say and trust who controls the data and who controls the rules about how we build the the system right openness also uh allows optimization and dem democratization and That's the way to to local AI actually. Uh you can specialize the model. Uh so you fine-tune a larger
model and you specialize for a specific task. So it will become very very good at doing something very specific. You can reduce the size of the model also. Uh with some optimization techniques that we will uh explain just in a few minutes. uh and then it means like you can run the models maybe on this laptop on a in a container on your local VPS on on
your local enterprise or organization infrastructure right so democratization just the same as uh open source right and that's the way to the second part local inference so I started to to analyze what do we need to make a model run and to use them to exploit it uh on a PC or in in a local uh uh architecture, local infrastructure or private infrastructure, right? And because I'm
an architect, I always do that kind of of things. Uh I uh started to uh draw some boxes and layers because there's a lot of names that pops out and you don't know and and I speak with people who are in the hardware part, they don't understand what people are doing in the in the top level on on this part. So the purpose of this second part
is just to give you some hints and some global map so you can easily navigate then when you see a new technology say what kind what kind of solution it is and where does it uh does it uh stands and how does it articulate with the others right so if we okay with that let's move through through quickly through those different layers with some examples and just
remember what are what are they providing to the global stack right [snorts] so the first one hardware So hardware of course we have open hardware but we're not there yet right we have two different kind of hardware equipment uh when speaking about AI you have the generic one with CPUs uh and and RAM CPU is for general purpose computing and the RAM is like for multitasking right
but uh you also have specialized hardware that are more adapted for the matrix calculation that is uh required for uh inference of artificial intelligence model like GPU or TPU it's more parallel computing right and uh video RAM that is more oriented for performance and bandwidth but on top of this hardware level then you have this accelerator acceleration layer so it's a piece of software that will uh
translate the instruction of the model to uh pure instruction to the hardware so it's a it's an intermediate layer right um And um it's optimized uh they they try to optimize the most of the the specific hardware like for instance GPUs and VRAM. Um you have uh different kind of libraries in this category. You have NVA CUDA that is quite known but which is not open source
by the way. There is an open source project on top of the of that that is called uh what is that? Tensor RT that is a simplification of CUDA but uh the CUDA itself is not open source. And then AMD and Intel with open veno, AMD with rockm are doing the same kind of thing but in an open source way because they're trying to d differentiate because
Nvidia is the leader in the in this uh domain. Right? So you have this hardware you have this layer on top of the hardware to optimize things. Then on top of that we will execute the the models. So you remember the this open models we have the the architecture you have the neural network we have the weights. So how do they run? So that we call inference,
right? Uh you have two kinds of frameworks. You have the generic frameworks that are uh built to do the inference but also the training and many other AI and machine learning uh processes right. So we speak about like for instance PyTorch of or PyTorch so which who is was a MIA project and it was donated to the foundation is under a BSD license. uh or you have
uh TensorFlow from Google who is under the Apache license, right? They're very very generic but then they're not really optimized for inference and most of them they're in Python like PyTorch, right? Uh and then you have a second category of uh of uh execution uh layer that is uh optimized for inference only, right? And most of them they are in C or C++. So it's a low
level. they can have some Python bindings and and parts of it but the core is in C and C++ and there uh you find uh a lot of uh solutions and specifically in the open source world you have a project like llama C++ maybe you heard about it it's that kind of stuff right so it's specialized uh for inference and uh it's very very uh efficient and
performant because it's doing just that right the open source way just do one thing but do it good and then with the architecture view we will articulate and compose the solution with different bricks, right? But you also have a VLM that I will speak about just a little bit later. you have the inference the the frameworks to do the inference then you have the models the models
we spoke about the open models in part one uh they can be optimized for inference. So you have different techniques like pruning. Pruning is the the the a way of removing all the unnecessary weight for what you want to do. So you simplify the basically you simplify the neural network. So it's less uh big it's uh smaller. Yes. And then it it requires less resources to to
do the inference. So pruning is one technique. Quantization is the reducing the accuracy of the weight. So it's a trade-off between quality and performance. You have also distillation. Distillation is when you take a uh you train a small model that already exists uh to imitate a larger one that is more capable. And then it's a mimic is mimicking the large model. So it's called transfer learning, right?
Or you can also have the low uh rank adaptation which is an adaptation of a existing model but you just simplify and change the upper layers of the the neural network. So you don't have to retrain everything. You just retain just a part of the neural network. And with all these techniques, it's very easy then to fine-tune to reduce to adapt uh existing model if they are
open uh to run on the local machine or in the local container right and you have a lot a lot of uh uh solutions and models that exist in this domain. Some of them are well known like Gemma which is a openweight version of Gemini let's say uh by Google. Uh you have Schwen from the Alibaba. uh you have Mistral D that mentioned and you have also
small language model from like 54 from Microsoft right [snorts] um when you have the model so we have this stack right the hardware acceleration layer uh the framework then the model then we we want to use it so how to serve the model and that's another kind of component that's uh serving the models what they do is beyond providing APIs that uh that the the your the
human or the the the the system will be able to consume. Uh by the way it's done now in a de facto standard that is the open AI like API right even if it's not ISO or it's a de facto standard. Uh they also provide other features like performance with integrated caching or dynamic loading and unloading of the model. So you can load unload the model based
on rules or very easily right. Uh they also provide uh features or capacity for developers like uh batching streaming the responses uh logging everything some observability uh uh features for people to to to learn and know what's going on actually. So basically it's to package the model to be used in an architecture where you will compose it with other thing right and here in this domain you
have numerous inference servers open source one of them that is known is lama maybe you heard about it right so it offers that caching dynamic loading and loading and stuff is very interesting also and a very dynamic uh community doing that so is the license is MIT VLM is Apache 2 so you see it's open source and permissive licenses and there's a lot of different actors contributing
for instance on VLM right [snorts] VLM is seen like the kind of the new docker for models right but you also have other other inference server like torch serve twon tggi from aging face uh and you can also serve your model locally with some uh desktop application like lm studio maybe you heard about it well it's not open source It's a freeware let's say uh but you
have Yanai that I love. It's a very interesting project coming from the east east Asia. Uh they they build everything in the open. It's a very powerful solution. So you have a lot of things. So you can serve you can optimize and then when you do that then you start using the model right. So uh first uh by users with users interfaces and then you have two
different kind of uh users. You have the hand users. So they will have their own version of chap GPT let's say so a UIX with a prompt and then I I have access to a model uh it can be web or desktop based so I spoke about Yani Thunderbolt maybe you heard about Modilla Thunderbolt it's a brand new project doing just that uh uh you also have
uh integration of the model within the browsers [snorts] now we can run some small model within the browsers there are uh standards that are emerging regarding ing this aspect uh and you can and you see in all the the tools of every day now AI is is infusing everywhere you know so the because you can have those specialized model uh not that big you can inject AI
everything AI is injected in my phone now uh I can run some models here and that's [snorts] right and but they're also targeting uh developers so for instance you know Jupyter notebooks which are coming from the data science uh world now it's uh It's a very very adopted way of sharing both the code, the result of the code, the execution but also documentation and of course you
have a lot and that's a domain that uh we spoke about a lot here until today in this conference about coding assistant and AI agents like FIA CI or cursor or open code and there there's a lot of solution for developers right now. Okay. So that's how to consume AI itself. Now how to build system with uh AI blocks, right? So two two two layers you remember
this. The first one is orchestration and augmentation and then we'll speak about agentic, right? But first orchestration and augmentation. Orchestration basically you have a frameworks like longchain whether it's the Python version, the type TypeScript JavaScript version or the longchain 4J version, right? or uh no code or low code solution like flow wise or kind of N8 but uh N8N sorry uh to build workflow based on AI.
Basically what they do they they help you orchestrate things including AI AI blocks like LLM agent and stuff. Uh they abstract the model so it's easy for you to design the system and then change the model change the the building blocks. They also implement the rag patterns retrieval augmentation uh augmented sorry generation. Uh so they provide a lot of connector to uh inject uh your own context
within the your own data private data within the context and then let the LLM give you an answer that is more contextual right uh and then here we have a lot of solution for vector stores. So when you want to do the semantic search that is required most of the time to do the rack pattern. So you have a chroma quant but you also have the the
usual suspect like elastic search that who added the vector vector store capaci capacity and semantic search or even posgressql right who supports that [snorts] and because we are we are we are speaking about building solutions uh distributing things uh we need protocols to integrate right so tools first with APIs so how to let the the the LLMs uh know about what is the context that he can
ask you to ask for him on his behalf, right? But let's say he calls it even if it's not really the case, right? But again to grab some new information and inject it in the context for you to have an answer that is uh uh relevant uh to your to your environment and now with the new kind of protocol like MCP that's standardized, right? So here is
the first generation of application that we we we built with AI component and then we move to Agentic AI. Aentic AI. So what is an agent? A lot of people have a lot of definition but basically it's an autonomous program that has a goal that you give give him or maybe another agent gave him, right? Uh and who will be able to plan toward to to achieve
this this goal, right? uh you build and we build most more and more multi- aent frameworks where you do orchestrate agent and you distribute the task uh between the different specialized agent and who say specialized agent with those SLMs with a different way of using AI you can have an agent that is using Gemini another one what is using F4 because they are not doing the same
thing and you don't want to spend the same money or or expose the same kind of data because it's private and you want to share it right there's There's a lot of framework to to to organize that. Uh some of them are the evolution of the previous solution like uh lung graph which is in longchain in the longchain ecosystem. But you also have um uh what what's
the name? Yeah, autogen for from Microsoft ADA from from Google and crew AI which is interesting because it's uh more distributed and the agent are auto organizing by discussing between themselves. It's it's not that that there is an orchestrator that is uh uh orchestrating everybody the agent are doing some kind of choreography they dance together right that kind of stuff right and for that you also have
new protocols because and when you have new protocols emerging in the ecosystem for me it means that it's mature it's maturing right because we need to integrate so we need the protocols so you have protocols like MCP we saw uh A2A to to allow two agent in different technology to exchange and to interact uh between themselves or even data for format protocol like tune optimized uh it's
not it's more optimized like JSON or CSV uh it's token optimized you know for to exchange the data with uh with models and inject them in the context and you can also have a specialized uh protocol I work in a payment company there's a lot of things going on right now about how to make sure that we have a standardized way to let the agent make some
payment on your behalf I don't know in which world we'll live in theuture few months but there are things going on right there. Okay. So that that's the local difference and when we have this openness this way of uh having a a way to easily run a model on this very particular box or in a container or in my private infrastructure then I can do what I
do for a long time as an architect design architecture and since the internet that we we all build together uh we know how to distribute things right. to enter the world of distributed and we'll see decentralized AI but let's start with distributed distribution so distribution we talked about it about the all the open source framework to do the orchestration the augmentation and uh the agentic workflow now
[snorts] and as I said you have some interoperability interoperability uh protocols that are there to uh help us to do it in a proper way and in a durable way right there's also the distribution of the execution itself so the inference or the training but let's focus on inference right [snorts] the the idea there is to overcome the hardware limitation all right instead of having a big
data center in Texas that that we will call colossus and try to build some nuclear plants around it because it requires so much energy maybe we couldn't distribute things right just like we did with the internet uh so there are different techniques to do that offloading offloading is a way of I'm running something on local but there is some mechanism that allows me to offload on the
cloud certain task or you know just like the scale out of the cloud of the cloud and you have solutions like Ola cloud so the amma community now is offering this also a cloud so they're going on the SAS uh business also uh but you also have a docker offload so you know you have docker uh model docker model runner that allows to run a local model
on your PC and then with uh the docker offload you can offload some some of the load on the cloud and you have other solutions but you can also uh have uh and for me as an architect it's interesting to see this the pattern of uh distributed computing so you start to have some gateway so light LLM for instance is a smart gateway so as soon as
you have that kind of component you can build your system right so it's a gateway and then it's smart so b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b based on on some contextual information, it will dispatch
your request to this LLM or this LLM or this other LM, right? You also have some the bus so the messaging uh pattern. All right. Okay. So on top of camel now with Wanaku you have this project that is dedicated and specialized about how to exchange between systems AI systems but with the bus pattern right and then you can also distribute the inference. So it means that
the inference won't be done only on one machine but on several machine right and you for that you have LLMD which is a very very promising very dynamic uh project. So basically it's the way to distribute the inference of a model on kubernetes right so there is an operator kubernetes operator and everything right uh you also have a ray uh ray uh which is also very interesting
and it's only not only limited to inference it's also to uh for training for instance and then uh the the weights are distributed on different nodes and then the inference is done or the training is done on a cluster right so we do distribute uh inference And we do distribute AI already today. And even the Yama C++ project which is very community based uh is just about
to release or just have released it the RPC feature of it which means that several instance of Yama C++ will be able to to collaborate. But this is distribution. So this modularity that we have we can distribute things. So distrib modularity is about how a system is organized right decentralization the last part the most crazy part of the of this talk now is not about that it's
about who controls and verify the system that is distributed right uh so the first one is distribution of task it's a technical act and the last one is a distribution of power and trust it's political who controls the system who controls the distributed system that's how some people tried in the past I don't know if you remembers you something about like the the hype for three four
five years ago right to decentralize things right so ah our superyama is time traveling again so now it's one year back uh 2023 because in this year uh I made a lot of series of talk about how it was urgent for us to redentralize the internet right it started like when I discovered the internet and uh I I helped build the internet. We all helped build the
internet at least with our own data, [laughter] right? Uh it was uh a very decentralized network. That's how it it worked. That's how it succeed right because we had the protocol to do that the low-level protocol. Then my children they discovered a very centralized internet with some platform whether they're business uh platforms or social platform uh that capture all the usage all the data that is uh
produced on the internet. So I speak about the the big tech giant the Google the Instagram the whatever. So uh they have they they they they decide what are the rules regarding the value uh that is created by us on their uh platform right and the decentralization and the all the decentralization uh initiative that we we observe with blockchain with web three web3.0 zero uh all that
kind of uh buzzwords yes is about how to create new protocols to make sure that even this value the rules and the business value it's managed by protocols so again that's why I use that the return of the protocols right now you understand that I my my I was young for for the first episode right so the pillar of decentralization and I won't go into crypto and
stuff don't worry it's not a talk about uh blockchain and that because everybody will be a little bit overwhelmed. the pillars and that that's what we'll find in what some projects are trying to do with decentralized network and AI are it's open to all so it's permissionless you don't need permission from anybody because if it is there is a central body who is distributing the permission right
uh you have actual ownership of your own data on your own component right and uh everything because we don't trust anybody we don't need because we do verify everything right so that's where you have all the crypto Uh so a token and here I I don't speak about the AI token, I speak about the crypto token. A good way to understand how useful it is is don't
think about money, right? Think about value. A token is something that bear some value. The value can be money. The the value can be power like voting right for instance in the decentralized autonomous organization. It can be an incentive. It can be used for something. You know it's a value and it's also a program a programmable tool the famous smart contract right and the token is the
basis to implement the key feature of decentralization which are incentivation so you do you play by the rules you earn some token right you cheat you lose the token right uh honesty you can stalk you can uh uh slash the tokens and uh consensus which is the the very essential uh protocol on how the different nodes of a decentralized network do agree of what is the truth.
Right? That's enough for now. Okay. What I want to do and that's the last stack of the the last part of this talk is what about decentralized AI. And then when I opened this Pandora box, uh I got really dizzy. I mean there are crazy things going out there. So I close it very fast and then I say no no no I have to look again but
try to have a map again. Right? So as an architect I did that. we can decentralize compute the data the intelligence itself verification and action let's go through uh those layers very quickly decentralized computing the idea is to reuse mutualize existing GPUs right instead of building new GPUs there's a lot of GPUs out there let's mutualize them right but how can I trust somebody I don't know
by using his GPU and why should I share my own GPU you remember incentivation all the token I don't go through all those fundamental mechanism of decentralized network. I just want to illustrate with two project. The first one is io.net. It's a generic marketplace where anybody you can rent your GPU, right? And they say it's minus 40 to 70% less expensive that a centralized solution to to
rent some GPUs. That's the first example. Another one is Ether. Ether is more aiming the enterprises. So they take SLA, they take commitment and security and have more homogeneous uh uh set of of GPUs. They can they even have some deals with data centers to use and and run their GPUs when they're not used, right? But today, yes, we can decentralize the GPUs and behind that you
you have all the they have their own token, their own incentivation mechanism and stuff and stuff to make sure that everybody get the value and redistribute it, right? So decentralized computing then on top of that the data right uh the what is at stake here is how transparent the data is how transparent and fair and proven is the sourcing of the data right because we know the
guys who took all the value from the internet they took all the data well we gave them the the data we we did write the the agreement and we did did agree about that that's why now they can build AI I mean they scrapped the old internet but They also had the a lot of data set from their own platform, right? So how how to make sure
that the the the the the sourcing of data is fair, right? And how can privacy be respected? How to trace the provenence? You have several project here. The first one is grass. Grass is a way to scrap uh data but uh decentralized. uh it's like uh for the end user it's like a extension in your browser where uh you can scrap the data and then you get
paid on the the the quality of the data that you just scrapped and that you shared with the community right and this ensure freshness of the data also you have VA is more like for privacy it's a safe uh where you you you put your own private data and you decide what are the rules uh who can access the data to do what and you can share
the safe with others so starting this decentralized autonomous organization for instance right and you have um zerog which is on onchain storage with I band bandwidth and parallelization so they they they found a way to distribute the storage on a decentralized uh network for AI right I go quick I don't need to go in all the details because you'll go too easy it's just for you to
have this global vision right now decentralized intelligence right improve the value created by models so it's not only about inference people who are trying to do that they they want to make sure that when we do decentralized AI, we bring some value to humans, right? Uh so the inference or the training. So how can we reward utility rather than calculation? Calculation is the the low level, right?
We we're in the level of intelligence. So yes, okay, you did the job, but what is the value? Can can can I prove that there is some value generated by by this AI? One good example of that is the Bit Tensor ecosystem. Maybe you heard about Bit Tensor. We could make a one-hour talk on Bit Tensor and still don't really understand what is Bit Tensor. It's a
a world ecosystem. It's a Darwin Darwinian ecosystem of sub network running on the on the blockchain right uh who has their own objectives. So some sub some some sub sub networks are for inference. So some are for embedded for reasoning for finetuning for serverless AI you name it. And between this this uh sub sub network there is the competition based on the token the in incentivation and
all the processes from the centralized AI but between the sub networks you also have a competition so only the best one will survive. So bit tensor with all this Darwinian crazy stuff is trying to let new users emerge to bring some value for AI with decentraliz decentralization network. I stop here because it's kind of easy, right? Another one maybe simpler is prime intellect. Maybe you heard about
intellect one. Intellect one was a model that was trained on a decentralized eterogenous set of GPUs. Uh [snorts] and they use some compression and uh uh exchange statistically important elements. So there's a lot of cryptography in all this stuff. They create their own protocols as I said, but they were able to uh train a model on a decentralized infrastructure. Pluralist uh also is a emerging uh solution.
They create node zero or node zero. Yes. Uh and they have something very interesting that they call multi-party computation where basically the the raw data used for training is not revealed. So it's heavy cryptography also. So you can use some data to to to to train a model in a collective way but not reveal the data itself. So that's crazy crazy mathematical cryptographic stuff. But there are
people doing that and there is a decentralized autonomous organization dealing about how the project is going and where it's going right will accelerate. So decentraliz decentralized validation so you have validation in every layer of the protocol because that's the very essence of decentralized network but there's also specialized uh network only for that for decentralization uh validation of things. So it's zero knowledge principle. How can I trust
the compute but I don't trust the the compute provider. Two example there aura any validator can uh challenge the the work for another node in the decentralized network. Uh so it's optimistic machine learning but if you get code you you lose all your tokens. So you lose a lot of value let's say money right. Uh Jensenis which is more emerging is interesting. It's a statistical game. So
the more nodes you have cryptographically the more difficult it will be uh to cheat right. So it's more and more and more difficult. So and this is just another layer that you can compose because you understood all the layers that I presented you can take some part of them and then compose your old solution with that kind of now the crazy part decentralized action. So we speak
about agentic AI right? What about agentic decentralized AI? So agents has autonomous economic enterprises because here they have this level of autonomy they have their wallet so they have their own token so they can decide with the token to play the game of the decentralized network right so that's another level of autonomy so how to prepare the machine to machine economy just two example to illustrate that
Olas uh who provides some autonomous services like models tools agents uh they have this uh this consensus uh protocol that is called proof of action. Prove that the the task that was assigned to you as an agent is done and then when it is then you get rewarded with a token, right? Uh and then you can have your this catalog of agent whether it's your own you
have a marketplace with your own agent that are deployed on this blockchain or uh you have a bazar where you can rent uh agent from others or or lend your your agent to other people. Right? And the last one, virtual protocol is interesting because it's a virtual entities that can act autonomously online. Uh so people co-own uh these agents and this agent will be able to create
social media content, create video, create gaming assets, uh that kind of stuff. And uh the people who co-own the the agent will share the value that is created by the agent. I don't know if you heard about this this uh this woman, the her name is Luna. I think it's from Korea. She has a number of follower that is completely crazy and it's autonomous. Uh it creates
values on the network. Well, value depends on what you find valuable, but some gamers and people in the entertainment industry thinks that there's value there. So, you see the world we're building. Yeah, that's interesting. Kind of crazy stuff. So, it moves like it's not an employee that you task with something. It's an enterprise that has a goal. You give us uh some token and money and autonomy
and then it will try to to to achieve the goal but in a decentralized way and by uh acting and interacting with other autonomous agent kind [snorts] of dyes right to uh finish and wrap up. Uh maybe is it a good idea to speak about blockchain and AI? I mean sustainability you spoke about that at the beginning of the talk. Well the answer is not that simple.
We speak about blockchain. Bitcoin, which is the the the worst ever uh decentralized network uh energy wise, right? And it allows also to optimize uh hardware existing hardware. So maybe to stop this crazy crazy race that we have like creates all these geopolitical problem that we observe today, right? It's also a good way maybe sometimes to optimize energy because some protocols will uh promote the nodes that
are cheapest and they are cheaper because they use energy where is cheaper to to collect like uh the green surplus not in the towns next to the waterfalls and that kind of stuff right as an epilogue. It's a little bit dizzy. I know I went very quick. I hope I gave you some keys to understand what's going on in both local AI openness and now decentralized AI.
Uh we are next to the cliff but you know this Yamama has a long long neck so we can all grab it and he's he's used to this those hates and he's not scared about that right [gasps] uh because that's the the the proposition of there's a lot of people who are trying to break down this monolith uh with distributed architecture with open stacks uh with new
protocols with decentralization principles with openness but for me it's not monolith versus distributed it's more if because I mean the world is multiolor and that's what I like as an architect to have the choice and depending on my use case I said yes this one I'll use a flagship model but for this particular uh use case I'll use this local AI and SLM and stuff and I
will decentralize thing maybe in the future if I dare to do that right what is the next t physical AI I mean AI is autonomous but can't act on the physical world there are things going on there already with robotics We speak about physical agent, mobile revolution where AI is trained on edge uh uh tiny ML and ambient intelligence. So you have always running uh powered by
battery of your of your sensor that is connected to the sensor and is analyzing the the signal uh on the fly you know and open robotics to do that we need some open uh models we need local and we need distributed to to build this world. So just to finish I asked the llama. I say you have a superpower. Can you go in the future because I'd
like to do where we're going with all this craziness with decentralization and stuff. He say yeah there are many possible futures. It depends on what you decide to do to today and what do you uh delegate from your own humanity to to those systems. And uh he he told me yes I can go in 2036. I say yes you can because I wrote a book of of
that. It's a sci-fi a sci-fi tour that I I co-wrote with the help of AI. You can scan it if you want to see. That's not the case when it's go very well, let me tell you. [laughter] Okay. Thank you very much. Uh you can scan this one if you want to uh stay in touch with me and uh have more details or more longer explanation of
some of the topics that I touched here. And thank you for attention. I don't know because I don't have the time here. It's it's there's nothing. It's finished. We have time for questions. Yeah. So if you have questions, I know it's a big subject but >> as you wish. >> My question is regarding the model weights which are being compiled on the hardware accelerators. Model weights are
hardware agnostic. Yes. >> So if we truly want to like democratize the local AI, how will the model weight compilation look like on different hardware like you mentioned? >> [snorts] >> That's why that's why you have those frameworks. They do abstract. As soon as you have frameworks layers in architecture, it's an abstraction. That's why you're able to have PyTorch or TensorFlow or those kind of network, they
will abstract the the the the hardware part. And the acceleration layers that I spoke about will abstract also the specifics of uh Nvidia, Intel or AMD GPUs for instance. So that's why we only ones which are open source like you mentioned four categories right only for the ones which have made their weights public >> well when it's when it's closed you you don't know what's in the
box >> you you just don't >> okay >> right okay thank you very >> [music]