Small LLMs at the Edge are the engine for open source scalable AI agents
About this talk
This talk challenges the assumption that AI requires large, proprietary models to be effective. The speaker, a Chief Technology Officer with experience in AI, discusses the evolution of AI architecture towards specialized, open-weight models that can be deployed on edge hardware. In 2026, it will be crucial for enterprises to adapt to new regulations and focus on models that offer high performance with fewer parameters. The session highlights the move away from the scaling race of larger models to smaller, optimized models that can effectively execute specific tasks. It also covers the importance of model distillation, synthetic data, and the operational impact of compliance regulations on AI implementations. The speaker provides insights into various models available today and emphasizes the efficiency gained by running AI workloads on local hardware.
Full transcript
[music] So welcome everyone. Uh I'm the only one standing between you and the launch. So I will try to be [laughter] very interesting and uh starting fast. In the next 45 minutes, I want to challenge one very common assumption and the assumption is okay, you cannot do AI without uh using frontier very big and proprietary models. I'm here to challenge this assumption and to show you that
there are plenty of use cases and solution that can be used to leverage that. Then first of all, uh who is the person standing here? I am Italian. Uh as you can probably guess from my accent I am chief technology officer at Mesa and at Lydia and overnight I'm not doing three jobs in one. It's just one job but we are part of a group and within
uh overnight and Lydia uh we work a lot of on AI especially with Lydia we had the very uh super cool opportunity to challenge the before under assumption and to understand how to move from big expensive proprietary models to something more open and this is the story that I'm telling you today uh by the way I'm also AWS hero cursor ambassador and genai Italy and open Milano
meetup organizer which means that I'm not adverse to public cloud or to sovereign to US-based cloud but in very narrow and interesting and useful cases we need to make better decisions then uh the assumption is that in 2026 enterprise AI architecture it is something different from what has used to be in the past. It is no more about the number of parameters. It is not no more
about the uh huge size of the model. It is a specialized openw weight model running at the edge hardware. And we will see that being able to run at the edge will solve a lot of issues that are differently difficult to to solve on external cloud-based proprietary models. Then uh what we are talking about today uh just why uh it is so important now and why it
is feasible now because we are in the middle of a market shift and then the uh small language model ecosystem and how we can address that and we can leverage that a technical architecture uh which worked for us and which could be uh elevated to a blueprint for everyone uh wanting to approach this uh this kind of solutions. Then uh a few words about sovereignity and regulation
because this August uh something is going to change uh and we have to account about that and industry evidences because uh theory without something working something put in production is just dumb theory and action points that you can take your own and start thinking about or start resonating about how to improve that. Then about market shift uh we are uh coming out from the scaling race the
so-called scaling race period which had been from 2023 to 20245 and it was just can be stated as larger models massive cloud clusters a lot of computive power on the table no optimization just bring me your bigger model and the new release will be bigger than the old one. Now we are shifting towards specialized uh less large models. Maybe not small language models but less large language
models for sure. Uh we have found out that models even frontier models needs to be optimized in order to achieve the best throughput, the best computing per token, the best quality. uh just making your model bigger won't fit uh for every kind of problem that you can face and uh every architectural decisions that we faced in the very first phase has flipped. We started thinking about bigger
models from 70 billion to 500 billions and more. But now we have sub 1 billions more that models that work uh quite well and less than 30 billions models that can run on on commodity hardware that work quite well as well uh and can be an equivalent substitution for specific tasks. And we will find out that specialization is something which is the clever point and it is
the sweet spot between frontier models and small language model. If you can specialize, if you can focus on a narrow set of tasks, you can move from large large models to smaller one. Then hosting hosting uh wasn't perceived as an issue in 2023 and it became a relevant compliance regulatory issue in 2026 and this is something that we need to consider. Then weights closed vendor proprietary weights
models demonstrated showcase that they are not suitable for uh optimization. They cannot be easily optimized. They cannot be easily distilled. They cannot be easily changed and transformed into other specific models. And this is a beforehand on open weights models that allow uh basically everyone even startups that start from this model to engineer new language models. Then interoperability uh we started with proprietary SDKs and then we moved
uh and this shift was driven by uh MCP and uh skills and all the stuff that has been uh released open source and then contributed into the Linux foundation and it is something that actually it is enabling a lot of use cases and then data and then unit of econ economics we are not measuring more uh as per token we are uh understanding how to move something
on and optimize that on a given hardware uh which can translate a not affordable opex into something more sustainable which is quite funny because it is an inversion of tendance an inversion of direction compared to the reason why we move to the cloud we moved to the cloud to mobilize capex into opex and now we found out that some kind of of offices are too expensive ensive
to be affordable for uh common use cases and then somewhere we are shifting back we are moving back these workloads into something that could be a capex but could be more manageable uh even because we changed the scale the factor is that we cannot be effic economically efficient at scale maybe we will be able to do that in one year or three years but right now uh
generative AI is not efficient at scale scale but it can be optimized on a lower scale and now we are facing three forces that are pushing in the same into the same direction. First one uh we are we have the availability of models that are small models and can match on given benchmark on public benchmark bigger models then uh the enforcement is really close. is the 20
120 days away because starting from August 2026, the AI uh act will be enforced and we will have to consider which kind of workload we are moving on AI and if it is something that is uh stated as label as higher risk workload, we will need to consider and to provide accountability, auditability and we will see that it it is challenging if you do do not own
the knowledge about the model that you're using and first of all we have hardware availability right now this Max MacBook Pro which costs say about a couple of cases uh or a Jetson Nano which cost less than $250 are hardware that is suitable to run language models and if we can optimize if we can distill language model uh to be able to run on this hardware we
and just run them with no extra expense, with no extra cost. Then the 2026 small language model ecosystem changed dramatically and a consistent shift arrived just a few weeks ago. Uh then now we have uh sub 1 billion parameters and we have techniques that can make these models uh suitable uh with uh quality that can be comparable to frontier models and especially I want to focus on
three techniques uh and three three factors. The first one it is distillation. We will see what distribution is and it is something that that can uh make a model a small model able to perform with the same quality of a bigger one on a specific focused set of tasks. Then we have the availability of synthetic data which are paramount. If you want to uh say if you
want to fine-tune the model you need a a data set you need a key value pair data set that can be fed into the model and now we can generate such kind of synthetic data using frontiers models of course uh and then we have methods that allow us to post-rain uh our train model. So we can align the model with something that could be our brand guideline
or uh something that could be our unwanted behavior or wanted behavior or wanted ton of voice. And we we have a very proven technique uh which is reinforcement learning and a lot of derivatives from reinforcement learning. The uh most the easiest one it is reinforcement learning with human feedback. But we have also something much more uh complicated than that that allow to have a model aligned to
have a small shift in the in the model weights in the direction of being suitable for Uh just just to be completely honest, I'm not selling you anything. Uh this uh advancement as every advancement cames with a trade-off. that hydropha is we are not uh using frontier models and their general purpose capabilities and then we are focusing on task specification. So if you if you need a
model that can basically tell you the recipe for a good carbonara and on the other side debug your code and on the other side uh be a good uh coach to your children probably small language model are not the best solution. But if you can extract a single domain specific set of tasks then you can narrow that the kind of behavior that you need. you can narrow
the capabilities that you need and then you are able to train or fine-tune or distill language model into them. Then uh we have seven different categories of models that are suitable uh to be used for that case and they are open source. They are available on on lama. They are available on open router. You can they are available on on game phase. You can download them and
you can run them today. I'm not talking about the missing parts and the missing parts are large language models which can be open source such as quen 3.5 397 billion but they require still require a huge effort on the hardware there. It is an interesting field maybe more for a uh solution provider, a big system integrator or a telco operator. Uh we have a solution we have
sovereign solution leverages these models but they are out of scope for this talk. What I want to show you is what we can do with model that have say less than 30 billion parameters and in many cases less than 10 billion parameters. So that can definitely run on the specific hardware. The quantry family is one of them. Uh we have different sizes. We have bigger sizes. The
quen 3.5397 billion is definitely the biggest one or one of the biggest that is available in the market. uh but uh we also have very small variance and this very small variance has an have an interesting context window of uh size which is 128 case which means that it is not too far from say GPT 5.3 contest window it is not too far from uh standard plain
set 4.5 200 uh kto tokens content context window. We are not dealing about uh 1 million. We have techniques to achieve uh pretty much the same the same result with a low with a lower context uh impact. Uh they are techniques related to such as say uh context compaction or context summarization uh that can be leveraged. They are out of scope for this for this talk. But
just to let you know that 128ks context window is definitely not too small uh to be usable for something. Then we have Gemma 4. Jamma 4s was released a couple of weeks ago. It is one of the most interesting solution that we have right now because it can run even on mobile devices. Uh I'm running actually while we are speaking uh I'm running Gemma 4 on my
laptop because I've installed open claw on that laptop and I'm running Jamba 4 and Gemma 4 in the uh E2B or E4B uh flavor. It is super fast on this m MacBook Pro. Uh it can output say 1K tokens per second. The E4B uh version of Gemma 4. I can also run there on this AR. I can also run gem of 31 billion which is a model
that embodies a very good thinking set of thinking capabilities. It is quite good for uh a lot of narrow tasks and can be leveraged. So, Gamma 4 is one of the best that we have and it is also multimodel which means that we can use this model for uh a lot of tasks that are similar to the same task that we used to leverage uh the frontier
model on. Then we have a special mention for F4 which is a very very small model that can be used to match specific needs such as uh PI uh reduction or uh say offus data opuscation or data reformatting. uh it has been built by Microsoft and it is super interesting because you can just run that model at scale on commodity server clusters and use that to anonymize
data then we can have we have deepse [snorts] R1 uh it's able to implement reasoning and uh reasoning of one class or one style on commodity alor and it is due to the fact that with deepse PR1 you uh can choose the kind of distillation. We um we will spend a few words about what distillation is in a couple of slides. But you can have different flavors,
7 billion, 8 billion models, very small models that are able to implement the same reasoning techniques. And when I say the same reasoning techniques, I refer to the fact that comparing on specific and public benchmarks, we have a result that is close to frontier models and in some cases it is even better than them. Then a special mention to RCAI. RCAI uh it is a startup US-based
startup uh and they released a lot of uh set of open source models uh their approach is something that is worth mentioning because uh the found one of the founder it is known for merge kit and merge kit is a tool that is able to take open weight models and merge them into a new large language model or small language model that has that retains uh some
kind of capabilities from the uh parent models. It is interesting because it doesn't require new training of the model. So you have your quen model trained and you have your deepsek distilled model and then you can merge this uh their weights into a new set of models without the need uh to have available GPUs or to use GPUs and they released RC agent 7 billion then uh
supernova medius 14 billion and virtual small 14 and 7 billions. Why so different models? Because right now we are talking about focused models. We need just one frontier model to be general purpose. But if you want to be able uh to address specific models, specific use cases, we need some kind of model with shaped down behavior. So we need a specific model that is able to follow
in instruction. If our goal is to have an agent model, we need a model that is able to generate text or summarize text. If our goal is on that focus, we are losing general purposeness, but we are achieving a lot of benefits in terms of cost than compliance and accuracy. uh just to give you a glimpse of about how uh the behavior or the approach changed. We
are not uh scaling parameters anymore. Neither the biggest frontier models are scaling parameters anymore. We are optimizing uh the models. we are shrinking down uh to prune parameters in and letting into the models all the parameters all the ways that re that are really useful. So we are just removing uh the noise and uh the we have a set of techniques that embody this kind of behavior
and we have shown we have been shown that they improve the quality of any kind of language models large small not so small whatever and the fun part is that uh uh this is the proof point on m 400 a three billion parameter llama uh running test times uh test time scaling outperforms 70 billion parameter dense model 23 times smaller. Oh sorry 23 times smaller. Uh you
can deploy a mod a memory efficient uh model locally and it can allow you to beat a very big not local deployable model. It is very difficult to deploy a 7 billion 70 billion model on this MacBook with A6 chips and whatsoever. And it is practically impossible to deploy that on Raspberry Pi or on Jetson Nano. But you can deploy a three billion if if uh on
specific Matt task uh math tasks uh the quality uh it is the same basically the same and you have you have a 23 per uh gain and uh you can deploy a number an array of them. It is very difficult to build an array of cluster. It is very expensive but it is pretty simple to buy say a 100 of just nano. Uh and then uh this
is something that I want you to remember. It is a paper that was released at the end of last year. So it is December 2025 and they were able to fine-tune a 350 million parameter op uh to achieve uh 77.5 uh 55% pass rate on tool bench outperforming chp chain of tooth tool lama and clo chain of tooth uh from up to 75% point. The interesting part
is that um the ability to call tools to use tools it is something that is for fundamental to build agents. So uh we are at the beginning of a paradox where something which is the point that generative AI field AI domain is shifting to which is agents and capable agents. It is something uh for what uh the bigger for biggest frontier model are not the best or
the most optimized solution. They showcase that tips that the fact I I want to um introduce two techniques uh that they used and then go back to this paper which is really a gamecher. The first one is that supervised fine-tuning it is the the cheapest way to turn a generalist model into a specialist models. Uh it em it embodies a task a teacher student model. Uh it
does sorry it doesn't embody a um teacher student model as distillation does. uh it's uh just define a set of tasks and the data set belonging to these tasks and then fine-tune the model on that. Uh it has a very uh narrow distribution. It is something that can be suitable and it at a very low budget. Oh, okay. at a very low budget such as less than
uh hundred dollars and it can run uh with quantization on uh RTX uh uh really on old RTX uh it can be trained on just one GPU Nvidia GPUs so it is a technique that can dramatically lower the cost of training uh of the model because yeah because you are not training the model you are just fine-tuning the model on a specific tasks specific task data set.
Then we have distillation. With distillation we have the capability uh to train uh using the teacher student approach frontier models say 400 billions uh one ter billions one terarameters uh down to uh 1.5 billion uh 32 billions 7 billions. So very small models and we can without losing uh too much without losing their capabilities. And here the architecture of the model is something that it is super
useful because a lot of uh modern models lot of models such as DeepS or uh Gemma uh they use a special techniques uh which I think that was invented by Mistral and it is called mixture of experts and with mixture of experts you have a formally statically big model say 300 100 or 100 billion parameters. But then you activate only a subset of this parameter uh when
you are do performing inference. And since your computing power is tied to uh the amount of parameters that you activate, not the amount of parameter that you have in your model, you are really running on a very small uh set of parameters. It is an a super interesting techniques that dramatically change the shape of the the landscape of the generative AI uh two years ago and now
basically every kind of models open source model is using that and probably even a lot of closed source models are uh leveraging that technique. But going back uh into our uh to our benchmark into our paper uh the basic uh behavior of an agent is okay I receive an instruction then I have to think about what to do. So I require a language model. I need a
language model that has some kind of thinking capability. Then I choose uh amongst my available tools to use that tool. Then I invoke the tool. The tool is basically an API. Then I collect the response and then I iterate. I understand whether or not that response was compliant with my initial input and then I send back the response to the user and then rinse and repeat. So
this basic uh workflow is the basic standard workflow of basically every agent harness out there. So being able to optimize that means that we have been able to uh increase the possibility the quality of uh agents calling agent calling tools. But we needed a work a test a test bench in order to measure that. and in 2023 and then revised in 2024 uh they published a set
of different tools calling benchmarks that can we that we can use to uh check our models and to identify whether or not our model are working. Then back to uh that experiment a single uh that model uh was just a single epoch uh supervised fine-tuned uh deliberately undercooked. So it was not uh overtrained. It was not trained for many many epochs. Just one one epoch. And then
they tried and they found out that this was the result. The uh 350 million model was able to score a 77% pass rate on that test bench. compared to the other model, it was an improvement, an astonishing improvement. uh if you if you compare that to Chad GPT chain of tooth uh because the chpt version of the model that was available at the end of uh uh
2025 uh to be used to invoke tools uh it is a triper improvement and uh the model it is uh say two order of magnitudes uh smaller than that one. we the one one thing that uh uh a lot of people argued when this paper was published is okay uh you just hit a single lucky category. So maybe they are good for say invoking APIs for structural
debt extraction or maybe they are good for invoking web books. So they repeated the test and the test was consistent among uh different categories uh basically all the categories of the paper. Uh this means that uh something was wrong with the generalist LMS. And then a study a paper from Entropic find out that uh LLMs tend to suffer from parameters delution which means that you have a
really a lot of parameters. you are spending a lot of tokens and a lot of uh you are using a lot of your parameters just to think just to discuss whether or not you should call that tool and uh uh but in most of the cases you don't know uh the kind of data that the tool is returning. So it is completely pointless to spend a lot
of time a lot of time into uh trying to figure out or in in some way or some others. You don't have this kind of complexity. You just need to route the request to the tool and that was uh uh an explanation about the this kind of very strange phenomenon where where uh three uh very small model was able to outperform really bigger models. uh then starting
from this starting from this point for this results uh we should understand how to uh actionize them how to use them. So what seems for me is the uh the the final question that we should answer to and if you consider which are the layers that uh we have to pass through uh when we build an AI solution for our company for our project whatsoever. uh at
the most basic layer we have uh the hardware and in this case it can be edge hardware. So it can be uh just work just nano or workstation such as this MacBook or an MPU or even uh Nvidia GPU uh and this is the silicon that you own. Then uh we have the model and runtime and we can use uh shrink down models. we can use efficient
runtimes such as VLMs to run these model or lama uh to run these models. Then we have the need to make these models uh able to interoperate to work with other systems to invoke our system and here is where MCP cames into the play because MCP changed that layer. Then uh we have the problem to orchestrate agents. Uh so to direct different agents if we are uh
focusing agents on specific task we need to orchestrate them and at the topmost layer uh we have the application your application basically in your code. a few words about quantization. This is not a talk about quantization. So I'm not hosting a two-hour lessons about that uh because a lot of details of quantization are very difficult and uh very difficult to understand even for me. But I want
you to have uh just a few glimpse few insights about what quantization is. Quantization is a method to reduce that you implement to reduce the precision of the uh floatingpoint weights that you have in your uh in your model in order to achieve a better size of the model in order to achieve a better performance of the model. So you move from standard front uh 16 floating
point sorry standard uh FP16 uh down to integer. So you are losing precision uh or even down to four bit integrals and in some cases just for such as gamma uh e to uh e2 b uh you are going down to uh uh in to two integrals. So you are really reducing the uh the size of the model. uh and the question is okay how much quality
I am letting on the table how much quality I'm dropping because it is a drawback and we can see that in some cases uh the uh loss of quality it is definitely upsettable uh just to give you uh a glimpse uh and something that uh that we found out last year say and near and last ago it was that quantization it should be hardware specific. If you
run at full precision, you can run basically on any kind of on a set of different hardware. Uh but if you need to reduce if you need to lose precision, you definitely need to run on specific uh hardware. And uh we actually we have hardware which are such as Apple silicord or GPUs that are uh optimized and if you can see uh we are losing just 8%
compared to the full precision to the not uh not reduced model on MLX on Apple silicon we are losing just less than 10% in uh in quality but uh we have a consistent gain in throughput. Uh in some cases such as the Nvidia GPUs, we got also uh an astonishing uh an astonishing 741 uh tokens per second throughput gain. Uh and the reason is pretty simple. The
model is smaller. The model is smaller. It can run more efficiently. It can run at more speed on the same on the same basically the same hardware. Uh just to give you a comparable example, the uh full precision uh 4.16 compared to uh for uh int quantiz uh this int 4 is able to run on consumer laptop or on uh Jetson or in nano or on RTX
uh 3060. much smaller and it can implement it can provide a significant and improved uh token per second uh figure compared to the uh 14.16 baseline. This is super interesting because if you can uh give up a few point of percentage com in terms of accuracy say we have seen uh seven eight uh point uh points in accuracy you can gain almost you can almost double the
number of token per second and the performances of the model um I don't know if you if someone of you has has been used has been using open AI codec The codex spark model uh just basically uses this approach. They they did not disclose a lot of details about that but they let us understand and let us know that uh uh they are basically uh using a
quanticized version of uh GPT 5.3 and the codeex spark is able to deliver more than 1k per second tokens. uh okay it it is not the most accurate model but sometimes is not required uh so significant u level of quality then in this scenario MCP it is something that allows these kind of specific tailored models to invoke uh external uh external data uh we are not deep
diving into MCP uh three points are the governance it has being the transport uh it can either be HTTP uh standard IO recently uh they have also implemented uh out 2.1 in order to have MCP servers handle directly the authentication so your language model can invoke a remote host a remote service through MCP and then it can challenge you back uh about say give me your credential
then uh following Oout 2.1 uh standard uh it can uh it can uh authenticate then retrieve the data or perform the action then uh I would like to suggest you to propose an architecture that can be implemented in order to leverage all the different models that we have seen. Uh this is the same kind of architecture that we have implemented uh with overnight and Lydia and uh
it embodies a supervisor agent uh that is a model uh that has only one duty and it is to fire sub agents. Then through MCP it can communic communicate and exchange data with specialists models or models train specifically on given tasks that you want to accomplish. Uh these are a set of quite standard and common uh specialistic models. Uh the librarian something that can retrieve it can
implement and perform rag operations uh on a remote uh vector database or an analyst that something that can run SQL uh statements or it can uh write and run uh Python code. uh but uh in some cases it not it is not our case with Lydia because with Lydia we stick on the uh first three uh agents but in some cases especially in industrial cases uh you
need vision you need to classify uh data you need to classify images you can do that you can use specialized agents for vision for classification maybe using gamma 4 uh that we have seen to be able to implement some kind of multimob modality and uh uh going farther just really a very few words about that uh putting all together if you want to take away one uh
disruptive statement for this session is this one we have been told that it's just a matter of having our service within the EU borders is not true is not sufficient because a physical location uh it is irrelevant when you have to comply with the law And the law it is related to corporate domicile. It is related to where uh the company that is giving you their hardware
uh is located where is the their uh main headquarter. uh we have we are facing some kind of very harsh uh issues related to uh the enforcement of cloud act and GDPR articles and things are getting even worse within a couple of months and the reason why they are getting even worse is that when AI act becomes into becomes legally binding uh if your use case it
is uh labeled as high risk uh you have to provide by technical data governance, technical documentation or you have to record uh pro record keeping and logging transparency and you also human oversight and it is very difficult for closed source models. It is very difficult even to not have data retained uh when you are dealing with this service. Even in some cases uh the providers are not
able to assure you that data is never retained or never uh acquired by external uh pe people external from your organization. And in some use cases is not a matter of designating someone to be your data processor. It's just a matter of something that you cannot do. If you are handling healthcare data but also if you are handling educational data or if you're handling employment data you
cannot you cannot share this data with people external to your organization according to the AI act. The AI act it is stricter than GDPR in this sense with GDPR you have to clarify which is your legitimate interest and then you can provide that data with AI act is not possible. Uh which means that uh uh if you leverage on edge deployments you can uh use these three
uh built-in compliance properties and to end lage you own the uh all the life cycle of your data. Then the processing it is uh and even your IP it is within your company boundaries and you can implement closed loop logging because you can log uh all the data going through your model and uh you can see even which kind of weights are activated and you can even
uh understand how your model performs because you can you own the complete life cycle. [snorts] uh just in 30 overview of uh some kind of use cases that we are are really uh really seeing fraud detection. Threat detection is something which is critical because it handles personal uh account data of financial account data of people and we have a latency target which is really constraining then uh
healthcare data patient data I've been working with a lot of uh healthcare institutions and it is really a critical point you cannot share your data you cannot provide the data you cannot contribute patient data because patient data is something that cannot change from time to time is something very critical. Uh then in some cases even industrial uh cases you cannot afford to have uh um internet uh
failure of some of any case. uh owning the silicon and can end the scarcity argument and it can provide you up to uh 250 bucks dollars per 7 almost 70 uh teraflops of data which means that starting from today you can just take this picture of this uh this chart which is the last one and uh you can focus uh by design on sovereign edge uh model
deployment and then treat every uh frontier public cloud deployment as an exception default to sovereign unless you can explain why sovereign solution are not suitable to you. This is a conservative approach is something that work in practice is something that we have been using with Lydia which is a legal tech legal tech domain and with our healthcare uh scenario healthcare uh companies [clears throat] then the intelligence
layer must operate within the boundary of your company and uh under your regulatory compliance. It is something that I've been seeing in the last uh say one year and a half uh getting more and more sustained more and more relevant for a lot of companies for a lot of use cases and now we have the tools to start so first to start edge first and then understand
if our use case may be in some narrow uh particular uh sub agent or uh aspects could be delivered to be uh deferred to external frontier public corporate obscure models. So start uh edge first then move cloud later. Thank you very much.