Keynote: Gemma 4: Compacting Intelligence for the Edge - Léonard Hussenot
About this talk
This talk explores the evolution of large language models (LLMs), highlighting the shift from pre-training approaches focused on next token prediction to instruction tuning techniques. The speaker discusses the importance of few-shot prompting and the introduction of 'thinking mode' for improving model output accuracy in reasoning tasks. Key advancements in model capabilities are showcased with the release of the latest Gemma 4, which demonstrates how smaller models can outperform larger counterparts in various benchmarks. Additionally, the speaker emphasizes the significance of accessibility in AI development, advocating for cost-effective, user-friendly, and versatile models that can be integrated into diverse applications, especially on devices like smartphones. The talk concludes with a vision for the future where smaller, intelligent models enhance user experience through effective instruction following and tool integration.
Full transcript
Hi everyone. Thank you so much. I'm incredibly excited to be here today and not only about you know, the topic which I'll get to at one point, but mostly because of the crowd and to explain you why I'll do a bit of history. Um if you roll back to 2020 um models were trained with what we call pre-training, right? Next token prediction. So you were feeding in
Zidane is the greatest footballer of all and it was outputting time because it was a good model. And you know, LLM users figured out that what they wanted was actually the model to follow instruction and not do next token prediction which in itself is pretty useless. Uh so they figured out that if they give an instruction and then show instruction answer instruction answer instruction then the next
most probable token was actually the answer to the last instruction. Well, this was framed as few-shot prompting. And it was somehow the premise of instruction tuning, right? So what people building model changed is that they introduced a conversational format and they started to train on instruction and answer pairs. So you know, then you could ask in this new format who is Zidane? And a good model would
answer the greatest footballer of all time. Right? You know, for a couple years we were training exactly like that, you know, solve this equation in the conversational format and the answer was, you know, the answer to the equation. But LLM users were like, "Cool, but I just realized that if I add think step-by-step to my prompt I'm getting a more verbose answer. Like the model is going
to spend more tokens to output an it answer, but it's going to get a huge boost in accuracy in all reasoning tasks, right?" And people building models are like, "Okay, let's me let me introduce this in the training actually." And we introduced what we call thinking mode, right? So now if you prompt it to think it's actually trained to think for longer before answering with reinforcement learning,
but not only and it will give you massive boosts in accuracies by spending more of these tokens. Same, you know, a couple years later you know, the models are trained to make a function call, you know, to support their answers. Typically it would be a Google search, right? To make sure the answer is factual. So models are trained to do one function call, two function calls to
help you know, the answer quality. And how it ends up being used is actually in harnesses like anti-gravity, cursor, deep think, deep research where there's hundreds of the tool calls that are made, right? Um and you know, then people building model are like, "Okay, cool. Let me actually introduce more multi-turn, more multi-step data or more multi-turn multi-step environments in my training." And now you can do stuff
where you know, you give access to bash to your to your model and when you ask it to debug it will, you know, check the logs, build tests, run tests, etc. etc. and that's the whole agentic systems that you're familiar with nowadays. let's go back to my point why I'm excited to be here and because of the crowd and it's because the users of LLMs are driving
the LLM development just as much as researchers, right? And that's really exciting and that's why developers, students, you know, researchers sitting in this crowd, they are also the people that are co-building these advancements and that's why I'm extremely excited to be here today. For me this calls for one thing and like if we want this to keep going, we need intelligence to be accessible and accessibility is
somehow the critical bottleneck here. And I broke down accessibility in three parts. Maybe you can tell me if I missed something. First there's the cost of intelligence, right? You need intelligence per byte to be as cheap as possible. Second, you need it to be easy to use, right? You want multi-modality, you want agents and you want this natively supported by the model. And third, you need to
be allowed to use it for your applications, right? So let's go a bit one by one and I'll try to explain you how we try to tackle that for Gemma 4 which was released last week in case you missed it. So we had the chance to test two of our models on on LLM CS. As you can see they are landing in a very sweet spot outperforming
models 10x their sizes. I know LLM CS isn't perfect, right? No, actually don't trust any benchmark, right? Check them out yourself, but it's still very very interesting signal to see that we have uh you know, either a 31B dense model or this 24B with only 4B activated parameters you know, outperforming models 10x their sizes is just amazing. And I think you will have to revise your priors
on what you know, model size you need for your tasks because this is just you know, crazy. And I'll dig a bit further in how crazy this is in this slide. So on the very right here you get this Gemma 3 27B model. So a year ago this was an amazing model and throughout the year it was still a pretty loved models and people at Hugging Face
they can tell you how how many times it was downloaded, but throughout the year it kept being downloaded and used for many tasks. Just look at these columns. If you nothing is tickling your mind there's an issue here, right? We have an effective 2B parameter model outperforming on being on par to this 27B model from a year ago. So I used to think, you know, anything below
9B is pretty, you know, it's a gadget that has to be fine-tuned on a specific tasks, cannot really generalize, it's a bit annoying to use. This here has completely shifted my understanding of that. Like the cost of intelligence has shrunk has shrunk has shrinked, I don't know, but you'll figure it out. Probably Gemma knows better than me English. It's just it's just ridiculous how now a 2B
model can do incredible stuff, right? And of course at all model sizes this growth is amazing. So like these 31B dense and and the the MoE are are just amazing and these benchmarks. I mean, look at the Tau 2 bench like very agentic, very trendy. You go from 6% a year ago to just 86% now. Things are just easy and natively supported. You know, it's it's amazing.
Just try it out. >> [clears throat] >> To make sure people realize what's going on here. So these evals we didn't run ourselves, so they were not, you know, used during model development inside the team because they were run by arena.ai. And I want to show that because I want to show that with we you know, people think that you know, if you do small models they
have to get you know, be specialized, they have to be good at one at Gemma we kind of believe the opposite. We're like, "Yes, if you have very very small models, you know, you can get better at one task if you specialize on them, but still the role of the Gemma team is to provide you with an intelligent model and intelligent at just everything." And you can
see here how at equal size roughly, you know, between Gemma 2, 3 and 4 the model just made progress on all topics possible, right? Um and this makes it just an amazing base model to work with to either try zero-shot or fine-tune then on your applications. So you know, this makes it both cheap and super easy to use, right? On the ease of use I wanted to
show you know, how on day one people are able to you know, use skills on their phone locally with a 2B model, right? And it's just insane, you know, this is a 2B model running locally on your phone. 2B runs actually with good quantization not even on a high-end phones. So on on high-end new Pixel phones you'll be able to run the 4B model, but the 2B
runs kind of a on 3-year-old phones nowadays. It's crazy to get access to all this new technology you've been seeing, agents, skills and you know, being able almost to develop with your voice on your phone without access to the internet. This is opening so many so many use cases, right? And that's why you know, people always eager to showcase you know, demos with their bigger model, right?
So it's always tempting to show you you know, what the 31B can do amazing. But you know, making demos with a 2B model locally on a phone, that's actually what excited me the most. So I want to go with a second actually a second demo. Um so here I'm so let you we're disabling the Wi-Fi just to show you no internet is used here and it's the
model is run directly on WebGPU on the browser, right? And so it's plugged. So you're talking to the model on the right and it's outputting the JSON that controls the robot, right? It's pretty easy task. You see that it's a very small JSON that has to be, you know, that is controlling you know, the the robot. But still, you know, it's a 2B model running directly on
WebGPU on your browser, right? And you ask it to walk forward and non-sanctioned, right? It needs to be just understands format of JSON, understand the task at hand, ask, you know, understand the instructions, output the right format, and that's, you know, we tried it out yesterday and it just worked, right? And that's, I think, what I mean by ease of use, uh, you know, we're putting it
in your hands and it just works, right? Uh, and and that's why I think you need to really update your priors on, you know, what size of model you need for your task because, yes, bigger models are getting stronger and stronger, but smaller models are just going amazing at the moment. Um, okay, I could not finish by still not showing you the 31B uh, also acting, you
know, best places for this food in Seattle under $30. Uh, you know, you're all used to this kind of task now. Multimodal, agentic, access to Google Maps, the model figures out it's a ramen ramen, not so hard, I agree. Uh, but also check out the ramen place in Seattle, check out the prices, check out the routes to go there that are the fastest from where the user
is located. And you have access to that locally on your laptop, right? 31B runs on your GPU. Um, I mean, it's just just amazing to have that, uh, in the in the hands of the users. Uh, of course, you know, you're the open-source community and I I can't imagine the creativity you have, uh, and you also plug that, you know, with models on the cloud and how
they're going to interact together is going to be just so amazing to see. I can't wait to see, actually, what all of you are going to build with that. and, you know, so I had cost of intelligence, I I showed you how this shrunk, right? Uh, I had ease of use and the license was my third point. It's the first, actually, uh, Gemma generation to go out
under the Apache 2.0 uh, license and we're super excited about that. I know uh, the community is, too. Uh, that's pretty nice. >> [applause] >> Uh, so a lot of work to get there, uh, but we're super super excited. And now that you're getting these three ingredients of of accessibility, uh, I really hope uh, you'll be building amazing stuff on top of Gemma. Um, I want to
insist on one thing. Um, I my job as post-training lead is to look a lot at benchmark and eval, right? Um, and although they are, you know, the most important thing that we're we're doing, right? Our work is, you know, the measuring is the key to making progress, you can't trust any of them. There's not a single metric that is a good representation of the quality of
a model because we're looking at general intelligence, right? We're looking at something that is supposed to be able to virtually anything, right? Now, not only output text, but also actions that, you know, activate agents, etc. So, you know, I just encourage you to try it out, basically, you know, try it out on your tasks, make your own tests, uh, try different prompts, make sure you're using the
proper formatting of the model you're using, etc. because, uh, you know, benchmarks are cool, uh, but we're really trying hard to make the model feel great and easy to use, um, and that's a bit of the magic behind behind Gemma. Uh, so, just try it out. And I'll wrap up with a couple of thoughts. When [snorts] we released Gemma 3N something like 8 months ago, we targeted
that phones only. Um, Andrej Karpathy, that most of you probably know, tweeted that and it really resonated with the way we think of models at Gemma. Um, so, first, he said, you know, models have to be natively multimodal. Um, they have to have reasoning, use aggressive tool, and he started to make a very interesting distinction between knowledge and you know. My point here, uh, is you don't
you can't always expect a 2B to store the whole knowledge of the world, right? Uh, and what I like with what's going on at the moment in agentic system is that people are figuring out that this is the case and they don't need these 2B models to know about everything anymore, right? Now, people are used to plug their models with the web, with Wikipedia, uh, and with
all the tools needed, right? And so, if the model doesn't know about William the Conqueror, uh, but it knows that it can Google search it or Wikipedia it and then answer it correctly, that's actually much stronger, right? And so, we we're seeing this trend with like smaller models that are getting smarter and smarter, uh, but also needing to accept that they don't know about everything and it's
the role of the community here to plug them with the right tools to really get the most of Um, so, like, that's why I'm I'm saying, you know, really be careful how you use these models in the sense that they can reveal themselves extremely powerful if you plug them with the right tools and if you use them correctly, right? And at this point, you know, when you
get this kind of intelligence on your phone, on your Raspberry Pi, on your Jetson Nano, etc., what matters is instruction following, right? If it follows the instruction right and you plug it to the right tools, you can just build the most powerful, private, fast, accessible applications in the world. And I think you're the crowd to do that. Uh, so, here you go. Uh, we released Gemma and
I just can't wait to see what you'll be building with it. Thank you very much.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17