PyTorch Conference Europe 2026

Keynote: Community Led Open Source RL - Joe Spisak

12:13 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk is presented by Joe, a significant contributor to the development of PyTorch, and now a key figure at Reflection. He discusses the advancements in training techniques for large language models (LLMs) that have evolved over the past year and a half. The conversation highlights key concepts such as reinforcement learning from human feedback (RLHF) and new training methodologies, including the integration of various environments and harnesses. Joe also introduces Open Enve, a platform for developing and training these models, mentioning successful community engagement through hackathons. Finally, he shares insights on Reflection's mission to foster open intelligence and superintelligence ecosystems, focused on sovereignty in AI deployment for companies and governments alike.

Full transcript

Good to see everyone. It's a far cry from Midway and Dog Patch from 8 years ago. Pretty cool. Welcome everyone. I'm Joe. I'm as 1 week old at Reflection. So, I was formerly at Meta. Um Yeah, a little bit about me. I spent a lot of time on PyTorch. I worked on the the project for almost 8 years, actually over 8 years now. Um now at at

Reflection, I helped found this foundation and and still sit on the the the core maintainers. Uh I also do a a fair amount of angel and advising. Hopefully, you'll like her and all the companies there and I also invest with Sequoia. Yeah, so I mean, today PyTorch, I think you you saw Ed's talk like a lot of stats, a lot of great, you know, it's it's kind

of the foundation. It's the open language as I call it of AI. Um you know, if you look at research, it's underpinning a lot of the the cutting-edge researches out there. Uh it powers top labs, Meta, OpenAI, my own company uh new company is is of course using PyTorch. Um we have a great ecosystem of cloud providers, hardware vendors. The ecosystem is super broad. And of course,

we have a a number of new projects. We voted ExecuTorch in to the core. Helium's now part of it. Uh we helped bring vLLM in and Ray um and of course uh DeepSpeed. So, it's a it's a pretty awesome core set of projects. So, but today I want to talk uh briefly about what's going on with training. So, you know, Ed kind of talked a little bit

about what's going on there. Um what's happening in in frontier labs actually quickly changing over the last I would say year, year and a half. So, of course, the early days, we had this kind of like compression stuff of pre-training where you had unsupervised training that was happening. Large corpus of of unlabeled data. Moving to SFT. We did a lot of SFT in the Llama um and

it helped kind of tailor the the tone of the models and and the safety and and so on. Um we brought in RLHF. So, this is human preference data. And of course, we're doing a lot more RLAIF uh these days to generate that. Um and then of course, like we're bringing in these RL environments. And then I'd say in the last year here, the biggest change is

we've included a harness or multiple harnesses uh for for the labs that are trying to generalize their their models to this kind of universe of of use cases And so this has incorporated a lot of a lot of a lot more complexity and it's kind of changed how training and inference actually it's not really a separate step anymore. It's much more of a continuum. And I would

actually argue that because of all these changes like training actually is kind of like a special case in a lot of ways when you kind of have a reset button in training but you wouldn't have that obviously at inference time. So what is an agent? So very quickly it's you can consider it basically an LLM plus some tools and some memory. And this tool use you know

you think about calling you know so like MCP or some type of search or something like that taking some actions you have these environments basically where you have some reward and some state. And those generate some trajectories that then generate experience and you train over the the RL and then hopefully your agent learns. And so you know really when you think about the customization of these models

and these systems you're kind of going from you know things like prompting or in context learning all the way to like something really really high effort like pre-training which requires a huge cluster of GPUs. And but but really what's what's happening these days is you get these new capabilities are happening these models. So for those who remember SFT and like the days when we were just kind

of fine tune over some label data we could do a lot right? We could actually change the tone of these models we could actually get it to learn different languages we can have task we could get it to teach it to to you know do things like call tools. What RL gives us is actually the ability to do things like multi-step planning and long horizon tasks. This

is like computer use and you we can actually even reward shape safety which is pretty interesting. So instead of just having you know basically telling the model to do things you can actually learn in in kind of safety sandboxes. And so, you know, where we are today is is we have tooling and infra that's actually exploding and I'll talk about this in like the next slide, but

we're actually lowering the barrier to actually using RL across the the ecosystem, which is really exciting. So, my own work and and the team's work at Meta and of course what I'm planning to do at Reflection. And you can kind of see like some of the the tools and the algorithms and things that are open source, whether it's RL frameworks like TRL from Hugging Face or Varro

from ByteDance. You can see even you know, Titan and I think there's going to be a talk on Titan somewhere as well as Forge, which was our original RL library to some of the environments and work like Open Enve, which I'll talk about in a minute here all the way to the algorithms themselves, which underpin a lot of this. So, about those environments. So, actually who was

at the PyTorch conference in San Francisco just a few months back? Okay, a few hands and saw Lisandro and myself up on stage announcing Open Enve. Pretty cool. Well, the project was actually in development I think for only about three and a half weeks before that launch, which was kind of exciting. Thank you Cloud Code, that was super helpful, but of course you have to you have

to have an idea and you have to have you know, you want to do something, but yes, Cloud accelerated the living hell out of us, which is pretty pretty cool. And so, these environments that we've seen, these are just kind of a snapshot of some of the the ones we've seen. You know, we're basically incorporating these into like you know, high compute mid training or an RL

in in kind of mid training all the way to post training and it's allowing our models to generalize and learn and and think about things in in more longer horizon terms, which is really cool. Um, and so, we introduced Open Enve and I removed the word introduced from the slide because it's out there and it's exciting and we just did a hackathon about three weeks ago or

four weeks ago in San Francisco and I think 160 environments were generated over that weekend, which is pretty cool. So, now basically we have if if you haven't learned about what Open Enve is, it's really two things. One, it's a spec, so it's a it's a GitHub repo which you can go to under Meta and allows you to basically see like what is, you know, what is

the interface and so on. And then there's a hub that sits in Hugging Face. And basically with the API that's that's on GitHub and there's a CLI, you can actually push to the hub once you have an environment. And so starting with the spec basically, we use a very very much an open approach. Davide Testuggine from Meta is largely writing a lot of the RFCs and they

get published. And so things like, you know, supporting the the API which is kind of a gym like, you know, a gymnasium style API. We support MCP as a first-class citizen, reward pipelines. And most recently we added agentic harness integration. And this kind of, as I mentioned, this actually becomes much more of a continuum. You can imagine you know, Open AI plus a harness plus something like

Torch Titan or Megatron underneath and you start to get agentic development and not just not just doing training and deploying for inference. So these things become much more dynamic. And then of course the hub itself, I think we have at least hundreds of of Big shoutout to Ben Burtonshaw from from Yeah, we have I would say more environments every day and you can create a collection and

then you can actually train over those. Um and you can see like what would happen if you actually pushed an environment into the hub. You actually get this human agent interface which actually allows you to even play with an environment directly. So you can actually prompt the environment like a sandbox, see what the result is and kind of reset the environment, get the state, etc. It's actually

really cool. Um so yeah, go go get started today. If actually you want to hack if you're in India by any chance in a couple of weeks here, we actually have I think 40,000 people planning to hack We're we're working with Scalar, Hugging Face and of course PyTorch, Meta. And so there's $30,000 in prizes there so you can sign up for that if you happen to be

in India. Cool. And then there's there's the link to the repo. Okay, so I'm going to switch gears. I have just a few minutes. I want to talk about Reflection, my new new gig. I'm very excited. Um so, who's heard of Reflection? Okay, more than I expected. So, it's actually a pretty stealthy company. Uh so, we've we're very shy company right now. Uh so, well capitalized. Um

and so, we're building open intelligence, open superintelligence. Um and this is something that of course if if anyone knows me, I'm very passionate about it. I worked on Llama and led the open source effort for Llama for Meta. And so, when we think about open intelligence or open superintelligence, there's three, you know, pet pieces to it really. There's this research community that we're going to be building

this ecosystem. There's builders themselves and there's obviously countries. Um I'll talk about this in in a little more detail. So, the way we we think about Reflection is in the early era of superintelligence, you had OpenAI with consumer, of course ChatGPT. Uh a lot of us had accounts or have accounts. I had an account. Um there's cloud natives like Anthropic, which a lot of us use cloud

code. And then what Reflection is going after, which is what we're calling sovereigns. And a sovereign in this case is really um anyone who wants to control their intelligence. Doesn't necessarily need to be a government or something like that. It could be for example someone like Cursor who wants to control the underlying model and and systems and agents underneath um their kind of scaffolding that they use.

It could be large enterprises, but of course it could be governments as well. And um uh Reflection announced the a deal with South Korea just a few weeks back. And so, why do you want to control your intelligence? Why not just like prompt a closed API forever, right? Well, I mean there's obviously autonomy. Like you want to be able to control the intelligence. It's important. It's strategic

for many companies. It's actually strategic for all of us. Uh you want to be able to customize. You want to be able to do even, you know, a deep level maybe a post-training, but even do things like continual pre-training. Um cost obviously, you know, when you you start to get hooked on APIs, you're kind of buying not only that token, but you're buying all that infrastructure behind

it. And so that margin can can be a lot for for for a lot of companies. And of course having, you know, kind of the data uh residency and security piece is is quite important as well. And so shipping your data and shipping especially sensitive data um off to to close APIs can be prohibitive for many regulated industries or just companies that that want to protect their

users. And so if you look at the the current state of the models, this is from artificial analysis. Uh I grabbed this uh just a couple days ago. You can see there's not a lot of US-based models that there are western models that are out there. A lot of them are are Chinese. And that's a problem um for a lot of companies and a lot of a

lot of folks out there. Um and so, you know, I think this is something we we hope to solve. But you can see basically uh you know, over time the open frontier models, you know, they kind of you know, between the China and US and I'm just showing kind of GPT OSS here uh versus when, uh there continues to be kind of a competition between uh the

countries. And I think there's a big wave of uh of western models that are definitely coming. Um and of course from the from the US government perspective, there's a lot happening. Uh we work our our company, part of it, works significantly with the with the US government um and exports So when we take a step back and actually look at what we need actually to be able

to deliver kind of super intelligence or open frontier intelligence at this scale, number one, we need a team, and the team has to be exceptional. Um it has to be at the top. These are people that have actually built um intelligence at this scale at other companies like OpenAI. And I'll talk, you know, Anthropic and and Google. We need a strategic backer. Um and every kind of

super intelligence lab to date has had a strategic backer, whether it's OpenAI historically with Microsoft or Anthropic with Google. Um and of course with uh with Reflection, we have uh Nvidia. And there has to be of course a commercial way to, you know, make money ultimately. We can't just give everything for free. We actually ultimately need to have commercial structure that supports this scale of investment. Um

and you can kind of see like who's who's part of this. This is, you know, folks that worked on everything from Gemini and AlphaGo to ChatGPT to, you know, Claude and and so on. Um and you can kind of see uh and recently uh we we raised our Series B at $2 billion. So, we are super well capitalized. It's still a very small team, um but we're

we're backed by some of the best VCs in the world. Um and of course we are hiring. If you are interested, uh shameless plug. Um it's a it is a great team. It's we're based in California in San Francisco, uh New York, as well as London. Uh so, please uh check that out if you're interested in Reflection. Thank you. Appreciate it.