Linkerd: Reliable Production in an AI/MCP World - William Morgan, Buoyant
About this talk
This talk is led by William Morgan, director of the Linkerd project, who discusses the role of Linkerd in achieving reliable production in AI and multicluster platform environments. He provides a brief introduction to Linkerd, a service mesh designed to simplify service communication in Kubernetes, emphasizing its focus on security, reliability, and observability. The speaker shares insights into recent features and improvements, including protocol declarations and integration with GitOps for multicluster linking. He also addresses the project's unique approach of utilizing custom microproxies in Rust instead of the widely-used Envoy proxy, highlighting its focus on simplicity and low overhead. Furthermore, he poses questions about the impact of generative AI on platform reliability and discusses the potential for integrating AI capabilities into Linkerd.
Full transcript
Okay, good. Uh, when I came in, the nice gentleman back there said, "I'm the star of the show." I prefer to my to think of myself as a you're your jailers. So, if you could lock the doors back there, no one is going to escape this room until the presentation is done. Okay. Thank you all for coming today. Uh, my name is William Morgan. I am a
director on the linkdity project which means I'm in the open source uh uh role but you know all I really do is edit the read me and and the website. I'm not really writing a lot of code uh myself these days. The title of this talk is reliable production in AI and MCP world. I'm I am going to talk about AI because you basically have to uh
to be here but uh I'm also going to give kind of an introduction to what linkerd is. I I'll keep that brief uh and go over some of the recent features and changes that we've made and and things like that. Um, so hopefully there's a little bit of everything for for or a little bit of something for everyone in this talk. Um, so before I get started,
you know, maybe we do a quick show of hands. How many people are new to linkd? You're here to learn. Okay, good. Gosh, you know, it's really encouraging to So, thank you for for coming to this. Really encouraging to say that. Uh, sometimes I feel like I've been talking about linkerd for 10 years, which I have. Uh, and I often say the same thing over and over
again. Um, but I'm always encouraged by the fact that there's new people coming to the community and and new people uh who are here to learn. How many people have been using linkerd in are currently using linkerd in production? Let's put it that way. Okay. All right. Good show of hands. And uh how many people have been using linkerd let's say in production or or not in
production for over uh over three years? Okay, good. Show of hands. How about over five years? Wow. Okay. Gold star for you and for you. Yep. Hi, Benedict. Uh, what about in production for over 10 years? Okay, good. You're truthful. You're all truthful because linkertd is 10 years old. So, over 10 years would be logically impossible. Okay. Um, I put together this nice set of like other
linky related talks and I realized by the time I did this set most of these talks already done. So, it goes too late. you're too late. Um except for uh we do have one which is not strictly linkertyd related but is from a linkery maintainer from Zahari about GPU profiling that's happening at 3:00 today. So if you're interested in learning how to profile GPUs with eBPF please
go check out his talk. Uh and then in terms of like maintainer representation uh at the conference um we have a uh a booth in the the project pavilion. We have a link booth and there's a couple maintainers and and other kind of linkd related folks hanging out there. So please if you are in that conference hall with 10,000 companies trying to sell you stuff and you
feel like you're you don't know where to go and you need a safe place go to the go to the linkerd booth in the in the or kiosk in the project pavilion and come say hi. Okay. And then of course you know you can check out these uh uh these talks afterwards as well. the supply chain security one uh was very well attended and that's based on
a bunch of work that we did in uh in a version of linkerty. So there's cool stuff in there. Okay. Uh so very quick overview for those who are new. Um it is linkerty is a a service mesh. Service mesh means that it sits uh on top of Kubernetes but kind of underneath your application. And we handle the networking between your your services and we try and
elevate that from you know kind of like the classic concerns of networking which are how do I establish a TCP connection from point A to point B to to a higher level of abstraction where you can say I'm service A and I want to talk to service B and can you just make that happen and can you make that so can you make it reliable and can
you make it secure and can you make it observable and I don't really want to have to care as a developer whether that's you know in the same cluster whether it's across clusters whether I'm talking to one endpoint or whether I'm talking to 10,000 end points whether some of them are slow some of them are fast you know I I just want you to make it work
and to make it secure and to make it observable so uh we were actually the fifth uh project ever uh to join the CNCF they created a whole new category for us called in uh called inception which later became renamed incubation um so I think we're probably the one of the few projects ever have been in the inception category uh and you know we've been around for
uh for 10 years and nine of those years have been in production, you know, at least in one place or another. If you're a big history buff, we actually was a first version of linkard. I'm not going to talk about that. That was on the JVM and written in Scala and came from like our Twitter history. Uh many many years ago we rewrote that entirely to the
modern implementation which is based on Go and on Rust and I'll uh you know I'll talk a little bit about that in a few slides. Lots of Slack channel members, lots of GitHub stars, lots of contributors, um the first service mesh to achieve graduated status, etc. Okay, but you know, so I talked about like what a service mesh is. Really the way that I view linkert's job
is we're trying to make Kubernetes into something into a platform into a true platform. And I think one of the beautiful things about Kubernetes is that it is a it's like a Lego base plate and you can build projects, you know, on top of it um and and and you build a platform and and link is a critical part of doing that. So that's our job is
to make Kubernetes a secure, reliable and observable platform. We have a design philosophy. I already lost one person from the audience. So hopefully they were just so astonished by my uh incredible slide they couldn't take it anymore. I thought we were going to lock these doors. Okay. What is Linkert's design philosophy? Uh you know in a word of simplicity. So we want Linkerty to just work. We
want to you know for you not to have a big config burden. Um we want config to kind of scale sublinearly to the extent that we can with like the power of the feature. Um, we want you we want to be small and quiet and to get out of your way and not to like consume lots of resources. We want to be understandable. This is a big
part I think of anyone who's in an operational role is something's going to break at 3 in the morning. And where linkerd sits, we're at the intersection of the network. We're at the intersection of the Kubernetes API and the application. And so linkerty in many ways is the canary in the co coal mine. Soon as you start seeing a problem that link is reporting, it's usually not
linkerty. It's usually like, oh, the application is doing something weird or the network is doing something weird or the Kubernetes API is overloaded. But you'll see those problems uh emerge first in linkerty. It's a canary in the com if the canary is also doing like the mining I guess. Um and then finally, you know, we want to make security easy or at least as as low barrier
as possible. So we do mutual TLS by default for example with zero config and uh you know you can add config and do fancy stuff on top of that but we want you to at least get that working without having you know any barriers in the way. So that's our kind of philosophical approach. Um, you know, and then kind of on the technology side, I mentioned this
a little bit uh when I talked about the history, but what makes linkerd unique is we are the only popular service mesh, I'll say those fighting words, that that does not use envoy. Only one only popular one that doesn't use envoy. Um, and there's a bunch of reasons why, but you know, one of the reasons is because with these uh uh instead of envoy, we build these
what we call microproxies in Rust. They're very small, very lightweight, they're very fast. They do a tiny fraction of what envoy does, tiny fraction of what kind of general purpose uh proxies do, but that approach allows us to uh to do things the way that I think is is the right way to do them and to to accomplish our our goals of of being simple and and
secure. So, Rust, of course, allows us to avoid a whole class of memory uh vulnerabilities. Um, we were one of the earliest adopters of Rust way back in the cloud native space way back in like 2018. The language had just hit 1.0. 0 and like the networking libraries were barely there and like idiots we decided to jump right in and like build a production grade proxy on
top of that. Uh we had to invest pretty heavily in in the Rust ecosystem to make that work. So if any of you Oh, show of hands. Is anyone writing Rust code either directly or through through your agent? Okay, that's pretty good. So if any of you have used libraries like Tokyo or Tower or H2. All right, great. Great. Yeah. So, a lot of the early work
in there um uh was uh thanks to your dear friends and team Linkard um funding those projects and and you know trying to get that ecosystem to the point where we could really rely on it. And you know, we kind of have a a little bit of an internal joke that really link is just a thin wrapper on top of uh tower and and and H2. Um
well, you know, the engineers don't say that, but I Okay. Uh ultra fast, of course, because Rust compiles to native code, state-of-the-art networking stack. Yep. talked about all those libraries and again from the operational aspect we want the proxies to be an implementation detail. So if you had to manage envoy you'll know that sometimes envoy becomes uh its own beast that you have to tune and and
and uh you know understand as much as possible we wanted to avoid that. Of course, there are situations and I'm sure you know the longunning linkerdy people here have seen them. There are situations where you do have to tune stuff like ultimately, you know, we can't be one sizefits-all. Different workloads have very very different characteristics, but we try and make it work as much as possible without
without having to to tune and without you having to think about the proxy as like a a thing that you have to worry about. Okay. Ah, here's my controversial slide. Buckle up. Linkert is a sustainable project. One of the things that I'm very proud of is we try and be really explicit about how linkerd gets funded. Who pays these maintainers? So we do not use volunteer maintainers
as a of course of course like any any like any uh uh project we're open to contributions and please send us PRs. We would love to have them and please volunteer. We won't pay you for that. But all the maintainers who work on linkerty have full-time jobs just maintaining linkerty. uh there's a company called buoyant that I run that uh you know that does that funds all
of that development and that is profitable by s by selling an enterprise version of linkard and I want to be really upfront with that like that is how this project works so we're not at the mercy of VC funding we're not at the mercy of you know uh kind of big players deciding that link that service mesh is cool and then you know being uninterested and and
funding dries up we are a self- sustaining uh project and for I think for a lot of projects the fact that there's a vendor behind the scenes is something you kind of hide away you know you're trying to talk about it I want to do the opposite I want to be really explicit about that that's how this project works we've been around this way for 10 years
my goal is to make that 100 years and the only way I can think of doing that is by having an economic engine that powers this project so controversial opinion I'm happy to uh fight it out with you Okay. Uh, and then just a quick note about kind of like how this thing works because it's a little quirky um, from kind of like the open source consumption
side, but we basically always roll forward in terms of the artifacts that we're producing. So, we produce edge what we call edge release artifacts. I should have updated the slide to have bigger numbers. So, edge 25.4.1 is the first edge release that we produced in April of 2025. And those edge releases have all of the latest code, all the latest bug fixes, security patches, everything is intended
to be production ready. But of course like sometimes these bugs only rear their ugly heads in production. So we will uh produce kind of like um these retrospective blog posts that are like okay these releases seem pretty good. These ones actually we found this issue so don't use that. So usually you know if you're doing this in production you'll wait you won't use the latest latest. You'll
wait for one of those reports or you wait to see kind of like how things are. Of course if you want to run the latest and give us bug you know uh bug reports that al also is great too. Uh, and then every once in a while we bundle these up into kind of major version, you know, announcements that represent a chunk of functionality and docs. Every
version has um uh a course, every one of those big versions has a git tag, has a corresponding edge release and like all that stuff. So we try and make this as as reproducible as Okay, so far so good. what has happened over the past year. What have we been up to? Um I'll just run through a couple features um that we've added and and things we've
changed. Uh so one is we have updated all of the the default TLS libraries to be post quantum ready. I think that I'm not a I'm not like a super crypto expert. I think that mostly refer to the key exchange algorithm rather than like the actual cipher suite. Um but we moved internally from ring which is a you know we've been using for a long long time
to AWS LC. So if anyone here works at AWS thank you for that uh project. And uh we've also added a bunch of observability about the actual cipher suite and and key exchange and stuff uh into the metrics. So if you are trying to prove to someone that you're actually doing like postquantum crypto uh you know you can at least get that from the And this all
happens by default. So you know as long as as long as you upgrade uh you know this should all magically happen. Um another big change and this one actually does have kind of an operator facing impact is the introduction of protocol declarations. So very for a very long time linkerd has detected the protocol automatically for you. So we have to tell we have to know when a
request comes in we have to know is this an HTTP request or is it not because if it's HTTP then we need to you know do a bunch of logic right we need to look whether you have routing setup that's specific to HTTP we have to know whether you have um you know security policies in place where like maybe this route is blocked in order to do
that detection or in order to do that that protocol determination we since the very beginning of time or beginning of the project what we've done is we have inspected the first couple bytes of the connection. Okay, well that's great unless a connection gets opened and you don't have any bytes. Okay, so then we add a timeout. So if you don't get any uh, you know, bytes on
that connection for the first 10 seconds, then we're like, okay, we'll just fall back to treating this as a TCP connection. So that way you don't have to have, you know, any configuration. Um, unless you want it. And we've given you ways of of specifying the protocol. Um, but it's not there's all sorts of like weird ways where that goes wrong. Obviously, you know, you can have
some some types of communication that don't send bytes from the client side, right? Like SMTP, for example, like waits for the server to respond before it sends anything. Um, okay. So, you know, you have to like not use protocol uh detection for that. Um, but where this really started to bite us at scale a couple years ago, um, was if the system is under load, then that
10-second timeout can fire even for regular HTTP traffic. And then you have this really like what I would call the opposite of a simple situation because now you have behavior that's non-deterministic. Now you have behavior that depends on like how much stress the system is under. So in order to ameliate that we added a um uh you know protocol de declarations which basically means in service records
in your service like CRDs there's already this app protocol field that has was introduced long ago. So now we just read that and that's a place where you can say hey I want you to always treat this as HTTP and don't do any protocol detection on it at all. And basically our recommendation is for systems that you expect to run at at high load where where you
know what the protocol is you should just start adding this. you know, it's an it's like an opt-in kind of feature, but it'll just avoid a whole class of like complicated situations. And if you see these protocol detection er Oh, has anyone Let's do a show of hands. Has anyone seen a protocol uh protocol detection timeout error in their linker locks? Okay, just one, two, two and
a half. No, just two people. Okay, good. Lucky. So, this is one way of addressing that. Okay, so that's a it's a very long description um of uh you know a very short feature. Uh but boy, we had to suffer a lot uh to get there. Okay, GitOps compatible multicluster linking. This was another big change that we made recently. Um so for a long time you know
one of the most powerful features of linkerty has been the ability to link uh across clusters and you know so you can have uh traffic from A going to B and B could be on another cluster and like or you could shift that traffic dynamically between clusters application doesn't have to care about that. Um we implemented this in kind of a world where people were uh you
know this is like the early days of Kubernetes. So you would have like if you had you'd start with one cluster and then if you had a second cluster it was like a big deal like maybe you acquired a company and now you had two. Nowadays you know the more modern approach well I shouldn't say that we're increasingly seeing approach where people have like 10,000 clusters well
you know hundreds of clusters okay and they want to manage them through GitOps. So we've updated the way that you do this multicluster linking so you don't actually have to run a command and you can just do this all through through GitOps and make it purely declarative. Okay. Okay. And then we did a bunch of work which is incredibly uh it's like the combination of incredibly boring
but also incredibly uh uh finicky to get right um around decoupling ourselves from the gateway API. Um so we've upgraded various you know support for various uh CRDs. Uh and we've gotten to the point where we're no longer for a long time we bundled the gateway API uh CRDs as part of the linkerd install and then we went to like a halfway point and now we're we
don't bundle them at all. So the onus is on you to manage that as a separate dependency. Otherwise it gets too complicated because some clusters have the the gateway API types in there by default because like the cloud provider put them some don't. Some like are you know are in this halfway state. So this is now something that you know we've offloaded you know and documented etc.
But like we have to offload this to you. You have to the onus is on you to like have these CRDs in your cluster. Um and we do depend on them right now for configuration of basically you know many many interesting features. Okay so that was a big change. Okay I'm doing pretty good for time. So that's where we are today. What is next? AI. AI is
next. Should we add AI to linkerty? Show of hands. Who says yes, we should add AI to No one really. Come on. One brave soul stand up. Stand up for AI. Stand up for the robots. No one. Okay. Who says no? We should not add AI to Linkerty. Show of hands. Okay. One, two, three, four, five, six, seven, eight. Okay. Nine. Yeah. Okay. Who uh who is
too scared to raise their hand because they think this question is a trap? Yeah. Okay. Yes. Yes. Very good. Okay. You have learned you learned. Who is uh just too scared to raise your hand at all? Show of hands. No one. Wow. Gosh. All right. No. Answer is no. Good job. Uh you know uh linking needs to be fast. It needs to be lightweight. It needs to
be predictable. And generative AI is is many things, but it's not any of those things. So, we're not going to add AI to linkerty. I mean, you know, it'd be fun to to think about what that what that would even look like. Okay, next question. Our platform problems are they now fundamentally different since we have uh ar you know this this massive wave of of generative AI?
Show of hands. Who says yes? Platform problems are now fundamentally different. Interesting. Who says no? Platform problems are still the same. Same Different day. Okay. Good. One, two, three, four, five, six. Okay. I think I probably agree with the nos here. Although I think you could argue, you could make an argument for the yes, you know, but I think where I landed on this is, you know,
ultimately the platform still needs to be reliable, still needs to be secure, it still needs to be observable. Generative AI doesn't change that. I mean, it changes some of the the the surface area of how how we get there, but the core the core requirements are still the same. Okay? And yet, you know, we if I could ignore this thing, that'd be great, you know, but you
know, we can't we can't link users are facing AI challenges and and we see this in a whole variety of ways. We've talked to a lot of people. I would love to talk to you anyone here in this audience who has uh you know feels like their life is different now as a platform owner, right? Because most of the AI usage that we see is very developerheavy
right now, right? the devs get their cool uh you know fancy ideides or they get their army of agents or whatever it is and they're cranking out code. What are the result you know what are the results for the platform owner right? Do do we have like you know all this sketchy vibecoded software hitting our production environment and like we can't trust it anymore. Uh are the
devs now so productive because of the miracle of AI that they have 10x of deploys and our CI/CD system can't keep up. Um are we serving inference from the cluster? I know that is a a big pain for people who are doing that. Um, do we have MCP like in our prod environment now? We have this new protocol that we have to deal with. Is anyone running
a gentic workloads that are doing their own thing in prod and and looking for ways to delete your database? Like these things are all happening to one degree or another and if they're not happening now, they're going to be happening another 12 12 months or another six months maybe. Um and and I think linkerd has to we have to be there. So as much as I would
like to ignore this and hope it all goes away um I don't think we can. Okay. So right now we're in experiment explore and prototype phase. One prototype that we demonstrated uh at KubeCon in Atlanta of last year uh geez just a couple months ago was adding MCP protocol parsing to the proxies and wiring up you know the metrics and um and security policy. So you could
do things like hey show me you know what are all the tools being called and the and the and the resources and the prompts or you know whatever you know MCP has like these kind of core primitives show me the latency of those things show me the you know the success rates and the error rates uh give me just give me a catalog of like all the
MCP usage that I'm seeing um and then give me the ability to define policy so I can say yes this tool call is allowed to this client this one is not allowed it's pretty It's interesting, might go somewhere. I don't know if if this is if this is useful to you or you think this will be useful for you, then please come talk to me. Um, another
thing that we're experimenting with pretty heavily is gateway. Is this an ingress gateway or an egress gateway or both like TBD? I asked AI, this is one AI image I use in here. I asked it to generate an image of the link maintainer and this is what I came up with. um you know does this handle ingress MCP traffic or is this handling egress you know is
this like a egress MCP gateway are we helping to serve inference you know are we doing fancy routing between models these are all things that we're actively actively exploring so if any of this stuff and I'm basically at the the final slide now um if any of this stuff is relevant to you or you think it will be relevant I would love to hear from you because
we are spending a lot of time trying to understand which parts of this are real which parts of this actually affect platform owners and which parts are are it's going to, you know, pass us by. Okay, so that's the end of my linky update. I have three ads for you. The first is, you know, for those especially for new folks, we have these uh, you know, online
courses that are free. So you can go to um, learn.boyant.io and, you know, or buoyant.iosma. I think they both go to the same same place. Um, we've got a bunch of online stuff and we also have recordings of a whole lot of webinars and and things like that. So, we we've got a ton of linkery content on there for you. We have a fancy new guide. If
any of you folks are on EKS or thinking about moving to EKS, we've got a cool new guide here. Um, you know, you could probably you don't have to scan this QR code. You probably just search for it. Um, and then tonight we have an escape room party sponsored by by and uh you know, there'll be a bunch of linkery related questions there. So, if you quickly
learn linkerty, you will come a ahead of the crowd to this uh to this party. Okay. Sorry. All right, let me let me go back there so you can scan that. Scan that or you just, you know, search for it. Okay, that's it. We have uh 3.9 minutes remaining. I am happy to take any questions. Two two minutes remaining. Okay, I'm Benedict, nothing. Okay. Yes. Yeah. So
the question was uh I mentioned when we started exploiting the r uh rust or using rust that the asynchronous like network ecosystem was was very uh early. Are we seeing a similar thing with MCP? Yes, we are. But the surface area of MCP is much much smaller than the surface area of like how do I do asynchronous network programming really quickly in in user space. So I
think we ended up I I don't know if we wrote our own MCP parser for this but like MCP is just it's you know from the protocol level it's JSON RPC and like plus some weird stuff plus you have to manage state for some reason that makes no sense. Um so yes but it wasn't it wasn't that hard. Yeah. Thank you. Good good question. All right. Any
other questions? Windows support. Uh, it is it is available, but it's behind a a payw wall. But yeah, let's talk. I knew you were going to ask that. I knew you were going to ask that. Yeah. So, I I haven't talked about it in this talk for that for that reason. Okay, time for one more. I saw you raise your hand. You moved you moved your hand
like that. No, no, not you. Okay. Save by the bell. Yes. One last question, please. >> Thank you. Okay. So the question was uh best practices for MCP are focused on the intent not on the API. How do we plan to integrate that into our API gateway? Is that right? >> Yeah. The agent. Yeah. >> Yes. Yeah. That's right. Uh actually I I would love to hear
your thoughts on this because I don't I don't know what the answer is. But if you have an opinion I'm I'm happy to hear it right now. Like I said, we're in we're in explore and prototype phase. So, a lot of open questions. That's a good one. Yeah. Okay. Well, everyone, uh again, if you have anything you know that you want linker to help you with any
burning AI problems, I would love to hear about that. Otherwise, thank you very much for your time here today, please use linkerdy if you're not using it already. Uh and I'll be around for a couple minutes for questions. Thank you all.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32