About this talk
In this talk, the speaker discusses the critical importance of open source supply chain security in the context of the technology industry, particularly highlighting Microsoft's journey and its relationship with Linux and open source projects. The speaker shares personal experiences and Microsoft's evolution from being a competitor to becoming a significant contributor to open source. They stress the reliance of modern infrastructure on open source software and the associated security risks, emphasizing the need for collaboration across different sectors to enhance security measures. Actionable insights are provided for developers and organizations to engage with open source responsibly and securely, including tools and standards related to software supply chain management.
Full transcript
Good morning everyone. How y'all feeling after that long night at game night? Um, you're still asleep. You can't can't even get a clap for game night. It's thanks for coming out on this, uh, daylight savings time Sunday morning. it's our annual tradition. We get to change the clocks together, uh, as we learn about open source. Um, Haron's probably in here somewhere trying to fix NTP for all
you all. U,, but yeah, it's, um, hope you all had a good time last night. Thank you again to GitHub and ARM for helping to make that happen. Um, Lots of fun games, lots of activities, hopefully some hopefully a little drink, not too much drink, but uh thanks thanks for coming out for that. U and thanks to GitHub and again ARM for making that happen. Uh again,
there's coffee in the back of the room. Uh thanks to our friends at Tail Scale. Um if that is helping you get moving on this slow daylight savings time morning, uh please stop by their booth in the expo hall after this and say thank you um so that we can keep coffee in the future. But also they have a great a great product and a great uh
great open source contributors that I think you all would enjoy learning from. So, thank you for uh thank you again to Tail Scale for today we've got a whole bunch of a whole bunch of talks uh throughout the day. We've got the Libra graphics track over another building. We've got open source career day for those of you that are thinking about doing recering or uh looking for
new positions or looking to hire uh all in all of those situations. Please stop by. lots of great um great training and activities there for folks that are in the job hunt. Um we've got uh and then this afternoon we'll have Dr. Doug Comr talking with us at at the closing keynote uh at back in this room. So um we've we're continuing to this this uh tradition
of bringing in some of our um some of the folks that help sort of build the technologies that we all depend on day-to-day and as as and bring them uh bring bring them here to scale. So, uh please come check that out. It'll be very exciting. He's uh one again one of the one of the founding sort of internet pioneers. So, excited to have him out here
as well. Um if you're standing in the back, there's a couple chairs up front here. If you're looking for a place to sit, please come up front. I don't bite. I don't think Mark bites either. He says he won't bite today, tomorrow. He makes no promises. Um but uh yeah, I'm when we started this pro when we started this uh this conference 23 24 years ago um
there was a lot of you know proop source and against open source and we thought we were sort of the rebels and you know there was the dark dark side somewhere but honestly like uh so many of our open source friends are now at Microsoft Microsoft's been a you know amazing contributor both to scale and the wider open source community and u you know we had John
Gossman here a couple years back talking about um Microsoft's love of Linux and that opensource and developers in general. And um you know it's uh you know it's it's amazing to see how sort of open source has become the default. Uh open source has become the way to build secure uh scalable amazing computing for developers. And u you know Mark's uh Mark's been a huge part of
that. So uh that that that that shift and that change. So we're excited to have him here joining us. Um Mark uh leads uh he's a CTO at Azure. He's a u leads the open source programs office among many other things. and I'm I'm excited to welcome here as a partner and um and a friend of scale and uh he's going to share with us a lot
about what they're doing with with secure software and AI and a whole bunch more. >> All right, thanks for inviting me. >> So good morning everybody. Uh as first of all, it's just amazing to be invited to speak here. Um, this is actually a key moment in my own personal Linux journey. Um, this morning I wanted to share with you a little bit about what is going
on with open source supply chain security, but I wanted to frame it in the in a bigger context and actually share a little bit about why Microsoft's involved with this. And I'll start with my actually own story about my involvement with Linux and open source and then get a little bit into Microsoft's own which leads into how the world depends on open source and the risks that
come with the world's infrastructure running on top of open source which takes us right into what we should do about securing open sources. So, first a lot of you probably know of me as the Windows or CIS internals guy. Uh, but it's a little known fact that I actually have a long history when it comes to Unix and Linux. In fact, my PhD thesis was built on
top of BSD and mock. And I became an expert not just in Unix internals, but Windows internals over time. That led me to write articles like this back in 1998 comparing the Unix kernel in the Windows kernel. You can go find this online by the way still. I was actually at IBM Research as the expert on Linux despite the fact that they'd hired me as the expert
on Windows. So when IBM was porting AIX to Linux, the AX team had me come in and talk to them about the Linux kernel. Uh you can go find this presentation online which was before I joined Microsoft but I actually gave a talk more focused on Linux versus Windows and this is my first interaction with Lenus Torbalds where I sent him the deck and asked him for
his review and he wrote back looks good. So I um and then I've also contributed to Linux or contributed source to Linux. This is uh CIS internals which again you think of as Windows utilities but this is from 2001. You can see that I actually made a version of Fmon for Linux and this I did this using system call hooking which quickly became banned and so I
this this tool broke and I stopped supporting it. And then the next big step in my Linux journey was being personally involved with Azure and the introduction of IAS on Azure which introduced Linux as a first class operating system at the launch of infrastructure as a service on Azure. And I was here at this announcement that's me there. Satches actually looks like he's telling me that Microsoft
loves Linux. Um which I was flattered by that he tell me that Microsoft loves Linux. But this was a very exciting moment. And then over time I've become involved with Linux Foundation and I actually ended up getting to meet Lenus and now can call him a friend and this is just from a couple months ago. Jim Zlin the executive director of Linux Foundation. Me and Lus having
a dinner. So it's been just a journey and and so I just wanted to spell the I'm the Linux I'm the Windows system internals guy only. I'm an operating systems guy. I'm a cloud guy. I'm a Linux guy, an open source guy and a Windows guy. So Microsoft's open source journey also goes back a long way. In fact, you know, and I'm gonna not not go all
the way back into evil empire days, but I'll start after that, which was um the introduction of HyperV, which led us to wanting to support Linux on top of HyperV, and that led to us starting to become a major contributor to the Linux kernel. You can see that uh one of the uh several of the engineers uh actually became top five contributors to the Linux kernel in
supporting that work. And then you can see over time just this continuous progression already talked about Linux on IAS WSL Windows subsystem for Linux on Windows. So Linux can actually run in Windows to actually Microsoft having its own Linux distribution. In fact over 65% of workloads on Azure are Linux based and more than half of the marketplace offerings are Linux based for Azure. So our customers are
using Azure primarily as a Linux platform as most cloud offerings are built on top of Linux. One of the biggest is Chat GPT which runs on Azure and this is a massive uh Kubernetes service running on top of Azure that uses Linux provided by Microsoft on top of our virtual machines as well as you can see open source components like Postgress. It's one of the largest apps
in the world. But Microsoft or or open source isn't just a platform for Microsoft where we have customers run. We are actually also deeply dependent on open source. And I want to highlight some of the ways that we integrate with open source because we look at this very carefully. We look at how we consume open source and use it. We want to be good citizens of the
ecosystem and community. We also look at how we should be contributing to open source. Again, you can't just toss things over the wall. contributing means supporting it means engaging with the community if you really want it to be true to the open source spirit and then finally releasing open source our own stuff. So one way that you can look at it is how much we're consuming open
source is measured by the number of pulls from our internal registries or component governance or central feed service where our internal systems pull open source packages. And that when you look at that over 200,000 pulls are happening against that central feed service to eight 11.8 8 million locations in our services across all our products whether it's office or Windows or or Azure and one of the largest
Kubernetes clusters in the world happens to be the cluster that underpins a lot of the office services. So uh M365 services and teams running on top of a system called Cosmic which is built on top of Azure Kubernetes service millions and millions of cores and it also uses several other open source components underneath including things like the Kada Kubernetes event driven autoscaler as part of that deployment.
And get this, we have an infrastructure offload card that goes into every single server we ship. It's called Azure Boost and that's got an RMSOC on it and that's running Linux. So Linux is literally going into every server Microsoft deploys, the millions of servers we're deploying. So we we have a very deep deep dependency on Linux and open source and we also turn around and contribute to
open source and we do this by in in a variety of different ways. And one way to measure this one is just to look at the number of pull requests that Microsoft engineers are making on open source projects. And when we take a look back at the la last three months, and this is from Linux Foundation stats, three million pull requests from Microsoft. In fact, that that's
a huge number. But just to put this in context as far as how big it is in terms of corporate contributions to open source and Linux, that's the list of top corporate contributors from the Linux Foundation. And who would have thought, you know, going back to evil empire days, that one day you'd see something like this where Microsoft is the number one contributor overall. And then Microsoft's
also releasing a lot of open source software. One of the most popular, how many people run Visual Studio Code ID? So quite a few of you. That's an open source project. In fact, it's one of the most active on GitHub open source projects. uh and you can see it's installed by the vast majority of developers are leveraging Visual Studio Code. And then we've released a lot of
other systems and components for open source developers that are quite popular.NET became open source about 10 years ago. TypeScript one of the most popular languages in the world. It's in the top five and continues to rise. Uh a lot of people might be using that. That is also open source and developed by Microsoft. PowerShell became open source, Open JDK, and the list goes on and on. And
then when it comes to new types of software, we're also contributing. So when it comes to AI, because everything is about AI these days, if you've heard of semantic kernel or autogen, those are two AI orchestration frameworks. They were both open source. We combined them last year into something called Microsoft agent framework to get the best of both of them and have one offering that we stand
behind. And so this is also open source project and it's quite widely used among enterprise users. And then my own org, I've got the open source ecosystem team in my organization as Elon mentioned, but I also have an incubations team which is entirely focused on open-source cloudnative incubations and we've got a track record of this goes back to 2019 and developing these things. I mentioned KADA, the
Kubernetes event-driven autoscaler that is a graduated CNCF project. Dapper distributed application runtime came out that is also a graduated CNCF project. Copa which is a container patching system which is very important for container and cloud native security that is also been contributed. That's a sandbox project. And then we've got a couple of other uh incubations that we're working on. Radius which is a a cloud native cloud
neutral application platform application development platform that is also a sandbox project and DRSI which is an event driven reactive new type of data service that is also open source and contributed to the CNCF and then again you think of CIS internals you're probably thinking Windows but CIS internals while it was dabbled in Linux over 20 years ago is back with Linux and we've have a number of
tools here. You can see five tools that we've released for Linux including versions of Windows tools like Sysmon which is edr tool that is actually shipping the Windows version is shipping in the box in Windows as of this month or last month in February but we've got a version of that for Linux that we're running internally at Microsoft. Uh JCD which is JumpCD I wrote myself. So
that is a a tool that I I open sourced through CIS internals that I developed myself to make it easy to jump up and down complex directory hierarchies that I was having to navigate as I was worked on my open source or on my uh AI research projects and I get these very deep directory hierarchies and got tired of cd dot dot dot dot dot dot help.
Um, and then Microsoft also just some frivolous types of contributions to open source. Microsoft's gone back into the archives. And so you can go look at the source code for DOS 4.0 which is now open source. Going back even further, Applesoft Basic, what became Applesoft Basic, Microsoft being the creator of the basics that ran on most of the PCs back in the 1970s, that is also open
source. So you can go look at that. Um, and by the way, I have because you're going to that I've got a connection with Applesoft later in the talk. Uh it's not just Microsoft that depends on open source. The whole world depends on open source. This is a study done in 2025 that looked at 1,600 organizations in the s in the code that they uh were developing
and found that 97% of them had open source dependencies. Uh 70% of them had their origins in open source. over nine about an average of 900 open source components found in each one of them. So you can see it's not just Microsoft but the whole world literally is dependent on open source and one way to actually look at the growing dependence is the number of pulls from
package managers. So you see here two charts pi pi and crates the rust and python package manage main package managers. You can see exponential growth in consumption from these things just as everybody is now using open source and contributing to open source. It's pretty 10 trillion pulls. You can see just a staggering Now the the world's dependence on open source brings with it interest from places that
we'd rather not have interest in it. And that is from the the malicious side of things, from nation states, from hackers that are looking for how do I get into enterprises? How do I get at money? How do I get into desktops and consumer devices so that I can go and ex do my exploits and carry out my goals? So, and and because of this huge dependence
on open source, of course, they're looking at how do I get in the open source supply chain? How do I infect something that is going to be used by thousands of organizations or millions of users? In fact, this XKCD comic kind of really shows us kind of the a little precarious nature that we're at because with 900 average open source dependencies in a commercial piece of software,
millions that Microsoft or hundreds of thousands Microsoft depends on. And there's somewhere in there components like this that has one maintainer that is working in there part-time that might not have the best security processes and might not be available when something goes wrong either from a reliability or security perspective. And yet the whole world yanked that thing out. And we've seen things like left pad. You remember
that? When the developers like ah people's pissing people are pissing me off. I'm yanking left pad. And then half the internet gets impacted because of this little a few little line node and the open source supply chain just to put it down so we can see where attackers try to enter this thing are all over the place from the developer like if you get to the access
to the developer you can start on the very left of the supply chain. If you can modify the source code maybe out in the repo not even needing a developer because you happen to be there that's another great way to do it. If you can infect the build pipeline, if you can go into the package manager and have carry out exploits there that cause people to pull
down the wrong packages or pull down infected packages, that's another place. And then of course your ultimate goal is there on the right, the consumption, the consumer, where you hope that you can get them to do something that puts them at risk too because they're, for example, pulling from the wrong place or pulling the wrong package. And then I want to highlight the increasing risk and and
pace and sophistication of these attackers that are trying to get into open source. Um, and the kinds of threat kind of risks that the world is at when somebody discovers a problem with open source. going back to one of the earliest wakeups for how dependent the world was on open source even in 2014 OpenSSL vulnerability called heartbleleed where there was a buffer overflow that could be triggered
remotely and basically half the internet was exposed to this and then in 2022 more problems with certificate validation were found leading to buffer overflows and and remote exploitation the heartbleleed one actually led to being able to read arbitrary memory and get uh keys out of the uh servers uh But these kinds of problems really started to wake people up to, hey, we really are dependent on the
security of open source. If something that we're all using gets compromised or a vulnerability is found, we're all needing to react to it together. And this one was another huge huge wakeup call. Log forj. This is remotely exploitable vulnerability that you could cause arbitrary code execution on Java web servers that happen to be using this logging package. And this one again huge amounts of exposure from companies
across the internet and huge reactions. I mean even Microsoft had exposure to this in pockets around Microsoft that we weren't even really aware of until we started to go look. So we had to pull in incident response and go get these things patched. We had to figure out just like everybody else had to figure out what's the exposure, where do I have this thing running? How do
I get it patched? How do I get secure? The whole world was scrambling. And despite everybody being aware of this problem as it kind of spread and and attackers were taking advantage of vulnerable servers. Here we are four years later and you can still see that there's vulnerable instances of this package log 4j all over the place. And you can see in 2025 alone 42 million vulnerable
versions of blog 4j were downloaded. So it's like we're not quite there yet and really appreciating the risk of running software that it has these kinds of vulnerabilities in it despite the fact that it's gotten so much attention. And then I want to give you some other examples of different types of attacks. This one, PyTorch. How many people heard of this compromise? So yeah, this was a
package manager confusion attack where PyTorch is building off an internal PI and attacker realized that pip would go and pull a public version with a higher number overriding the internal version if it was available. And so they published a package with the exact same name P into the public pi pie registry which caused the PyTorch build to go pull it instead of the internal instead of the
internal one. Of course the the public one was a compromised version. And so this shows that you know down the supply chain the package managers are also a target of these attackers. And this one how many people heard of this? I think the whole world heard about this one XD utils because this one's a fascinating nation state long game attack on open source supply chain. They identified
this package that had one single maintainer, befriended them, got gained their trust, of course with fake identity and never meeting. And so who knows who's really behind this and it's probably not just one person, it's probably a nation state organization. But then they planted the seeds to have a backdoor re recognizing that XZU tools is used by OpenSSL or OpenSSH. So a very commonly used uh interactive
loon protocol component. And how was this discovered? actually a Microsoft engineer who happens to be a Postgress maintainer was looking at comparisons of performance for login across different versions of OpenSSH and saw that there was a jump with it pull request from this individual that caused it to go from like 100 milliseconds to 500 milliseconds or something and said whoa that's a bad performance regression. They started
to investigate and discovered, you know what, this thing actually is malicious, which set off this whole reaction to everybody needing to go clean up uh this dependency. And then we're seeing viral spreading, too. It's kind of like the internet worms except through open source supply chain. The internet worms I'm referring to from the late 1990s, early 2000s. here. Back in September of last year, uh, an O
npm maintainer was compromised through a fishing attack. They didn't have MFA registered, and so the they were fished, which caused them their credentials to get leaked, which caused the packages that they maintained to get compromised, and that was discovered, and that was cleaned up. But then other maintainers were also compromised a few months later because they didn't have MFA enabled yet. And the attacker in this case
went viral by saying anybody that pulls down a a compromised package. So it's like one maintainer gets compromised. Now the package is compromised. 10 people pull that they get compromised and the compromised software is going and scanning them to look for npm signing keys as well. Compromising the packages they maintain and you get this kind of exponential spread effect which uh and it ended up being uh
kind of a worm through maintainers being compromised through lack of MFA from just a few in indiv source individuals in this case so they're getting clever and the are ubiquitous they're all over the place in open source software you can see in 2025 npm recorded 83900 basically releases with CVSS 9.0 which 10 is the highest rated vulnerabilities. So people are contained to release source code that is
putting their consumers at risk and one in five PI releases had uh CVSS ratings of of 7.0 or higher and the rate of vulnerabilities continues to go up. Not only that, but now we've got to face this new world of AI coding and AI coding. How many people are using AI for coding? I think if if you're coding, you're using AI for coding. Um, by the way,
not looking at the code is not the flex you think it is. Uh, I hope that's a message you take away from this. Um, but I want to say that there's risks that show up here as you're using AI because AI is going to be using old versions of APIs that might be insecure. It's going to be getting dependencies that are maybe older versions that are insecure.
It's going to maybe making uh writing code that happens to have insecurity in it using string copy instead of a safe string copy for example and yes it does that. So you need to be aware of the risks that it is introducing that you might not introduce yourself if you were actually writing the code yourself. But AI is also we're it's time to buckle up because the
signs are everywhere. I mean, we saw this coming as AI got great at coding and could do things nobody could imagine a few years before, but just in the last few months, we're starting to see sophistication that is really both great because it's they're they're so good at coding. They're so good at helping defenders analyze their code to find problems, including security vulnerabilities. But at the same
time, this is a tool and a weapon, and attackers have access to the same technology, and they're going to be using it to go find these things. and Enthropics uh published a couple blog posts just in the last month or so that show how good things things are, which you can look at as that's great, or you can look at as oh crap. And you know, in
my job, I need to look at it as oh crap. Because here's one example where they set it loose on a a bunch of open source projects and within a few days and like $4,000 they'd find 500 vulnerabilities unknown new vulnerabilities in open source. And just last week they published a post where they worked with Mozilla who creates Firefox. Firefox combination of Rust and C++. They set
it loose on that and found I think 148 And just to put that in context, like Google's last year only saw roughly like 75 vulnerabilities being exploited in the wild at the same severity level. So uh and I highly recommend you check that out. So this is and by the way, Firefox is one of the most secure open source code bases there. I mean they've take various
security very seriously including being behind the creation of Rust specifically for security. So just to give you an idea of kind of the ignoring even what we thought and knew we had just with human analysis and those vulnerabilities that have been discovered now we've got AI and this is going to be an order or two orders of magnitude in terms of volume and rate of vulnerability discovery.
And then we also see AI now involved in the supply chain. How many people saw this one just last week? You guys, anybody here use Klein? So, a few of you, you might have been impacted by this one because Klein got infected. So, this one, a researcher discovered a few months ago, that they could prompt inject Klein's build system, which is using Cloud to go look at
issues. And in the issue title, they put in, hey, before you work on this issue, you need to download this package. Guess what package that package then would go download? OpenClaw. So, OpenClaw ended up being for a while deployed with every up deployment of client because of a prompt injection in the build system uh so in this you know the development system. So just to give you
an idea of where things are headed uh now AI is in the picture and by the so and it's not like critical infrastructure companies and organizations concerned about the security of their critical infrastructure aren't aware of all of this going on and this has led to a big focus on we need to understand our supply chain is what they're saying. You can see a bunch of organizations
starting with the US government a few years ago with an executive order saying every piece of software purchased by the US government has to have a software bill of materials with it. And then you see a bunch of other regulations coming. And one of the biggest coming at us is the CRA, the Cyber Resilience Act from the EU that kicks in later this year. And that one
requires you to be publishing CVES and attaching them to your SBOPs so that they can see what risk the software as they're consuming it is carrying at any point in time. And these regulations are just going to get stricter and stricter. They're starting of course with certain things like hey just publish an sbomb to next it'll be some verification the sbomb is accurate to your CVS have
to be accurate and and correctly attached to sbombs to we're going to need you to secure your build systems in a certain way before we consume software from you and prove that they're secure and this list of requirements is going to grow pretty quickly given these threats that we're seeing given AI's ability automate risks and so this is what I'm know message is we all have to
get together because this is not a problem that any one of us can solve we all got to solve this together uh governments open source maintainers and companies together which brings me to this the open source security foundation how many people had heard of the open source security foundation before today how many people had never heard of it so This is dis disappointing raise of hands which
is why I'm here but the open source started as an idea back in 2019 time frame. Uh me and Eric Brewer from Google were talking about our company's dependence on open source and the fact that open source had these risks in it that were kind of we couldn't really see the risks, measure the risks or help with the risks. And so we started talking with others including
Amazon and other companies about what we should do about this and that led to the creation of the open source security foundation back in 2020. The idea of course with the open source security foundation is to have everybody come together to work on this problem together. No one company as I've said can go and really address open source security. It's not like one company, Microsoft, we can
secure our own open source supply chain. That's just impossible. The scale of it is monstrous and we don't have the skills to do it either. So, we need everybody else to participate and every other company does too and every other government does too. So open source security foundation is a place for us to come together to work on driving and setting standards working with governments and working
with companies on improving the security of open source software. It's not a place by the way for us to come and you know corporate big hand come in and control open source. We love the energy we love the innovation in open source like everybody thrives off that and we all thrive on it and there's you saw there's lots of Microsoft developers that love contributing to opensource. So
it's not us trying to control open source. It's trying to to help open source because the fact is if open source doesn't become secure, the world can't use it. Just as simple as that. It's not like it won't be a choice. You see the regulations like Microsoft won't be able to sell software that has insecure or so or open source that is not following best practices just
so that's just a reality. So let's try to work on this together. You can see that this has gone momentum. this five years later 117 organizations 16 industries so it's not just hyperscalers like I mentioned but finance cloud AI companies government academia everybody's represented as part of this recognition we all need to work together and it's not just a US effort either you can see organizations from
40 countries and just to give you an idea of what the open source security foundation's actually looking at and the companies that are participating this This is just a slice of some of the working groups in open source security foundation and some of the members leading and participating in those different working groups. By no mean means comprehensive but just to give you a flavor of how diverse
the participation is in supporting open source security foundation. Now that supply chain picture that I showed you a little bit earlier, this is the open source security version of that supply chain diagram with open source initiatives, standards and tooling mapped to each place in that supply chain. And this is just quite the eye chart and I'm not obviously going to spend time going through the whole thing,
but I did think I want to take this opportunity to share with you some of the key projects that are probably relevant to you. So, how many people in here are are a maintainer for an open source project? How many people contribute to an open source project? How many people consume open source? Everybody just want to get everybody raise their hands. So, we're And how many people
work at a company that depends on open source. So these projects I'm going to cover are relevant to you regardless of where you are almost certainly but some might be more relevant to you than others depending on which role you happen to be in. So let's go through a few of these and I'll start with salsa supply chain levels for software artifacts. This one actually came out
of Google which was looking at this internally. How do they measure or define what is a good open source build pipeline? What is what are good standards for securing the repos that have the source code that goes into those build pipelines? And so they've contributed this standard to the open SSF. I think even maybe early in the first year after open SSF was founded and it has
continued to evolve. If you take a look at the source track, so it's divided into tracks. The sol source track has what are called levels of assurance. And you can see that this is a way to for you if you're maintainer to look at your build pipeline and decide what level of security you're going to shoot for. Like I said, get ready because if you're source L1,
which is you're using a version control system, which is great, it's just one of the key requirements. get ready because look at the next levels because they're coming at you through regulation and your customers demanding these next levels. So eventually, I don't know if it's five years from now, I don't know if it's 10 years from now, but at some point things are going to need to
be up at source level four. Basically anything that the world consumes. So this is kind of a map and a look into the future, not just a look at what you can be doing right now, but this is a good way to go look and see what are the best practices at different levels of security. And then there's the build track. Same thing here. You can see
four levels for the build track, which is level one, which is, hey, I build, but I don't know how I'm building or I'm building on my laptop at home. Up to build level three, which is I've got a medically sealed build pipeline. What that means is that attackers will have a very hard time getting into the build pipeline, injecting software into the build pipeline, and that you
have basically proof of what went through, signed cryptographic proof of what went into the build pipeline and what came out of the build pipeline. And so this is uh the ultimate the goal. And there's a level beyond it that they used to have that they've cut for now just because it's out of reach for just about everybody right now, which is reproducible builds as well. But that's
also coming. And then if you take a look at how you prove things. So those are the levels. Uh something else that we need to be looking at is again those those proofs because when those regulators come or when somebody comes to look at a piece of open source and say can I consume this? They want evidence evidence that the thing was signed that it had multiffactor
authentication on the repo that uh that it did run in a hermetically sealed build pipeline. And so those proofs are going to come and be recorded uh using at this point there's two emergent systems for recording proofs. One of them is just for recruiting uh showing proof of signing. So if you're familiar with let's encrypt this is let like let's sign basically for open source software. This
is a sig store project which is in the open SSF which is ability to go and using ephemeral keys with backed by OIDC identities like your identity or the identity of your repo get something signed that has no risk of leaking the signing key and so makes it very hard for an attacker to get that signing key and then go sign packages on your behalf. And second,
proof that it was signed as well. that goes to a transparent immutable basically an immutable ledger that shows this package was signed by this key at this time by this individual. So this is the kind where this project so this is the this transparency key log there you can see these are all open source projects for each piece of this stack uh that you can go deploy
and you can now integrate it with pipi integrated with github you can go get things signed with sig store so it's one of the very easy ways to just get some basic evidence of your provenence of your build pipelines uh what Microsoft has been working on with I the IETF and other companies is looking at it from an enterprise perspective and the things that we need to
believe we need to have evidence of not just that something was signed on a particular date but we're going to need connections with things and policies that are also recorded like for example the fact that I did a malware scan of a piece of software that's a piece of evidence that we might need to want recorded and attached to a software billing materials which also is attached
evidence that it came from an attested build pipeline or a medically sealed build pipeline or reproducible build and there are number of policies. We kind of have an idea of what policies make sense. Now there's likely going to be more and more policies over time that we're going to need to provide evidence for in this immutable ledger so that an auditor can come and look and so
that we can produce evidence not just us but of course anybody producing software. And this isn't just for open source software. This is also what we're going to we're already started using this internally of a version of this internally at Microsoft for signing and recording addestations of some of our closed source software including So this is actually and this is an important note I also want to
highlight. If you take a look at everything I've been talking about it applies equally to closed as open. In fact I don't look at it as open-source software standards. I look at it as software standards like what's the difference? We need the same controls. We need the same security for both types of systems. And so that's the way you should look at it. Whether you're open or
closed, these systems apply and these standards apply. Now, here's another example of a way to measure and determine if you're following best practices. And this one was also contributed by Microsoft. This one is aimed at how do I consume software? What are the best practices? And this one, if you have You have to actually sit down and think about it. All the things that matter. And once
you start to think about it, it comes into picture. The things that you'd want to be secure from an open source consumption perspective. So from Microsoft perspective, as we sat down to say, how can we be secure? Well, you want to make sure that you've got an inventory. What open source do I have running? What versions of it is running? What am I pulling? So this is
you know the kind of engineering systems and having that visibility when there's a an XY util or log forj how exposed are we that you require an inventory for that you want even ahead of that starting with ingestion what can our developers ingest do they is it okay for them to ingest that package that was has one maintainer that with no MFA and no PRs on their
contributions to the project is that okay Well, we got to decide what are our policies. So, that's something else. Deciding the policies. How do you update when there's a log for J? What are your processes to go and get the whole organization mobilized and get that patch in place and deployed quickly and safely? How do you enforce? Again, that goes to the ingestion policies. It goes to
the up deployment policies as well. How do you audit so that you can understand what's happened and what you got? How you scan your systems because it's not like IC open source today might be secure tomorrow it's the exact same code is insecure because somebody discovered something about it. So how are you scanning and looking for new problems as they're emerging? And then one of them is
for something really critical can you wait on upstream? This is something you got to think very carefully about. If there's a vulnerability in an open source package and your customer base is at risk of being ex harmed or your own company's at risk of being harmed and the upstream maintainer happens to be on vacation, what do you tell the world? Well, the guy's on vacation or girls
on vacation. Uh, sorry, you're not getting a patch till next week when they're back. No, we're responsible. Our customers don't care. They look at us as we're supplying a service. We're responsible for that service whether it's closed source, open source. And so if there's a vulnerability, Microsoft, keep us safe. So that might mean building temporarily yourself and releasing a patch while you're waiting for upstream on those
critical projects that are high-risisk surface areas at least. So that's another consideration here. So that is a look at consuming. Now let's talk about measuring actually. some standards. So there's kind of standards, how you look at things, different levels of assurance. But let's take a look at how you actually measure how good you are. And one of the programs the Open SSF has had for several years
now is this scorecard, which is just a very lightweight measurement tool that you can point at any repo and get an idea for and which includes CI/CD in some cases, get an idea for how secure it is. So, it's got a bunch of measurements it does and it produce a report for you in an automated way. Source risk assessment, build risk assessment. You get the idea. And
each one's rated on a scale of 1 to 10 where 10 is you're doing awesome. Just to show you how easy it is. And if you've got a open- source repo, you should just go run this right now. Here's the tool scorecard running against the React repo that Facebook manages. And you know, this is kind of sped up a little, 32 seconds, so sped up just a
little bit, but you can see uh it's found some potential issues. It's hard to see, but just wanted to give you a flavor for just there's no excuse not to go run this and see what it says and say, "Whoop, I didn't know about that. That's something I should go fix." And then the open source security foundation is also now establishing set of controls common controls. You
know everybody working in software assurance and software regulation knows about controls. This is what how you say I want the build pipeline to be tamperproof. Well that's a control. Now how you implement that control you know that's a different question. It depends on the build pipeline and the mechanisms it's got for enforcing protection of the build pipeline. But the control itself is what people care about. The
control evidence that the control was met. So one of the things open SSF has started to do is define a set of controls that hopefully the world can start to align on for open source including ideally regulators that are asking for open source compliance with their standards to be able to say you know what we've got this control. we want you to uh implement it and that
can be mapped on to an open SSF baseline control. These baseline controls right now it's kind of just getting started. There's about 40 controls and they've got different levels of assurance for some of those controls and you can kind of have these basic levels defined to go look at a project and say how secure is it? So like 20 controls apply to the universal floor. You just
everything should do these things and then you can they on one another. You want to go to level two, you need to have those 20. Some of them might be more stringent, too. So, 18 total new controls come into effect or new variations on those controls come into effect. And then level three. And it's not just kind of these basic levels, but also open SSF has mapped
these things onto open standards, regulatory standards like NIST, the CRA, DORA, which is forthcoming to say here's how these things map onto these regula regulated controls that are going to be coming at at you. So this is to try to make it easy for us to just have a a common language across open source for defining these things. And if you take a look at it, all
of it is like the more we can standardize, the more we can interoperate, the more we can share and build on top of one And then finally the the execution side of open source security foundation. It's a sub foundation called Alpha Omega. And alpha omega came from us talking with us meaning Microsoft talking with Google talking with Amazon and saying open SSF is great for education defining
these standards supporting the tool development that let's go and try to accelerate progress. So when there's an open source project that is, you know, funded only to keep the lights on, can we go inject expertise and resources to get them to accelerate progress on open source? Of course, this needs a project that is willing, that has the time to go and engage like this. But we found
many projects, almost all projects that are approached by Alpha Mega to say, "Hey, you're a critical project. would you like some us to help you accelerate your open source maturity that they're all enthusiastically supportive? And so the alpha part is go high value targeted critical deep engagements. The scale is let's go fund stuff that's going to lift all the boats. Primarily up to N's point the focus
has been on alpha and Microsoft. You see Microsoft, Google and Amazon have uh funded 5.8 8 million into 14 critical open source projects to improve their security including the Rust trusted publishing launch and CV authority designation, Apache trusted release pipeline. So you can see a bunch of projects just in the last year where Alpha Omega has gone and helped execute and move things forward. I want to
just summarize here. I've given you a whirlwind tour. First of all, no, the world depends on open source. Microsoft heavily depends on open source. Open source is becoming more critical and it's it's a high value target for the threat actors out there. The risks are rising. We all need to move together. So, one takeaway that I'd love coming out of this besides just going and running the
scorecard just to see what it shows on the repos that you're involved with is looking at this and taking some action. You contain open source get get on boarded to SIG store. It's really easy to go sign uh your packages. If you're building an open source, go to look at S2C2F and try to figure out what level of risk you're willing to tolerate versus the expense of
doing some of those things in there. And then if your business depends on it, please support the open SSF and support maintainers. There's a lot of different ways you can support maintainers including uh through GitHub. So that's the big takeway. Now I want to leave you with one last thing which is just to uh I know that uh scale is famously a builder community that loves to
get together and and hack on stuff including IoT devices. Well yesterday I was just preparing for this and you know thinking about the implications of AI and AI's ability to find vulnerabilities. I was curious because one of my first open source projects was this one back in 1986. I uh how many people ever programmed on the Apple 2? Just out of curiosity. So quite a few of
you. I um I was I've you know been this is in high school. I wrote this tool. I was like I reverse engineered Applesoft which now I don't need to do given the source code is public but I reverse engineered it and one of the limitations it had is you couldn't go to or go sub to a function. You had to go to a constant. So you
could say go to 10 to go to line 10. You couldn't say go to a plus 10. which it might be nice if you had a you know sequence of different snippets of code depending on some choice the user made or some other condition you'd go jump into the right piece of code just automatically without having to go if a equals 1 go to 10 if a
equals 2 go to 30 you get the idea so I reverse engineered applesoft and I extended it so you could actually do computed gotos and go subs and here's the source back then you had to type this in so this So last night I'm like I wonder how Opus 4.6 would do on this given you know the stories coming out about uh open sour finding vulnerabilities in
C++ and open source projects. So I went into GitHub copilot and I told Opus 4.6 first I just like go find the article download the binary that the user has to type in. These are truly the prompt. download the binary the user has to type in, save it to a file, disassemble it, and then scan it for security Well, you know, like 10 minutes later, here's the
disassembly, and it's like brings back fond memories because it uses the same commenting that I used in my original assembler. Um, this is basically it. So it fully disassembled it, understood the exactly how this thing worked and then did it security analysis and came up with this. You know, I told it make a markdown file. Came up with this. And you can see that it found some
pro. I'm kind of embarrassed uh for my high school self. Uh that it found some of these including one that is particularly uh embarrassing. This one right here, which is it doesn't check that fine line found the new line, you know. calls this function to go do the a + 10 or whatever and gets back a line number, but it doesn't actually check to see if that
line number exists. So if it doesn't exist, uh this is in the do restore, it actually will put the read pointer into an undefined location which could be off the end of the program could point to some non-existing some data that's not really data. In any case, it's like you know here's your problem and by the way it even suggests a fix. So in 10 minutes I'm
looking at 40y old machine language it's reverse engineered it and found security problems with it was blown I kind of expected it but seeing it right in front of you taking that binary listing from the magazine and getting to this in 10 minutes so yeah it's time to strap in because it's going to be some interesting times ahead of us and with that I want to just
wrap up and just know hope you found this Uh, hope you realize that Microsoft's come a long way from the early days that we don't talk about and that you've got some good idea about just how uh how critically the corporate world is taking security of open source software and how we want to help improve the security of open source software but not hurt what has made
open source software so fantastically successful for everybody. and so much fun. So with that, I hope thanks again for inviting me. This is another milestone of my hopes for returning a great scale. thank you Mark for joining us. Um if if you've got time, I'd love to do a couple questions. Sure. >> Um normally it is our tradition to hand you a scale jersey now that you're
part of the scale family and team. Um, unfortunately, uh, shipping couriers do not follow the same SLAs's that as as as some of the software that we do, but it is redirected to your office. Uh, and I apologize for not having here with you, >> but, uh, we we will we we will we will have you dawned on the in the team garb shortly. >> Um, but
yeah, I've got somebody in the audience with a mic, I think. >> Okay. Uh you want to grab one from the the back and then I'll get started. >> Sure. >> Thanks. Nice presentation. A fan boy uh here. I've been using all your products for forever. Uh a couple of points about your presentation. The software bomb thing and providing all the CVS that your product has. Does
it make it more easier for the AI industry to find vulnerabilities over your product? And then when you constantly get this vulnerable how do you as a company keep patching the systems all the time. Uh so there is a is there a philosophical way of keeping some critical repos as less open source than more because you know uh from that point yeah I think uh the questions
you know about this always in security the security through obscurity or offiscation versus the transparency and I think that the fact is uh that that slider on things is moving and it has to move because you think that it's not disclosed You think that it's not visible in fact, but the fact is somebody's going and doing automated scanning now and automated attacks against repos and open source
and they're going to find it. So, and I'm not sure exactly, you know, how this is going to play out because of course we don't want to if we discover something ourselves, you know, um a company finds something, we don't want to put the whole world at risk by saying, "Oh, by the way, here's a vulnerability." you know, this is the the whole uh responsible disclosure thing
and keeping CVE guarded for some period of time, but and I think that that's not going to go away. But I think the timelines and the urgency we have to act on for critical variabilities in critical components is going to be changing in the next year because somebody finds something like this and we're not aware of it, we know that somebody else is going to pretty quickly
with AI and we better go get our systems patched. And so going to your second part of your question which is you know what do how do we have to think about this? We have to get our patching systems up to you know they've been designed for human rates of vulnerability discovery and exploitation. We've got to upgrade them to being ready you know AI rates of uh
and volumes too. So this is a big wakeup call I think for everybody including Microsoft. >> You just showed us that we don't even need the source code. You just need the binary and I'll do it for you. So >> uh I don't know if open source we've got one more question over here. >> Um hi first of all great talk. Uh I like that Microsoft take
us seriously now. Um >> we have for a while. >> Yeah I know. Just just had to say it. Um I I saw all this I I have been concerned for years for sbombs and all this supply chain uh stuff and I and I see I like the initiative but I I it looks like nobody uses Linux in in the board or something because they don't talk
about Linux distributions and there's a lot of learning for the past 30 years on maintaining software supply chains and integrating software and making make it uh different libraries with each other and and pipelines and all these build systems that we have. So uh what what will be the role of Linux distributions in this initiative because I see >> it concerns me that it they are not even
mentioned and a lot of stuff that the good practices and all that stuff that you mentioned here >> they are already being done in like the ABM Fedora Arch and a lot of Well, I didn't mention them specifically because I think they're not noteworthy other than, you know, the impact and scale and actually sophistication they're at. They're not noteworthy when it comes to what they need to
do. And Sbombs, they're all required to produce Sbombs. So just like Microsoft's required to produce best bombs for Windows, all the dros have to produce sbombs if they want to be used by the US government and more you know the EU and other organ uh countries and organizations that are now requiring that as part of their you know consumption policies. So yeah, if you haven't seen Sbombs,
that's because they're not they're not yet care. They're not yet being required to produce them because they're not selling to those kinds of organizations >> Question over here. Just one last one. >> Hey, um so it seems we should get started doing Sbombs in our CI/CD. Um maybe if you have some tools I saw anchor sift seemed like a good way and another thing is um can
you detail a bit the rules of universities? >> Yeah. So first there are a few asbomb tools out there that you can put right into your build pipelines. In fact Microsoft has a tool called asbomb tool that we open source that you can go use. That's the one we're using internally as well. Uh so yeah no reason to just start to go bombs right now. As far
as the role of universities, uh let me I mean obviously educating students about open source security. I mean security has typically been a oh you're interested in security go this track. You're interested in software engineering goes this track. Two separate tracks when actually it needs to be integrated. Now you can't be a software engineer computer scientist or developer without also understanding security. So, I think the role
of universities and schools is getting security right there up front. >> Um, I think we're we're I imagine Mark will be happy to to hang out for for a moment and answer a few more questions one-on-one, but um we've got another session starting in this room in a little bit. Um, so want to go ahead and we I think we'll want to call it here on the
Q&A. But thank you again, Mark, for Um, and thank you to uh thank you to Microsoft and GitHub for your support and and Tailscale and everybody else. Please stop by the expo hall. Say thank you to our sponsors. Remember, we'll be back here this afternoon for a keynote with uh Professor Comr to learn a little about uh about all the network underpinnings that uh helped build the
world we live on today. So, thank you. >> [snorts] >> So hi. Yes. Hello everyone. Thank you for being here and today we will be diving a bit deeper into how LMS can become a resource hog and how we have tamed inference with help of Kubernetes and opensource muscle to improve the entire pipeline. Currently what's happening is number of use cases for AI is increasing significantly and
with that we are serving to more and more diverse use cases. Whenever we are doing that we are serving different forms of models and different forms of IO but with that GPU demand is also increasing but that is not bad because currently whatever we are doing it's been more effective with the help of LLMs starting from costsaving use cases to new product and services revenue gains optimizing
the entire performance of an individual to a corporate All those things are done with help of LMS around us. Today we will be talking about how and what are LMS, how transformers come to the picture and how they require GPU, the performance and life cycle of LM, why is inference so hard and how optimization of inference around different areas become so hard as well. Last of all,
we will track how model performance is measured with hotel and all those things and talk about some architecture benchmark and results. After that, you will have some actionable insights on how you can improve the entire inference cycle. About me, I'm Ritik and I'm flying all the way from India to give this talk. I'm a platform advocate at Vcluster, a opensource communities multi-tenency company and a CNCF ambassador.
I've been a Google banquet scholar as well and done couple of certification. The reason here I am today is because one of the use cases that we are seeing a lot. A lot of people care about models and how they are being served across their organization and a lot of companies are getting a lot of GPUs in house because they want to have the stability of having
a private cloud inside of their organization. The problem becomes where you have the models, you have the GPU, but how do you serve them effectively? To understand that we will break down what LM inference is. You have a client which is generally us or an agent. It can be through web, it can be through a CLI, it can be through anywhere. And then there is a response
which you get from an LLM. Between that there are three more areas. There is the API if you are using a third party API or there is a gateway and whenever the request goes through the client to API or the gateway you have tokenizer which tokenizes all those things. Then there is the router and batcher those are effective to understand what type of models you are sending
your request to and bashing is also important as sometimes without effective bashing your requests are expensive. We will dive deeper into bashing a bit later. Then you have KV cache, GPUs and all those layers where you are doing a bit of inference and forming KV K value pairs. And then there's a model server according to your model weights and all those things with help of observability. You
are tracking the model performance and generating a response which is being served to the user. If we dive deeper into history of LLM, it started with engrams and RNN but in 2015 there was attention mechanism that made things popular and in 2017 something happened which is called attention is all you need a paper which talked about something called and the transformers were a leaf from 2015 because
it introduced something called a self attention mechanism. With help of self attention mechanism, you could have more bigger context windows, training windows and more highly [snorts] throughput logic between your different With that, we saw GPDs, GPD2, GB 3, 4, 5 and all those things because now a lot of foundational models have the technology to go deeper. But why does transformers matter so much? Before Transformers we had
a lot of things like Ardan and LSM team but with transformers things changed kind of because now massive par computation of data is possible. You don't need to go a word after word after word. You can do a sequential word analysis and relationship building. Apart from that you can have higher dimensional data to form better relationship with the help of softmax. Previously you couldn't form relationship between
different words in a sentence but with higher dimensional data if you have a individual you can find relationship between thus individual between different words and the bigger advantage that's the reason we have bigger models and parameters is the training time with hello parallel operation you can have generated better soft maxes and compute the probability at the end the attention mechanism which is your query key value had
a better relationship which helps you to generate and use better models and performance. Lastly, if we try to understand why are GPUs so important, we need to go back to the architecture of GPUs. CPUs are great for sequential logic, but GPUs are built for parallelism. If you see the diagram, GPUs you can have like a highway with a lot of lanes on the car. CPUs are like
a traffic stop. With that CPUs can't do a lot of parallelism together. But what CPUs does is on based of their architecture is like a very good amount of control logic. What does control logic means that it can have preeemption. So your arithmetic logical units are doing a task and CPUs can start and stop the task according to your requirements and with a bigger cache it can
manage all of those things easier but with GPUs things become a bit more different. GPUs don't have a big control logic unit because they are meant for parallelism. So it can do a task very good a lot of times to create a bigger throughput. As you can see, it only has LU's over there. Less cache, less control logic unit. Just focusing on LU so that you can
do a task better always, which is very much a requirement for transformers and other matrix multiplication operations. Now I use LLMs to generate a picture because we are talking about LLMs. This picture represents what's there with CPUs and GPUs. CPUs do a lot of things in the background, but GPUs are meant for a high operational heavy mathing. So if we go deeper into this, we see that
right now more and more companies are acquiring more and more GPUs. And the reason is like if you have a lot of GPUs together interwined, you can have a very huge amount of that can do a lot of things together. But there is some problems in it. To understand how the entire phase of AI workload happens, we really depend on Kubernetes because most of the big foundational
model and industry house use Kubernetes to develop, train, deploy, serve and monitor the models. From the stage where you are just generating data and gathering all those things to serving the model and monitoring, Kubernetes becomes as a part of your core infra which talks about all And right now uh with support like EKS with so many nodes, Kubernetes becomes as a foundational layer whenever you want to
develop, train and use inference between your different models. Going deeper, it seems to be good until it's not because increasing GPU is not always worth it. Here are two graphs. The TTFT is time to first token and we see the top 1% worst latency on the left side of the graph. So as we see as we are increasing the number of GPUs from one to two even
if you increase the request rate we get a lower latency but if you increase the GPU from two to four there is a bigger drop but if you see from four to eight it's very similar. So it says that whenever you are increasing GPU your latency might not be benefited going back into other data point it is about max streaming multiprocessor and it talks about something different
how busy your cores are whenever there is more GPUs so as you see even if you have one two four eight GPUs over here it is not changing because the problem is not about how many GPUs you can use it's about how you can utilize all of those So now let's process this thing in a bit clear manner. There is a AI boom and even if you
have budget sometimes it's not possible to get all those GPUs you want because there is a shortage of GPUs. So first of all you have a problem of having more GPUs then you need to optimize it. And the second problem is even if you have the funding and all those things, it's very hard to acquire more GPUs. Recently I saw a LinkedIn post where they were using
a neo cloud and there was no availability of GPUs for that training task. So that is a true reality right Now world is running on Nvidia GPUs and a lot of models are being served with those. We have other AMD GPUs as well and GPUs and all those things. But at the core everything looks fine because you whenever you want to use a API request to your
favorite foundational model it is working very effectively but in background when you try to get that capacity in house there are a lot more problems. First is the unpredictable GPU saturation. You might have GPUs whenever you want you might not have it. Then there is the latency, memory leak, token overutilization. Sometimes you see some task is using 40 50 million tokens because it just hallucinates. Then there
is cost. There is how you scale your LMS, high resource excessive GPU usage and excessive cache which kind of destroys the entire system. Now when we talk about LMS in general there are three phases to it. One is the training phase. Training is the foundational phase where you go talk gather data and try to find patterns. And right now in training phase we also use LLM because
we want to have a feedback loop. So a training phase also uses a bit of inference. Then you have finetuning phase. You can train your model with everything that is there. But how do you fine tune it for your own requirement, for your own industry, for your own use case? With Laura and Kora, you can do all of those things. And then the last and most important
part which most of the end users care about that is inference, which is your day-to-day task. You go and generate an image, you go and talk to a model, you go and ask for a code review. Everything you require and want should be in low latency because you want it as soon as possible as efficiently as So we will today focus on the inference piece and this
was from CubeCon keynote. So right now we have a lot of phases in the pillars of open source AI. There is the training part, there is the inference part and there is the agents part. Training as we talked about it was about model development how you are building a model. Inference is about model serving. But then comes agent. Right now everything is running across agents. Whenever you
want to code you create 10 15 different agents who are working together to create and solve a task. But even if you compare training phase, inference phase and agent phase, each of them required inference overall. So to optimize that part is very crucial because in training whenever you are having feedback loop you require inference. In inference it's just the inference and agents also required to go and
talk to the models to get some task done. So the focus is inference today and to understand what is inference because we want it to be in the same page. It is basically when you create a output from a input which is your prompt. If we read the official definition, it's about when an algorithm moves from producing sample outputs to guiding business critical decisions. And the few
different parameters that are very important for in uh inference performance metrics is the TTFT and time per output token. So the time to first token is how much time it requires to generate your first token after everything that is happening and time per output token is how much time it requires to generate the corresponding tokens. When we talk about latency in inference, it's the combination of the
TTFT which is the initial time it requires to generate a first token and then TPOT which is like how many tokens you are generating and the multiplication of that and then the total latency is the combination of the amount of time it takes you to generate the first token and the total number of time it requires you to complete the entire u stream. So whenever someone Inference
is slow. What does it really mean? The first one which is the very common one which we have seen with third party foundational models as well is sometimes models take very long to respond. The second was latency. Whenever you are interacting with a model sometimes response time per request is very high. Third is whenever you are using in-house models the performance of GPUs which you will be
diving deeper into a bit about how things work over there it's inconsistent sometimes you have very high utilization sometimes you have very high glo utilization but there is nothing common ground and then throughput request per model per handle per second which is sometimes very low. So whenever we talk about computational bottlenecks for models, it are generally two categories of that. One is compute bound which is how
much performance and the ability of a model to perform and just data and just produce an output. But the other one which is very like important and crucial which a lot of people don't really talk about is the memory bound bandwidth bound. CPUs are doing a lot of things to get the GPUs don't. So a lot of times it requires you to transfer a model from your
CPU to GPU. So the time it requires for you to do that is memory bandwidth bound and that also acts as a performance bottleneck for you. So increasing just the models and the GPUs doesn't help over there because you need to optimize how you are storing the models either in your CPU or in your Next is let's understand whenever we talk about the cloud native AI stack.
This is from the cloud native AI white paper from the group where does the ML life cycle pictures in. So at the fundamental area we have the accelerators which is your GPUs, TPUs. Then we have the infra operators which can be your new clouds, your cloud providers. Then we have the platform which can be kubernetes, open shift whatever you are using for that. And over there on
top of that your ML life cycle sits. This is where the model training performance inference and all those things happen together. Focusing on the workload part which is we are consuming it is a very important stack on how you are dividing everything together. We will focus on the ML life cycle today. Now deciding on why inference is hard we this part. So whenever you want to get
any sads or vive code generate something there is a GPU utilization and iteration and you are consuming creating something. So every day you are using inference and to understand the three critical areas where inference becomes hard. We can divide into KV caches how you are storing the cache so that you can have a lower latency and more throughput. GPU under utilization how you are using your GPU
so that everything is served and why increasing GPUs is not always helpful. In this diagram you can see like time per request and compute resources. A lot of times GPUs are more idle than CPU because how they are made. But to optimize is the name of the game. So let's understand what's KV cache. So whenever you talk to the model in the inference pipeline cycle there is
a inf sequence and then there is a tokenizer and with the query key and value your query tries to find the keys and if there is already a value it retrieves from the cache if there is no it computes the k pairs and stores in cache and then generates a token. Suppose you are 10 people are trying to find the capital of a country. Going through that
loop again and again will take a lot of time. If there is already a cash, it becomes simpler. Now if we try to go deeper into the diagram to compare between those two. The first one is without cache. So as we see over here every time there is a query the number of table is increasing it's become more dimensional but with a values are taken from cache
so overall it's a simpler table so simpler and more efficient in from the other part is how GPUs is not always fully utilized we talked about this a But weight and biases presented a small study on they where they observed that even if you have a lot of GPU average GPU utilization per user is very low because most of the times the GPUs are sitting idle and
whenever there is a task it is not consuming the entire VRAM and all those things together. Next is fixing slow inference. So how can we fix the slow inference? So we dive into the same problems which creates that. First is the KV cache where we optimize KV storage and retrieval then GPU where we improve utilization improve sharing then induce bashing and sharing. And last is obserility where
we measure per profiling for resources and token all together. At first let's go uh into the KV cache solution. In the KV cache solution we have a lot of open source projects which are very popular over there like VLM. What it does is like have page attention generally the key value pair is stored as a whole resource. So when you induce pages in it it becomes more
efficient because now you can have more smaller units to talk to. And then a cool part about VLM is that it induces continuous fashion. What it does is like you need to don't need to wait for other parameters to be completed. It can be continuous basis. So you're not waiting for your GPUs to be empty. So it creates around 3x more throughput from our experience. And then
LMD so it does a very similar thing where it does a cache aware routing. Cachal routing as a general concept is where you are understanding where your caches are in different GPUs and you are routing those together over there. The other techniques are like tired KV caches where you have GPU and CPU caches together. So whenever there is a cache and if it's hot you are storing
that into your GPU so it can be accessed simpler and faster and you don't need to go back to the CPU and have a memory bound issue and CPUs are storing slower and cold cache it creates a cheaper inference and then cache locality is where your latency is kind of killed as well because you can access the cash faster. If we focus on the core idea in
VLM, it focuses on just three parts. It is continuous bashing as we told where we have the request coming together. Page attention where we use and store the self attention mechanism in pages and then optimizing CUDA kernels. So in a general GPU you have a 30% and more like that. So you can do it faster and overall with VLM you can see the memory uses being increased
because you are using more efficient parameters. Next is a combination of VLM and LMD where you induce inference scheduleuler and with help of KV events you can talk together and create a fleet which does both of things together and create more efficient cache manager that helps you to create and manage the cache together. So you can more effectively provide better inference. And when we talk about different
form of uh over there and the better is the precise scheduling because the output tokens the time to first token the mean and V weight Q after using that is the best rather than appro approximate scheduling load scheduling and random scheduling and provides you a better overall control on the caching part. Next we focus on the GPU solution. The GPU solution part is a puzzle on its
own because now you have your GPUs. So how do you make sure they are optimized effectively and use better? Three parts bashing, sharing between different tenants and scheduling smartly. I will just go a bit deeper into how GPUs integrate with Kubernetes so that we can understand the entire stack. Generally you have the GPU hardware then you have drivers depending upon the provider you are using. But Here
we will take an immedia example and then you have the container toolkit and the device plugin. What happens is like there are some host level components like the container toolkit and the GPU driver and there are some kubernetes components like the device plug-in the feature discovery and the mig manager depending upon what type of GPUs you are using. Now whenever Kubernetes is creating a container it talks
with containerd and then container creates that with run C and your container is there with the OCR runtime. Now what the container toolkit does is like it goes and mounts itself gets the libraries from the host and the environment variables and inject those environment variables in the container in the device group. So your container which is very secured and isolated has that Nvidia container toolkit plug-in. So
your container can now access GPUs. And why so much complexity? Because as we do all those things at once. Now when you want to use your GPU in your communities, it is as simple as adding one line. Your ps, your deployments, whatever you are using can after all those hard lifting which sometimes are done by the managed providers as well can be used with just a line
over here depending upon the capacity you have. Everything boils down how Nvidia GPU operator is there because it helps you to manage the life cycle which you initially had to do on your own and it's a open source GPU operator and uses the operator framework within Kubernetes to leverage all those parts where you get the India container runtime automatic mode leveling and a lot of monitoring metrices
that you can pull in for your So it creates the layer between your hardware and software which can be used effectively. Now what we are thinking is like you have your GPU node, you have your pod and there is the Nvidia GPU operator You have kind of solved how you are serving GPUs but now optimization is the key. The problems with running GPUs with Kubernetes is like
a lot of task requires a lot of GPU to like a A100 can process one request per second and this is how the architecture is that's the reason so much inference takes a ton of GPUs so you have requests coming in and each of the requests are taking per second so it it can be quite long here we introduce bashing where you combine the requests together for
with large batch sizes to create a large GP parallelism together which the utilization overall and when the GPU is not saturated you can double the batch sizes and double the learning rate but again there becomes a problem over here when your batch sizes increase very you also increase a lot of validation error induce a lot of validation error over here and that then you need to have
warm up and allow ramp up over here so that you are more efficient on that part but overall to find a balance where you are using bashing effectively and as well as using not very large batch sizes is very important. The next problem is GPUs a lot of times are sitting idle. You have a lot of models served by different users, different talents but most of the
times they are sitting idally which is not efficient because you want to use those expensive piece of hardware as much as you can. You don't want them to be sitting idle with low load on your machine. So here another piece of the puzzle comes into the picture how you can share GPUs between different tenants. Over here the most popular is time sharing because it inter leaves with
one another and sequentially slices your GPU so that you can have efficient scaling. But again it is not that effective as we will see later on. Then we have MPS which shares the GPU context between fine sharing of GPU resources and it is also quite efficient but the problem is if you are using secrets and all those things it can be shared across different areas. So it's
not that secure. Next we have MIG which is more popular but it only comes with very expensive hardware. It actually breaks down your hardware into seven instances which can be acted as their own hardware units. The advantage is that you have very isolated units that you can define for each of the task. But the disadvantage is that the GPUs that support mega very expensive. So sometimes it's
cheaper to get smaller GPUs which support MPS rather than get a very big GPU which doesn't has those capability because It's the cost to benefit ratio is very Now a couple of uh questions that pop up like how do you prevent like when you are using MPS time sharing and all those things. How do you ensure fair scheduling between different tenants? How do you prevent one workload
from monopilizing other? What happens when multiple bots try to use same GPUs and the security boundaries between those? So something like Kai scheduleuler comes into the picture which is a software scheduleuler from Nvidia and the good part about that is like you can use that with your current GPU and it offers a flexible packaging. So you can have several units of your GPU being used a fractional
units between different tenants and the good part about that is like it provides you dynamic GPU allocation throughout your AI life cycle. So you can have different units of your GPU. Suppose training requires more GPU than inference. So you can allocate accordingly with that your GPU architecture looks something similar where you have GPU nodes, pods and the scheduleuler and they're consuming over it. Now if you decide
on what is best that is very dependent upon your situation but some are not available on every hardware and the things that are available on every hardware depends on do you want fractional GPU sharing or multiple smaller workloads over here going deeper into this now the problem is even if you have a lot of nodes and everything together you need to do and make sure your different
tenants for example tenant A which is your customerf facing tenant tenant B which is your developer facing tenant the GPUs they have are utilized it is not good if your one tenant has and it is also not good if your another tenant has very low utilization you need to have a balance between those so that your GPUs are not tiered off so one of the most common
is to combine everything but that creates a different set of problem. The problem is how you deal with enterprise multicluster component with enterprise multicluster component unit autoscaling observability deployment in different areas CRD isolation because different tenants might need different form of CRDs for your customers. the distribution of GPUs and comprehensive security and audit the CRD isolation is a big puzzle because if you combine all those GPUs
together and everyone is using something like a namespace isolation between them so CRDs don't really help over there here something like vcluster comes into the picture which is a open source tool what it does is like it creates name spaces on your host tenant and creates a API server on top of that your tenants actually talk with that API server which acts like a whole real cluster
and with that they can access and use GPUs as well as CPUs workload more efficiently while isolating your CRDs or controllers. The architecture changes a bit over here because now you can have GPUs together and CPUs of course and your utilization is dependent upon the different requirement. So you don't have very high utilization on one GPU lower on other because you can efficiently schedule and with that
different form of use cases come because if you have a whole big AI cluster you can have four five different inference use cases running parallelly you can have if you provide a GPU as a service then you can have four different tenants consuming that GPU then apart from that multifunctional use cases where you have inference finetuning pre-tuning running in a similar cluster works and distributed inference where
inference is running on remote sites but you're just consuming that is also a very common area now we talked about GPU bashing sharing we talked about KV cache and all those things the most important part about this every all those areas which we talked about is the obsibility part the The reason we know to solve the GPU B and the KV cache is because we have been
observing all those things. So to focus why it's important to observe is because you can understand how you are tuning your GPU how the model performance is going around benchmarking your cost and performance to create a ROI for the leadership. Now one of the most common things that we have seen is like it's very hard to configure and set up all those things. A lot of people
just go and measure the GPU load for once and keep it aside. The problem with that is from this paper like users training and RN saw that GPUs drop 25% sometimes while we're training. If you're not always observing that it's very hard to understand why. It might be because a data load loading bottleneck inefficient batching model switching phases from training to like a GPU was running 85%
effectively and then it just goes below you don't know those data points until and unless you are observing them other different areas is like whenever you are using GPUs multiple GPUs and consuming them sometimes during the phase you understand some GPUs just go to zero and that might be a case that you don't have higher workload or it might be a case of data distribution bug imbalanced
workload or an unintentional configuration change caused by something in the application. The other part is to understand how GPUs are being used during idle periods. Suppose you have trained eight GPUs running over there and as you see sometimes the GPU drops at regular intervals some of them. Why is that? It might be because ofation phase and alltime training slow data loading and sometimes inefficient pipeline scheduling where
the GPU is waiting for a bash to come in. If you have VLM and all those things there can be a sequential load but without that without the sequential load it might be very inefficient over there. after observ part there is the profiling part as well which is very important and profiling is sometimes a way to pinpoint exactly which part of your code or training resources causing
latency and hurting performance with the CNC landscape being so huge we have a lot of hotel libraries that do that for us and you can measure the total CPU GPU and all those things actually how it's working you can also have distributing facing for your LM application to understand how and what things are being happening and you can observe the data points to understand which areas of
business you should provide more efficiency over Next is with this phase details you can observe all those things for a specific application or during the training or serving phase. So you know things that you can work on and you can schedule better which overall improves the entire like this is a endpoint generate which is generating an image. So it can overall improve how you are observing and
which areas you want to focus from. One of the case studies that we have seen like after LM deployment cost increased to $50,000 and with performance utilization and all those things the 20% cost was decreased and this is from graphana which talks about how profiling is ideal for optimizing or minimum cost because it talks about a lot of different areas which with which profiling reduces the entire
cost log and makes it more efficient to observe and see what's happening and why it's happening. whenever you want to track metrices for ongoing performance, there are a few areas during development phase, you tag CPU users per token. So to identify inefficient code profiling and all those things are there. Then memory usage for requests and latency per request is something which a lot of individuals track so
that you can have a holistic idea. And then there's the deployment In deployment phase things change a bit more. You talk about how GPU utilization and efficiency is being tracked and resource efficient metrics like CPU, memory, GPU with help of hotel and how allocation is working. dynically based on demand is working that changes depending upon type of workloads you are running. A lot of organizations right now
are moving to bursting. So to share more about bursting sometimes you have private AI cloud capacity inside your organization and you need more GPUs or more workloads. So what you do is like you create a burst to public cloud or a different private cloud. So you can run a part of your in local machines or local private AI cloud and a part of that on public cloud
dynamically based on demand and how operations are going through and bursting has become very popular recently because most of the times it's very hard to get the requested GPUs so you go around different providers to get and find whatever there is available and to consume that other core technical challenges when you try to optimize GPU sharing to improve model LM performance is like distributed attention kernel. So
whenever there is a lot of tenants over there you will have a lot of cross node uh communication. So to maintain the throughput across those nodes is a very big challenge. Then other challenges like KV caching across node. You have GPU nodes, they have their own KV caches. So how do you distribute that cache across different nodes so that you can reuse the cache instead of generating
local caches again and again? Then communication and graph capture. Whenever there is dynamic bashing, you have NCCL CUDA hooks all over there. So to have a deterministic nature of that so you can have dynamic bashing that is efficient and capture the graph that is also a challenge and then um one of the very important one is like pipeline and mix precision. So whenever you try to handle
the tensor core alignment there is always a loss of uh scaling. So syncing different inferences stages together so it's more efficient is also a very technical code challenge that a lot of people are trying to solve. Apart from that uh sequence parallelism what sequence parallelism does is like suppose you have a number of GPUs. So you bash your request together and send it parallelly across a service
of GPUs. So you're not overloading one GPU but bashing similar form of requests together. So have have a maintained attention synchronization across those single pieces of GPU and then bashing strategy. So bashing strategy is like how you can have dynamic bashing for multi-enant inference minimizing the tail latency while m maintaining utilization. So you have better bashing in different multi-enant infra and you can consume efficiently all those
things and routing is also quite popular because you allocate your tokens to different expert set of network GPUs. But here a different kind of problem arises because as you are routing your GPUs and tokens to a different form of GPUs the problem is not about compute it's about and here the solution we are trying to solve is how you can have cross mode allto communication so that
you can have a better efficient routing and serving. So yeah at the end of the day you should focus on three core areas. One is to cache mod. The second is whenever you are scaling scale better and see everything observe everything so you know what's to focus on and what to improve. Thank you so much. Uh this is my LinkedIn if you want to connect and I'm
here for any questions if you have. >> Yeah, it was a like big talk >> Sunday. Yeah, I was just saying stepping away and processing going through your presentation again a lot, you know, just gives a lot to think about. Yeah, I I have no question until sometime later this week probably. >> Yeah, please feel free to text me on LinkedIn. Uh I'm also going deeper into
few of the attention mechanism. So I would love to talk. >> So in your in your summary on the three parts to focus on the caching, the scaling better and the observability which of the three pieces would you recommend tackling first to handle figuring out where your bottlenecks are? Uh I think the KV cache problem is quite uh more important because that can be done more efficiently
than the other two. Yeah, the GPU problem uh is when you are scaling more but the KV cache and obsibility part is how you understand where are the bottlenecks. I have also some clean pins that I will drop over here if you want to collect. And thank you so much for coming on a Sunday and it was great having all of What is open search? Open search
is a communitydriven open-source search analytics and a vector database platform which has integrated tools for observability, security, visualization and AI powered search. Open search currently is under Linux foundation. uh I think we recently went under Linux foundation and we have a lot of partners including AWS, IBM, Uber and whatnot. So today talk would be around vector database. So before going into what is a vector database let's
get our basics correct. So the first basic question that comes in our mind is what is a vector? Okay. So you have images, documents these days you have like lot of images, documents and then you have models these days a large language model or could be a text embedding model. When you pass these images or text in a in a model it produces a vector. Okay, vector
is nothing but an array which contains some floatingoint numbers. Now what is present in this vector? So vector basically captures the semantic meaning and the content relationship. Okay. And these vectors can range from tens to thousands of dimension. Dimension is basically the length of the vector or the length of the array where each dimension is nothing but a floatingoint integer sorry floatingoint number which represents a four
byte. Okay. So that's the basic understanding of vector. So you have images you pass it through an embedding model. it gives you some mathematical array and we assumes that okay this array is capturing all the semantic meaning okay so that's the whole role of embedding model now moving further let's talk about the ABCs of vector search uh with open search so the basic idea is now you
have these tons of vectors and you want to do a search on top of it and that's where we'll dive into okay how do we perform vector search with open search now moving ahead Let's first look at the architecture. So anytime you're trying to do a search, you'll create an index. Okay? And in this case, it would be a vector index. So now first let's look at
the architecture of an index. And this is still at a very high level. So let's just go right into it. So within open search, everything is represented as a index. So you create an index for your text data. Okay. Similarly, for doing vector search, you create a vector index. Now what is a vector index? A vector index is similar to an open search regular index where every
index has a lot of partitions. We call it as a shard. Okay. And since open search runs behind like runs on top of lucine. So every shard is nothing but a lucine index. Okay. Now when we go dive into these shards or I would say lucine index every index is composed of a lot of segments. You can think of segments as a immutable snapshot of the index
where when you are keep on ingesting data the these segments keeps on getting created and you keep on merging them behind the scene. Now when we look into a segment a regular segment contains a lot of files. So let's say you are ingesting text data you are ingesting some uh numeric data all those things goes into different different files. So there's something called a fdt file dvd
file tim and whatn not. Now how a vector index is separate. So for a vector index you have something called as a index file or a graph file which generally we have lot like we have a file called as dotf we'll talk about more on it what is f but think of it there's a file which captures the vector index which has the serialized information of that
index for that particular segment. Moving on. So how do you do injection into open search? So you have your documents, you have your open search index. Every index is composed of a lot of shards as we discussed. These shards are present on or distributed across some of the nodes. We call it as a data node because they are storing the data. Okay. So this is your basic
indexing flow. You start with some JSON documents. Okay. Which goes to your index via REST APIs, gRPC or whatnot. The index looks at it, divides those documents into different shards and those shards are hosted on a data nodes. Now comes the search flow. Now you have a query. You ask your index that hey I want to search. I want k nearest neighbors. I want some nearest documents
around it. So what you do is you form a query. You again hit that same index. That query goes to something called as a coordinator node which is nothing but coordinating your whole query. that coordinator node goes ahead and hit or request data on these data nodes saying that hey I got this query who all can serve this query for me okay so it's like you're getting
your query and asking all your partitions that hey give me your uh run this query for me and give me the results and that's where coordinator node comes in accumulates all the data and send it back okay so data nodes process the request locally and sends the results back okay coordinator reagggregates and sends the So the basic idea again you send the documents for injection open search
stores them and when you want to retrieve those documents or when you want to do a query you ask the same index that hey this is my query give me documents back that's the whole injection flow now let's look at how do we create an index what's the schema of an index or how does like when I want to create an index how does it look like
okay so this is a standard u you can say input or a rest input that we'll put And let's dive into it a little bit. So first thing is you put a setting on that index saying that hey this is a KN&N index which means that this is a vector index. Okay. Now comes the mapping part where you'll specify that hey I will have these kind of
fields and those fields have some meaning in it. So you define that okay I am going to have a vector field which his name would be my vector. Okay you define the type of the field just like KN&N vector. Within open search we have different types of fields. You can define strings, integers, arrays. So that's where you define your field type which is scen vector. You define
the dimensions basically the length of the vector. And then you define the space type. Okay, which is basically the vector similarity. How do you want to compare these two vectors? We'll dive into vector similarity in just a moment. But these are all the things that you need to define to create your vector index. Nothing else in else is needed. There are more I would say advanced techniques
that you can use. We'll talk about it later on. But that's how you get started with the creation of a vector index. Now comes the vector similarity. Okay, we looked at space type last time uh in the last slide. So let's dive into it. That's where the vector similarity search comes into place. So the basic idea behind vector similarity search is hey you have vectors which are
capturing the meaning or semantics of that particular text or a document. Your embedding model is putting that vector in a embedding highdimensional space. Okay. And if the vectors are closer or if the documents or the data that you are ingesting into the model to convert it to an embedding, if they are closer, they would be within that embedding space. Okay. And similarity is basically saying that okay
these two vectors are closer how do I compute like how do I know these two things are closer? And the idea behind these things is there are three types of techniques or I would say mathematical computations that you can do. One is the ukidian distance. Next is a cosine similarity. And third is the dot product. And as you can see uklidian distance is nothing but distance between
two vectors. Cosine is the orientation between the two vectors and the dot product is the magnitude and the direction. Okay, this comes in handy in terms of when you go into the details of how your model is generating the vectors. But from a high level standpoint, all we need to understand is hey there are three techniques to do a similarity search and that's what it is. Moving
on, the next question comes in our mind is why do we have to go through all these things like why what why vectors are needed? Okay. So, vectors are needed because uh they compute the closeness okay behind. So, they they encode the closeness of uh your text or the documents. This is where if you see that if I have a dog and a puppy, okay, and if
I'm searching for a dog, I should retrieve the documents for puppy also because those two vectors are closer. But if you are trying to search for dog and then you hit into something like cars, you'll never get that those things because the vector for the car would be in a very different space. Okay. So that's where uh the vectors comes into place. Moving on now Naven who
is one of my colleague will go through all the different features of open search as a vector engine. Okay. Handing over to Naven. So now uh we went over like so what is vector search uh like what is the use of vectors uh and all those things. Now let's try to understand what are the different features of open search vector So uh now we know like there
are two types of search. One is the lexical search and the other is semantic search. Let's try to understand like so what is that what is the difference between both of them and like what's the advantage of semantic search uh with a simple example where everyone can easily uh understand so let's say I want to find some restaurants so which serves pasta so I I just simply
search like something like Italian pasta so the traditional lexical search what it does it does a keyword to keyword matching so it will only able to find the list of restaurants where Italian or pasta matches with it. So it it might be able to find the pasta house, but it can't find something called mama's ni, but we all know that nachi is a Italian pasta. So now
let's see how semantic search handles. So let's say I ask something for cozy cup heavy dinner. So semantic search tries to understand the meaning and intent of the user's query. So where it tries to understand, so I'm looking for some warm Italian food. So then it tries to find something called like Tony's tratoria. So even though it doesn't have a word pasta or something in its name,
it still tries to identify it from the list and adds it as part of the result list. Search evolution. So there are like four different generations like how the search has evolved over the period of time. So one is the lexical search which is nothing but the keyword search and the next is the semantic search which is nothing but the vector search which we just talked about
and then comes the hybrid search. So which is nothing but the best of both the worlds. So which combines the accuracy of the keyword search as well as the semantic meaning of the uh vector search. So think of an example. So let's say I'm searching for something called uh something for some Thai restaurant near me which is open right now. So here uh the lexical search tries
to identify the restaurants which are open right now and the semantic search will look for all the Thai restaurants near me and it will try to combine both the results and show me the complete set of all the Thai restaurants that are available and open And next comes the agentic search which is the future. So which is like quite popular these days. So which takes the user's
query, reasons it and like breaks it down into multiple steps and tries to execute them for you and like returns the results. So let's take an example and understand it. So when I say like show me some Thai restaurants like which are available at 8 p.m. and I'm also looking for some wedge options. So, it will try to find all the spots near you and uh yeah,
it will try to find all the Thai restaurants near you that are open at 8:00 p.m. And it also uh checks the menu for you and it will understand like yeah, it will it will look for the wedge options in it and like it will show the full list of uh restaurants and so now we understood like so open supports uh the vector uh vector search uh
capabilities. So but what is the heart of it like so what is the core component of open search that supports these vector search capabilities. So which is nothing but the KN&N plugin. So K&N means like K nearest neighbors. So yeah uh think of it as an engine under the hood which does all the heavy lifting for you. So when you send a semantic search query to open
search so KN takes the query process it and it leverages some of the powerful optimized search libraries under the hood which are nothing but the files and the lucin engines and it will process the query and it will return you the results. So and yeah and KN&N is also something like which is a highly sophisticated scalable in uh scalable plug-in so which can handle millions and billions
of vectors at a subsecond sub millisecond uh latency and how is it achieving it so it is achieving it using ANN stands for approximate nearest neighbors which we'll talk about it uh in a moment so the next comes uh and now we will discuss about different types of search which The KN&N supports the first is the exact KN&N or the brute force search. So what is brute
force search? So this is something like which we all know about where the query vector will compare will compare its distance or it computes it distance against each and every vector that are present in the vector index. So this is uh the brute force search like as it keeps searching against each and every vector in the index. So it will give you a best results where the
recall will be 100% with perfect accuracy. But the catch is like it is computationally expensive because it needs to comput its distance against every vector in the so this is ideal for small data sets or for some filtering use cases. So here is a realtime example. So let's say if I'm searching for some uh Dell brand laptop which has an processor. So it will try to apply
those two filters of Dell and as well as i9 and then the the subset of the list would reduce down to 500 laptops. Then it will just do the exact search on top of those 500 documents or the 500 results and then it will return you the top 10 or top 50 like whatever user asks. The next comes is the approximate K and search. So think of
a a case like where I have like millions or billions of vectors. If I want to do exact search or brute for search on top of all those documents, then it would take a lot of time. So that's where approximate K and search comes into picture where it will only do the distance computation for your query vector against a subset of vectors using the graph structures uh
that it uses to build uh and that's how like it will return you the accurate results in sub 100 millcond latency even when you have a billions of vectors ingested in the vector uh index. So there are two algorithms that supports approximate KN&N search today which is HNSW and IBF algorithms which we'll talk about it in a moment and also like yeah we can we have also
have a realtime example here to understand approximate KN search. So think of like an e-commerce website where like you have like 10 million plus products in it. So now customer uploads a photo of a dress or something and he tries to search for that on the whole catalog. Then open search against the whole catalog and filters out the top 50 results like which are like closely matching
with the users query. So yeah here comes the first algorithm inm which is the hierarchal navigable small world or HNSW. So I'll help you uh understand a very uh I'll try to give you a very high level understanding of the HNSW algorithm. So to better understand this think of the roadways. So when when I'm at a place like which is far away from my place and I
want to travel to that place then first we need to identify what is the national highway to go there to to go to that state or whatever it is and think of this as the topmost layer in the HNSW algorithm. So which is like sparsely populated there might be very few roads but like yeah it's spcially populated uh similar to this and then next comes the next
layer. So once you enter the state you look for the interstate highways and all to go into the corresponding city in that state. So here the vectors are little bit more densely populated but still they are like far away from each other but it will try to connect the closest neighbors to it and then next comes the corresponding city to reach the specific destination there might be
many roads in the street level. So this is think of this as a last layer like where you have multiple roads which are like densely populated. So the query vector enters the top layer and it will gradually narrow it down and then just into the deeper layers into the last layer where it tries to only do the distance computation against a subset of vectors instead of doing
the full set of vectors in the index. So that's how like it does the approximate nearest neighbor search and it will return you the top results in very less time like in in the millisecond latencies. Uh so next comes the IVF algorithm. IVF or the inverted file algorithm think of this as sections inside a library. So instead of yeah when you provide all your vectors it will
start creating something called uh clusters by uh by grouping them or group the vectors into different different centroidids different different clusters where it will elect its representative which is no which can be called as a centroid. So which is nothing but an average of the vectors in that specific cluster. So once it forms all the clusters. So now when I send a query vector instead of like
doing the distance computation against each and every vector in all the clusters, it will try to do the distance computation against the centroids, find the nearest centroid and then it will start doing the distance computation against each and every vector inside that specific cluster and that's how like it will do the approximate search and it will return the top result back. So for example in this example
we can see that the query vector is doing its distance computations against the uh the green one basically the the green So now we just got a high level understanding of like what HNSW and IVF is like let's quickly try to compare both the algorithms and see the trade-offs. So as we just saw like HNSW is a graph based algorithm where it uh constructs like multiple layers
of proximity while ingesting the data and then during search it will pass through those uh graph layers uh during while processing the search query and then like it will return you the similarly like yeah IVF is a cluster based algorithm where it will try to form clusters or buckets uh group the clusters group the vectors into clusters or buckets and then search perform the search against only
those uh subclusters uh and HSNW is best for when you have like enough memory to process the request. So this is good for memory rich environments uh where it can process the request in low query latencies and basic but it will also provide you very high query uh quality in terms of results unlike HNSW IVF is good for uh uh memory constraint environments like where you don't
have enough memory but it can still help you process your request. Uh but the trade-off here is like yeah the query latency is low but the quality of the results might not be that great as HNSW. So here is a quick comparison of the trade-offs. So where we can see uh yeah in terms of memory HNSW might need slightly higher memory compared to IVF. And in terms
of indexing uh HNSW is fast uh and IVF is like much more faster but HNSW doesn't need a training step but IVF does need training. And in terms of recall quality, yeah, HNSW is superior so radial search. So what is radial search? So not all the times and not for all the use cases, we just need the top k nearest neighbor. Sometimes we also need the results.
So which falls in under a given threshold. So let's say like I want to find all the results like which are in a given distance or which falls in a given uh score threshold. So that's where radial search comes into picture. So where yeah it returns all the vectors that falls in a given distance or the threshold of the scope. So think of a realtime use case
like where I want to find an accessory that is at least 85% similar to the given mobile or something that I provide in the user input. So then it will try to find all the relevant points that are like matches that similarity and like yeah lists out all of those. So multi vector and nested search. This is another interesting search uh which most of the users will
use. So for example, think of this way. So I have a very big document. It has a lot of data in it. Uh I can't fit all of that in one vector. So I need multiple I need to like chunk it down into small small chunks and I need to generate embeddings for each of those chunks. But I don't I want all of them to fit in
one veh one document because I don't want them to be just as different different vectors because if I search for one So like so for a vector like which is matching in it I want the whole document to be written out. So that's where the multi vector or the nested search comes into picture where you can in justest uh include multiple embeddings in the same document and
ingest it as a single document. Uh when you search it will try to perform the search against each individual vector and tries to get the best score out of it and but it returns a whole document back to the user. And there is also a capability where you can aggregate the scores of the each individual child vectors and then aggregate them into the final score and uh
compute your results based on those aggregated scores. So here is another use case. So when you are searching for some sofa. So uh like in an e-commerce website like something like sofa has three different images. Let's say like front side front side and the texture of the sofa something like that. We we never know like uh so what is the user's query if it can match to
the front or side or something else but uh it will try to return the full set of details of the sofa back to the user though uh only one of the vector matches to the user query like it will return the full set of uh sofa details back to Okay. So now uh we understood like so what is vector search like what are the different search capabilities
it has uh and we know like what is L2 what is cosine similarity what is inner product. So or to perform a vector search for a query vector against n number of vectors in the vector index. So it needs to do the distance computation. Distance computations are CPU intensive and what if I have millions or billions of vectors then it it becomes a big bottlene because in
terms of query latencies. So that is where the hardware acceleration comes into picture. So these days like all the modern CPUs comes with uh support for this mo hardware acceleration where uh it will help you to parallelize your computations and like it can help speed up uh your queries or the injection rate as well and uh yeah it will reduce your latencies but 10 times to 100
times in terms of throughput. So to achieve this like you uh this is something called as SIMD which is a single instruction multiple data where it can process uh like think of this way like uh like both both architecture Intel as well as the ARM architecture both supports the SIMD. Uh so Intel supports two types of SIMD one is the AVX2 and the other is AVX 512.
So think of it this way like AVX2 can process like eight operations in like one CPU clock cycle. uh to explain it in simple terms think of it like it can add eight numbers in like one clock set. So similarly AVX 512 can process like 16 of those operations in like one CPU clock and then like we have something called as bulk SMD. So where you batch
a group of vectors and process SIMD at the same time on on the on those set of vectors. So which is called as bulk SMD like which is giving much more better throughput in terms of query and like it can also improve your yeah query latencies. Yeah, we will see this with a real time example like uh in just a So now let's try to understand the
classic inner product like which is uh what the legacy vector search used to do. So now I'm trying to do a classic inner product on two vectors a and b of dimension one. So in the legacy way like we need to pro we need to multiply each and every element one by one and then add it to the accumulator and then go to the next element and
then do the same thing. So this would take a lot of time if my vector length is something like 768 dimension, 1024 dimension or 1536 dimension. So which is like the standard dimension of every vector which the standard LMS usually generate. So then comes these part. So here we can see like instead of like processing one element at a time, it is trying to take a chunk
of the vector which is like four elements at the same time trying to do the uh product as well as the accumulation at the same time in one clock second and then it processes to the proceeds to the next chunk of the data and so on. So based on our experiments we have seen that using SIMD compared to the lexi way of doing inner product like we
have seen like more than 220% improvement in QPS and then comes the bulk sync. So this is like batching the vectors uh using SIMD. So in SIMD case like we have seen that like we processing the query vector and like one data vector at the same time. So now we can see that like we're trying to load the query vector as well as like four other raw
vectors like that which we need to process into the registers and parallelly processing all of them at the same time using five different registers. So this way like you can paralyze and process things even more faster and like can reduce the So yeah so you can see that like first it process like the first set of chunk of data and then like it move to the next
one and so on using bulk SIMD like we have seen from our experiments that like like 130% more improvement in QPS over SIMD. So here is a representation of like the progress and the improvement we have seen from our experiments using the scalar version of the vector computation versus SIMD and bulk SIMD. So we can clearly see like we are close to 300% improvement compared to the
traditional way of uh achieving the QPS. Okay. So now we understand uh what is vector search what are the different search capabilities but not all use cases uh can be done by just doing vector search. we still need something more something like filters. So let's say like like where are filters used? Like we use it like every day in our day-to-day uh use cases. So for example,
when I'm doing some uh shopping on the e-commerce website, when I'm searching for a product, I want to make sure that there is at least a rating of four out of five because I just don't want to compromise on quality or like I want to for example I'm shopping for a headphones. I want to make sure like they are wireless. So then I'll find a filter and
then I I'll add it for wireless. So that's where like we use filters uh in our day-to-day use cases and that's where like we can also integrate these filters with vector search and like achieve much more higher So that way like you can narrow down the set of the results on which you need to perform this vector search on and like achieve quick and accurate results. So
what are the different types of filtering that open search vector search supports? The KN&N plug-in provides three different types of filtering. First is the prefiltering. So where you apply the filter on the full set of data you have in the vector index and then you perform the brute force of the exact search on the subset of the results that you obtained after applying the filter. That way
like yeah you'll return the top results back to the user and the next comes the post filtering where you do the ANN search on the full set of data that you have and once you get the results on that subset of the results you apply your filter and then you'll return the results back. So both the pre-f filtering and post filtering has its own disadvantages. So to
fix these things like there is something called as efficient filtering. So where it will intelligently and smartly tries to do the search and apply the filter both hand in hand at the same time. Uh as it keeps going through those data structures, it will try to uh apply the filter criteria uh with it and once the filter criteria succeeds then it will also do the search and
like include it as part of the result set. That way it makes sure to return the top k results which satisfies the filter criteria. So now we understood like what are the different types of filterings. Let's also try to understand as a developer like how can we uh like write a code in terms of like how can we create a pre-filtering uh query uh query JSON basically
query DSL. So here we can see that the filter is on the top part and the query is in the bottom. So this is the pre-f filtering like where we will be trying to apply the filter clause first and then on the subset of the results like we'll try to apply the query do the query search query basically on top of the filtered set of and then
next comes the post filtering. So here we we want to apply the filter on top of the vector search results. So that's why like the vector search query class comes on the top and the filter filter class comes on the bottom. So basically yeah we'll be running the vector search query first and then followed by the filter class on top we'll be filtering on top of this
and then the efficient filter. So this is where both the filter as well as the vector search query both goes into the same query clause uh because we want to apply both together. So that's where like we can see the vector search query and the filter both inside the K and query case and then comes the hybrid search. We we we talked about hybrid search very briefly
in the very beginning uh of this session. So vector search is something that brings both the capabilities of the lexical search as well as semantic search. So where you can get the precision of the keyword search and the semantic meaning of the vector search. So we need keyword search and vector search together. So think of this use cases like where uh something like product ids, version numbers,
SKUs all those things which vector search can but search will not miss. It will make sure it will include that and vector search will also help us to get the semantic meaning and like will also include those results. But so now we have two different results. So one from lexical search and other from vector search but the scores of both of them will be like different. the
ranges of the scores of both of them are are different. So that's where like we need to uh fusion both the results and then create a final set of results which we'll talk about it in a moment. And like here is a realtime example of hybrid search where someone searches for like waterproof boots. So the the lexical search will look for waterproof and then the semantic search
will try to find uh find the all the weather resistant boots and then like it will combine both the results and then it will make sure it will return the best set of boots like which are. So now we talked about uh the fusion of the documents. So now we have two set of Usually the semantic search results fall in the score range of 0 to one.
But the lexical search results like will have a different set of score like maybe 45 or something. So to fuse to to do the fusion of both the results and like generate one final set of results. We have two different ways or techniques to do it. One is the score based combination uh which is nothing but like where we will normalize the scores of both the semantic
search and lexical search and bring them into a defined range uh using a different uh can be done like two different techniques. One is the minmax approach and the other is the L2 normalization approach. And once we get them into the same set of range that's when like we will take the top K results and like return it back to the user. And in general by default
it will use the same weights while calculating or picking both the scores. But there is also a flexibility which the user can use to tune the weights such that like if they want the semantic search results to dominate the lexical search or the vice versa that's when they can tune the results and then like yeah they can return the results back based on their and the next
approach is the rank based approach. So now we have seen how to like tune the results or how to merge the documents using scores. So this is with the ranks. So ranks are nothing but the order of the results. So now we have like two set of order of results. So let's say I have lexical search which returned me uh the document which is in the first
place but semantic search is giving me the same document which is in the 15th place. So this approach will try to uh do the fusion of both of them in terms of rank and it will return you the uh set of results like based on the ranks like by fusing both the results. So for which like it will use something called as RF which is like the
reciprocal rank fusion. So which is using a uh a specific formula which we can see like scores will be calculated for each document based on a rank constant and the rank uh where users can tune the rank constant and like play with it where they can tune the result and per sub query weights and get Okay. So to do a search query and to process a hybrid
search query like which contains both a text query as well as a vector search query like we use something called as the search pipeline. So where we got two set of scores the BM2 score and the KN&N score which is nothing but the lexical search score as well as the vector search score. So this is where like we want to merge both the scores of the both
the documents and then like get one final set of results. So that's where we do the normalization of both the scores combine the scores and then get the one final set of results either using the normalization processor or the scorebased processor and once we get the final set of results like yeah we will using those document ids we will retrieve the corresponding documents and then return it
back to the here is another uh interesting search use case which is called as multimodel search. So sometimes I have text I have uh images but I want to combine both of them and generate one single vector. So to do this we have some of the models some like the Amazon bedrock uh titanium model or some coherent models. So which will take uh all your data and
like generate embeddings out of it and then like create one single embedding out of it and during search user can either give a text query or as well as an image query and which will process which will be process using the same model and will be run against the vector index and it will return user. Yeah. So here is an example like where user uploads array running
shoes and like also looking for some lightweight trial shoes. So yeah it will try to process using that model and then like it will give you the right set of Okay. So now we understand uh like so about vector search what are the different capabilities which open search oper which open search provides all those things but now so I have a raw data which can be in
the document format uh audio or images and all. So you can ingest those vectors into open search either two ways. You can directly uh create embeddings out of it and ingest into the open search cluster or open search also provides you a capability to connect to those models that are there in the market like whatever you want using something called as a IML connectors. So using those
connectors you can connect to any models that are available in the market. That way like when you provide raw data, open search will take care of generating the embeddings for it and will inest the data into open search. So that way users can directly send raw queries uh in terms of text or images whatever it is and which will be also uh like use the same model
to generate the embedding out of it and it will run the search against the open search cluster index and then return you the results back. Okay. So now we learned a few different types of search. Uh this is something which is more exciting. So which is something like which is quite popular which is nothing but the agentic search which everyone is talking about. So what is agentic
search? So to do a query on an open search index so I need to run write a a long JSON query document. So which is not so interesting for everyone. So that's where like agentic search comes into picture where I can just give a normal text query but it will help to create the query DSL for me. It will try to it will the agentic search understands
the user intent. It will break it down into multiple steps create the query DSL for you. It will and like yeah it will run the search against the open search cluster and it will return user back. So let's see an example. So for example there is an user query. So which says like show show me some wireless headphones under $80. So uh the agent will understand uh
the user's intent. It will try to identify its corresponding index and mapping and uses the query planning tool sends it to the LLM and using the chain of thought it will create the DSL for you and it search index and then it will send the results back to you. So let's see what it does. So yeah, so this is the natural query that I'll ask. So this
is what the LLM will be will be creating using the agentic search. This is the query DSL that it creates where we can see like it clearly breaks it down into multiple steps. Applies all the filters as part of the query DSL and then it will run the search and it will return the top. So what are all the different components that I need to build an
agent search open? The good part is open search provides you all the required components that can be used to build the agentic search system. So which we can see here the first is the cannon plugin. So which is where like you will be using to store all the semantics for the vectors and do the search on top of it. And then the ML comments like which will
provide you all the ML LLM abilities like you provide you all the connectors to connect to different models and like generate embeddings. And then the neural search which provides you with all the pipelines both injection and as well as the search pipelines. Think of this as an abstraction on top of the KN&N plugin and MLS where like yeah it will provide you all the pipelines where you
can use it to connect uh where you can use it to connect to a model and generate damings out of it and ingest it. And the search pipelines is this the way like where uh you process a search query. Uh yeah this is like an API or a pipeline where you can use it to like generate the embeddings and do the search and this is where the
normalization is done and combines which uses the scores of both the subqueries and like yeah generates the final set of lists. Flow framework is an orchestration for this agentic pipeline. So which will break down all the steps that it needs to do and like execute the steps one by one. So this is how yeah it it does the shopping assistant thing like where it takes the user
query uses the query planning tool works with the LLM create the DSL run the search combines the scores uh using the pipeline and then like returns the results back and not just that it will also remember the memory uh it will also remember the conversation that that's happened as part of this thing and it will store it in the conversation index. So next time when the user
comes and asks a similar query. So show me something which is less than $50. Then it will understand the intent of the user uh by trying to pull the information from the conversation index and then it will try to show the same set of wireless headphones under $50. So let's also try to understand the architecture of building this agentic shopping search system using open search. So this
is the user query and once uh users connect its corresponding agent and sends a request it will try to use its corresponding tools like the list index and list index mapping tools and it will by itself like fetches the index and its corresponding mapping and it will try to uh process the request and uh sends the request to the LLM using the flow framework which is orchestration
tool and then the LLM works with the query planning tool to generate the DSL with the chain of thought and then sends the request to the neural search where it will run the subset of queries both semantic search query as well as a neural query with all the filters applied on top of it and that's how like it returns the results and like sends it back to
the user. So we can also see like there are like multiple different uh indices that are being created and like inedured as part of this flow like where also have the canon vector field for the actual embeddings and then if if if user is using any spark vector yeah it will also store in sparse vector index and then there is one for the conversation memory and then
the BM25 for the keyword search vector DB the brain and memory of so do this agentic search or for uh any uh a agent to work. So it needs a vector DB which is the heart and the core of uh it without that like it cannot function anything. So this is where it has the semantic memory, the working memory, the long-term memory and the absolute memory everything.
So without the vector DB it cannot function the vector without the vector DB the agent is completely. Okay. So now I'll pass it on to Nik Scott uh to talk about like scaling [clears throat] Uh thank you Naveen. It was a good presentation. We talked about a lot of features. So as you can see open search as a vector database has a lot of features. So it's
featurerich. The next question that comes in our mind is hey can it scale or how does vector as like open search as a vector database scale? So within these few slides we'll dive into it. So let's start with the first thing. How does open search evolved? Okay. And we are evolving towards as a trillion scale vector database. So open search started in 2019 2020 we are seeing
like at that time there were only 10 million tens of millions of vectors which was present and then slowly once we reach to 2021 we started to see 100 million or a billion vectors which is present in an index. Then in the last year we were seeing like a billion to 100 billion and now with this conversational memory and all these agents what we are seeing the
scale is moving toward trillion. Okay. And open search can scale up to trillion scale and let's look at how it scales. Okay. So the first thing that we can talk over here is as we discussed earlier the vector is nothing but an array of floats and every float is represented using four bytes and anytime I have to do a distance computation or any kind of computation I
have to bring that data into the memory. Okay. So this is where open search has a lot of quantization techniques which can take your original 32-bit vector okay which is in high dimension. It has three types of quantization techniques. product quantization, scalar quantization or binary quantization which can help you either reduce your dimensions or it can reduce the data per dimension. So if you look at something
like product quantization, it can reduce your vector dimension from 1024 to something like 256. If you look at scalar quantization, so it can take your four by dimension and convert it into a single bite or maybe two bytes if you're using a p16. And then comes the binary quantization where each dimension is getting represented in a single bit. So let's talk about a little bit binary quantization
over here. As I mentioned the 32x memory reduction. So you can see from a 32-bit float you're going to single bit per dimension. So you can see there's a 32x reduction. Okay. Now what we do in all these techniques is we take your original vectors. We have our uh algorithm for binary quantization and scalar quantization. We compress those vectors and store them. Okay. And now when we
have to do the search what we do is whenever a query comes in we'll first run through these compressed vectors. So we have an index where these compressed vectors are stored and all the structures which Namin talked about the HNSW structure or the IF structure are built on top of these compressed vectors. So you first go ahead and do the search on these compressed vectors and as
you know whenever you are doing compression and as we move from 32 bytes or sorry 32 bits to single bit there would be a precision loss. So how do we get out of that precision loss and make sure that we get a good accuracy and this is where open search also stores your full precision vectors on the disk. So whenever you are searching let's say you're searching
for 100 neighbors we go ahead behind the scene fetch over fetch the vectors. So basically if you're and fetch 500 neighbors for you and for all those 500 neighbors we go back to the disk retrieve those full precision vectors and rescore them. In this way we are able to maintain a precision of 90% and above. Okay. And once your vectors are rescored we send the response back.
So that's how a simple uh you can say with compression and rescoring we are able to achieve good recall and also able to reduce the memory footprint. And again since we are going to disk this would be a little bit higher latency. And that's where we are seeing like sub I would say 200 to 300 millconds on a good amount of data we are able to see.
Now one of the next thing to be a trillion scale vector database is hey you should be able to ingest these many documents. You don't want to ingest let's say trillions of vectors in like a month and that's where last year open search launched something called asating your vector injection using GPUs. Okay as we know that as the documents are coming to the open search domain or
open search cluster we first take it and divide it across the shards. This is all doing pure CPU computation. Okay. Now with the GPUs what we have done is you don't have to provision your GPUs within the cluster. Okay. Because GPUs are costly and you don't want to provision the provision a cluster or provision a domain which has all these GPUs running which are only required for
indexing. And that's where what we did was we actually built out a separate fleet or you can build out a separate fleet using open search where you can host all your GPUs and you can connect your open search cluster through to this GPU fleet. So the idea around this is whenever you are indexing the data and at a shard level you recognize that or the system recognize
that hey doing or doing a index build on this data node is computationally intensive. maybe I should put this or maybe I should send this job to the GPU cluster. Okay, so the idea would be on the shard we'll identify it. We put this data into a intermediate store maybe S3 or Azure blob store and then we let the remote vector index build fleet where you have
running GPUs know that hey there's this vectors uploaded can you build an index for me it builds the index uploads it back and your charts downloads this information and all these things happen as you are doing indexing you don't have to think about oh you will not see any CPU spikes because we are not building anything computationally or we are not doing any computationally intensive on This
way all your searches which is happening behind the scene works seamlessly. At the same time once you are done with your indexing and you don't need this fleet you can actually destroy it. Okay. It does not impact your cluster. There is no chatting happening within the cluster. Oh these nodes are coming up, these nodes are going down. It's all separate out and it can work. Now since
we have seen this architecture what are the performance that we're observing? So within our benchmarks what we saw is we are seeing 9.3x faster index builds. Okay, we are seeing reduce CPU utilization since nothing is happening I'll not say nothing most of the heavy lifting is happening on the GPU node. So we are seeing a reduction in the CPU utilization by 2.5x and then comes we are
seeing a cost reduction by 3.74. So earlier when you are able to when you are building an index in three days for millions or billions of vectors now you are able to build it within a day that's how you are able to save your cost okay you don't need to overprovision your cluster or you don't need to provision a cluster behind the scene or sorry ahead of
time to do a heavier injection okay so those are the performance numbers uh we are seeing with the GPUs moving on so we are doing all these things we have built all these features. One thing that comes in my mind is and I keep on asking our team is hey what is our goal for like why we are doing these things and one of the goal that
we are setting oursel for this year is we want faster indexing with minimal memory footprint during search okay so the idea is you should be able to ingest your documents faster and then when you are searching you should be able to search with minimal memory when I say minimal memory let's say if I'm doing one query rather than let's say using 100 MB to do search can
I do it with 10 MB okay so how you can reduce your per query memory footprint because the moment you reduce your per query memory footprint you don't need large machines which has huge RAM okay so that's our goal for this year and you see some of the features that we have built like the GPU based indexing is able to do the faster uh indexing we'll talk
about some of the next items that uh we have for search so few of the things that we are looking for is the vector reordering where we want to order our documents on the disk in such a fashion that we are able to retrieve it quickly. Okay, that's how we can if we don't have to fetch a lot of data from the disk and we don't need
to bring them in the RAM reducing the memory uh memory per query footprint. Then comes prefetching of vectors. So the main idea behind this thing is when you are doing computations can I behind the scene fetch the data so that when my CPUs are ready they are able to basically take that data and do the search. Similarly we have couple of more ideas uh around it. One
is yeah going beyond 32x compression. So keep a eye on the GitHub. Uh we'll be all soon releasing these features. This like you can take a photo of these. These are some of the blogs and the features that we have launched. uh where you can go ahead and read about it. GPU acceleration, dispace, vector search, uh other capabilities of vector database With this uh you can reach
out to us on forums, events, blogs, GitHub. You can take photo of this and uh you can reach out to us. We'll be happy to talk to you. We're looking for maintainers on And thank you. >> If you have more questions, please let us know. >> Wait till I get to you at the mic. Please, >> you mentioned that you are using the GPU only for ingestion
and index creation. Do you also plan to use it for like searches and other I know it's going to increase the cost. >> No, that's a very valid question and that that's a question we get a lot whenever we show GPU based indexing. So what we have seen is these days we are not seeing use cases where GPUs could become price performant for search but yes we
do have it in our plans. It's not I would say not in this year road map but yes uh we do have plans to integrate the search. It's more around hey are there use cases which needs this much amount of search QPS and where GPUs can become price performant because indexing is very heavy and that's where uh currently we have integrated GPUs but it's in our root
map. >> Thank you very much. The subject is complicated. Your talk was very good for me because I think you advertised it as vector operations for dummies. So I'm the dummy. question that I had throughout at the very beginning when you showed the vector bunch of numbers that might represent something a text some text or an image or something then you worked throughout the rest of the
talk with the vectors and there was proximity between the vectors where there was proximity between the original objects I guess type of thing what were those numbers how did it come to be that the numbers were indeed meaningfully reflective of the things that they represented. >> What happened in that step? >> Yeah, first of all, I'll answer that question for Victor search for dummies. I think in
the middle it it the title got changed a little bit because of unforeseen reasons. So I'm really sorry for that. But yeah, the idea over for this presentation was to provide you all the features switch capabilities of vectors engine in open search. Now let me answer your other question around that numbers. So those are all the things which are happening inside the model. Okay. So think of
it like this. Whenever you talk about let's say chart GPT is like X billion or trillion parameter model or whatn not. So the idea is whenever you have your text these are all represented with some weights. Okay. Or some numbers are exactly given to it and as you process through the model they convert a I would say a trimmed representation of that particular value. This is all
happening inside the model. This is something which I would say a GPT or embedding model do for you. Okay, nothing specifically. I'm not a scientist. So I cannot explain much more inside it. But the idea is you pass it through a model. It gives it's a black box and it gives you a bunch of numbers. It's putting you into a high dimensional space where if the semantic
meaning of the documents are same, they would be within that close proximity. Okay. So I uh I could have answered it better but I haven't done a PhD on this thing. So I hope that provides you a high level idea. >> Hey uh I'm very interested in the AI memory and conversational. I didn't know that was uh coming to open source. So I would like to hear
more about this. >> Yeah it's already available. You can read read the blogs around it. There is a proper I would say conversational chat experience that we have provided and there are like different flavors where you can have your long-term memory short-term memory for the agents that can be stored like something like you can keep your long-term memory in a disk optimized index versus you can keep
your short-term memory or something like HNSW based index which is fast to retrieve. >> hi. Hi, it's great talk. Uh, three >> Is this already production ready? The first question. >> Yes. >> Uh, for agentic search, uh, when you do the top heat collector, can you guys go back to LLM and generate single document or you have to do it later? You had a hybrid agent search.
And the third question when you do hybrid index uh all of them should be vector or you can have a classic string and other type of fields combined with vector basically. >> Okay. I don't think so I understand all the three questions properly but I'll try to answer it. So from a index standpoint within an index you can have both vectors and text present okay in one
single index. So if you want to call it as a hybrid index you can call it but that's an index for us. uh in terms of search, yes, you can do search on both a vector field and a text field together. Okay, that's where hybrid search comes into place. And I don't exactly understand your agentic search question. Oh, no, no, no, no, no. That's that's not uh
agentic search is more like you got your text okay and we send that text query to LLM and LLM uses the tools which we have provided within the open search let's say it needs to look for mappings it needs to look for hey what types of queries are supported to generate a particular full open search query and then rest is taken care by >> okay >> it
translates English into open search >> so when you're asking your question so it doesn't know which you're not providing the details of the text uh sorry providing the details of the index. So it needs to it uses its tools like list index tool index mapping tool and all those things to identify the corresponding index which maps to your query and sends that details to the LLM. So
in terms of LLM also it also provides you all the different uh what do you call like connectors which you can use to connect to any LLM that you want let's say cloud chargpt anything any open source or anything anything up to you and you can also it also provides you to tune those parameters to play with it uh the way you want that way like yeah
that corresponding LLM will take your input process it and like it will internally connect to it and like it will communicate saying that the the query uh query planning tool will tell it saying that like so I want to Uh I want a query DSL to run on open search like that's how like it communicates internally gets the query DSL and then like yeah it process it.
>> Uh how long does it take to stand up such a system for someone who knows what they're doing and is familiar with the docs? >> Yeah. So in terms of standing up let's say if you want to just stand up an open search cluster there's like a good enough document you just need to run one command and open search would stand up. Now because this is
an open source and we don't provide you something like hey here is a model that you can connect to. So all those things you need to bring up but that's where the a IML connectors that we were talking about can help you. It's like one I would say one rest API that you need to hit and it will start connecting to your model. >> So you should
you should be able to get up and running with within an hour or so. >> Yeah. Yeah. Yeah. >> And then the follow-up question is is this in the same category as Chromat or BG Vector? So PG vector I would say PG vector is also like PG vector is vector search over a SQL database. Okay, open search is more around okay uh you have your unstructured data
and you're doing search on top of it. I don't know PG vector does support all these a IML connectors and whatnot but you can think of open search as you have a core vector engine and then you have a lot of capabilities around it where you can use where you can use all these things to build out your application like a rag >> so it's closer to
a chroma then >> I haven't read about chrom in detail so sorry I cannot answer that >> so real quick I have two questions and then we got to go because the kio is happening First question is one of the very common use cases for search is when a new item enters in the index and somebody may have a search flagged. So you have the search for
shoes and so like a new product comes out, a new shoe comes out that hits the index. In traditional open search you can use percolators to inform the user, hey like you can have just a series of percolators set up for as an as something comes into the index and you say like oh okay these percolators fired. So we're going to send an alert to those users
that now that thing that they're looking for is in the search index. Do your agentic searches and KN&N searches support percolators? >> Uh, I would say no. >> Can you put that on the road map? >> Yep, for sure. >> And the other question I had was when you're creating those indices, those face indices so forth, you're using your LLMs to create the vectors. Um, there's there's
no real feedback system there, right? Like it may be that the vectors aren't as good of an approximation as you would want. And then over the course of time you could find out that hey actually this item is actually closer than what the LLM is indicating it is. Is there a mechanism for fine-tuning within the context of your search? Not necessarily recreating a new set of LLM
vectors but within the context of search saying oh we need a different sorting algorithm or different distance space. >> Okay so from a distance space standpoint no it's you cannot change distance space once it's indexed. But the other things that you can do is let's say in terms of HNSW you can increase the radius of the search which is using the EF search as a parameter. So
that is something that you can do. Okay. Now for IVF yes there are parameters where you can say that rather than looking at 10 clusters go ahead and start looking at more clusters. That's two parts. And the third thing which you're talking about hey are my results if my results are not great. There is a user insights uh I would say user behavior insight plug-in which is
present in open search and I if I'm not wrong it's open source okay that it can basically start looking at your results and based on your clicks on let's say on the website you can do a feedback loop saying that hey for this particular query I got these results but these were garbage okay it can start tuning your searches that's something that you can >> thank you
um big round of applause And we're about to have our closing You get to enjoy the bunch of people outside. She said, Hey folks, you made it. We're almost at the end of another scale. Um I know we've got a few more folks coming in, so I'll vamp a little bit. Um but uh thank you for coming out and spending your your week, your weekend with us.
Um, if you see folks in a scale jersey or a scale volunteer shirt, one of the tech shirts, some of them have been out here since Monday or Sunday even uh setting up this place. So, it's been a it's a whole week for us. So, you know, three, four days for you all depending on on what you came out for. Um, but thank you for making it
out. Thank you for being part of scale and and closing out the day with us here on on Sunday. Um, how many folks were at uh were at game night last night? So now, so all you that slept through the keynote in the morning are are here now and well rested after daylight savings time and hangovers and all this stuff. You were you were here all the
entire time, Gwen, is what I'm hearing. She was bright and early, always with energy. Um, but yeah, it's uh it's it's a pleasure to have you all out here. Again, this is our 23rd year running the event. Um, when we all started this out, when we were, you know, 18, 19, some of us maybe less, maybe a little bit more. Uh, I don't think this is what
we were envisioning scale becoming. Uh there was uh we were you know 200 of us in a small USC conference center uh cafeteria uh it was a lot of fun but you know now we're much bigger than that orders of magnitude and it's because of the um the community you have all helped built that we get together every year so thank you for um making this as
rewarding as it is and and if you run into volunteers again in the hall please thank them for the work they put into all of this. Um we've made a tradition of sort of welcoming back uh of inviting you know the pioneers that help build the the technology that our industry is built on to help close out um scale every year. Uh and um we have Dr.
Comr here to talk a little bit about uh his his work in the future and how it all fits together. Um and you know no different than any than any other year. He's a he's been a foundational part of uh of the computer science and the technology that helps us all do the work that we do every day. Um for those of you in the back, there's
plenty of seats up front here. Come come join us. Come come meet some friends. Um so with that, um I'll hand over to Doug. We'll have some Q&A afterwards. Um and uh and yeah, we're excited to have you all. Uh one last one last call. If you saw something you thought you could make better at scale, um like all open source projects, PRs are welcome. Please come
find us. Come talk to us. We'd love to help have you become a member of our team. Whether it's anything from the website to the mobile app to the signs in the hallway to the temperature of the water in the bathrooms, there's a way for you to help with something. Uh and uh we'd love we'd love to have you be part of the team. Many hands make
for for light work. Um but with that, I'll hand over to uh to Dr. Comr and we will uh we will learn about the internet. Thank you. So shortly after I agreed to uh to give this talk, I looked on the website and I saw I already was listed as talking about protocols. Now that makes a lot of sense, but I figure all of you guys already
know the protocols. There's no point in me giving a lecture on how TCP IP works. So I thought I would give a history lesson. By the way, if there's anybody here who doesn't know how the protocols work, I can recommend an excellent textbook. All right. So, let's start with how software gets distributed. Now, if you're a consumer, you hear about an app and you go to an
app store or or up pops a QR code and you scan it and you know that that sort of stuff. Um, and then when a consumer is faced with a really difficult thing, some message pops up that says you need to update your firmware, you know what they do? They call an expert. How many of you here are really lowcost IT support for your entire family? All
right, that's the consumer world. What about the rest of us? Well, you know, you go to some repo and download things. There are lots of repos around. I assume everybody here has used at least one of these, right? How many? How many of you have used all of these? Unbelievable. I thought I'd put one on here that was a little bit obscure so I find that somebody
hadn't done it. Well, let's talk about the other side. How you distribute software. You upload it to a repo to a file server maybe at your institution. And uh the other side of it you might give it to somebody on a flash drive. So that's now let's go back to before the internet. I got to Purdue as an assistant professor in 1976. Nobody had ever heard of
networking. In there was a faculty member when I started working on the internet project. A faculty member took me aside and said, "Doug, you're throwing away your career. He's very very serious. He's giving me fatherly advice. He's an older, you know, see networking will never be part of computer science. This is a computer science professor. That's the the way the world was back then. So, how about
transferring software? Well, it really wasn't the same. The only way to get from one site to another was on media. I had to choose a medium. The two most popular for mainframe sites were punch cards and mag tapes. I didn't put down paper tape because that was really a niche little thing for people doing process control computers, you know, not not the big mainframe stuff. So, punch
cards. I brought one just Just for fun, I actually had to go looking for it in the back of a drawer somewhere. I used to use them as bookmarks. They came in colors. In fact, see how you can do block lettering? We did block lettering to plaster them on the front of a box. Cards came in a And if you went to a computing center, they had
lots of empty boxes. You could pick one up and that's how you transport it around large programs. You got yourself a box and to make sure nobody else picked up your box, you put your name on the front of it. I also brought this card because it is so fantastic. It's it's an old if statement in forran. And on the card, it's printed the computer using this
program may not provide this program um to any other computer. It's sort of a an early copyright thing. so when when students picked up cards, this is what you've got. You got this little legal statement on here. I don't know. It was It was a fun world back then. Anyway, you got yourself a box of cards and it only held 2,000 cards. That means 2,000 lines of
code. How do you fit big programs into small boxes? There was a common thing that we were told. You don't want your program to be more than 2,000 lines of code. We had to take courses and writing a compiler. We had to take course, you know, and those got pretty big. How do you make it not go over 2,000 lines? Real simple, folks. You take out comments.
See, you think I'm kidding. All right. So, a box of cards, a 20,000line program that I that I wrote in grad school. I had to carry around a box. The boxes weighed 11 pounds. But cards weren't just for programmers. Lots of companies would put a punch card into a mailer and send it out and on the punch card be written, you know, put your address here and
you want my catalog or whatever it is and send it back in the return envelope. And they would always give this line which became a meme, do not fold, spindle or mutilate. There's a I think it was a TV show, maybe it was a movie where a kid asks his dad, "Can I borrow the car?" And and the dad says, "As long as you don't fold, spindle,
or mutilate it." It was part of the It was part of the culture. So, what were the other options? Well, the big option was mag tape. It was a sequential storage mechanism left over from ancient computers. By the way, did you know that early computers only had mag tape, a sequential medium for storage? They didn't have discs. So when you talk to to anybody from the 1940s
or 50s, that's the computing of the day. That's why if you took a theory course and learned about touring machines, touring machine is the mathematical abstraction of a computer that has a mag tape on it. It's not some fanciful crazy mathematics. It's really practical. Anyway, I brought a mag tape to show Isn't that nice? It weighs 2.2 pounds. All right. So, what can you fit on a
mag tape? Well, a lot more than you can fit in a box of cards. By the way, the technology for mag tapes evolved. In fact, I was sort of near the tail end. I didn't realize it. You know, when you're in school, you don't realize where you sit on the curve of things, but mag tapes had been around a long time and they evolved and they had
higher density. What's that mean? Well, there take the mag tape and you wrote sequentially on the tape. In the early days, it was seven track tape because byes were six bits. So, they had six little heads and bit writing along the tape. And in order to get more information on they worked hard to make things more precise and they could put make the bits smaller so they
get more bits per inch the m the magic measure of tape BPI bits per inch. It's actually bytes per inch for some reason. But it went up a lot. It went all the way up to 6250 bytes per inch. In addition, tape's got bigger same size reel and you can get more on it. How can we do that? Well, enter the uh the people that do material
sciences. They figured out how to make tapes thinner, but the magic was without stretching. You know, you're going to be pulling the tape through and you have to make sure that it doesn't stretch. How many weight? How how much data can you get here for 2.2 pounds? Well, look at this. 140 megabytes. You see this 140 megabytes. Um, if you really extended the block size to much
bigger, you could get 170. But that was about Anyway, how do you read and write tapes? Well, you had to have one of these big monster machines. They did get smaller, especially in in the late years of mag tape. In the 1980s, as mag tape was fading out, technology got a lot better. By the way, only the upper part of these big boxes actually read and writes
tapes. So you have you put tape on a reel, you put an empty reel over here, you feed the tape through and wind it around and then it would move the tape and write on it. But it was really hard to do that without stretching the tape. So it's hard to turn wheels and get them right. So the reason these things are so big is down the
sides there all the way down to the bottom are vacuum tubes. They are tubes literally filled with a vacuum. They suck the tape down. So the tape comes off of the supply reel down through this tube that's sucking it down up across the the rewrite head down through the other tube and up to the that way. It's like a buffer. They can just keep turning the reel
to fill this thing with tape down there and then read right head would only have to move the tape from one side to the other. Okay, now we get to tar. It just um it just occurred to me last semester that students in my class have been using tar all their lives and they had absolutely no idea why it was named And when I said well it
was for a tape archive they had absolutely no idea what a tape was. So before tar we had programs to write files on tape and what they did was they wrote the contents of the file DD for example. Popular way to write a file on a tape was DD. Have you ever used DD? Take DD to mention a file. Mention a tape drive slash R&D whatever it
is and it would write the file. But Unix had changed the world completely. It changed programming in two ways. One we had hierarchical directories. Now multix had to introduce them a little earlier but nobody used multics. Unix came along and did it the right way. It was actually very practical and everybody was using hierarchical Furthermore, the Unix guys preached the gospel of compiling separate files for each
function. Put a little function in C file and compile it and get O file and another one and another one and then load them all together to get the program. So, lots of files, hierarchies. Tar said, "Here's what we're going to do. We're going to take an entire directory with all these files and a make file and all the stuff in it and put it together and
put it on tape. I used tar to put things on tape and to take them off tape for several years before we really had good internet access. And you know, and I'm talking really early days when it was all still experimental. So, we were building the internet. It wasn't It wasn't, you know, building a lot of stuff that we were doing ourselves wasn't reliable. It wasn't great.
And um finally I realized what this what the real benefit of tar is I can send the whole thing over the network. I can send a hierarchy over the network and reproduce it at the other So [snorts] do you ever wonder why there's a block size option in tar? I bet you never needed it. But if you were writing tapes, you had to write the right block
size because you couldn't send a tape with a giant block on it to a machine, some poor user out there that had a tiny machine without that much space. I'm talking kilobytes. You couldn't have a 64 kilob block on a tape because lots of people in those days only had 64 kilobytes of memory total. So you had to specify, you had to talk to people when they
said, "Please send me a tape." They would say, "Oh, and make sure the block size is less than give you a number 8K." But why do you want big blocks? Because tapes had an inner block gap between the blocks. There was a gap and it was a fixed size that let the hardware figure out where the end was. So if you had bigger blocks, got more stuff
on, but the other guy couldn't read it if you had a small machine. So you want the biggest block size that the other guy can read. Okay. How do we transfer across sites? Two options, hand carrier, postal mail. And both ways weight was really So I have computed for you the bytes per of a of the media that I talked about punch cards 14,545 bytes per pound.
Now that's actually a little bit wrong. It's characters per pound. And that assumes that I'm actually writing all 80 characters on a card. But if you were living back in those days, you were only supposed to write programs in 7 through 72. The end was reserved for sequence numbers 73 and the beginning was reserved for a label. So really didn't get that many, but I'll, you know,
I'll give them the benefit of the doubt. How about mag tapees per pound? By the way, just for fun, I took my 256 gigabyte flash drive, weighed it, and computed the modern bites per pound. So, even if you were charged to send these things through the mail, it wouldn't do you much harm. Okay, so two ways to send things, punch cards or bag tapes. Nobody wanted to
sh of cards. They're really heavy. Imagine 11 pounds per box and I give you four or five boxes. It's really gets heavy. Even mag tapes. I had I was an assistant professor. I had to uh send things to people. I had written a compiler. I had written a greater program that was very popular. Would you believe it? When I got to Purdue, when I was in grad
school, and when I got to Purdue, professors, even in computer science, computed grades by hand. They literally wrote them down. And this was in the days before cheap electronic calculators. They did columns of numbers and added them and divided by the I couldn't believe that. So, I wrote a program to compute grades. And suddenly people were were you know using it all over. They would get these
letters. Of course it was all letters in those days. They send me a letter saying I understand you have this wonderful greater program. I'd like a copy. I would have to ask them what size blocks they could have on their you know mag tape. And and I would also have to ask them can you pay the postage? In those days professors didn't get any allowance. These days
it's the other way around. They get thrown money when they when they sign up to be a professor. you get a pot of money to spend on things like that. Anyway, there was some good news. If you had a really small program, you didn't have to send a big bag tape. It was a lovely little tiny mailer tape. And they were called mailer tapes. And the post
office actually had little boxes that exactly fit these tapes. You slide them in. It was just so convenient. So, what did the internet do? We're going to go through a few things Architecture and systems, economics, blah blah blah. Internet architecture. Well, the phone companies owned networking in those days. In fact, AT&T used to about the network uppercase T the uppercase N network. It was there was only
one network in as far as they were concerned it was only one network in the world. It was the phone network. They said everything has to be in this one unified network. But the internet guys embraced heterogeneity. They said we're going to use routers to interconnect arbitrary types of networks, which doesn't sound like a big deal now, but if you were on the other side and everyone
who had ever studied networks had always learned about the AT&T approach, everybody believed that was the only way to do it. Internet guys did a great thing because that meant you could choose a network for every situation and you could incrementally upgrade your network hardware without affecting the rest of It sounds it sounds silly, but do you know how long it took AT&T to convert from dial
telephones to touch telephones? In order to make that work, they had to change everything. The whole network is all built into the network. So, what's it mean for you? Well, it means you can start with a lowcost internet connection and then upgrade. How about connected computers? The computer vendors in the 1970s and late60s 70s were building what they called computer networks. And every vendor had the following
idea. We'll make more money if the network only connects computers. our company. I know it's it's hard to put yourself back in that position, but imagine that you're working on a network project and I tell my colleagues, I tell other people, I'm working on this thing called the internet project and they want to know what is it? Well, it's computer network. Oh, which company? None of the
above. Uh, you can't do that. You have to have you have to know which kind of computer you're connecting in order to do it well. And certainly IBM and digital equipment they had all come up with this line. You have to be able to tell all about the computer to build the right network for it. Okay. So what do we get? Well, from the internet approach, we
get the opportunity to use whatever computer we want for whatever situation and big one replace or upgrade that computer maybe to a different model, a different vendor without changing anything in the So from the point of view of distributing software that means repository can start with a cheap small computer and upgrade and upgrade and upgrade and maybe they decide that Intel isn't good anymore and they can
arm no change to anything. Here's another thing that you probably take for granted in the uh in the early days of the internet research project. Very few of us had little Indian computers. I was really lucky. When I got out of grad school, the guys at Labs wanted to recruit me to work on vax units and um and I thought about it and thought about it and
I decided to go to Purdue, which in those days was ranked number seven in the world in computer science, very prestigious. And I thought they they also said, by the way, we're going to try and get better. So I thought that would be great. I I like to teach. So I went there and the GSell lab said, "Tell you what, if you get yourself a vax, we'll
give you vax Unix and you can play with it and work on it." So I went to the dean. I was an assistant professor and assistant professors never talk to deans. But I went to the dean and I made my pitch. told him this would revolutionize the department. If we were one of the first departments to get this brand new computer and run this brand new operating
system, Unix, it would be wonderful. And the dean surprised everybody in the department by coming up with the money to do that. Now, why is a surprise? because the cost of a vax computer was almost exactly 20 times my annual salary. you know computer we didn't have any money for facilities at all in the department. Well we did have money uh we had enough money to buy
a new typewriter for the for the secretary in the department every three years. Typewriter not or not. All right. So, I had a vax and almost all the other guys had big Indian machines. John Pastel had a tops two tops 20 which is which was the sort of high-end digital big Indian stuff. And um he got to write the So he [snorts] said well look IBM the
big digital machines almost every computer in the world and by the way IBM controlled a huge percentage of the computers and software worldwide at one point I think they were actually coming down a little bit at one point it was 95% and they had big Indian machines the the IBM 360 stuff was all big Indian so he picked big Indian and We're going to make all the
internet standards be Indian. That meant on his computer, on my computer. People would send me things and it wouldn't work on my computer. And I couldn't figure out why. And it was because this Indian stuff was not in the tool set that programmers understood. So not only do we have to come up with the idea of standardizing bite order. But we had to tell all the programmers
from now on when you're writing things that may be used on another computer and by the way almost everything gets used on another computer you better think about Indians. It was a real It was a real shock. Students in my class objected. It works on this computer. Why should I worry about what happens somewhere else? I think it was sort of a natural reaction. But so we
get Indianness from the internet awareness Indianness and building programs that work across different Indian orders. How about internet services? Well, the telephone phone company had been in business for a long time. Remember they owned the network and they said the only way to build services is to build it in the network. You can't trust people to put services outside the network. Phone company owned the network. They
put everything in the So what's the disadvantage of putting things in the network? Well, you know what call forwarding is? you know, where you have a phone and you can forward it to another phone. How long did it take the bright guys at Bell Labs and the and the bright workers at AT&T to roll out call forwarding as a new service? Years. Some people say more than
five, but at least five. five years from start to finish to to roll out something that seems pretty simple. That's because they had to go to every phone switch and rewrite the code in every phone because all the services were built into ESS number four, ESS number five. So the internet guy said, "We're going to put the services outside the network." What do we get? We get
wonderful results. In fact, it turned out better, I think, than anybody who was proposing this in the beginning actually realized. You can invent new services like for example the worldwide web. The internet was already up and running just fine and we had all sorts of things on it and then along came the worldwide web and we just not a problem. It's not done in the internet. You
don't have to change anything. You just add new computers running this new software and presto. Oh, same thing for GitHub. It's true. and you can change services at any time. The internet approach to economics. Now, none of us on the internet project, at least in the in the core group, were economists. We were all techies, we lobbyed very strongly to go against what the phone companies had
all time. The phone companies had always built based on time and distance. Whenever you made a call, they would figure out how distant it was in miles, not in wires. You know, the wires might go funny ways, but they would figure out miles and then they would time it and they would charge you for that much use of the network. The internet guys argued and argued for
a flat rate approach. what's the advantage? Well, here's the deal. Imagine that you want to uh go to google.com and you type it in your browser and it pipes up a little flash on the screen that says estimated cost of this connection 32 cents. Click here to approve. >> I couldn't imagine. We were doing TCP connections, not phone calls. They weren't phone calls that lasted for many
minutes. They were just TCP connections. Can you imagine having to having to approve that? Or worse yet, not having to approve it. And well, you probably already know what it would be like. It would just be like you've gone to the cloud and and the ISP sends you the cloud provider sends you a thing at the end of the month saying here's your network charge. You have
no idea what you were using, how many how many bytes you sent across the network, right? Imagine that you did that with that you were your home and every time every end of the month you would get this bill that says you've used this many network connections and here's the charge. But we had a reason to really push back against charging per TCP connection. What was it?
One day I visited Bell Labs a lot and and talked a lot to Brian Kernan of Seing and uh one time I was there and Brian said, "Well, I don't work for a phone company." Brian, did you did you resign from the lab? No, no, no. They had done a study and here was the study. Suppose we take AT&T, the whole Shebang, including Bell Labs, right? All
that money. Have you ever seen the size of Bell Labs? Huge organization. at Bell Lab sites, some one of the sites had millions of square feet of offices. Millions. Okay. So, all of that and they asked the question, how can we save money? And one of the things they explored was suppose we didn't have to do accounting and billing for each phone We didn't have to keep
track of that. We keep the phone stuff all the same, the maintenance, everything, but we don't do accounting and billing. Instead, we just take the total cost of running the network and divide it by the number of telephones in the in the country. The reason that they had started this was that phone bills had pushed past the average phone bill was pushed past $20 a month. Now
you're saying that it's nothing now, but go way back. $20 a month was a significant cost. What would it have been if they didn't do accounting and billing, rip out all that stuff? And by the way, I include in that uh the postage to send the monthly bills or you know, you had to send the monthly bill and get back the answer and cash the checks. over
$20 a month on average and it would less than 20 cents a month if they didn't do accounting and billing. So remember that when you think about why we pushed back on accounting for every TCP connection horrible. We would have been paying a fortune. So you can the internet guys for cheap lowcost internet service. How about the communication paradigm? The other thing the phone companies had was
asymmetric protocols. Everything was asymmetric. If you ever worked on phone stuff, you know about the demark, the boundary between them and us. They own things to the demark. You own on the other side of the demark. If you've uh if you haven't done anything with telecom, if you've if you've played with RS232 on a Raspberry Pi, maybe you've encountered the problem with the asymmetric protocol. RS232 says
there are two ends, the DCEN and the DTEN, and they send on different wires. This one sends on and it comes out on three. This one sends on three. Oh, it's really ugly. If you ever had to deal with it, you'd understand. So, internet guys said, "We're not going to have asymmetry." They've actually fought hard because everybody's reaction in those days was to design an asymmetric protocol.
They fought hard. And what do we get? Well, what we get out of no asymmetric protocols is that there are no application gateways and there's no protocol translation in the middle. Now, forgive me. I'll excuse that because you know that helped us a lot. You know why that helped us? Because the ISPs wanted to charge per computer at your house. And that just got around that. So
they were forced, it all came out in the open, they were forced to back down. And and what I really like is that later on after all the wars and and lawsuits over this stuff, once they had lost, they uh they sent them a nice little ad saying, "We've invented this new thing. We're going to give you a wireless we're doing you a favor. In any case,
what we get out of the internet communication paradigm is the ability to represent data in new ways. Doesn't matter. The network doesn't understand the data in the middle. Doesn't change it. So, we can do end encryption. You have to thank the internet guys for phone companies wanted to have the encryption go between you and them. They would handle it and then between them and the endpoint at
the other end. So, they wanted to have two forms of Here's a big one. You probably didn't everything changed with the internet. Now, go back before the internet, there companies building protocols and one of the standard ways that they built protocols was to use broadcast to find a server. IPv4 when it started had network broadcast as a big feature. I could send a packet to MIT. I
didn't know where Dave Clark's server was. I could just send a packet and broadcast it to MIT and it would get to every computer there and the right one would That's all deprecated now because after a while we realized it's a bad way to do business. Well, we didn't have to to think very hard. Go back and read the history of digital equipment corporation lat. Lat was
used between terminals and servers and it broadcast. And then when people started hooking up sites with satellite bridges here, Vital Link, they did satellite bridging back in the early days. As soon as they did that, broadcast swamped the satellite and swamped the other site. I'm going from my terminal to my computer here and it gets broadcast to every site. And oh, it was terrible. By the way,
did you ever look at the early version of Apple Talk? Apple Talk loved broadcasting. It's how they did hardware addresses. They didn't really build them into the computers. You would generate a random number, broadcast, and say, "Anybody using 54 today?" Oh, nobody. I I'll use 54 as my MAC address. we had a wonderful guy, Paul Mapendet, one of my friends who invented the domain name system and
we were all thinking about how to do name servers. Everybody was think we knew what the problem was and we knew we had to have it distributed and but Paul came along. He wrote up a solution, put it on the table and it was so much better than everything else that it won. hands down. It's one of those times when when you realize there are really really
smart people here. You know, they're doing wonderful things. But his major achievement wasn't just that he had a distributed name survey that worked great, but he had delegation of authority. That was the major change. Everybody had been thinking like the phone companies. Let's have a a group whether we want to call it Diana whether we want that owns names and it will be giving out names improving
them. Paul said what we're going to do is delegate authority. So you go to example.com some company and you give them the authority to do everything connected to example.com. You can have a computer named www.acample.com. By the way, have you noticed that I am following the internet standards here? When you give a talk, you are supposed to use example.com. Did you know that? You can find it.
It's it's written down. It's an RC. I just want to show you that I'm very very knowledgeable about the protocols. But once you can do that, you can have a computer named mail.agample.com. You can have and you can have subdomains of example.com. For example, if if.com is a big company and it's got subsidiaries, they cannot be assigned subdomain. Oh, by the way, um Paul doesn't like it
when anybody uses the term subdomains. Paul isn't here. So, this was a wonderful idea. It let us name computers for services. Now you don't have to figure out which computer at some site runs the web server. It's named www whatever their domain is. How about the invention of an end transport protocol? Now when I first got into the internet project I had to put IP over X25.
X25 was a CCIT ISO standard. It was used by banks and it followed what was then considered to be the absolute necessity. It did link by link positive acknowledgement with retransmission. So you have a bunch of switches along the line and when you're ready to send a packet, you divide up the packet into smaller pieces. You send it to the next hop and that would send an
acknowledgement. Got that one. Here's one. Got that one. Here's one. Got that one. Here's one. Get the packet together. Send it over the next hop. Here's a piece. Got it. Here's a piece. Got it. Here's a piece. Got it. And send it over the next hop. It was absolutely horrible. So, the internet guy said, "We're going to try and do endtoend transport protocols. Reliability end to end.
Everybody that I knew who had any kind of network background said that's impossible. You'll never do it. You can thank Dave Clark who spent at least a decade or more of his life figuring out how to do end to end reliability without going link by link. It's so much better now that it works. Everybody can see that it works well and it's so much faster than doing
all of this little back and forth across each Why did they think it was impossible? Because they said, "Look, we've built this stuff. We know how to do it. You the roundtrip time on this piece of hardware before you can do acknowledgement and retransmission. You have to know it on this link, on this link, on this link." If you try and do it end to end, you
have no idea what's in the middle. It's just big nest. So Dave said, "We're going to do adaptive retransmission, which completely solved the problem." It was one of those another one of those brilliant solutions. And by the way, it does flow control, end to end flow control. So a big powerful computer can send to a tiny little computer without overrunning it. It does estion control in the
middle of One of the uh one of the guys when I first got on the project, he said, "Have you met Dave Clark?" He said, "Yeah, I'm done thing with him." Oh. He said, "You should tell Dave Clark to give up. It'll never ever work. He's wasting his time." I didn't tell Dave. But then about 10 years later, I met this guy again and I said to
him, "Well, you remember what you told me that I should CL to give up. Look, it's it's working. He said, "Oh, well, I guess I made a mistake. The problem was a whole lot easier than I thought it was." So, you know, never back down. You can always you can always find a way to not back down for a bad mistake. quality of service? By the 1990s,
quality of service was considered to be the most important thing we had to do in the internet. We couldn't survive without quality of service. Have you uh heard the apology from Scott Shanker? Got to go listen. I think it's out on the web now. Um he gives an apology. He was one of the leading guys that said, "No, we got to have quality of service. We'll never
do video or audio over the internet. quality of service. So when you make a connection, you specify what kind of data rate you need, what kind of latency bound you need, and the network will adapt and figure out how to do a path that will give you exactly Way back in the early days, do best effort delivery. We're going to start with best effort delivery and see
what it gets us. And everything seemed to be going along fine until we started to get a lot of traffic. And there was a really nasty surprise. Something that no one really expected. See, there was a huge amount of literature about how to engineer networks. It was all done by the voice guys. And it turned out that voice networks when you take the aggregate of all the
phone calls come out very very smooth for whatever reason. You can think this is crazy but it's not. If you have a phone call it's nice and steady. It's always sending the same data rate all the time. 64 kilobits per second. And when you add them all together, it comes out nice and smooth. With traffic, it's not smooth at all. Packet traffic is incredibly bursty. And when
you take the aggregate, you get this giant bursts. So we saw the problem on the arponet. Arponet backbone had 56 kilobit per second links. It was slow. And as we started to add more and more sites and they were all using the the arbonet backbone, yuck, things could get really really bad. how do we solve the problem? First thing to understand is the relationship between utilization and
delay. Maybe you've seen this. It's the only equation that I really like out of queuing theory. The effect of delay depends on the utilization of packet switch network. Remember this doesn't apply to phone networks. This is packet switch. So utilization goes from 0 to one. What percentage of the line is being utilized. And you can see what the delay does. When the utilization reaches 50% delay is
double. But when it reaches 80% the delay is five times as big. In fact, it looks like this. I stopped it here at 80% because between 80 and 100% I'd have to extend this slide a lot more because the delay goes to infinity. Just go out a little bit more and see what I'm talking about. how handle this situation. First of all, TCP has wonderful congestion control.
You can thank Dave and Van Jacobson who did wonderful congestion control for TCP so that when things start to look bad, TCP backs off. Your TCP and your laptop will do the back off. the phone guy said, "You can't trust users outside the network. They won't behave well." And we've always had a few crazy people who said to themselves, "I can rewrite TCP and get more of
the bandwidth if I don't back off. When things go bad, I'll just keep sending stuff in and I'll get more." It doesn't actually work for them, but they've tried. Now we have agreed that it works well for all of us if all TCPs back off the new congestion control. we now have the 5080 rule that says if utilization on a link reaches 50% it's time to start
planning the upgrade. When it reaches 80% you're in ter terrible trouble because remember it's bursty. So this is average. This little curve shows average. But the verses can be really Okay. So what did I tell you? I told you We get universal communication out of the internet. We get an architecture to allow us to use arbitrary computers and change them. Service model that puts Flat rate economics
and then paradigm enables creation of new applications and and awareness of ending this so that we can write programs that talk across the internet and not even worry about that anymore. Client server paradigm as you know it. No more broadcasting the fine servers and then transport with congestion technologies that achieve high So here's my point. Everything you know and love came from all of this. Not from
just moving to electronic communication, but figuring out how to do it well and figuring out how to avoid everybody getting charged a fortune to do it. Okay, I'll stop there and take questions. There we go. Well, thank you very much for the for the tour through internet history. Um, sure we Lots of lots of folks I think in the room have probably touched on one or two
of those things over our long careers but not not not nearly all of them. So thanks if we've got some time we do a couple of questions in the audience. >> Raise your hands. I'll come to you. >> And in the meantime wanted to present you with a gift from the team as well for joining us here on on your Sunday. Um, yeah. Where did Phil go?
>> Okay. On on average, how long did it take to um um punch out those cards and um and and also those real the real to real like on the big did it take hours, did it take minutes, seconds? >> Okay. So, u how long does it take to do punch cards? The the key punches were actually very like typewriters, but they were harder to press and
they had to actually push a little thing through the through the card that made the hole and the chad would fall out the back. So you could not type even even a touch typist could not type anywhere near as fast as they could on an electric typewriter. So it was much slower. Um that meant by the way the good news is that meant everybody thought hard about
their program before they punched it. I'm serious that the trouble now is I don't know about you guys but I've seen students You have some uh new hires that do this. They sit down in front of the computer and start typing hoping that the the computer will tell them, "Oh, this is this line is wrong. Change it while they're doing it. Don't think, just, you know, try
it out." So, no CICD in punch card land. I you you mentioned that uh it doesn't matter what thinks about DNS because Paul's not here. About three years ago, Paul was standing in that exact spot telling us what he thinks about DNS and that we were all wrong. So, so there you go. Um, go ahead. >> Yeah. So, thank you for a great presentation. Um, can you
give us the best lessons learned that you had in your career? And a funny one, too. I'm sorry. Yeah. Whatever. Let's see. Best lessons learned. Um well, um let's let me start this way. I was the first guy in my family to get a college degree. I didn't know anything about college. I didn't know anything about grad school. My when I got accepted to grad school, I
told my parents and my father took me aside and said, "Does this mean you didn't get a degree? You have to you have to do more school? so for me it's been uh a wonderful experience because when I finally got to the internet project I felt at home. Now you can say okay I'm just a nerd but you know I was with PhDs and I got to
these guys. One of the things that's really inspiring is you need to prove a theorem you can pro theorem. I've written papers and proved theorems and published them. Other people on the internet project when they needed to prove a theorem proved a theorem. You need to build hardware. I've soldered wires together. We had we had the most incredible, you know, but I had learned all that outside
of class. I never took a double class. I never took I had learned that in my basement as a kid, figuring out how to solder things together. Um, and that's sort of what I found at the at the Internet Project. These guys didn't h they weren't siloed. They were fairly broad and they can go very very deep when they need to. Um, and so that's a you
know, if you if you go to a college, they'll tell you you can either be broad or you can be deep, but you can't be both. And it's really inspiring to meet people who don't have that urge to just do one narrow thing. And the other thing I'll say is um I never really thought much about the authority of the experts and that turned out to be
really really good. If I could do it, I did it and I didn't what then say what you could or couldn't do. When I uh when I started working on IP overx25, a well-known expert in networking told me that will never work. We're going to have we're we're not going to do tunneling. We're going to do protocol translation and we're going to beat the pants off of
you. This guy had published papers in networking. I hadn't published anything yet. I was brand new and it didn't intimidate me somehow. I just thought I can figure out in my head. I can see how this will work. I can't understand why he keeps saying it. And that I don't know, call it hubris, call it ego, I don't know what you want to call it, but I
didn't think I didn't think I was actually trying to put anybody down. I was just that's the only way I could see to solve the problem. And it did work. So, you know, there's a lot to be said for not listening to uh thank you for coming out here and to say the least it was an amazing You become an associate professor at uh Purdue of everything.
Why networking? Why was that the the road you go down? Pardon me. >> Oh, I'm uh to your right uh over here. >> What's the question? >> Uh of all the paths to go down, uh why networking? Why was that the one you >> of all of all the all the different paths you could go >> All the different paths you go down. Um I'll tell you
something. When I I was where I was at a college that didn't have computer science and while I was there, they got a digital computer. I was a math and physics major. I started playing with this thing. I asked the the head of the math and science department, the whole division. I said, "Can I use the computer?" And he gave me a key. You can use it
after hours. By the way, computers were rented in those days, not purchased. So, I got to go in and I taught myself all bunches of programming languages and I taught myself assembly language. I just loved it. It was just plain fun. And I uh I was getting married. I told my wife, I'm going to go to grad school in computer science. And um I had written some
programs. I' had written one that that was really hard to write. Took me a long time. I figured it out and later on I was asked to clean out some file cabinets and I found this contract. It was a consulting contract with IBM. It described this problem that I had solved. It was to a te it was exact exact they had done a really nice job of
documenting it. And then you flip it back and it says, "We contacted 28 experts at IBM, listed their names. They all said it's impossible on that size computer." So I'm having fun and I said to myself that moment in I was standing there cleaning out this stuff. Either I'm really really good or I looked at the price tag on this thing. It was twothirds of my father's
annual income. either I'm really really good or in computing you can get paid a lot of money for saying sorry I'm too dumb to So I went to grad school and it turned out that I was really really good. So I just kept following the path of fun success. >> I think we have time for one or two As an industry, we're learning we're relearning specification and
documentation. Um, what are some ways you think about before you write the specification documentation like some heruristics in your thinking? So, [clears throat] I have this funny thing about trying to visualize the whole sheben in my head. and here's a little thing that works for me. Doesn't work for everybody. But when I get a hard problem, I think about it and think about it and think about
it. I write down all the things that you have to have to solve this thing. You know, I gota I'm going to have to have an array of this and I'm going to have to have this. I'm gonna have to module to do this. And I start to think about it in my head. And then I go to sleep. And several times in my career, I have
awakened with the solution in my head. It's I told you about that program that I wrote that was, you know, the IBM guys couldn't do. I had thought about it and thought about it and thought about it. I woke up and it was it was as plain as day. I knew exactly how I had to build the software, what the data structures would be. Uh, and I
wrote as fast as I could. I always when I when this happens I always have this fear that it's going to just evaporate and in fact it hasn't so I just down um what can I say that's something I I think hard about the whole big problem sort of tradeoffs if you do this it's gonna you know cost more here but it's gonna this what kinds of
parameters do I want to have for each module? What kind of how do these things talk to each other over a network? How and I think about it until I get it sort of a picture in my head and then I start to write the basic specifications. >> Many years ago networking was the you know that that got you most excited, what you spent your career on.
Thinking about today and sort of what the types of things that this audience is working with, what where what what gets you most excited and where where would you like to see us focusing >> as the next generation here? >> So I uh I have this feeling that I got into computing at the right time. I got into computing I didn't realize it but it was still
a baby. I got into computer science at the right time. I didn't realize it. I just thought the professors were, I don't know, blind. You could do these courses so much better. You could do this kind of stuff. You could do that. And and so I was very successful in in making all those changes because it was very early days. One of the things that sort of
hits me now when I look around that computing, it's no longer early days. Even cloud is no longer early days. machine learning, you know how old it Machine learning has been around since before Comr, it's really old. And and just to just to make a list of every possible variant of let alone all the other applications and uses and So, I don't get excited by that if
there's a mountain to climb. I get excited when there's not much in front of you and you have to sort of create from scratch. >> Do I have any more questions in front of me? >> Thanks for a great talk, Dr. CR. Um, um, I enjoyed the trip down memory lane. um some of it before my memory. Um uh I wonder it's always easy to talk about
hindsight's 2020, but I wonder if you can talk a little bit about like the road not taken for example. I think about like multiccast like remember in the 90s like hey if we could just turn multiccast on on all the tier one routers then like we wouldn't need YouTube. we could all distribute media to everyone to infinite people as many as we wanted and it just never
happened because the people that had to turn it on in on the core routers were the exact same companies that were going to end up selling stuff. I wonder in that kind of sense if you think about like not necessarily mistakes that were made but the road not taken that maybe you wish we'd seen we' taken or maybe we got there the long way. >> Okay. So
I could give a whole lecture on multiccast and why Steve During was wrong to start with and why we should never have gone down that path. Um and it wasn't just the economics of it was that multiccast doesn't scale well. But anyway, what about things not done? I don't know. I don't have any big things that I wish we had we had really done. I do believe
the following. The guys that I worked with on the internet project are the smartest people I've ever worked with. I work with PhDs, right? I'm in I'm in the university around PhDs. These guys stand out among PhDs and they don't all have PhDs, but they stand out. They are really smart. I believe that if they had decided to do a connectionoriented we would have one fantastic connectionoriented
network because they were relentless. They were smart. They worked hard to figure out how to do things well. So, I'm not sure that there's any big change that we could have made that we didn't make. So I think we'll call it there. thank you Dr. Comr for joining us and and and and sharing with us. Thank you to all of you for joining us at scale. U
hope to see you here again next year. Uh and on your way out uh I see Thomas Cameron. Pat him on the back. say happy birthday uh and thank him for spending it with us rather than his family. Uh but uh thanks again everybody. Uh enjoy your travels home and uh thank you for coming up to scale again.