SCaLE

Room 107 Sunday Mar. 08 - SCaLE 23x

3:56:31 · 05 Mar 2026 – 08 Mar 2026 · YouTube

About this talk

This talk explores the evolution of programming languages from low-level languages like COBOL to modern higher-level languages, framed through the significant contributions of Admiral Grace Hopper. The speaker discusses Hopper's pioneering work in creating the first compiler, which transformed how programming could be approached, making it more accessible for a broader audience beyond just mathematically inclined individuals. He emphasizes Hopper's famous assertion that institutional resistance to change, exemplified by the phrase 'we've always done it this way,' is a persistent challenge in technology. The speaker reflects on the current landscape of AI and coding, drawing parallels with historical innovations, and stresses the importance of maintaining rigorous engineering practices despite the advent of AI in software development. He argues that while AI has the potential to simplify coding, it still requires careful oversight and planning from engineers.

Full transcript

Yeah. Great. So, we're gonna we're going to self-s serve. I my uh my hobby that's quote unquote outside of computing is technical theater. So, I hope I can turn a microphone on and make it work. That'd be ideal. Although, I do do lighting, not sound. I feel like if anybody does technical theater, live sound is much much harder. Lighting is much more like a script that you

run. You get it all the work done ahead of time and then you just hit the button. Um whereas live sound is is very hard. Uh but I can I am capable of turning on a piece of equipment. So we got that working. Uh yeah. So thanks so much uh for coming everyone. Uh today I want to talk about from cobalt to cursor which is a little

odd because I don't work for Anthropic and I work for a company that kind of competes with cursor but that's okay. The the alliteration works really well. Um and I've got the slides here if you want them. Uh they're also posted on the scale website. Um but you can go to brennan.fyi/ccobalt uh if you want the slides. Um, and yeah, we're going to we're going to just

kind of kind of walk through a little bit of a history of how we got to where we are. Uh, and we're going to talk a lot about um, Amazing Grace, Admiral Grace Hopper. Um, and so that kind of motivates this whole talk of bringing uh, bringing a language first view to uh, what we do as software developers. Uh, and I think it's a really interesting parallel

to where we are today. Um, but I think we've got to kind of lay that groundwork first. uh as um those that may or may not know all of the amazing things that Amazing Grace uh contributed to our field. Uh we'll walk kind of through that and then like what that first revolution that she kind of led uh looked like in terms of shifting from uh very

low-level arithmetic only programming of uh computers to kind of a higher level language. Uh, and then we'll talk about her least favorite phrase in the English language, which is the way we've always done it. Uh, and how there's always institutional resistance to change. Uh, and I think we're seeing a lot of that currently in in software. And then we'll actually dive in and say, okay, so what

now? What are like the practical implications of this, Brendan? It's a great looking slide deck, but like what do we do? Okay, so first let's talk about Admiral uh Hopper. I like to call her Admiral Hopper because a that was her official title. Also, I'm from Annapolis, Maryland. So, I'm a Navy uh fan, born and bred. It's kind of weird being from Annapolis when you didn't go

to the Naval Academy because the Naval Academy becomes my like hometown college, right? Which like everybody's hometown college. You like get very, you know, pumped up about. But then if I'm wearing like Naval Academy stuff, you all can see me. People kind of give me like a [laughter] like, no, no, no, no. I'm just I was from Annapolis. I swear I played uh backup nose tackle for

the football team. No. No. Um yeah. And so I we I I alluded to it earlier, but Admiral Hopper's uh famously quoted as saying that the most dangerous phrase in the English language is we've always done it that way. Um and I think that this is this is really true. And if you've ever heard any of the talks she's given, if you haven't, I highly recommend. There's

a couple of great YouTube talks um that have been recorded of her her giving um lectures on her very different way of looking at the world. Um and I think you will not be surprised if you've if you've ever listened to those that this was a kind of something she focused on a lot. Um, and I think that, you know, it's important because revolutions or even just

like change in general always face this kind of reaction, right? This, hey, we're kind of the status quo is working for us. Uh, you know, we've always done it that way. It's not possible to do something another way. Um, and you know, that's just kind of the natural human default, right, of hey, we've always kind of done it this way. It generally works. Why why mess with

it? Um, but I think we're going to talk a little bit about how she used this phrase to overcome a lot of resistance to change uh in her career and hopefully that can again form inform us what we can do today uh in resistance to that. Um and you know I think that's because disruptive change always looks very messy to incumbent players in anything right like this

is forget technology forget even engineering as a general practice just in again in human nature um disruptive change when you're the incumbent and you've been in an industry for a long time or you've worked the same way for a long time or you've thought of a religion for the same way a long time all of these things in human history uh disruptive changes look very messy to

the incumbent folks there. Um, and that's something that I think I've seen over my career. So, I I've been in software since 200 uh six. Uh, and I started off life in in healthc care software and I spent 10 years in in breast imaging and uh breast cancer detection software. So, a whole different world than what I'm in now. Um, but that eventually led me to spend

five years at a company called GitLab. um that's an open- source uh company uh saw that grow exponentially uh and then left that to go to a couple other startups and now I'm at a startup called Kilo Code which is you know a focused on uh an AI an all-in-one AI agentic engineering platform. Um and so that's very different than the worlds I've been in before. Um,

but I think that, you know, this this current phase we're in, if we look at Grace Hopper's, you know, 60-year career, uh, shows us how these this move towards language first interfaces is going to just make computers and technology and programming and all of these things that we do much more accessible to many more people. Um, and so that's a little bit about me. Let's talk a

little bit more about Grace. Um before before we dive into what she did in her career, uh one of my favorite stories is when she was 7 years old. Um she took a aart uh seven different alarm clocks to try and figure out how they worked. Uh her mom was pretty upset with her about that. Um it's uh after my own heart because it's actually going away

now. Um but I have this little scar here on my hand that's from taking apart a phone that wouldn't come apart. And so I hit it with a hammer and the plastic, old plastic of a phone and a hammer had, you know, an interesting reaction uh that left me literally scarred. Uh she then went on to Yale, got her PhD from Yale, uh taught at Vasser and

um what didn't apply but was asked to join the Navy when she was 37. Uh, and she was actually 15 pounds underweight for the Navy, but had to get like a special exemption uh to be uh um uh admitted as an officer to the Navy. Um, and then she retired uh from the Navy in 1966 and 6 months later or no, sorry, yeah, like two months later

was recalled for a six-month additional post-retirement assignment that turned into lasting from 66 to 73. Uh and during that time is when she was promoted to uh admiral And the the first thing she worked on for the Navy was the Mark1 computer which was the first computer the Department of Defense ever had. Um and you know that was a computer that was programmed just in binary uh

code. I think the Mark1 was with punch cards. She would go on to use uh be very famous for using the UNIVAC which was actually programmed with magnetic tape. Um, and then very famously, if you don't know, this this um picture here on the top uh is a log book from the Mark1 computer, um, which is actually in the Smithsonian now. Um, because that little kind of

yellowing area there is a moth that was, uh, blocking a contact in the computer. like people like to say, oh, that's that's where the term bug came from. It's not. We were using the term bug in engineering before that, but uh but Grace actually wrote first case of an actual bug in a computer um because the uh the actual the problem with that they were experiencing was

this moth um breaking contact. Um and so at that time right again these these machines were all programmed um just in straight binary through a punch card or like like I said or through magnetic tape and there was no such thing as a programmer, right? There wasn't such a thing, right? That's a role that didn't exist. Um because there was no portable like readable programming language, right?

There was not a concept of a programming language. Um and she actually ended up writing the first um log book for the UNIVAC because that's all she got handed when she worked on the Mark 1 was just a log book for this is how the computer works electrically and good luck. And and so that kind of moves us to this this first revolution that she really she

really uh was was an advocate for. Um and it was both her technical and social mechanics of change. Right? So in 1952, you know, we talked about it's a nice stack of punch cards here. Um again, technically the UNIVAC was published on magnetic tape, but that doesn't make for as good of an illustration. Um so you can see you can see the punch cards here, right? And

for those who haven't, I mean, I haven't actually ever programmed a computer with punch cards either. Uh, but you can see we've drawn like a line, a diagonal line so that if they get out of order, right, I can put them back in the right order because if they're in the wrong order, you're going to have a bad day. Um, and then there's also like other revisions

that have happened, right? So like in the middle here, oh, I can I can illustrate with this probably. Like in the middle here, right? Oh, wow. That's way too big. Um, there's there's additional lines. Oh, the eraser doesn't work. Okay. Well, now it's all just a mess. Um, yeah, I thought there'd be an eraser, but there's not. Okay, this is great. I love that for us. Okay.

Oh, and it's still there. It's good. That's a fantastic fantastic feature. But the these these uh these like other lines, smaller lines here, right, are revisions where like I had to repunch all of those cards and then put another revision in. Um, so if you think Git was hard Oh, um, that's real hard. Oh, I hit the wrong. Now I' I've figured out how to kill it.

Okay, hold on. I have a plan. Just reset it up. Cool. I think I'm this way. Yeah. Mouse, maybe. Oh, yeah. All right. Now we're back. that that was the kind of the state-of-the-art, right? Was was programming like that. Um, and Admiral Hopper had this idea of like what if a machine could take human instructions and then translate that into machine code? Um, like what if you

could write something that was more like math or more like even English words and then the computer could have a program that then figures out how to make that into binary code itself. Right? This this concept of a compiler um was what she came up with and she um you know was told very quickly you can't do that because computers can't understand English. They only understand math.

There's no possible way it'll ever happen. Uh and so what does she do? Well, she builds a compiler, right? uh and she takes uh that she builds it's called the A0 compiler. It was the first compiler ever made um to take uh at the time what would become flowatic and then become cobalt um and turn it from you know these English words into uh compiled machine code.

Uh and it's kind of this you know it's kind of interesting because that that leads to this whole world that we're in today right where we have this concept of you know machine code that we're writing directly or maybe even assembly code which is basically just machine code but in a different form uh to cobalt right that that Admiral Hopper ends up inventing uh to you know

now this you know what are we doing now right we've got I've got I'm writing nextJS and I've got a node modules folder that's bigger than the whole project um And now we can have even you know LLMs do it. But the reality is these kind of filter into three areas right machine code and assembly are very similar. Like I said they're they're really you know one

is a binary representation one is just a pneumonic representation of the same binary. But really everything after that sometimes today we have these arguments about what's a highle language and how lowle is a language. But really everything that's to the right is a high level language in the sense that it's it's an abstraction built on top of um on top of the uh the actual assembly or

machine code that's going to be written. And again the problem she faced at this at this point was um she was told it's not possible. She actually wrote the compiler and no one used it besides her for three years. Um so she she was the only programmer for like three years. uh because she was the only one with the compiler that that decided to use it. Um

and you know she was told again over and over again that it wouldn't work. She and she said she kept pushing right she went past that um a zero compiler ended up making multiple versions of it and ended up making cobalt from that. Um and people said but then Grace like everyone's going to be able to program and and she said exactly that's that's the plan. Um

she wanted you know programming to be something that could happen closer to natural language not just mathematical notation right eventually there was another programming language uh called forran which was much more focused on like the mathematical notation right the formula translation of the math that we're doing in the inside the computer versus cobalt which was this um common business language that was say hey how can we

turn that into English instead not just symbols right words that like business people could read and understand. Um, and I think that we see this shift over and over again, right? She she said, you know, what I was after was an English language programming to bring a whole other group of people into being able to use a computer easily. Like she continued to call for user friendly

languages. And she was even quoted in in a uh 60 Minutes interview in 1983 saying that was her like number one goal uh as she started building this. Um, and it's it's crazy to think that, you know, user friendly was a word that she was thinking of in the ' 50s, um, when, you know, the computer is in in the entire uh, taking up an entire room.

Um, but, uh, yeah, and I think this is something that we see kind of like shifting all the time, right? Um if we you know she had this problem of going from you know even assembly languages were at first kind of decrieded because they were abstracting away the zeros and ones uh to this pneumonics then you know we move from assembly to cobalt and other highle programming

languages and we get a lot of like this isn't real programming um you know we see python and we get to make fun of it a lot and call it not real programming because like you know it's not object you know it doesn't doesn't conform to the way we think of programming. And and now I think we are seeing the same thing again where you know we're

saying oh well if you're just prompting an AI agent well that's not real programming uh either right and so we've seen this kind of as we move up the stack and abstract away um the the incumbents kind of say well it's the way we've always done it and uh this is the best I could do but she also would keep a clock on her wall that ran

backwards um so whenever you came into her office there was a clock running backwards which I just flipped it and said the numbers are backwards. So, sorry, but I had to get it to go counterclockwise. Um, and it was it was a reminder when everybody came into our office that hey, like we've always done it that way is not going to be accepted in this office. Um,

and so let's talk a little bit about these these modern parallels um that we're seeing in the in the age of AI coding, right? So, here's uh two different uh tweet two different uh internet things. One is Andrew Carpathy introducing the term vibe coding. Uh probably one of the most famous tweets uh in in our industry in the last couple of years. It's got I think I

took this picture yesterday six and a 6.8 million views. Um and then there's uh our you know Reddit CS career questions. Um you know AI slop code is hiding all of the all of the uh the real code. Um and it used to be obvious when there was a problem. Oops. Sorry. Did it again. um used to be obvious there was a problem, but now it's all

AI slop, right? And so those are kind of the two different you we're seen as these like there's only two different ways of thinking about it. It's either slop or it's, you know, all vibes and don't ever read the code. My goal, if I can get my computer to function, is to tell you why it's not either of those things. Let's just do this less complicated. [snorts]

Um and so why why is there this disparity? like what what is what is causing this massive disparity between you know people that are theoretically in the same position as software engineers um but seeing the world of AI coding so differently um well one is it can be c like AI is very good at being confidently wrong right um it is never going to tell you you're

wrong it's always going to tell you you're you're absolutely right and uh it's you know I if you just trust it innately You're going to get burned by it, right? You're going to have major issues, right? [laughter] I don't know what happened. I vibe coded this whole presentation and now now look at me. >> I started talking bad about LLMs and they got me. I don't know

what happened. I'm still connected to the HDMI. That's [snorts] my goal. >> Well, so I unplugged it. It didn't react. >> Yeah, I think so. Maybe. >> I think there's a signal. I hit the table. I don't know what I did. [snorts] >> No, >> wait a second. >> Right here. >> It's solidly in there. Is your computer maybe stop sending a >> So, it thinks it's

sending a signal. >> Okay. Well, I'll keep talking. Hopefully, we can get it back for at least the demo, maybe. Okay. I'll keep talking. We don't need the slides so much for me to talk, but uh hopefully we can get it back for the demo. Um so yeah, we were where were where were we? We were talking about how it can be confidently wrong, right? And I

think this is something that's interesting because it's not that different than you know a junior developer, right? Who's very well read and maybe never really um implemented code in production before and can be very confidently wrong, right? And just like you wouldn't, you know, necessarily trust uh that junior developer to, you know, go and and just push code straight to production, you shouldn't do the same for

AI, right? Because you'd have to trust but verify, right? Read the code, understand the changes are being made, and we shouldn't hold our AI agents to a different standard. Uh it should be the same standard we hold any contributor to. Uh and then next would be that it's like you know the next argument against would say oh well we're making engineering a commodity if we make engineering

a commodity then we're all out of a job right uh and again I think hopefully we've seen through this these multiple abstraction layers that have come into uh software engineering throughout time um you know it didn't reduce the need for programmers uh in any of those abstraction layers it increased the need right um every time we've built uh you know solved part of the stack and brought

ourselves up to a higher level, we've been able to do much more incredible things, right? Like we wouldn't probably have uh you know, half the technology we have today if we still had to be, you know, programming everything through a punch guard. Uh so I think I think it the re the the revolution again can be scary when it's happening, but I don't think that we've yet

seen and I could be wrong. You could come back to me in five years and tell me how wrong I was. Uh but I don't think we've yet seen a revolution where we've been able to make a complete commodity out of out of engineering just by moving up the stack. Uh and then the the next argument would probably be that you can't trust it for anything important.

um you know if you you know like if you're if you're going to have AI code something and it's production grade or it's it's really important you know what is what is the the um you know the reality of being able to then ship that into production can I really trust it to ship it in production um and I think again if we think through this this

history of um through you know moving up the stack it's like when you run something through your compiler, do you read the machine code that comes out of that today? Right? Like do you go and inspect that and and read it? Um do you look at like the, you know, just in time output of the runtime that you're running and really inspect the V8 engine, the JavaScript

that you're running and look at how it's actually interpreting your code. Um and then again to like the silly example I made um of myself and bringing in lots of known modules, it's like we're shipping lots of code we haven't read. if it's in, you know, a dependency, right? We're not reading every every line of code in a dependency. Uh, and so I think the question is

not necessarily like, can I or can I not trust this? It's not kind of that that black and white of a question. Uh, the question is like where do I draw the line? What what can I trust? Um, and you know, that line's always moved, but that line is always the important thing um that that we had to to um to understand. and then okay so uh

I'm I'm just trying to buy some time for the demo. Um so again another thing that I wanted to talk about um is when Admiral Grace was uh interviewed in 1983 uh on 60 minutes which again is available on YouTube if you want to take a look at it. Um it was interesting because they said to her like oh well you've seen all this advancement since the

50s. Um, and you know, do you think we're we're really, you know, is the computer revolution over or, you know, have we have we reached the peak? And her statement in ' 83 was that um, no, we're at the Model T. Um, yeah, go ahead. Go ahead. Do what you need to do. Um, and so it's it's interesting to me that her her assessment of of the

time between That was the button I hit, didn't I? >> Thank you. Hey, thank you. Now we know why I did I hit the table. We know um yeah so it's interesting to me that her assessment of the progress made from you know the early 50s when she was involved to 1983 when she was of course retired was like okay that got us to the model T

right um and that was surprising to Mory Schaefer who was interviewing her u but I don't think it's surprising any of us that have now seen what's happened from 1983 until now. Uh, and I think that that's interesting to to again think about in the days in in the age of AI where you know the models are the worst that they're ever going to be like today.

Um, the our understanding of how to work with it is the worst it's ever going to be today. Um, and we've we've seen that all of these these changes technology adjust over time and we we learn how to adapt uh and make it work Great. So just in time for the for the demos the or the uh so what now Brendan right so how do we if

if I'm you know sitting in the audience I'm a software engineer I've been a software engineer for a long time how can I work with AI in a way that you know doesn't make me um you know doesn't have those problems that we talked about which is what can I trust to ship uh how can I avoid it being confidently wrong uh how can I make sure

that it's actually a partner in engineering and not not a deter not uh making it harder for me. Um and so there's a couple of things. First of all, I think you have to be very explicit, right? Like just again if you were working with a junior engineer who's never seen your codebase before and you just tell them, "Hey, there's a bug with the button on the

table. You hit it and no one knows what you did." Um you know that that doesn't give them very much context. If instead you say, "Well, the button's attached to uh some sort of HDMI splitter below, and if you hit the button, it's not obvious that you're now looking at something else, right?" Like that that now gives them enough information to solve the problem, right? Um and

so being very explicit, right? Um that, you know, the ongoing joke on the internet is like make no mistakes. I mean, that's helpful. It's a good idea, but you know, that does not a good prompt make. Um and then also, I think providing context. So we've seen a lot written about context engineering um you know providing to the the the large language models the patterns that exist

in the codebase already examples of what you want to be um and building kind of good habits around keeping the context to be just about the thing that you're working on uh and how you want it worked how you want it to be solved is really critical. Um, and then setting the right constraints, right? So, we have to you have to it's like any new tool, you

have to build an intuition for how does this tool going to work with me? How do I when do I need to stop, start, restart, or push it further? Um, and so that kind of brings me to like this general concept of all of this, right? From from Dr. from Admiral Grace's uh you know concept of cobalt to today you know software engineering's always been about expressing

the right intent of the engineer right um and that's been true since you know for any programming language and it's going to continue to be true if not even more true in the age of AI because expressing very clear intent and having a clear fundamental understanding of what that intent is is going to be very critical and so how do we actually apply that in practice Um

there's been a lot of again a lot of different uh frameworks out there. Um I don't claim to have the the the like monopoly on frameworks, but the one that I keep coming back to is this RPI framework that that folks have been writing about called research, plan, and implement. Um and so this kind of solves those same three problems we were talking about before, right? If

I first research and spend a lot of time understanding understanding the scope of the problem uh and where you know the problem's going to get fixed uh I can then plan a lot better right I can then build those patterns and examples in to my plan and then when I'm implementing it helps me build that intuition and I've already got a plan in place um that that's

going to be successful. So that's where I want that's what I kind of want to demo here. Um oh well first first let's talk about research specifically. The idea of research is understanding the system that's in place. Uh finding the right structure for what we're going to be building and then you know understanding hey I as a human I'm going to going to ask it to do

some research but then I'm going to review the results. Uh so the example that I'm going to give um in this demo uh is something I actually encountered yesterday when I was on my way to the airport here. Um, has anyone heard of OpenClaw in the room? Okay, a few people have heard of OpenClaw. Uh, I've been playing with it a lot. Uh, Kilo even has a

product out now for hosting it. Um, but one thing I wanted to do is I built um a thing called Pinchbench. So, the idea is like what if how can we benchmark this like most wild of AI agents, Open Claw. Um, and so I created a a a benchmark that has a bunch of tasks uh that try to mirror like what people are actually using OpenClaw for

and then let's run that against all the you know various models and get a score. And mostly it was for fun, but then I'm literally on my way to BWI. I'm from the East Coast. Didn't mention that. And uh this tweet comes out yesterday at uh yeah 7 am our time here, but I was I was in the airport in Baltimore. Uh, and it's Peter, uh, the

creator of OpenClaw, and he tweets, "That's a really interesting benchmark. I just found, Pinchbench." Um, yeah. So, uh, this caused, uh, quite the spike in traffic. Uh, he also DM'd me and said, now that he works for OpenAI, he DM'd me and said, "Hey, why is, uh, GPT 5.4 not in there yet?" And I was like, "Okay, yep, I'll fix that." Uh, so so I had this

need to go and implement because I had I had just run the benchmark myself. I had built some kind of virtual machines, a little bit of orchestration. I hadn't really planned for this thing to take off, right? It was kind of just a fun side project. Um, but like I said, it it did take off when someone uh, you know, when Peter took notice. So now I

had to kind of say, okay, how can I automate this whole system and bring it, you know, like like orchestrate it? Um, and and so let's let's talk about how I I kind of did that. Um, whoops, sorry. if we're apply Oh, and let let's just do a little demo here of Pinchbench. So uh again, the idea is it's relatively simple, right? It's it's inspired by the

simplest uh benchmarks I could find out there. Um and you know, each one of these these is a set of tasks that get judged. Some of them are automated judged, some of them are LLM judged. Um, suffice it to say it's like a lot of Python uh getting run to orchestrate uh uh open claw against it and you get a lot of like cool graphs. Um, right

like cost versus uh accuracy. Um, love it. It's great. U but now I need to automate it because it had just been my laptop running a bunch of vulture vulture in instances itself. Uh, so step one would be to research, right? And so uh you know one key I think is to think about these modes differently. And so that it happens to be Kilo there's lots of

other uh iterations of this but in Kilo we think of it through these different modes. And we have uh code mode which is like the typical demo you see of like spitting out a bunch of code at you. But that's not what I want to do. I want to plan this out better first. Uh so I'm going to go to ask mode and I'm going to say

I want to create a project to automate the generation of the new benchmarks. I'm in the benchmark uh repo here, right? This is where all the tasks live, right? Like the weather task, um like what the expected behavior should be, how it's supposed to work. Um and I'm asking the agent to say uh you know based on this, how would I what would I need to understand

in order to create an orchestration layer on top of this, right? So this is where I want to research it. And we can see the agent uh here in ask mode actually cannot write any files. All it can do is read. It's read only and okay so it created this uh I have a uh understanding now um let's just pull this into a markdown oh still writing

it almost done uh live demo it's going to be with me wasn't with me earlier with the projector but it's going to be with me this time all right so it can't write this file but I can I'm still thinking >> that's a good question yeah I don't know Let me I wonder if I can switch to light mode and it'll help. There we go. Okay. Uh

that'll help a little bit. Okay, great. So, it created this this file. Um Oops, that's not what I want to copy. I want to copy this. >> I can. I'm trying to get it up in a preview so we can see it better. >> Yep. Just give me one second. Just copying all of it. Okay, I'm gonna copy all of it so we can see it a

lot bigger. >> Yeah. So again, it it's now um given me this entire uh concept of and I it didn't save it as markdown or it didn't it didn't come over as markdown quite right, but it has like a whole graph of how all of these things work together. Um all of the requirements that I would need to parse the markdown, all the things that I'm going

to need to create and manage agents, right? So this research in real in real life I would read and understand and make sure that we really um are you know all of the flowcharts and everything are exactly what I would expect based on my knowledge of the codebase. Um and I'm going to ask it to write it to a file because I think it'll be Oh man,

so big I can't even do it. try to be nice to our AI overlords and say please every now and again just in case it'll be easier for us to see the the diagrams if I can get it to write I can also go the the the the thing to always have in the live demo is a cooked turkey. So I can check out my branch cooked

And in that branch I have my benchmark automation guide. >> All right. So here we can see it's got these diagrams of what the input layer is going to look like, what the orchestration will look like, um what components we're going to need to build, [snorts] how task management works. Right? So again, we're I'm going to go through all of this as a human. uh and make

sure this research matches my my understanding of of the system and only after that you know very explicit time spent in research and in reading that am I going to then move on to plan right and so again I'm going to have the plan mode output the exact steps it's going to take include how we're going to test it be very explicit about how I want it

to build it and its output is going to again not code but be a human reviewed plan file. So if I go back to VS Code, I've got my guide and I can say in this new repository where I don't have anything yet. I can say I do have something because I'm on cooked turkey or I don't have anything yet. I can say you know hey I

want to plan the automated benchmark tool that can run live in the cloud come up with three to five different architectures so that we can talk through the implementation options. Uh this is another real key when working with AI is you can ask it to give you options and then that actually can help you um think through some of the issues there. U below is the detail

from a benchmark repository of how it works in one machine. So I'm going to put in that whole document we just got and this is actually going to take it much longer. So we are going to go right to the cooked turkey which is why I was on it. um because the plan file that comes out of there then is going to be a very detailed architecture

about what I'm going to need like what are the secrets that we're going to need uh in order to run it and then it's going to give me different architecture options right I could use modal.com so this is all Python so it it thought hey well moto modal runs Python well right what if you did that um and it shows you what an implementation would look like

gives me pros and Cons, right? Very easy to implement, no idle cost, uh con and vendor lock in, right? Like I'm going to be building against their API, estimated effort one to two days. Little does it know I'm going to make it do it after this talk um in the next 10 minutes. Uh AWS step functions, fly machines, io, you know, fly.io, um GitHub actions, and then

the nice thing is it actually gives me a comparison, right? So I can actually think through, you know, what is what are the the criteria that I need to think about before I execute this plan and how do each of the architectures um weigh into that. Now, of course, this still is coming out of the AI, right? So I'm still going to have to as a human

read this, understand it, make sure that I agree with its assumptions. Um but again, it gives me things that, you know, maybe I haven't thought about, right? Like um what about retry and error handling? Again, I wrote this code. There's going to be a lot of that, right? So like how is that going to work? Uh and so only after I again read through this plan and

choose from the architecture am I then going to go and say okay now I want you to go ahead and implement it. And we're not going to spend time on that from the demo perspective because of course again that's going to take a lot of time. But the key about doing those two steps ahead of time of the research and the planning is now when I go

to implement it's going to be very low context. Right? The context for that agent is only going to be those things that I really wanted to go do and I've already made a lot of the decisions for it. Uh rather than letting it kind of just decide by the roll of dice what decisions it's going to make. Right? If I again we saw the different architectures that

it came up with. If I use that same model and said implement it without a plan, uh then it would have just chosen one of those at random, right? Like it wouldn't it wouldn't have necessarily thought through those those uh different uh you know trade-offs. And that allows me then to review every change uh and commit frequency frequently. And this human review at that research and planning

stage is the high leverage of your time. Right? That's the engineering portion of the software engineering life cycle, not the you know actual uh typing every letter in or you know just like we would have to punch every card in before. Um and Dexory who works for a company called um Human Layer uh has this great quote that I love which is AI can replace can't replace

cannot replace your thinking. It only amplifies the thinking you've already done or the lack thereof. Right? Um, so if you haven't done a lot of thinking beforehand, it will be great at amplifying that. But if you've done a lot of thinking beforehand, uh, you you can actually amplify that very quickly. Uh, and so this is, uh, that's a link to the YouTube video where he he talks

about this and he talks about his he has tw he has I had three steps for you. He's got 12 factor AI agents for you if you if you want to really dive dive deep. right I think because of that you know these old disciplines these disciplines of engineering and making strong decisions and thinking through architecture those are actually going to matter more in the age of

AI not less right um if writing code is cheap now which I think it is right we can all kind of agree with that every demo of something writing code really quickly has shown us that uh software is not right like like engineer software is not cheap and it's not going to be anytime soon because the engineering habits that we have around like it being an expensive

thing to write the code are gone. You can now just just build things, right? Um and so with that constraint gone, it doesn't mean that all of the constraints of engineering are gone. It just means that that those only come to matter even more. And then yeah so just some some uh three further reading if you if you want to um path.o.ai is a pragmatic guide for

engineers and leaders that I'm writing uh and open sourcing has a lot of the links that I used in the research for this talk there. Uh if you haven't looked at agents.mmd yet I'd highly recommend it. Finally, all of us crazy AI vendors are ho hopefully standardizing around a way to instruct and give the agents the right context to begin with and that's going to be through

agents.mmd. And then lastly, Simon Williamson, if you haven't read his thoughts on agentic engineering, uh I think he's probably the the best writer on it um that that exists today. Then I can be found on the internet anywhere at ori crew. Um, and yeah, I would just I would just close by saying like I I think that we owe um Amazing Grace a lot. Uh, I've got

on my t-shirt here, uh, Ada and Gene and Grace and Margaret, which is um, the Grace is Grace Hopper, Margaret to Margaret Hamilton, Ada Love Lace, obviously. Um, you know, she didn't write the compiler so that like we could like stop learning. Uh, she wrote the compiler so that more people could learn. um she wrote it so that we could learn different things, so that we could

do higher level things, so we could do more valuable things. Uh and so I think if we think about uh AI agents as another iteration of that, uh you know, it changes it can change your perspective on on how to use them and and and where we're at, right? And if we've got this kind of progression um towards natural language, I think that's exactly exactly what Admiral

uh Hopper was going for. Uh and so yeah, thank you very much. You can find me anywhere on the internet crew and uh happy to take any questions. I hit the button again. Happy to take any >> Yes. >> So what did the history going? >> Yeah. So, so that's a great question. Um, first of all, there's a really great map. I will tweet it uh after

the talk. Um, where someone has mapped out the full lineage of every programming language with each other. It's like 19 pages wide. And actually, if you print it and post it on your wall, they then post your picture of it on your wall back. Um, and so the interesting thing there is I think Forran technically predates Cobalt. Um and she didn't she really thought that it was

too mathematical, right? She wanted it to she wanted computing to be accessible to more than just the the mathematically gifted, right? PhDs. Um and so Cobalt is a common operating business language something don't quote operating language, right? And so it was really about expressing this you know business needs uh which for her at the time was the Navy but then later other businesses um and so it

was about like expanding using English to expand who can understand what a program is doing right uh and so forrans obviously much more formula based uh and so she really it was really about natural language uh and bringing natural language to it does that make does that answer sorry I don't yeah keep what she knew that she wanted to go from a series of developer operating steps,

>> right? That's a good question. Yeah, I mean I think I again I think she she invented this idea of like we could have some middleware program that compiles it down to assembly, right? Um, and so I think her goal was to make it was to make that. And I think it was just I don't know the exact steps how she got to the syntax that she

chose, but it was if I can create a syntax and then turn that into machine language with with program a compiler and then run it on the computer, then a lot more people would be able to be I think she was working with a lot of folks, a lot of very smart engineers who weren't necessarily electrical engineers in the Navy and she wanted to bring their knowledge

into what they were able to get done, right? So they were they were computing a lot of uh nuclear engineering um things at the time of course for for World War II. Um and so I think she wanted to bring that knowledge into the computing space and and make it easier for people to express themselves who maybe their first discipline was not electrical engineering or math. >>

Three out of the four names on your t-shirt. >> Jean um she is poof that always gets me. she worked for NASA. Gosh, what's her last name? >> Simon. >> What's that? Simon. Yes. Yes. Yeah. Yeah. Yeah. Worked for NASA on in the computing division. Yeah. Yeah. >> Yeah. Go ahead. >> Yeah. No, please. I don't know. I know very little about >> cobalt. the hostas. >>

Yes, >> that's true. It was Flowmatic was actually what she wrote. Yeah. Yeah. >> I think you had a question here. Yeah. >> How do you um How do you consider your privacy when you're using these AMR? And then um how do you keep yourself safe when you're using open? >> Two very hard questions, one harder than the next. Um yeah, so how do you consider what

do you consider privacy? I mean I think um I'm very excited about all the work that's being done in openweight models because I think that's going to enable us to um not rely on, you know, a a duopoly of companies to provide us with AI inference. Um I have been uh with Kilo now since April of last year. Uh and last year it was all about like

Opus and how you know great that is. Um I think this year is going to be the year that openweight models catch up and pass uh or at least catch up with the the premier labs. I think that's a really great way to you know run your own inference right. Um, I I have a Mac Studio at home and I can run um pretty serious models on

it now and I think that's only going to get better, right? Like I think Quen just came out with um can't remember what it's called, but I saw someone demoing it on a phone, like running the model on a on a phone CPU and GPU. Um, so I think if privacy is a major concern, I think there's a lot of options out there for running your own

inference. Um, I think that's the only way to guarantee it. Um, I think there's also a lot of inference providers that make guarantees to you, right? If you want to trust them about, you know, if you're using the API, um, then you're you're not going to be used in training data. If you're just logging on and not or not even logging on to a chat app on

the internet, I would assume you were going to be in the training data. Uh, and then second, how do you keep yourself safe with OpenClaw? That's a great question. Um, I actually wrote a whole blog. If you go to uh blog.kilo.ai, AI uh and search for claw, you'll find it. Um I when I first encountered openclaw, I installed it directly on my Mac because like that's what

everybody was doing. Um, and then I actually went to um, our company, we're located all over the world, but like kind of half of us are in Europe and we were going to Spain u for a company off-site and uh, it was like deleting emails of mine that I needed [laughter] and uh, like for my wife um like, "Hey, can you fill out this form for school?"

Like, "Oh, yeah, sure." Um, and so I I like I took it off immediately, right? um because I was getting really worried. Uh but then I' I've reinstalled it. I reinstalled it on a on a VPS and now actually I run it on a computer, Linux computer in my in my uh house. But what I did is I have this same model, right? We talked about this

model of like treating it like a junior engineer. It's like you wouldn't hire a junior engineer or an intern or even a personal assistant and just give them your passwords for everything the next day. You'd give them their own account. So, I actually have um like if we go to github.com/scuttlebot, which is the name of my openclaw, this is actually my my openclaw agents uh GitHub account.

Uh and so it has full access to this GitHub account. And then this GitHub account is a developer, not a maintainer on my repositories, right? you can see um you know it gets a little active again yesterday I was trying very quickly to respond to Peter's tweet and and so he had a lot of contributions uh yesterday um but this way I get to treat everything he

everything that the openclaw agent does as I would any junior developer and I it sends me PRs and I review them um and so I think this is kind of the way that that I found and again I wrote a whole blog on the details of how I did this. Um, again, that that's trading off some things, right? It can't read my iMss for me and respond

to them and right like there's this convenience versus uh security or or you know, trade-off. And I think that we're still learning what that the right line's going to be there for sure. >> It's pretty good at it actually. There's there's actually a paper someone wrote a paper on it. I don't have it off the top of my head, but somebody wrote a paper on it. Um,

I think it's pretty good. I I don't know that I would just go into a bank tomorrow and convert all the cobalt to Nex.js. Um, but uh, you know, again, I think I think it's it's pretty solid, right? But I I I don't think I would trust it to just do it. Um, but you know, these the semantical differences between languages is something that large language models

actually pretty good at at reasoning against. Well, not reasoning, but reasoning against. Yes. >> Yeah, it's a great question. Yeah. So, I I you know, I talked about this RPI methodology. Um there's a lot of other methodologies out there and you know, spec driven development is another one um that you'll hear people talk about a lot. Um, I've actually got this section on trends and patterns on

this path.ai where I talk about a lot more of them. Um, you know, I think I think that agents are really good at running tests. And so if you're already kind of in this test driven development methodology, um, I've seen a lot of people have success of I wrote the test and then had it write the code. Um, assumes that you're already really good at tests, which

maybe maybe maybe everybody in here is really great at tests and it's just me. Um, but I'm not really great at tests. Um, and so that's why I didn't like choose that as like the thing to talk about in this talk. I think RPI applies even if you were just saying, hey, I want to do it test driven. It's like you would still research like what should

the test be? Plan, hey, what is the structure of the test and how how what's that output going to be expected and then implement okay now the tests exist. let's implement against those tests. Uh so I I I think that's why I chose that for the talk, but there's a lot more information out there on folks that are doing specri development and test-driven development with AI. >>

Yes. Go ahead. >> You said open positive privacy concerns. Um, do you do you think that's enough, you know, as opposed to like open source? >> I I I would love it if it was open source and the training data was open and repeatable. Um, I think if it's the beauty of open weight is you can then run it on your own hardware and you know that

nothing's going outside of your hardware, right? Um, I I think it'd be better if we had the training data open as well. Um, but I think that open weight gets us gets us a big chunk of the way there, right? Like the the my concern on from a privacy perspective would be the code that I'm writing, the code that I'm researching and planning against is that then

going back in to be fed into the model's training data. Uh, if I'm running on my own hardware, I I can guarantee that that's not true. Um, and so that's why I think openweight gets you the 80 the 8020 I think is openweight. Of course, I'd rather it be open source. Uh, I've worked for open source companies for a long time, but I think open open weights

does the 8020 at least. Awesome. All right. Well, I'll be around a little bit longer if anybody has anything else, but thank you all so much. Thanks for bearing through my uh major HDMI issues. Appreciate it. 45. >> Okay. Testing. Good afternoon everyone. We're going to go ahead and get started. Uh this afternoon we have Dave Neri. One second. Phone locked. Dave Ner leads the developer relations

team at Ampure Computing promoting the adoption of ARM 64 cloud native processors. Please take it away. >> Thank you very much. Uh thank you for that introduction. And uh this is a new talk, so I'm going to be trying some stuff out. I'm hoping this will be a little bit interactive. I also just recently, and as anybody who saw my presentation yesterday knows, I talked about it

a bit. I recently got progressive lenses, so I'm still trying to figure out. I'm in the acclimatization period for anybody who has a progressive like sometimes I take them off, sometimes funky. But anyway, I'll give it a go like this. And if you see me looking around, I'm just trying to find that spot where the magnification is right. Okay. So today's talk, you don't care who I

am, you care what I have to say, right? So what I'm I'm going to talk to you a little bit about about what developers should know about computer architecture. As a rough agenda, we're going to go over uh why developers should know a little bit about computer hardware design in the first place? Like why is this relevant? Doesn't the compiler take care of everything? We're going to

talk about memory and disk latencies and how that can affect the performance of programs. We've got some examples in there. We're going to talk about the way modern CPU pipelines work. If you're familiar with that already, then you're probably got not going to learn anything because I'm going to come at it from pretty high level. My goal is to target high level programming uh highle programmers who

have never really thought about how the hardware works. And then we're going to have a section on developing for parallelism. What does that mean? How does that change what you need to do? What are some of the pitfalls that you can fall into in terms of uh kind of destroying the performance of things rather than taking advantage of multiple cores or multiple CPUs? We're going to summarize

with some lessons learned. So, let's get started. Okay, first off, why should developers care about what the hardware is doing? Doesn't the computer doesn't the compiler in the high level language, don't they already take care of that? At least that was my perception. You know, compilers are pretty good at optimizing code these days. So, why do I need to know, you know, how memory layout works or

how the CPU pipeline works? Well, modern compilers are amazing. Uh, they can do lots of work that you don't explicitly ask for, like they can unroll loops. They can remove branches that they can move to branchless instructions. um they routinely reorder or change instructions so that you can uh so that as long as it doesn't affect the output so that you're more uh more optimized in how

you use the CPU pipeline. Uh you can even automatic automatically uh vector code. So the cont the compiler will detect if you have SIMD modules it will detect oh we've got the opportunity to process this vector of data four items at a time and it'll do that for you. You never need to know how to do SIMD. SIMD is single input multiple data. Uh so it's a

way to um it's vectorzed. So anything any GPU program is going to use SIMD programming. You're you're basically processing vectors of data item by item in chunks. So modern compilers are amazing but they don't do everything for you. You don't they don't own your algorithm. They don't own your data structures. Um, for any uh Angelinos here, this is the St. Francis Dam. So, I don't know if

any of you know this story, you do. Okay. So, the St. Francis Dam, it's up up uh one of the canyons near here. Um, it was designed and built by William Mullholland, eternalized in Mhalland Drive in the Mhalland Dam, which was another dam built by this Irish guy. So, sorry about that. Um, so it was it was a British Irish uh self-taught architect who was very active

in designing the water system of Los Angeles in this dam was built in 1927-28 and the dam didn't really take account of the substrate, right? The dam did was not designed with an understanding of how the rock underneath was was made up. There were there were several tests and and and Mullhalland to his credit took absolute fully full responsibility for it. It ended his career when this

dam collapsed. It killed 431 Angelinos uh uh sometime in yeah it collapsed in 1928 right before the great depression. Uh part of the valley in which the dam was built was shist which moved a little bit under the pressure of the concrete caused cracks during the construction. It passed an inspection 12 hours before it collapsed. Um, so this is just a cautionary tale. If you don't take

account of what's happening underneath what you're building, you can have bad results. Hopefully, this is not going to happen with your code, right? But you never know. So, you know, algorithms, data structures, deployment choices, these matter for your application's performance. So it's important that you understand how what you're doing at the high level translates into what happens at the low level, particularly if it's going to dramatically

impact In general, you want to swim with the current. In this presentation, we're going to cover a number of hardware behaviors that might be invisible to you as ana application developer, but they can be the silent performance killer for your application. They can cause your code to run 10 times, 100 times, thousand times slower. Um, by gaining awareness of the underlying behavior, you can just write more

optimal code. You can take advantage of the hardware and just make it easy for the for the CPU to help you do what you want to do. So, we're going to cover three big topics. Uh, memory locality and data coherence, data access patterns, uh, modern CPU pipelines, and coding for parallelism in multi-core or multi-host uh, environments. Okay, so let's get started. Uh before we get into the

content, I do have to talk a little bit about Ampere, but it's only just an introduction to show what a modern CPU looks like. When I talk about coding for the cloud or coding for a modern CPU, what do I mean? Um this is the Amper 1M. It's our latest product. Uh it's up to 192 cores at up to 300 3.6 GHz. You get on each core,

you get 16K of L1 instruction cache and 64K of L1 data cache. We'll talk about those a little bit later. [snorts] You get 2 megabytes of L2 cache per core, 64 megabytes of system level cache, that's cache SC shared across all of the cores on a single socket and 12 channels of DDR5 memory up to 1 point 1.5 terabytes of RAM. And we also have some vector

units per core, two 128 bit vector units for SIMD calculations. Right? So this is what I'm going to use when I do latency figures here. I'm going to be citing the figures for Ampear. It's just to because I know what they are for a start. Um, but this is like when we're talking about modern CPUs where we're talking about coding for parallelism. This is what you get

in the cloud these days. >> Uh, I do not know what the TDP is. I do not know what the power consumption is, but I know that we're very very competitive in power consumption because that's kind of the core. We're we're an ARM 64 CPU. Uh so that's kind of one of the core reasons why we believe ARM 64 CPUs are are kind of the the question

was what's the power consumption of that of that CPU. Um power consumption is like uh power per core but even power per socket we're very very competitive with anything out there. It's much lower than x86. Uh, and that's one of the reasons why we think this is kind of the future of the data center is if you put a a a CPU that big that's an x86

CPU, you have to account for heat, you have to account for cooling, you have to account for air flow or water flow uh through the through the server. Uh, you're putting fewer cores per rack because you're you have to leave empty spots for heat. I don't know if anybody have have worked in a data center, but like the classic cold aisle, hot aisle, um you you need

to account for the air flow that's going through these things. With Amper CPUs, you don't have that problem. So, you fill up your racks. You've got you can have 42 servers in a rack. um standard size rack which which which means that you've got more cores in your data center which means you're building fewer data centers which means you're putting less pressure on the grid which means

there like there's a whole bunch of good stuff for CSPs but it's also good for customers because your CSPs are renting you those cores cheaper he asked I wasn't I wasn't going to go into the sales pitch but there you go that's that's why we think ARM 64 CPUs for the cloud are really important for the future let's get talking about uh some of the things that

you need to know as a developer. Um, so we're going to start with memory and disk latencies. Uh, so to make this a little bit more concrete, something that you can identify with, I said to myself, how can I turn, you know, nanconds, microscs, and milliseconds into something that's meaningful to this audience? So, since we're in Pasadena, I'm going to take a local landmark as a reference.

We're going to take the Rose Bowl, the Pasadena Rose Bowl, and we're going to start from the 50-yard line. And to illustrate time taken, I'm going to use light time, right? How many uh how many? So, a meter and a uh so one light second is three billion meters, right? Uh so, it's three billion meters per second, one light second. We're going to take access time to

uh different levels of cache or memory as light distance. And I'm going to ask you all to guess uh how far light can travel in the time it takes for data to load. And I'm going to give you a freebie. If you are accessing uh layer 1 cache, you can barely see the red dot in the middle there because from loads from L1 cache are fast. They're

about one nanocond. Uh so one nancond is 1/3 of a meter, 30 cm, right? About a foot. Uh so you're still basically at the halfway line. Maybe if if if we're doing this when uh you know there's a there's a an a football match going. You're and and the the ball is placed on the halfway line, you're about like the the the you can just about get

outside the football, right? That's the kind of scale we're looking for. But remember, you only get 16k of instruction cache and 64k of data cache. So I is where your program lives. data cache is where data structures can live that are accessed. So you've got small amounts of that data. Um what that means is that if your data that you care about on the fast path of

your application is less than 64k access to that data is super fast but you don't have a lot of it. Okay. Now moving from L1 cache to L2 cache. How far? Who Who wants to guess what? Like how far we get away from from midfield? >> Sideline. >> The sideline. >> Um, okay. Anybody else want to care? Want to guess? >> Middle of the city of Pasadena.

>> Stadium. >> Stadium height. >> Okay. Okay. So outside the field more than 100 yards um a 20 yard line I heard. Well, so those are all uh good guesses but they're a little bit pessimistic. We actually get about um so it's 3 to 10 nanose on average about 5 nconds. So we're talking about 5t 2 m 2 yards, right? Uh so light has now traveled between

1 and three yards. So let's say two yards. So you're you're surrounding your your center and your defensive linesman maybe, right? It's not too far, but your arms reach from the ball. So even though like layer 1 to layer two is slower, it's five times as slow. Um, and you now have two megabytes per core. So if you can fit your hot path data in 2 megabytes,

you're guaranteeing that your access time is going to be still pretty fast, right? Next up, we've got the system level cache. This is a shared cache by all of the cores. So you've got some bus latency, right? Because you've got mesh interconnect between all of these cores with their local memory. But if you're going to system level cache, it's a little bit slower. Anybody want to care

to anybody want to guess how many times more slower? >> Twice as slow. 20 nconds. >> Five times as slow. So that would be 25 uh 30 nconds. Uh and that's it ends up being somewhere between 10 and 40 nconds. So about 30 nconds is the average. Uh at this point you've got so we're we're gone past 10 yards away from we're we're now it's about I

think 15 yards give or take 12 15 yards. Um so we're now past the 40 yard line in either direction. Uh so again thinking about ampear ones you've got 64 megabytes of this per socket. So that's shared across 19 up to 192 cores. So you don't get a lot per core. But you also have the ability to define policy so that which cores get access and which

cores get fast access to the system level cache and you can deny access to the system level cache to certain cores. So it's um you you do have ways to tune this but you've got a lot more uh data. This is still, I think, pretty fast, right? Okay. So, now we're gonna we're going to go to the next one, which is uh SDR RAM, Are we still

on the field? Why do I not have an SD RAM slide? Yeah. Yeah. But with SDG RAM, you're about 80 yards. Uh so SG RAM, you're talking about 100 to 150 nanoseconds. So it's about five times slower than this. So where we're 15 yards, five times we're about 75 yards. You're in the end zones potentially. Certainly sideline to sideline. Um I've kind of buried the lead. Now,

if I now go to the next fastest storage medium, it's local attached NVMe solid state discs, right? Uh, okay. At this point, right, let's say we've got a 100 yards for for SD RAM. Um, how how far away from the center field are we with NVMe? >> NVME is basically storage is is basically RAM, right? Uh, it's a little bit more than that. um it's uh 20

to 120 micros secondsonds. So we've changed order of magnitude. We've gone from nanoseconds 100 uh 120 to 150 nanoseconds. We've now gone to 20 to 120 microsconds. Now we've got a lot more of it. We've got hundreds of gigabytes of this stuff. If you have money in [laughter] and bought some before the prices went through the roof. Uh that's also true of SG RAM by the way.

Um but uh latency times for fast NVME discs are in the tens of microsconds. So in terms of the landmarks from here you in the optimistic cast you can get case you can get to the JPL in the pessimistic case you can you you certainly have time to get to Burbank airport or the Hollywood Hills right we're still inside LA you've now kind of surrounded Pasadena right

oh I've gone a little bit too fast the next fastest I've got two options to to offer and I'm going to do a pop quiz which one of these Do you think is faster? Local attach spinning disc or local area network attached SAN running NVMe SSDs? >> The SAN. Yes, it is. And it's it's quite a bit faster. So SSD powered SAN uh we're now up to

hundreds of uh microsconds. It is another five times slower. Uh so we now have, you know, tens or hundreds of terabytes. We have storage arrays that can basically be arbitrarily large, but you're about halfway to San Diego at this point. Um, entire Los Angeles County is surrounded. Uh, you know, you're now in like Santa Barbara north uh west or you're in uh what's what's that town there?

Carl'sbad in the south. Uh you're you're you're quite a bit away, right? We're talking about um 40 50 miles. Um, next up, if I am using spinning discs, if I'm using SATA discs, how far Anyone want to care to wager how far, not wager, but guess how far light can travel in the time it takes me to read from a spinning disc? >> Las Vegas. >> Las

Vegas, >> Washington DC. >> DC, New York. >> Who said New York? Who thinks New York is right? Who thinks Vegas is right? >> yeah, it's you can you can get to Hawaii or or New York. You can get to the the um we're talking about I had I could swear these slides were Am I working on an old version? Anyway, we'll flow with it. We'll go

with it. You can get to uh you know uh the Empire State Building easily, right? It's we're talking about 5 to 15 milliseconds of latency. So we're talking about another order of magnitude, right? Another two orders of magnitude from so just the spinning discs you're actually you have platters, you have a a mechanical head that's seeking to read from a certain area. Um if you need to

go to on disk swap for example, your performance takes a hammering. [snorts] H at this stage the scales have gone from uh nanconds to microsconds to milliseconds. Uh [snorts] so by the time your data is loaded into your hard drive, light could have had the time to reach New York. And in terms of lessons learned from this, where you read your data from really matters. The closer

your data is to the CP CPU, the faster your reads are. So you want small, fast, and close to the CPU for hot path data where where performance is critical. And and in terms of orders of magnitude, DRAM is 100 times slower than than L1 cache. NVME is a thousand times slower than is a thousand times slower than DRAM. That doesn't that No, a thousand times slower

than L1 cache. And traditional hard drives are another 100 times slower than NVMe. So 10 100,000 times slower than reading from cache. That's a massive performance kit hit, right? If if if you have to read from disk, it's it's a performance killer. So what what do you do with that information? How can you uh make sure that you're using algorithms and data structures that that allow you

to have that hot path data be efficiently used as part of of your um program. So designing hot cache hot data to fit in cache processing large data sites in cache cache sized chunks using algorithms that allow you to do that cache by chunk by chunk that's going to be important for application performance if performance really matters and we'll get to that a Okay, next let's move

on to how data is actually moves around inside your system. When you read a bite of data you're not moving a bite of data into cache, right? Um, the smallest unit of memory transfer is a cash line. I've used a carton of eggs here, but it's you don't you don't move data bite by bite. You move data chunk by chunk. Um, the minimum amount of data that

gets moved around in cache is is a cache line. Cache lines are 64 bytes. So when you read from a memory location, that memory location is first in RAM, right? If the if the data is actually on disk and needs to be loaded into RAM, it's going to be loaded into RAM. That's called a page fault. That's a slow operation. [snorts] Once it's in RAM, you seek

to the specific location where that memory that you care about is, and you look for the uh boundary aligned 64 byt uh chunk of data, and that's what you read into your L2 and L1 cache. Um, why does that matter? A cache line is 64 bytes of memory. And if a cache line is in cache, reading it or any of its neighbors that are already in cache

is essentially free. It's one nancond for access time versus five nanconds or 150 nconds or uh you know a couple of microsconds. Um if a cache line is already in cache um if a cache line is not in cache the whole data the whole line is read from DRAM before uh data is loaded. And the way that works is you have a memory controller on your CPU

um activates a line of RAM. You have the way the way RAM activation works is you've got a first operation. It finds what row RAM is laid out as as a 2D array of capacitors. basically they're either uh capacitors are loaded in which case it's a one or they're empty in which case it's a zero um and it finds the right row that operation has a certain

cost and then it finds the column inside the row where the where the cache line starts and that operation has a smaller cost so if the row is already open you get an optimistic like maybe 50 nanose if it's not open then it costs about 150 nconds to read from from RAM so the your memory manager is also kind of smart. It looks at how your program

is accessing memory and if your access patterns are predictable, it will automatically read in lines that it expects you will need in the future. If you're wrong, that has a cost, right? You've loaded memory into RAM that then gets gets evicted and something else needs to be loaded in instead. Um what that means as a programmer is that cache lines um are prefetched prefetched based on data

access uh patterns and predictable access patterns mean that you're going to be using that data more efficiently. If you do sequential access of memory across contiguous blocks of of memory, then you're optimally you're you're basically your average read time is going to be closer to one nanocond than than higher, right? And data structur data structures really matter for this. If you're a C++ programmer, if you're using

arrays or vectors, those are fast, those are contiguous data blocks. If you're using on C or C++ or really any language uh kind of binary trees or uh linked lists, those data structures essentially jump randomly around in memory because you've got a pointer to the next thing is is part of the data structure. Um and every single lo every single read of an element of a linked

list is going to cause some kind of cache eviction and loading in a new line. Um so arrays and lists outperform linked lists or binary trees. Again, I'm not saying that you can't use or you shouldn't use those data structure structures. What I'm saying is if you're in an area where performance is critical, you want to make sure that you're optimally using all of the data that

you load into cache. You want to make sure that every like every cache line you want to get 64 reads, right? You don't want to be you don't want to be reading one bite out of a 64 chunk and then reading another bite out of a 64 chunk and doing that over and over and over again. That's optimally poor performance. So you want to get optimal performance

which means using that sequential data structure. So this is not just theoretical. Uh let's take one example. If you're doing image manipulation I used to work on on the GNU image manipulation program. Um commonly in graphics programming what you're going to have is the image is going to have three arrays or four arrays depending on how many color channels you're using. um with an array for the

red data, an array for the green data, an array for the blue data, and an array for an alpha alpha data. The reason why we store it that way is because you can optimally um do a lot of image manipulation algorithms uh by processing chunks of those data. You can do that SIMD because the data is laid out in memory already. you're just passing in a new

address and you're adding adding 64, adding 64, adding 64, whatever. Um, so game developers routinely store image data in an array per color channel because it it allows them to efficiently process that data using SIMD. Uh, this is called, by the way, in the game world, this is called data swizzling if you're looking for a if you want to want to look this up. Um uh so

laying out data that you really care about the performance in in this kind these contiguous data struct structures is is kind of something that's real life performance engineering. Okay. So we're going to check how much this matters and I'm it's going to be hopefully interactive again. I'm going to take two I've generated 1 million entry matrices 1024 by 1024. uh the so an A matrix and a

B matrix. We're going to multiply them together. So each entry in the C matrix is going to be an a sum of A I K * B JK. So we're going to go across the rows of the matrix and we're going to go down the rows of the second matrix. Um so this is what we're doing. We're multiplying two 1024 by 102 matrices, 1024 matrices. The the

code for this is super straightforward. Um, this is too small for you to read, which is uh my fault. But the naive method goes for I, we're going to go for each I J in the C matrix. So, we're going to we're going to traverse the C matrix row by row by row. For each entry in the C matrix, we're going to process a row of the

A matrix and a column of the B matrix. Right? So that's the the naive method is we go for uh for so the inner loop is K which goes across a row and down a down a column. The second method we're going to use is column first. What that does is we're going to change the order. Uh so instead of going across the C matrix first we're

going to do as our inner loop um uh the the column of the inner loop is going to be the column of the C matrix and the column of the A matrix. The outer loop is going to be um or J which is the column of the B matrix. Okay. Uh and then the third we're just going to change the order another way and we're going to

do row first. So we're going to change the order. So that we do the row of the A matrix, the row of the B matrix, and the column of the C of the C matrix. I I I'm messing this up. But like the point is the only thing that changes here is the order of the loops. I JK. >> Okay. Um Oh, I was going to ask.

So let's just let's just jump in. Um the consequences for performance are dramatic. um ensuring that the inner loops are following row order for all of the major matrices, right? Uh means that we're for each entry um uh that so in the naive case every element of the C matrix we're doing we're calculating it and then we're moving on to the next one. In that naive case

we were getting 126 seconds on a single amper one core. Um, if I do the worst case, which is the column first as as much as I can, it's actually not that much slower. It's 146 147 seconds. Um, so those were both getting really hit by uh cache line loads, right? Because every single uh every single entry in the C matrix that I'm calculating is loading different

cache lines every time I go down to a new row. Um, if I change just the if I only just change the order so that for each entry for each uh for each member of the outer loop, I'm going across two rows. That's the only change I made, right? Um, the numbers go down to 8 seconds total, right? So 126 or 146 to 8, that is a

20x uh timesaver, right? This is for a fairly simple mathematical operation, right? Um, by the way, there are faster and better ways to multiply matrices and we'll we'll talk about those a little bit later. I reference them, but I'm not going to go into like tilebased matrix multiplication, but I will talk about that um in terms of pointing you in the direction of places where you can

learn So far from being a minor concern, this has a dramatic impact on the performance of something like matrix multiplication which is key to as we all know AI right. So this is the kind of thing where you really care about the performance of So how you access data ma data matters in addition to um where you access the data from. Sequential access of contiguous data is

fastest. Cach line misses and page faults are have can have a dramatic impact on application performance. And putting that into practice, you want to make sure that for the performance critical and I'm I'm kind of repeating this because we'll we'll talk about premature optimization at the end. But for the performance critical aspects of your application, you want to make sure that you're using data structures with flat

data types. So everything's the same size. So you're not doing arrays of strrus. um that are contiguous in memory and you want to access those data types sequentially to optimize for hardware prefetch so that your uh your CPU memory management module can can like optim make sure that data is ready for you when you need it. >> Putting it though how you access it changes where you

>> how you access it does change where you access it indeed. >> Yes. So, so that this gentleman said, "Shouldn't isn't it like how you access it changes where you access it from?" And that's correct, right? If you access data in uh a random manner, basically, you're going to be going to DRAM an awful lot of disk potentially. Um, okay, let's talk a little bit about modern

CPU pipelines. Uh, it's possible to modern CPU pipelines. There's a a mult process process multiple instructions at the same time. Uh, so this may be this may be old hat or knowledge that everybody here knows already. Um, but instructions are loaded by the CPU into I'm going to get in trouble here because I'm probably going to be wrong, but I think the ALU uh, so they're loaded

from Instructions that are loaded from memory are being loaded from Icash, which we saw earlier. That's 16 megabytes per core. 16 kilobytes per core, excuse me. Um in another stage those instructions are decoded. So they're get got they're prepared for execution in terms of reading from registers that need to be read from. Then you've got the execution stage where the instruction is executed. Then you've got a

memory stage. So if the execution is for if the instruction is a load or a store or requires some access to memory, memory access happens in stage four. Um and and this is a simplification. There are there are CPUs that have different pipelines. This is just the a kind of a a standard risk pipeline. And then in the fifth uh stage, you've got write back. So any

results of that instruction are written back to registers so that so they're available and uh and then the instructions retired. So optimally you can have five instructions being processed at the same time in different stages. That's what's happening in the green column here. Um, however, each one of these stages can stall, right? If you're reading instructions and the instruction is not available in Icash, then you can

end up going to main memory um to to get the instruction that you can decode. And so you're going instead of a one nanocond read, it can be a 150 ncond read if your main memory is busy. Um similarly on memory writes or reads um if the memory is not available in your dcache then you're going to be going to uh main memory or disk and that

can be a very costly operation. So you can stall the CPU pipeline in a number of ways and one of the ways is uh branch prediction. So some of you may have heard of branch prediction or um in the context of spectre and meltdown. This was the site this is what we call side channel attacks right um basically you can access some kind of indicator to what

would have been loaded in the branch predictor so that you can see what was so that you can see something that never actually happened but that can reveal data about the what's going on CPU I don't fully understand it um I have several friends who understand it a lot more but this is the mechanism that's used basically modern CPUs are going to have like any any program

is going to have branch points. You know, if this happens, then do this thing, otherwise do this other thing. And because what you're doing is essentially um you want to keep that pipeline full. You want to make sure that you're loading what's the next in instruction that's going to going to execute. Um and inside the CPU, you have the branch prediction module. Basically says, oh, I see

a branch coming up in three instructions time. based on what I've seen in the past, I predict that this branch will be taken and you load that instruction and you're getting it into the start of the Now, if by the time that branch uh that branch instruction is actually being executed, your prediction is wrong, you throw away all of the work that uh in preparing that instruction,

in reading registers that are required, in getting it ready for execution. You throw all that away and any other instruction that you've loaded after that and you have to restart your pipeline. Oh, okay. I was wrong. Branch misprediction. I throw everything away. What's the next instruction? I'm loading it from my cache. Um, so the cost of the prediction if it's wrong can be up to 20 cycles.

On an Ampere one, we've got an average uh branch misprediction cost of 10 cycles, which is very impressive industrywide. branch predictors on current hardware can achieve 95 to 99% accuracy on typical workloads. Um, and we'll see some numbers in a in in just a second. The predictor maintains a history of recent branch branch outcomes and basically is very good at predicting what comes next. Um, so you

want to keep that prediction success rate as high as possible. 99% up is what you're aiming for. Um, however, programmers sometimes do stupid things and the the the way that you code, um, oh, sorry, I I'm uh the way that you code can can make it more likely that you're going to have a branch misprediction. Um, there are a few dependencies. There are a few things that

that happen in in uh in CPUs, right? You can have data dependencies where an n1 n plus1 instruction is calculating something that's needed by the n instruction and you can't execute that n instruction until the calculation is completed. That stalls your pipeline at the execution stage. Uh you can have control dependencies where the outcome of instruction n changes the instruction n plus one to be executed. Branch

misprediction. We already talked about that. I'm not going to spend any more time on it. So here are some of the ways that developers can break your If the CPU is pretty smart about branch prediction, what can I do about all this? Are there things that I do to make it better or worse? And you can definitely make it worse, right? Um, three ways that the pipeline

stalls. I've talked about backend stalls. That's when you uh don't have memory that's required already in cache. You have frontend stalls. That's when your instruction that's required is not yet incach. That happens for example if you have a lot of functions. You don't necessarily think of a function call as a branch. It's a branch. you're going to a new part of memory. Um, that can happen if

you have, for example, if you have a C++ code with a lot of polymorphism. If you have multiple definitions of a function that are rent that are rendered at runtime, that's essentially something that the CPU can't predict, right? It doesn't know what the next instruction is based on what you've been doing in the past. Um so you can cause instruction cast misses with um uh with bad

function etiquette especially on the high on the on the hot path but and uh branch misprediction you can make it hard for the CPU to predict what comes next. And we're going to go into an example of uh some of the ways that that one example of how you can make branch prediction difficult. Uh but yeah, so if you have lots of small functions that are called

through function pointers, if you have lots of polymorphism, if you have lots of that, these are these are things that can cause uh higher branch misprediction rates. So here's one example. This is a um we're going to generate a list of 10 million random numbers between 0 and 100 floatingoint numbers. Um and we're going to um count how many numbers are under 50. Right? If the number

is under 50, add one to the count. If it's over 50, add one to the over 50 count. Um, just to have a little bit more time, we're going to iterate that over that 20 times. So, what you're going to see here is um 20 * 10 million, 200 million. We have 200 million branch operations, and we're going to see how many successful predictions we have versus

unsuccessful predictions. On the first time through, we're going to do it on uh random data just straight out of like the random number generator. On the second time through, I'm going to sort the list and do exactly the same operation. Right? It's the same code being run on the on two different data files. One where it's random, one where it's sorted. I'm going to use Perf to

count the time and the branch predictions for for the Okay. on the unsorted data. Uh, so I delay 500 milliseconds. That's just to allow the Python VM to start up. And I'm I'm kind of not counting Python startup. Uh, I'm looking for how many cycles, instructions, branches, and branch misses I have. Uh, I'm going to pin this to a single core. Um, just so that it doesn't

move around like I don't want the the kernel to be doing anything smart. So what I get out of my count is you know um I don't know how many digits are in that 300 uh 36 9 35 * 10^ the 9 cycles. Um I end up with uh you know 2.5 instructions per cycle which is we were talking about you can execute up to five instructions

per cycle. This is pretty awesome like 2.5 is pretty good. uh we end up with um 100 million branch misses which is 1.92% of all branches there are in general and this is kind of expected we're running Python right this is the full Python VM there's a lot of code that that happens before my code gets even processed so we end up with uh I don't even

know how many five by 10^ the 9 five billion branches um 100 million branch misses but 100 billion is 100 million branch misses is kind of what we were expecting from our data, right? 200 million comparisons. We were expecting 100 million misses, 100 million hits. Basically, it's we're flipping a coin. Uh we end up with our total time being 11.2 seconds. Uh user time for processing this

list, right? Uh so branch misprediction 100 million time 11 seconds. On the sorted data, I do the exact same operation, but I'm running on a sorted data list. What happens is the first time through it sees something close to zero. Next next time through it's 0.2 or whatever. Um, and it's like, oh, I predicted under 50 last time. I'm going to keep predicting under 50. And then

it hits 50 and it gets it wrong once. Maybe it gets it wrong twice. And then it's like, oh, now we're over 50. And it just predicts over 50 everything every single time. Uh so we've gone from uh 100 million branch misses to 2 which is pretty much what we what we would expect. Um we've gone from 11 whatever 11.2 seconds to 9.7 seconds. So on a

relatively simple piece of code we're saving a minute a second and a half on uh on this particular algorithm. not by changing the code, by changing the way that we feed the data into it. Uh so this is all in Python. Um there's a reason I use Python and not C for this. I talked about how compilers are amazing, right? Um if you do the same thing

with C compiled with the right optimization level, it automatically detects there's an if else and it uses a branchless instruction. It will use a turnary instruction. It will use some kind of method to say okay I'm going to do a test and then I'm going to like it uses an instruction which which doesn't use a branch right which the which the CPU can then read sequentially. So

we get an identical result with and without the data the sorted data just because the compiler is awesome. So the yeah some of this stuff that I'm introducing to you here compilers do help you take care of it if you've got minus03 minus04 um the way it does it is pretty neat and it's a it's worth dwelling on a second uh the C compiler does better because

it can avoid branch mispredictions by avoiding branches. So for those of you who here has heard of branchless got two three folks. Okay. Um we can avoid branches altogether by using using branchless algorithms or language contrast constructs. Here are just two um two ways to to eliminate the branch in what we just saw. One is a turnary operator which gets translated into different assembly instructions. So you

go x less than threshold under plus+ or over++ if it's greater than threshold. Or you can use bitwise arithmetic. You create a boolean that is this data is this data element is less than a threshold. And then you add the boolean to both the under and the over. If it's less than the threshold, the under goes up by one. If it's and the over doesn't change because

the boolean is false. If it's greater than the threshold, the under doesn't change because it's false. And the boolean and and the over goes up by one. Uh so there are a lot of branchless programming techniques. It's out of the scope of of this uh talk, but if any of you have heard of the billiondoll challenge, anybody here for the billion dollar challenge? Uh I recommend you

go looking at it. Yeah. So the something that came out of the Java community uh last year, I think um so there was a a a reference file of a billion rows, no a billion row challenge. um a reference file with temperatures from weather stations um that had a billion rolls of data in it and you had to accomplish some task of the you know the the

min max and average temperature for each weather station and uh there was a kind of a reference implementation that was bog standard Java that took uh I think 140 seconds right to process the billion rows and by the end of the competition people had gotten it down to three seconds Um the way they did that was some of the some of the techniques were really interesting. It's

just like we're going to load the data with a mem and read instead of reading rows. Uh some of the techniques are uh you know we're going to um skip the the the slow part of terminating the process. We're going to trick the the program into ending the time and then we'll do cleanup after. uh but the biggest gains were by moving to branchless processing of the

data. So I strongly recommend you look at that. Um branchless programming is a way to deliver maximum performance on the things where performance really matters. Sometimes as a warning some of the people I've I've talked about to this are like yeah but branchless code you know sucks because it's hard to maintain. It's kind of confusing. Um, and that can be true, right? If any anybody who's seen

uh John Caramax inverse square root function, right? You know that like magic numbers kind of pop up out of nowhere and you don't know what's going on. You're kind of casting an integer into or a float into an integer, adding a magic number or anding it with a magic number and then casting it back to integer. Um, yes, there are there are sometimes situations where you're going

to make your code harder to understand, harder to maintain, but if you really care about the performance of that thing, if it's a thing you're running billions of times per second, then you want to make it as fast as possible. warning. Don't fix what's not broken. before spending a ton of time fixing something that will only have a small effect, you you need to do some profiling.

So, this is a a a graph generated from uh an Amper tool called the Ampear PMU profiler, which basically investigates uh CPU events uh using the PMU performance monitoring unit. Um, and it can tell you where you're losing time. Are you losing time in the front end because your instruction cache is being is is not being hit well enough? Are you losing time in the back end

because your memory is not being consistently or or or efficiently accessed? Are you using time losing time in your pipeline? Make sure that you're using tools that are available like Perf, like the Ampure tool suite to to see what's going on in before you before you invest a lot of time uh doing things like improving the performance of the idle loop, which is kind of the classic

uh example of premature optimization. Okay, so lessons learned. You can make it hard for the computer to predict the future. You can help or hinder your CPU branch prediction. Try and get out of the way. Getting getting out of the way means that your CPU you're you're basically taking advantage of all of the engineering that's gone into your CPU. Inefficient programming of of uh of that that

leads to poor branch prediction can slow down your program. uh using branch branchless techniques or algorithms that uh or or al algorithms is going to reduce branch misprediction and if you're doing any high performance programming if you're doing any GPU programming or SIMD uh using SIMD modules on the the CPU core um SIMD doesn't have branches you're just processing bunches of data so learning how to do

things with bit bit maps and uh uh and and processing data chunk by chunk has the benefit of automatically being branches, right? So, it does require learning some branches techniques. Okay, we're running out of time fast, so I'm going to just zoom through developing for parallelism. Why is this important? The cloud has kind of changed how we develop code and how we develop applications. Um, we now

have dozens of services, uh, services or applications. Each have their own components. Each are running on virtual machines that can have lots of compute, right? We talked about the Amper 1 192 cores. It doesn't make sense to run a a a singlethreaded application on 192 core machine. You're going to have one red hot core and 191 that are sitting idle doing nothing, right? Uh so the cloud

has fundamentally changed the way application development happens. So there are a few techniques that we can use to make sure that we're taking advantage of all of those cores that we have a at our availability. Coding for more cores is different. And here are just some of the ways. There are a lot of them. Um I really like this Ian Cold Water. Um [sighs] um there are

a lot of them. First is you're going to be running if you're using in a single application many cores. You're going to be using threads. Um I I like a joke where you know a programmer had a problem so they use threads. Now they have two problems. Um, threaded programming is hard because you end up needing to be careful about critical areas of your code where you've

got information shared between threads, right? So, you need to have the appropriate level of locking. Uh, you need to have the appropriate semantics around acquiring and releasing data. And you can do uh, you know, we'll talk about this in a second. Uh, you don't have to completely lock data. So it's only one thread accessing the data particularly if everybody is reading the data and you only need

to update it when it when it writes right. So having the right level of of locking is important for performance parallelizable algorithms. So we talked about the matrix multiplication earlier. You can do tilebased multiplication uh multiplication of matrices which means that you've got small cachsized blocks of your matrices A and B in cache. you're generating a block, an intermediate block, uh, and then at the end once

all of the inter intermediate blocks are available, you add those in a certain way and you get your your your matrix multiplication C. So having parallelizable algorithms means that you can ship little pieces of a problem to different cores and aggregate the results when they're done. But that results in data synchronization issues. You need to make sure that there's no data shared across different cores. you've got

mesh congestion issues where you where you're if you're shipping a lot of data between different cores, you can gum up the system basically and add latencies to your mesh congestion. I haven't talked about and I won't talk about beyond this NUMA. Uh so NUMA is uh it stands for nonuniform memory access. Thank you. Um so if you have multisocket systems, each socket is going to have a

fast track. I think I actually have a slide on this. Each socket is going to have a fast track to certain storage, memory, peripherals, right? Network cards. If a CPU core on one NUMA node tries to connect to a network card on another NUMA node to do network reads, the latency is going to be terrible because it's not on the fast path. But the main thing I'm

going to talk about here is cache coherency and false sharing. So I don't know who here has ever heard of false sharing one. Nice. Good. You're all going to learn something. Excellent. Before we get into cash coherence here, I want to just give credit where credit is due. None of this is original material. Um this is Scott Meyers um who is a an old-timer programmer. I strongly

represent the pres the rep uh recommend the presentation CPU caches and why you care uh by Scott. um he goes into this in a lot more detail than I'm going to. I'm going to give you a flavor of it, but um this is there's way more and I recommend that you look into everything that's involved in in parallel programming and I'm going to link to some resources

in the slides that'll be available for you afterwards. But let's talk about cache cohesion. When cache lines are loaded into the CPU cache, if you have a two threads that are using uh that load something into CPU cache, they're copies of the canonical uh reference data which is in main memory, right? There is only one reference in main memory. Uh all of everything else is a is

a copy. So multip if multiple cores are using the same cache line writing in one core needs time to propagate to main memory and to the other copies of that. So let's look at what happens in as it happens. Let's say we have two cores core zero and core 1 who are accessing the same cache line. Uh, and I'll I'm talking about cache line because again that's

the atom of how memory moves around moves around in the system. Core zero and one both get a readonly copy. It's marked readon. That's it's marked readon and shared. Um, if co if core zero writes to the cach line, core zero's uh copy gets marked for synchronization. Core 1's copy gets marked as dirty, no longer up to date. Uh, so core one's copy is now marked as

invalid. and an invalidation mesh uh message is sent to any other cores that are using the same cache line. Core one tries to read the line from memory but is stalled until it updates the line from either core zero or main memory. um both cores then both uh cache lines then return to this status shared right so basically when I write here I can't read here until

that write gets propagated and it's cache line is the scope I and I'm emphasizing that because you'll see why in a second so this uh there's a there's a kind of a seminal uh Dr. Dob's article from I think 2008 or 2007 uh by this gentleman Herb Sutter who describes this problem of false sharing. He describes a scalability issue that he came across. So again the code

on the right is unreadable. It doesn't actually matter. All we're going to do is we're going to take a task. We've got a large corpus of data. We want to count how many times something happens in that corpus of data. And what we're going to do is we're going to break that down into 16 chunks. We're going to send a chunk to to different cores. We're going

to count in each core how many times that thing happens. And then we're going to uh add up all the counts at the end. Right? Um the way that we do that is we have an array with 16 entries. Core zero, core one. Uh so count zero, count one, count two, count three up to count 16. Um, and in in inside an inner loop we have we're

going to in increment that counter every time we get the a match on the thing we're counting for. So when he ran this on a 24 core machine, not a 16 core machine, but like we're we're going to just run with it. On a 24 core machine, this is the performance he got, right? When he went from one core reference 1.0 0 or 100% to two cores,

he actually lost 40% of the performance. This is not what you're looking for uh in scalable programming. Um it took until he started to gradually improve the performance as he threw more cores at it. It took until I think he gets to 16 cores before he's above 100%. So he's using 16 times as much compute and he's generating what he could have done with one core. This

is not good. And eventually it gets to a point where performance starts to go down again when he gets to core 20 21 22 23 24. Um so this is obviously something is wrong He made a change. Uh so sorry before we get into the uh how we fixed it, we're going to make a small assumption. Our array of integers storing counts is stored in a single

cache line. [snorts] So here's what happens, right? We've got our main memory cache line with our with our 16 counters. Core zero gets a copy. Core one gets a copy. Core zero writes finds a match somewhere in the in the corpus of data and writes its entry into entry one of the cach line that marks the cach line dirty. Okay. At this point, uh, core 1's cache

line is is is out of date and is also marked dirty and needs to be updated with this new write. Okay, so at that point, the core one copy of the cache now gets marked stale until uh the core zero right replicates uh core core one has found a match and wants to update co count one. It wants to update the second entry in the in the

array and it can't do it because the cach line is 30. So the cache line gets propagated which takes a few milliseconds I think. Um and then now core one can write to count one. Okay it's written to count one and that cache line gets marked dirty again. Now note they were not writing to the same memory location. They're writing to two different entries in a in

a in an array. But because those two different entries in an array are in the same cache line, it completely stalls the So they lose all of the all of the benefit of um having multiple cores thrown at the thrown at the problem. You make a small change. Rather than using array entries, you have inside each thread, you create a local me local uh integer that's just

count for that thread. After you've processed the chunk of data that's allocated to that thread, you set the array entry equal to the count that you end up with. Okay? By doing that, you've removed all of the intermediate writes. You're writing once per thread. So, you're writing 16 times. All of those writes get synchronized. You add them up at the end when all of the threads complete

complete complete and you get the result. Uh what Sutter found is that he went from what we saw earlier that mess to this, right? adding 24 cores, he gets 24 times the performance. Uh so what are the lessons to to be learned from here? Be careful when you're sharing data against multip across multiple cores. Attempting to modify a shared cache line on multiple cores will end badly.

You need to have some awareness of of of um anytime you're sharing cache lines because uh as we've seen, it's a massive So it does need a little bit more care. So what can you do about it? Um we have some really good techniques that have developed over time for parallel programming and and and uh uh they're common lessons from microservices, right? Consider functional programming best practices.

So this is kind of generally speaking shared nothing immutable data um pure functions. So a pure function is a function that when you call it it has no side effects. And um like any changes it makes are are are only it's it's it's given a task. It does a task. It communicates with other functions through messages. There's no shared data. Uh you can use perf C toC

uh which I really like to see if there are more uh load hits on modified memory than you would expect. So this is a this is basically an analyzing cache line locks. so in general thinking about Oh I think I thought I had something on uh on uh memory. Um this may be in here. So part four conclusions and lesson learned. Where you read your data from

matters and how you access your data matters. you can make it hard for your CPU to predict the future and you can be you need to be really careful when sharing data across multiple cores. So those are the lessons learned. What can you do about it? Again, keep your hot path data this is I'm going to be insisting on this. Keep your hot path data small and

contiguous and process it in a contiguous sequential manner. Uh use your data structures with flat data types that are contiguous in memory. Access those data types sequentially to optimize for hardware prefetch. Avoid a lot of um I don't think this in the in the list but avoid a lot of things like uh uh um function pointers that are dynamically rendered at runtime or um polymorphism particularly in

in performance critical parts of your code. Use branchless techniques and inline code to reduce branch misprediction. Compilers can help you here. And functional programming techniques will help for parallel programming. And I I want to particularly call out to memory oblivious. I I think it's cache oblivious programming. Um cache oblivious programming. So if you're getting into the world where you're programming for performance with multiple cores or multiple

machines, cache oblivious uh programming is basically this this is a set of algorithms for common high performance tasks that allows you to operate locally on data. So you process data by chunks. You can do the do things in parallel. Um, obviously important for AI. Uh, and it's particularly important with, uh, generative AI because you've got multiple transformer layers, multiple attention heads, and all of the rest. Uh,

but most important, do not optimize before you're ready. Right? Measure twice, cut once. Benchmark and characterize your performance issues before you start making any major changes. The biggest changes are going to usually come from very simple changes. choosing the right algorithm, choosing good data structures, and that is it. And I've I've gone a little past an hour, but if I if I have the indulgence of the

the conference, I'm happy to stay for a few for a few questions. And I'll while we're taking I'll put up uh first, this is some resources on where you can join the Amper developer community. um some information. We have an awful lot of information on performance engineering. We have a very very high quality performance engineering team that communicates tuning guides about how to get the most out

of your computer hardware which um which is effective both which is both for uh like it'll work on any architecture but obviously particularly we give advice for ARM 64. There's sample code documentation examples and videos. Uh we have a lot of information there and I'm going to put up a list of resources at the end. So I'm happy to take questions. >> I have the mic here

for anyone who has a question and let's thank Dave for a wonderful presentation. For any of you that still are attending the keynote, there is a keynote at 4 p at 3 pm in ballroom D and E. >> I'm coming around with the mic for Thank you all for coming and I'm like I said I'm happy to stay for a couple of minutes for questions but if

you need to go to the keynote or want to go to the keynote you can leave right now. You don't don't feel like you're hurting my feelings. Um does um pair have any um have like uh custom ports of various like linear algebra libraries like blas or anything like that that's [snorts] optimized for your >> uh so there is an ampear performance uh so we've got the

ampear ai optimized uh library which is a set of um optimizations for ampier of uh pietorch onyx tensorflow lama cpp um I'm Not a Are we recording? >> Um, maybe I'm not. >> Okay. I I would I would like us to be >> everyone can move closer. [laughter] >> Yes. Yeah. I' I'd be happy to answer that question later with a little bit more detail. Um I

I I I think yeah we do have some Yes. >> Anyone else for question? >> Uh you mentioned there's a piece of software there's a PMU monitoring solution that has what is it called? >> Uh so uh you can see the link uh well you it's not a link amper performance toolkit. We have a set of five tools, one of which is only available to customers for

reasons that I'm not sure what they are, but we have the Amper porting advisor, which uh gets your code working quickly on ARM 64 processors. It basically tells you um which versions of your tool chains and what what where you might run into issues with dependencies. It even looks at code and says, "Oh, you're using this header and this library. Those are not available on ARM 64."

Uh so that's the Amper reporting adviser. We have the Ampear system profiler which is a curated set of reports that we generate using standard performance tools uh things like MP stat and and PF. Uh then there's the Amper PMU profiler which gives you those CPU level reports and finally we have the Ampear um um performance benchmarker APB which basically runs a set of curated uh micro uh

profiling um uh uh suites benchmarking suites uh to um and and encapsulates the the important part is it encapsulates all of your system configuration in the report that it generates so that you can share uh that with somebody else and and replicate a problem which like replication of performance issues is kind of the number one problem that our performance team has had like customer says this is

slow and they're like well it looks fine to me um this is the tool that allows you to replicate a a system configuration on another system and verify if you're seeing consistent performance. So those are the those are the Ampere tools. >> Thank you so much. Also, does it work on other ARM 64 devices or is it just >> Yes, they're all they're all ARM 64. Um,

they will work on every ARM 64. I'm I'm and and it's all open source, right? So, it's it's on GitHub. Um, uh, Amper performance toolkit. You can find it in the Ampere GitHub repository. >> Anyone else for question anyone or comment or you know, >> thank you very much and thank you all for staying to the end. I'm going to >> uh let you all get to

the keynote now. I've kept you long enough. Thank you.

From event

SCaLE

05 Mar 2026 – 08 Mar 2026

All event videos
Back to Watch