DevDays Europe 2025

Mary Grygleski: Enter the Brave New World of GenAI with Vector Search

49:20 · 20 May 2025 – 23 May 2025 · YouTube

About this talk

In this presentation, Mary discusses the evolution of generative AI, focusing on the implications and applications of vector search technology. She begins by connecting the excitement surrounding tools like ChatGPT to broader historical narratives, remarking on the transformative potential of generative AI. Mary explains foundational concepts in AI, including machine learning, deep learning, and the significance of large language models (LLMs). Emphasizing the emerging role of vector databases, she highlights their utility in enhancing AI capabilities through complex similarity searches. The talk also touches on various player contributions in the generative AI field, including notable models and advancements. Mary concludes with discussions on retrieval-augmented generation and challenges in the AI landscape, urging closer consideration of ethical issues and the importance of ongoing innovation.

Full transcript

[Music] he hello here we go we are live hello everybody and welcome to yet another fantastic session I'm so happy to introduce uh Mary H for this session Mary and I were on a panel yesterday having a great conversation about the use of AI and I'm very keen to hear from you Mary about the new world the Brave New World of gen with Vector search now Mary

Mary griesi is a Java Champions and experience a passionate developer Advocate software consultant you have done so many things right you work for companies like IBM data Stacks ER you focus on gen streaming system obviously Java open source a lot of things now without any further hesitation I'm really really looking forward to this session over to you Mary and good luck thank you thank you Stefano thank

you so much and thank you again to De days I yeah really appreciate so thank you everybody if you came oops let me kind of do this one thing I need to switch back and then um I will make it small so then I can uh I won't be like blocking everybody yeah okay it should be good okay everybody Welcome to my talk um my talk right

enter the Brave New World of geni with Vector search um and you know as such you know if you are a uh science fiction uh Enthusiast you will recognize the Brave New World title right away is by an English uh author uh Aldis Huxley and very famous book about Brave New World so I kind of when I was trying to you know come up with the title

for my talk I figure generative AI Chad GPT was took the World by surprise right everybody was all excited about it um but in elders elders h 's U novel too it's basically painted a very dystopian type of uh culture right that's resulted from the future you know of our world in which we are kind of you know in some ways too we have lost our individuality

and it's like controlled by some you know like world power all of these things so it's kind of dystopian right isn't it like all these sci-fi like to say you know like the people are suffering basically you have to like you know listen to a central power now the thing is too is that I feel it's kind of somewhat relevant to gen especially when it first came

out right it's kind of exciting you know and and basically we actually got to a point in which we were nervous that well geni can take over our jobs but right now you know as such you know since his debut right back in 2022 November it's you know the chat GPT 3.5 that came out we are kind of noticing that okay the hype is kind of lifting

now there're still kind of plenty of work to be done and I myself personally think that it will be here to stay but then we recognize there are a lot more work in order to make this technology truly usable truly productive but that's what it is right but let's uh kind of delve into this world especially I find that there are a lot of developers we have

been kind of more used to right we're experienced doing things in the traditional way traditional Computing way the Enterprise Way perhaps if you're a Java developer um so we didn't get into the Gen yet um so let me kind of introduce to you this topic and also specifically to focusing on a bit two on the vector search at at one section of my talk so okay let

me real quickly and move on to the next page and who is Mary right so this is kind of my um well kind of blocking a little bit but you can kind of make out of it here but this is the agenda for today I'll do a quick introduction and then giving you some brief background of AI especially I'm assuming that perhaps you know also from my

experience is that a lot of you might have started maybe coding some you know AI but it's good to kind of get an understanding of AI first of all and then understand who are the players in this new generative AI era and then also I then introduce to you all of the commonly you know that heard of terms like gpts NLP llms and um you know rack

and also like the vector database too and that should kind of fill the space and then if there's time I can do a very quick demo as well and then also share with you you know kind of remind all of us too there are benefits while there are benefits there are also like challenges as I kind of like I'm kind of like pointing out in here all

this different things so so that's the um quick introduction and let me kind of go straight into really quickly so who am I I'm a passionate developer Advocate and thank you to Stephano he already gave the quick kind of introduction but these are kind of pictorial um kind of explanation of my background kind of just the key points um I'm doing a lot of things too these

days I'm transitioning I'll be starting as a uh generative AI practice lead for a consulting firm here that's based in the US in Ohio very soon but anyway but this is how you can get a hold of me but I'll be sharing with you my contact information uh at the end of this presentation too but P primarily too I'm just very interested in Computing um I'm kind

of was doing a lot more like streaming stuff uh back at IBM and uh as a developer Advocate and then a data stack to streaming and then kind of led into this generative AI um but I'm kind of definitely kind of Die Hard I'm in started off interested in operating systems moving on to like distributed systems reactive systems iot mqtt real time AI ml those are kind

of fascinating to me okay so let me start very brief background of AI so um kind of if we kind of uh start taking a look right uh of this uh you know kind of AI Kind of Revolution Let's kind of we all probably aware that it's not completely new so as you can see over here this um uh you know kind of like the timeline and

by the way too oh I might be blocking it but that's the the the URL as where I kind of got this uh timeline from so I will be actually sharing with you uh my slide deck so if you miss it don't worry about it you you can get this full uh stack to uh after the presentation um so okay so let's kind of quickly take a

look at this timeline so actually one thing I didn't kind of point out in here is that if you kind of do some research on AI I mean kind of really strictly speaking too in some internet literature basically it says AI the revolution began you know in in the BC like before Christ kind of um kind of era so many thousands of years ago and it's basically

the world at that time it was uh designed by a Greek uh person uh Phil philosopher or called archus and he kind of uh did like a steam powered pigeon right and so so the idea basically is that we want machines to do our job we want the automation part but of course back then it was very primitive way so let's take a look at the 20th

century as you can see right we started off you know essentially almost a 100 years ago the 1930s time that was when you know a lot of science fictions came out right golden age of Science Fiction and then very soon too then in England uh the father of modern computer science Alan Turing came out with this title can machines think so it's kind of this whole Trend

right kind of moving towards you know encouraging a lot more research being done and all of these you know first AI program uh there Dartmouth has some summer research you know project on AI and all of these things right as as you can see I won't kind of get through all of these but it's is quite interesting and I want to also point out too actually at

I was just came back from uh Europe from jcon in Cologne Germany and uh one person point out to me was also a perceptron too that he point out to me so if you're interested in that it's also like a particular kind of research during that time in the 50s too so I I'm going to add that into it in my next slide but anyway so if

you can kind of like see all of it and uh you know major um Institute right MIT uh Carnegie melon um you know in England there um all the you know University Cambridge Oxford all of these and all over and also then in terms of Industry uh deep blue came out of IBM and as we all know deep blue is kind of a dedicated machine to play

chess and then it ended up uh beating you know um Gary casprov the the chess champion in 19997 is actually beaten by Deep Blue so all of these have sparked a lot more interest into it too and there's also like you know late 19 or late 20th century already came out with dragon systems was basically the first uh speech recognition software that's developed uh by them so

and this it's very interesting but um okay I see a question Stephano kind of point out to me question from the audience what about ethical stuff concerning AI stealing artist art that's a an excellent question and that's something too that um in my panel just previously too we discussed a little bit about like you know ethical security kind of stuff but let me address that a little

more maybe towards the end as well but that's a a great question because you know as such you know there there are also questions about all these things but ultimately too if you kind of think about it we want machines to be you know not just automate some mundane task is basically we want machine to be doing things completely automating maybe to the point of taking over

our intelligence however you know and but that's a danger in it too but and also too if we kind of look at it that way then I'd say you know the present uh stage of gen is really just at the tip of the iceberg it's just starting right it's it's far from ready for production right and I'll share with you a story later as well as to

why I say so but let's first take a look at some background so if we kind of look into this AI there's you know especially for folks who may be new to this topic area just try to get an understanding right there's AI there's machine learning deep learning and how do they relate to each other but I don't think it's very hard to understand right now if

you think of it like an onion um kind of shape there's a artificial intelligence like basically the onion itself is essentially right we try to mimic human behavior and also automate things and having you know the machines um having intelligence to even decide what to do and to do the stuff for us uh however on its own it can't do much so that's why we get into

the inner inner core of this onion in which you know there's machine learning so we getting into the actual the substance what gives it the power is basically machine learning first machine learning is the technique by which the computer will learn learn from who not from us training it is basically giving it data have the computer figure out what to do right and and learning the behavior

and uh and it's it's that's what it is about training the model from large and huge data sets it's it's called corpora of data the data set right or or Corpus of um data and the thing is too even machine learning itself is not sufficient to solve a lot of the intelligence problems so that's when we get into the deep learning space in which you know it

kind of get into the neural networks you know we are basically a lot of research are being done here in which is figuring out how does the brain works right and how how does our brain works and apply it to machines and and you know training to to let machines to train on itself using this particular techniques is the neuron networks so that's kind of like a

very quick kind of introduction to that and let's take a fascinating look you know fast forward to today the Gen AI era so let's kind of first understand too Chad GPT this whole category of kind of computing is called generative AI as we all know and it is a disruptive field in AI why is it disruptive okay you might say well you know there are already some

things being done and there were also predictive what is called like predictive AI so the difference is that basically predictive AI already been there is essentially to Based on data it will make predictions right of maybe business forecast um what would it look like in two years in three years for business or weather forecast in like what does it look like next month the weather pattern all

these are very predictive and it's also trained on basically you you feeded data so from the data it kind of deres you know what it should be and then give you the prediction however there's a time kind of sensitive sensitivity to it because once the two years is over for your business forecast and it's no longer valid now going to generative AI first of all is that

we all probably are aware of prompt engineering it's a kind of a new subfield of profession too in which you know we basically have to prompt you know your AI uh the robot uh to kind of get what we want so to speak right so these are you know you give prompts but the also one difference is that there's no need to follow uh what we call

like traditional rule of doing computers in which you know we all know so well if we write a script for example right bash script you know even a Powers shell script if you're a net person is you basically have to say you know when you do something you need to give it a set of parameters and you have to follow strict rules and so it's very basically

robotic kind of way of asking you know your your systems your program to do stuff whereas for generative AI there's no need to you can just ask ask this machine this robot ask it like the way you talk to another human being so that is a very kind of a you know kind of a revolutionary way but of course it's also not all new it's already you

know kind of came out with speech recognition by Dragon systems just now and that diagram and also um an Apple computer came out with Siri right in the early 2 10 kind of time frame so not all new but the difference is that right now you know it's really capable of understanding in you know human spoken language really well too so generative AI essentially take the prompts

and then based on that it does a lot of similarity searches and basically from his algorithms in the models is able to figure out what you need it's basically what is called like chat completion right so and it produce the contents for you so the creative thing about generative AI or I should say generative AI focuses more tends to be all more the uh creative side because

for example we want to write an essay which kind of have been wowing all of us you know in these days I think even like students might be using cat GPT to write their essays or us too you know we work in the in the industry we need to oh marketing need to write a promotion we can actually ask cat GPT to Pro provide us with the

content but very often though we know you can't just fully be depending on the chat GPT you still need to do some kind of fine-tuning to it but for the most part it does its job very well and also um basically too you can also like do text to image right you can describe a scenario have the have have a dolly mid Journey all these to generate

picture for you so that's kind of like the generative AI thing that's kind of exciting so really quickly too let's take a look right all of the you know how does it kind of came to being and we can kind of go back to 2003 the first feed forward neuro Network by yosha Benjo and his team uh developed that first language model that's kind of like pointing

towards the direction of this chat GPT and and as such I mentioned about 2011 Apple brings Ai and NLP assistance like n natural language processing assistance to the masses by releasing his first uh Siri uh on his iPhone 2 and then 2013 a group of Google researchers by Thomas mikeloff created uh word to vac and so that's actually an important kind of research uh being done and

it's basically formed the basis of the today's especially like vector database vector search because it's essentially looking at words and convert them to um Vector uh data as such so it's kind of formed the basis of vector database and then 201 17 it's basically there's also another paper at Google um it's called the Transformer now if you are have been following some of these uh papers all

of these you have heard of you might have heard of uh attention is all you need so that's where that came out of this paper from it's basically the Transformer too so it's basically the Transformer is a simple Network architecture that supports um all these kind of like essentially gearing towards the generative AI era too okay so now after kind of a quick understanding of gen AI

let's take a look at all the players I think it's also very important too who are you know kind of in this field now first of all with generative AI we also we have all heard of the llms the models too so these are generative models so I like to kind of point out to you too these are some of the more popular models too and of

course there are more too and and don't get me wrong you know these are just some of the more popular ones that we hear about so first of all open AI we all know is kind of like the defecto um kind of uh organization right that's kind of in charge of all these geni model and first of all it came out with in 2022 November that while

the world was the the GPT 3.5 U model and uh then after soon well then kind of also along the way it isn't like the first thing but if you kind of look over there in the corner is Bert it's interestingly Bert stands for by directional um encoding encoded representation Transformer from Google that actually came out before the chat gpt2 and you again some of you might

have already heard of it so it's sort of an early form of llm but I put a question mark there because at defox UK somebody pointed out to me that that's not really an llm but you know it's interesting when I'm doing kind of research and that's what they call it an early form of llm but yes strictly speaking if you kind of measure everything from the

chat GPT era then maybe berd is not quite qualified but it's up to kind of discr you kind of discussion too but I just want to still set it up there because it's sort of um kind of like you know the the the early form of it now if you kind of look further to there are also codecs um over here in the middle that's what kind

of formed the basis for GitHub co-pilot codex and then also then there are also uh from meta uh the you know Facebook is llama 2 llama and there's also llama 3 these are open source models too and then of course um there's also like mid journey I want to point out mid Journey Dolly stable diffusion these are text to image models too that came out of stable

diffusion came out of stability AI um and then there's also a GPT 4 as we all know that came out last year sometime in the middle of the Year 2023 is an enhanced model from GPT 3.5 and then there's also whisper which is an audio based uh kind of take care of audio type of data too so again these are just couple of kind of representational ones

and of course there are companies like mistol and also um and uh germani right that came out from Google that's multimodal um and also coher for example is another companies with kind of kind of well-known models as well and uh okay so that's that and okay and then Stephano too is asking questions says in your panel yesterday also with Stephano and the other speakers you mentioned about

prompt engineering for getting better results from gen do you think the future versions of gen will be more inst in intuitive right and S senent so that it doesn't need more context to be able to understand conversation and context I think exactly Stephano I think that's what I feel today's you know prompt require us to kind of know what we want and then ask it in the

right way so to me I think it's not truly um you know gen I mean it it it's not truly artificial intelligence because the intelligence part is that you want the robot to be able to kind of think on his own and really sometimes too when we need to decide on something what to do we basically want to ask the robot hey you know I can't figure

it out can you figure it out for me um so writing prompts become kind of a chore in itself in my opinion right shouldn't it so anyway I think there are plenty of uh research being done we I wouldn't be surprised you know very shortly we are coming out with even more um Advanced like prompts you know that can kind of you know essentially really moving the

machines towards truly thinking on its own so but great question Stephano thank you very much so okay here the next slide is about generative apps right so these are like essentially these are apps that are making use of all of the models and of course you yourselves too some of you already started working with llm apps you are already producing apps that are relying relying on the

generative models but here are some of the example I already pointed out for example chat GPT is kind of an apps that's kind making use of GPT 3.5 and also the 4.0 right and then there are also other ones too as you can see highlighting I want to say GitHub co-pilot um and which also is developers maybe a lot of us are using or maybe attempting to

use it and then also to wanting to point out through ger I and Bart they came out of Google as well German as such too actually I might have put German into the model side I you know sorry if if you kind of spotted it but anyway gerini the the amazing thing about it is that from Google is that it's supporting multimodal and I'll explain a little

bit what modality means too so and then all of these other ones they basically come from different companies too they are apps that are making use of generative models as well okay so here's the multimodel slide so I think that's what is kind of making this uh super exciting right is that you know when we first started off like geni it's basically text to text you want

something done you you talk to the chat uh chat bot and then give it some text like let's say you spell something like um I want to know how I can um you know kind of uh you know build a an automobile engine or something right it goes back and then it comes back with you a whole paragraph you know how do you build an engine like

that but the thing is though what makes it even more exciting is that well I want you to show me right um how do you build an engine and um have a video show it to me so that's text to video so that is actually the where the the true value of of this whole thing is you can do text to text no doubt about it but

doing it multimodal is basically adding even more value to it and you can also like show an image right like a picture and say well who are these people in the pictures and they should come back with you to you and show the text or it can say you know um show me you know this uh this uh actor uh you know what uh movie He plays

in and it should do image to video so you get the idea so basically there are kind of six main types you know of modes or kind of modes you can have text image code audio video or 3D so these things you ought to be able to do it so then you can do multimodal type of um kind of a transformation of data in between them so

that will actually make things even more useful for what we need right and um and basically too a lot of companies are also kind of into this um kind of a space as well you you may notice so okay so now what about the people right all of the players and don't forget we are people we're working with it but I wanted to point out right they

are data scientists computer vision Engineers they are the ones that truly understand the data what it is supposed to be doing but very often too they are the ones that who may not be developers or Engineers like maybe I assume most of you in the audience are developers and engineer so the thing is too if you kind of look into it on the on this side right

the other site and and that one is basically they're data engineer AI engineer and ml Ops too now I just came from a def Ops panel we talk about you know integrating emerging technology to me I think this mlr Ops to machine learning Ops which is to me is an extension of Def Ops into the machine learning space I think there's a lot of potential in there

as we all know too machine learning there are kind of two categories of activities right one is the pre-training like kind of broadly speaking I'm talking about there are many steps small steps to don't forget but broadly speaking you take the whole data you need to train the model so that's a huge task in itself and then go after you produce the model then you can do

inferencing and this is where you can apply this um you know what you're needing to search is you send the queries over to the model and get what you want do the chat completion is inferencing so each one of them too they take a lot of steps so there are repeatable steps you know we need to identify them how you know how to do some of these

operational stuff now not to forget too data Engineers they still focus on the data but they work very often work with data scientists to kind of uh you know do all of the you know so to speak like the the the workflow you know kind of like crunching data all of these but they are complex in itself too so data engineer would do that and then AI

Engineers kind of deals more on the AI kind of level but of course there are normal software Engineers too we still need to do you take the data you need to like you know uh scrub them and then clean them and also filter messages or filter all the data and do all the feature engineering for example so just kind of want to give you an idea to

so do not worry right there are people that said oh it's going to take our jobs no I think we need more people to kind of make this more usable more production ready too is what I think okay so really quickly then let me go into the all of the ter terminologies GPT generative pre-trained Transformer so basically this kind of describes right what it does is it

takes simple prompts right again in human language form as input and it does all the pattern matching and and that's what we call like searches similarity searches and it answers questions for the prompts and it produce content such as a new essay a blog post you know a computer program that's GPT and now real quickly is that GPT how does it came about right it's essentially in

2018 and that's when the first GPT paper is about a language model that came out from open Ai and that was by Alec Redford's paper and that actually capture a lot of attention because there a lot of innovative uh kind of like research kind of discussion in there and based on pre-training based on a large and diverse set of data and then in 2019 it's basically gbt2

came out and it expanded that that um all of the documents that it takes right it's gbt2 also came out of open uh Ai and then let's go to the next page and and so fast forward a little bit 2022 and um essentially chat GPT I kind of go a little bit ahead it releases like gpt3 3.5 and it's essentially wild the world since then right it's

really hasn't been like you know hasn't been quite two years right a year and a half so far um and essentially it takes the World by surprise and then but not to forget to I pointed out about stable diffusion and that's came out from stability AI so it's essentially a deep learning text to image model that generates images right based on text descriptions so it leads to

like mid journey and Dolly and I'm sure a lot of you too if you're trying all the images things out and that those are kind of very popular model too so as we can see they came about then and then 2023 then we see kind of more and more companies vendors coming out with um you know new models or new apps to like Bart from Google and

also Bing from Microsoft um and then um all of these things are happening as we speak so I won't spend too much time on that let's look move on to the next thing which is natural language processing and that's what kind of the magic that'ss kind of behind the scenes with the human language interpretation so NLP it stands for an interdisciplinary subfield of linguistics of computer science

so the idea is that it wants to process natural language data sets right and it takes you know these data sets are huge these are text corpora or speech corpora and basically it process that and then it uses like rule-based and probabilistic U machine learning approaches to kind of process you know prompts that comes in and helps essentially the llms to kind of produce what it wants

and essentially so essentially llm large language models will be using this technique underneath the hood and it enables computer to learn from the contents right including right and and these days too we have to kind of um kind of move on to a next kind of way of thinking is that you know we're doing like you know all of the contents that I produced it's not just

like you know what we used to we need something it's definitive kind of answer in this case too with NLP is basically it can come back you know with you know a context of information around what you're searching for um so essentially too um it included you know when you search for something it comes back with some contextual nuances of the language itself and it's essentially too

it can draw insights from the document so it's pretty kind of uh Innovative in that sense so now let's kind of quickly take a look into llms right I mentioned about it so llms large language models and that's what developers we work with directly right we interact with through all the libraries all the a apis it is a type of machine learning model it's also Foundation type

of model and we know that all the general like llms they cost a lot of money it wouldn't be for us to do it because it takes you know many gpus and also weeks of processing too and it essentially too what it does is it performs all the you know NLP task underneath and it generates and classifies all of the text and it answers questions right answers

all the prompts it's basically feeding into lmms to kind of um it for it to figure out what you want it completes all the chat just like a human right not only that it also analyze your sentiments too and also dealing with chatbot conversations now real quickly too if you kind of look at this diagram what picture is worth a thousand words so earlier I showed that

set kind of theory that onion so if you kind of expand it a little more to include llms and that's where it sits it's the Transformers and language Transformers it sits inside the Deep learning space right so over here too as you can see deep learning has all the generative AI Transformers in it too there's image gen and also llms too and that's where it sits and

as you can see llms kind of leverages on NLP kind of uh techniques and all of the algorithms uh that's kind of working with nlps because the prompts comes in in human language form it needs to figure things out so it mixes of NLP and as you can see this whole diagram also explains is that there's also data science that are being applied because AI is not

AI without the data so we need tons of data so the data science that's how it comes into play okay so now you may be asking that well how can I actually work with llms that all sounds good so I just wanted to introduce to you couple of these libraries so if you want to kind of start working with it these are uh open source libraries are

kind of open to all to use l chain and you might have already kind of heard about it or work with it already so it is Python and also has uh typescript and JavaScript uh interface to that too and that seems to be the most popular at this time if you search for some popularity kind of ranking Lang chain and then also llama index is another popular

one and then there's also from Microsoft is semantic kernel and that also has a Java SDK because I come from more of the Java space so it's really good to see some of the Java uh kind of like getting into AI generative AI too and then there's also pal and that's pal came from Google and that works with this vertex Ai and then there's also hugging phase

now this one um in my previous panel I brought up hugging phase because we talk about okay GitHub is a way of helping right in the def Ops kind of process because we produce tons of code we need to version them and track them all of these other things too not just tracking but a lot of things that can be done right with the source code so

how about like in generative AI space we're talking about all of these models um generative AI models so huging face essentially allows you to kind of um you know or kind of I should you should think about of it like it's a repository that houses all of the uh generative AI models so you can get an account there and download some of these models but bear in

mind that these models are very very big you know so make sure you have enough space and also like a large bandwidth to be able to download them but there are also some smaller models too that you can kind of play around to with right for example gpt2 has a medium-size one that that's kind of reasonable size to it you can download it easily now this one

too because again I came from java if you're Java developers and these are some of the API Frameworks I like to recommend there L chain 4J and bear in mind it isn't the same as L chain um these are different groups of folks you know kind of Open Source groups working on L lch 4J I say that one is is actually pretty good you know kind of

close to the you know kind of takes the concept of the uh Lang chain and and it's actually quite um yeah good to to work with if there's time I can quickly show you but otherwise go to the GitHub you can find out more information about them and again semantic kernel also has a Java SDK there's also a j Lama too which is like a Java Port

of the Llama um kind of uses the Llama uh models but actually uh because I know the um the author there Jake is name and so um yeah he he also actually supports other models like Bert as well as um I think there's also not just llama I think he he said something else so you can go there and take a look and J Vector wanting to

point out too it's basically a lower level libraries that kind of provides um kind of with a vector a API for Java uh if you're interested it's here and also llama 2. Java that's a direct Port from llama C okay so now um this one because I was kind of presenting to a Java crowd so if you kind of want to take a look at right and

maybe I won't spend as much time it's just a very simple Hollow world the idea is that you always want to get an open AI um API key so that one is also free of charge you can register for open AI um and then be able to then generate a key and with that key you can use and of course too now that's another thing we know

that GPT 3.5 is actually much more cheaper than it is if you want to use GPT 4 so that's what I've been using too so you can kind of use that API key but of course if you use that then you're only working with 3.5 but so in this case too you can just see you know interacting with the model you can just do a model. generate

hello world and it will print that answers out for you too okay so now real real quick too I I kind of realize I kept talking because this is such a you know a topic full of a lot of information but I still want to quickly point out there's a retrieval augmented generation rag um concept that you might have seen Al floating around in this space and

essentially too what is rack let's quickly kind of talk about it it is a hybrid framework right it integrates two components of the rack models so essentially to we have to think of data that comes in we need to retrieve all of the data so that's the retrieval model and then there's also generative model side in which you know it generates all these fancy things that we

are seeing in here so you know the idea of R is that it wants to produce text right that is not only contextually accurate but also information Rich too so what does that mean right so let's kind of really quickly touch upon this retrieval model think of it like like a a librarian right is a library it pulls in relevant information from some sources it could be

a database could be also corporate of documents or like even a stream like event stream that comes in a Kafka topic for example all of these you can use retrieval that model to kind of do things and um so essentially too with what you have pulled then you feed it to the generative model because the generative model is what kind of seems to be the producer of

magic in which it acts as a writer right in a textto text kind of situation it writes and and kind of craft out cohesive you know very coherent kind of information Rich text you know based on all of these retrieve data too so essentially retrieval and generative model work together in tend them to produce what you want um so one thing I want to point out is

that this rack pattern is only an architectural pattern but it it isn't actually the you know tells it how you implement you can Implement in whatever way you see fit but essentially it follows this model of having two components the retrieval and the uh generative model so now take a look at the generative model as such you know it is like a creative writer it synthesize the

retrieved information and then produce for you like relevant uh kind of text to it too so it it's built upon like large language models too so it's basically it can create text that can gives you grammatically correct and semantically meaningful type of uh output too and it's aligned with the prompt and so now going back to Stefano's question is that yes currently too you need to prompt

kind of set up a prompt that you know you you have to know what you want and you set up the prompt uh kind of like correctly however I think I see the future as the prompts you can be a little bit sloppy and it should still give you like very accurate answers too or maybe at least suggestions of what you can do with some things too

so essentially too if you kind of think of in the r model generative model is basically gives you the final piece of the puzzle right it it's basically producing the textual output that we interact with so that's rag now why is it necessary because if you kind of step back thinking right llms as kind of humongous it it seems like a huge monster it can do things

but the the input it right is some fixed set of data now I'm talking about llms that are more general purpose right chat GPT 3.5 4.0 whatever it is so it needs time to produce to train the model so as such for example the chat GPT 3.5 it came out in November of 2022 it only has processed data up to 2021 so you may be wondering what

happened now there's there's data data doesn't take a break right it keeps coming coming so all these data in between you know the time that you produce your your model and you know the time that it's cut off then you know what to do so that's when you want to apply a um kind of mechan mechanism is using the rank model to be able to feed you

know to your um um you know kind of to to augment that's what it's called augmentation or augmented model because it augments the answers because the LM doesn't have the answer but you want to feed in external data to it so that's when you use rag model okay so now finally get to the vector database similarity searches although I realized I only have five minutes so I'll

be real real fast okay so what is Vector database so now if you kind of look at you know llms is basically a static piece of humongous set of data but it doesn't know it doesn't remember things if you ask it something you just ask it and that's it and you know where where does it go right how do I store like very often in real production

you know kind of environment you want you know you're doing queries you it's like you're going into a store to ask for help you know somebody is helping you they don't not just help you and say bye-bye you the thing is too you need to buy something you have lots of questions back and forth you test something out you ask more question so llms can't do that

by itself so that's why maybe you want to consider some storage mechanism you can always do things in memory but that will be very limited because your memory no matter what it is it can't hold everything in the whole world so that's when you want to consider storage like a vector database so Vector database is a purpose-built database that serves up Vector data type for complex machine

learning purposes now remember I said that essentially too uh you know this generative time right generative AI is that you're dealing with context that have multiple dimensions and what does it mean right so essentially you want it you know the data to be able to figure a lot of things out it's not just returning to you like you know kind of like apple and apple orange is

orange but maybe you say you know what is it about an apple you can return a lot more things there are many varieties all of these things right so all of these kind of the capability is basically can be enhanced by using a vector database um Vector search because you take the DAT first of all you work with a vector database you take the data that comes

in as input you need to basically transform it into what is called embeddings so embeddings are numerical representations of the data that are stored in Vector database so if you kind of look into it it's basically they are arrays of floating Point numbers so in other words that you know you need to rely on a model so that's what is called embedding model that you want to

say I want to convert the data using this embedding model so with that it's store into your vector database and then when there's a query coming in you you may be saying that I want to look for an apple that's kind of delicious and all that stuff and it might kind of go back and kind of do all the Sear searches and say it returns you the

red delicious apple it's truly delicious and so it does all of these things it may suggest to you well pink lady is also delicious too but maybe a little bit sour taste to the red delicious for example something like that so all these kind of way of you know allowing it to be be searched and you know kind of giving you all the context is relying on

Vector database um really quickly too if you want to take a look at Vector data I'm using a graph to describe it you can have many dimensions too but for Simplicity sake I'm using a two-dimensional uh kind of uh space to describe to you so you can have data the data itself is not just the data but what is called in mathematical terms you know if you

if you are familiar with linear algebra is basically Vector data is has a magnitude and also a um value to it too right so so essentially too that's how you can plot it on a graph two-dimensional graph so V1 V2 and 33 so let's kind of take a look at the next example it can be a cat a dog and a house so let's say cat and

dog are already stored they are that's where they live in that space and you want to kind of search for something like what is in the house and it can come back to you what is closest is the dog in this case because usually dog are more loyal they guard your house they don't go outside like cats do during the day so it then maybe that's why

the dog can be um returned too right in your searches so again this is just really really quickly kind of describing to you the mechanics of it but here I won't get into all the details but if you are interested essentially to Vector search will not just like do that linear algebra just one dimensional thing right one simple thing um like that but it's essentially go through

different layers too it does a lot of computations between different layers and then it kind of sort them all of it and gives you the answers that you want at the end so that's what all the magic is existing behind the scenes okay so examples of vector database these are open source and I like to you know this I actually take from from an OSS ranking site

mil is appeared to be quite popular and PG Vector is an extension of postgress that's something I'm kind of working on a workshop too if you're interested and there's also chroma fast from meta or Facebook um also quadrant and we8 actually there's another we8 uh developer Advocate at this uh Daniel is his name if You' have been to him his talk that was very good too so

these are kind of example of vector databases that are open source that you can consider using and what a ctor embeddings being used for for searches for clustering for recommendations anomaly detection diversity measurement classification you name it right so these are kind of very promising type of works but of course you know it's not as simple as you think to use it but once you figure out

it's not bad okay thank you Stefano said one minute so just wanted to point out traditional database you can use it to but then it doesn't know how to handle a lot of the dimensions patterns and relationships as well as Vector database and the vector search is is using like linear algebra mathematical algorithms to do it so it can do things a lot faster in that sense

okay so this is just another example a picture of the of a rack application if you're interested um I'll share my slide Deck with you really quickly benefits as we all know you can see right you ask and you shall receive but make sure to ask wisely however make sure that you are aware of hallucinations right llms are not magic it can come back with wrong an

answers or sometimes we don't even know there could be ethical concerns there are no regulations to all of these usage of llms right now you can produce your own private llms as well but how about Real Time upto-date data as I point out you can use a rack pattern to kind of augment what llms are missing too so with that I think I'm kind of timed so

if you're interested this is the okay I I better like step back a little bit if you want to um take a a quick uh kind of picture of this QR code and also the bitly uh kind of uh URL to it too now if you miss it don't worry about it and there's also some sh sharing the open source ranking information for the vector search engine

if you're interested and also I have my own stream too if you want to join me I'm trying to get back to it my twitch Channel please follow me there and also multistream to YouTube as well and with that I want to thank you and let me kind of kind of walk away but you can find me on LinkedIn you can uh find me on Twitter or

the X and also my my Discord Channel too let me let me real real quickly uh kind of get myself uh kind of uh move move myself away so then I don't block it so okay so this is uh how you can get a hold of me but I ask you to stay in touch with me and share with me your projects there there really exciting thing

so thank you Stefanos thank you um de days Europe it's been really fun here even though it's like virtual it's been kind of fun to interacting with a lot of people and uh please feel free everybody um follow me and uh give me feedback and share with me your projects on gen thanks a lot bye yeah

From event

DevDays Europe 2025

20 May 2025 – 23 May 2025

All event videos
Back to Watch