About this talk
In this talk, Isis Rudis discusses the importance of human oversight in the development and deployment of Agentic AI. She defines Agentic AI as autonomous systems that can perform tasks and make decisions on behalf of users, distinguishing it from simpler AI like chatbots. The speaker emphasizes the risks associated with Agentic AI, including the necessity for context and memory to ensure that these systems operate correctly. She describes several use cases where her company has successfully integrated AI agents for tasks such as automating hiring processes, managing scheduling, and engaging with clients. Additionally, Rudis presents the potential dangers of miscommunications and errors in AI-driven actions, highlighting the importance of structured oversight, explicit guidelines, and the involvement of human decision-makers to mitigate these risks and ensure effective AI utilization.
Full transcript
[applause] Hello everyone. So um my name is Isis Rudis. I'm going to talk today about uh why AI still needs human oversight. So let me present myself first. Uh I uh as as as my colleague mentioned uh I'm teaching at Williams University. I'm a professor and uh it's been for quite a while and uh I also uh like eight years ago with my colleague started a company
labs where we uh decided to do all sorts of AI related projects in various uh uh topics including you know um machine vision, agentic AI and so on and so on. So what I'm going to talk today is Agentic AI what what it is the risks of Agentic AI and uh the oversight done right. So everybody's talking that agents agents agents everywhere but very few people are
talking about the dangers and how to uh make things right so that you get the benefits and you don't uh face the dangers that much. So part one agentic AI. So what is what is Agentic AI? Agentic AI we we we used to things like a chat bots. Everybody knows that you uh open a web browser enter chat GPT uh and you ask some questions and it
answers some questions. Later um the chat uh chat GPT became a little bit more clever. It started actually browsing the web and searching for information, not just uh telling everything from his own uh memory and later it started using some tools. It could write a little Python program, executed and so on and so on, but still sitting in the isolated box of the browser. Uh AI agent
is the next stage where you typically you are you have a lot of variety of AI agents and one of them is you install some tool on your computer and few few famous examples are open AI codeex or for example uh uh code their code and uh you can ask the agent to do a things on your computer. So read a file from some from some location,
you know, download some information, even install a software on your on your computer. So this is AI agents which you can use like read your emails, send send emails, you know, pretty much uh uh do a lot of things and autonomous AI systems. It's a completely closed loop where uh it's predefined and it does a series of things uh automatically and makes some decisions automatically. So this
is three stages, three different levels of uh AI. So uh let me show you several examples how we use AI in our company. So very simple thinguling internal meetings. So someone you know wants to meeting with several colleagues. It goes and uh searches for the calendar and also sends in emails to external people asking what would be their availability. collects the emails back and then reads them
and then proposes the final uh date for the meeting. Uh automated hiring when hiring people we more and more tend to uh uh look less at the CV of the person but actually at how the person is performing. We give them simple exercises. Even the hiring administrative assistance uh assistant, we give a set of tasks like create an invoice, create a simple agreement, do that and that
and that and um uh upload everything to the Google Drive. send us a Google Drive link and then you can send the agent to actually uh send the emails to the candidates uh read the results of the candidates, double check that we actually done this uh exercises and evaluate this exercises. Obviously at the very end you get a summary from agent and the decision is is done
by the by the human uh sourcing sales leads. So the agent uh searches for for prospects when you need to you know uh find a partners in some conference or something. Uh you you you send the agent to first of all search for information collect create a a database for you and then ask it to outreach this uh uh people. Jira management Jira is is a task
management system. If you haven't uh used it before uh you uh people who are working on different tasks they create create tasks in Jira and uh uh indicate that the specific task finished or not finished. So uh agent uh is able to connect to the Jira and make various u u various um uh actions you know uh change the priority log the progress you know keep the
boards updated automatically and other various other other type of things. uh Slack agents. Slack is another very popular platform for for chatting inside the company. Probably uh another very more popular platform is Microsoft Teams. Probably quite a few of you used it. It's it's pretty much the same thing. So uh agent gets integration with a with a with a chat with a slack and uh you simply
chat with your agent as with your employee you know like your colleague you know you send him a messages ask him to do certain things and so on. So we have uh uh at least two um managers uh like one is project manager who is overseeing uh one project and another agent is a claw manager just a general purpose looks after some tasks Um okay. So um
one more use case for for agentic type of work uh which was proved to be uh extremely useful. Lithuania launched a project to collect 10,000 hours of Lithuanian speech uh and um uh so make it open source so everybody can use it to train various uh models. So a lot of volunt a lot of hired people actually working for double-cheing that the speech is corresponding to the
to the uh to the text and then you need to verify this you know uh many many thousands of files. So uh you you launch the agent and you say you know go check these files uh do a speech recognition compare it to the to the to to the text recognized. Uh sometimes you do uh you you check the same text like twice and uh you never
know that the person who did the job did he do the good job or not good good job. So you ask the agent, can you just quickly compare the 20 files of this person previously before doublechecking and after double checking? How many changes we've done? Did we do any changes and so on and so on helps a lot and speeds up the process you know um from
500 you know minutes to to 10 minutes. Um uh so general u agentic uh everybody saying that we're using agentic currently the source of price water coopers says that about 79% of organizations they state that they use agentic AI and uh 88 uh companies are 88% of the companies are saying that we're going to expand the budget for uh Agentic AI for the forthcoming uh years. So
how to decide when you need to use the agent for your agentic type of work? Typically probably the first rule would be repetitive process. If you're doing the same thing or a similar type of thing multiple times, it's a signal for you to actually that uh you can try employing an agent. Obviously the first time the first time applying the uh the agent will be will be
uh not as uh you know smooth. So sometimes people even uh can start thinking that um you know I can do it faster myself than just trying the agent. But it's the same with every new thing. Yeah. Then you you know first time you you you you start driving a car you know you think no this is like a crazy thing. I would faster walk myself you
know. Uh so it's the same with with agentic type of work. When you first time start using agents you you you it's not it's not smooth. It makes some errors and you need to spend a little bit of time to get used to it. So first of all you have to identify the right process and then you have to start uh from the outcome. Instead of thinking
of steps you need to do, you can say that uh okay, I need to reach this uh point. This is my my outcome I'm I'm looking after and then uh see what the agentic type of work is suggesting you to to achieve this uh point. So the next part of my presentation is risk of agentic AI. So we have several risk. one agents know nothing about your
business. This is very typical error people do in uh in in in even with charge GPT you know they say I went to charge GPT I typed my question it it generated me some nonsense you know uh this AI is is is useless you know but the problem is that the person is not giving enough context to the agent so he can agent or AI to make
the correct decision. So the um contacts memory it doesn't the the agent has no no information about you. The agent doesn't know information about your business about your work about your family about nothing. It just starts completely fresh. Uh one uh by the way one way of of doing it is um with simple charge GPT the advice is create something called a master prompt you know um
if you if you if you check on the YouTube what is master prompt basically master prompt you spend a little bit of time by yourself putting all the information about you. So who you are, where you work, what is your family, you know, what are your hobbies, what you do on your free time and so on and so on. It takes time to put a lot of
information. But then you have one big file which is a context to any uh query. So you just uh create a project in charge or simply the drop the file with the context and then you ask your question and then you see how much time you save uh instead of you know like every time asking a simple question and then giving it a giving it and giving
it more uh detailed information. You give it a prompt the context about everything about yourself and then you give it concrete question. So again uh agent knows nothing about uh about your surrounding. Uh okay. So example uh the real world example the trade fair sheduling case the agent was adopted to find some some contacts for the for the conference. Basically I was in the conference in u
in uh Barcelona. There was Mobile World Expo not too long ago, about um a month or two months ago. No, probably more than that. And um I decided to try the agent to uh contact some potential customers of of of the company. So the agent went into the website of the of the conference, search for the various companies who are participating in the conf in the comp
in the conference selected uh like 20 or 30 potential clients and then I said uh the agent okay now go and send a message saying hello we are participating in the same conference we are sitting at the stand there and that in that location, come by, say hello, you know, uh maybe we can find some interesting collaboration opportunities. And the agent went and did this thing. And
then later I was browsing through the through the messages and I saw that uh there was one company who pretty much wanted to sell us something. You know, I checked the company and I completely didn't like it. So I was saying uh being polite just simply answered them, you know, uh sorry, we are not interested. we're not uh we don't want to meet and after this message
agent send another message again saying hello we're interested to meet come to our stand you know and so on so it was very funny situation when uh there is no context for the agent and uh uh that's how to I mean you have to be careful careful with with that um okay so the prevention how do you solve the the risk number one, you have to give
it a context as much as as much as as as as you uh uh as as as you can uh you have to maintain the agent's memory. You have to save the information about what's shared in this session. So you keep chatting with your agent by doing some task and the very end you see now remember everything we talked. Yeah. So in some uh file. So this
is a an example uh in in in black where the agent is given instruction and actually saving some memory uh some memory in some memory file. He's saving his experiences. Okay. Next we have a risk number two. Big instructions get executed literately. So we were trying to create um basically participate in some international uh project um there's a projects called like horizon type of projects and we
were searching for various uh uh participants who could be a partners in the project. So we identified a number of people and um asked the agent to go and send an emails to all these people offering uh uh you know would they be interested in participating in this uh uh project to be uh being a partner and a agent was so excited trying to send the emails
as fast as possible. So he sent uh 50 emails one after another and some spam blocking uh you know system basically blocked entire domain of of of our company you know. So so this is a real world examples how you have to be careful with with agentic type of uh work. Um so prevention you need to give a a strict rules. You need to give a strict
rules. Uh you have to have a memory. You have to explicitly say that for example some people uh already uh declined you know any communication and never email to some people so that the agent knows uh knows this information better next time. Okay. Risk um risk number three agents fail silently uh and you pay for it. So real world examples uh some company implemented some agentic type
of work uh two of the agent got stuck in the infinite loop. You know they writing something they're reacting on something and they it basically goes in circles and in circles and every time agent is working it's consuming tokens. Yeah. So you're paying for the tokens to open AI or to the entropic via API and you keep paying and employees actually noticed that the um amount of
money we are spending on tokens is growing and it was growing and growing and growing and growing uh for for for like in a loop for like 11 days and only then they realized that something is wrong because we thought oh our business is going well we have more customers that's why that's why the usage of API is increasing and and in reality it was a two
agents stuck in the loop and continuously consuming uh uh tokens. So to avoid risk number three you need uh uh uh prevention model layering. So um you need to um define the uh failures. The agent needs to be uh evaluated and agents need to know when and how to fail. The typical scenario says that if the agent was not able to complete the task in in three
times, yeah, it's a it's a point when you should give up and ask for human assistance. So uh the general statistic is let's say 70% of agent fails and you have to not only implement the success scenarios but also uh failure scenarios. Okay. Uh uh agent stopped uh stop until uh agent don't stop until the task is done. So another example which was published in in in
Bloomberg where the person the developer created a agent for himself to go and select various sources of information and send him like a summary of this this information at the beginning of the day what what's is happening throughout the the day and uh uh the agent got crazy and instead of just sending to him a via iMessage He sent this uh uh message to him multiple times
to his family me me me me me me me me me me me me me me me me me me me me me me me me me me me members to all his contact list like 500 different uh different uh uh messages to various people in his uh uh domain. So risk number four, how do you prevent it? Build for failure, not only success. uh prepare for
emergency cases and uh make sure that the human can ask can can intervene and uh human can ask some uh uh while agents keep working you know even in some uh deadly you know loop So you need to define various rules that human should be uh alerted. So let's look at this example. You have um you have uh some tasks MD description and the agent is given
this uh this uh descriptions because it gets a conflicting instructions. So you can explicitly tell to the agent if if if if you have one type of instructions and then someone from another source you get conflicting type of instructions stop and ask for the human uh feedback. Yeah. So this is example where agent is actually stopping and not processing until uh you resolve the uh the conflict.
Risk number five uh agent act without a judgment. So agent doesn't have a feelings as we have and the agent is not afraid to be to be wrong. We are trained to be as best as they could but they don't uh we don't have this uh feeling about about uh uh you know doing some uh wrong things. So example from pocket OS case it was very famous
and it was all over the news that the agent destroyed the uh production database. So basically unreversibly it deleted uh not only the database but also all the backups of the database as well. Um as I remember from this specific case uh agent was working on some problem and then suddenly he realized that there is a difference in credentials and he cannot log in to some specific
database because credentials are not matching. So he decided actually to to delete this database and recreate it again with the right credentials. and uh he didn't know that this database is as production database. So uh risk number five, how do you prevent uh things uh uh happening? So agents are not decision makers. Humans should be decision makers. Uh agents are aren't afraid to be wrong. Uh and
the simple thing is don't give a sensitive access to the agents. and also in the uh so one is a physical way you may not give a access and another is uh u like a software where you just don't give them permissions to do that. We've been discussing with a friend about this u security thing of the agents and we realized that probably the simplest approach to
the agent is actually the same as with the employees. So different levels of employees have different access. So for example you have a software developers and then you have database uh administrators. So software developer cannot go and just delete the database. Yeah you have it has to go and ask for administrator to delete the database if he wants to delete the database or create the new database
and then have to justify why he needs one. So having such similar hierarchy as humans have in in the real life to the agent is one typical way which is is is quite u uh useful. So going back to my uh prevention number five uh you should not give irre uh access to irreversible actions and um uh you should use some sort of uh model council. So
one AI is overlooking another AI's decision. So one AI is is deciding to make a risky decision and then uh another AI double checks that it makes sense and everything is is is fine. So to summarize my um risks and preventions uh is to uh say AI is powerful in turn not expert. Uh I want to remind you that when you first launch a launch an agent
is knows nothings about you. It knows nothing about your company, nothing about you. So you have to teach it. You have to give it background information. You have to teach it. You have to work with with with your agent uh as with a new intern. So let's say okay, we do this, we do this, we do this, and they say, "No, no, no, you didn't do it
right. Uh let's go back and do it again." And then you have to go through the process of of of doing several steps. And at the very end, you have to remember to tell the agent to remember everything. See, remember how we actually went through all these multiple steps and reached the final result and remember it in the file. So next time you you run the agent,
you see you see this uh MD file with the instructions how we worked last time. Read it first. Then he read it. And now let's repeat the same thing we did the last time. If you don't, you actually have to do exactly the same process again. Train the model uh train the agent you know to do certain things again. Um so yeah so people uh what people
tend to do they run the agent they expect some professional person who knows everything about the problem about the environment and uh uh they we get disappointed and we just forget about it. Uh going I saw there was a discussion about responsibility. So within European AI actu the person is still responsible for the his agent's uh damages. Yeah. So if you are the person who is using
a agent you are responsible for his u actions. Okay. Part number three the oversight um we've done several projects and uh we try to rule we we try to follow all these rules which I mentioned in the in the uh in my previous uh part of a presentation. So here's several um examples. Lithuanian municipality for social support assistance. So basically in municipalities a lot of people apply
for social support. Uh so let's say we have a new child who is born you know or we have some disabilities or some other we and the municipality is processing like 12,000 uh of these applications every year and one person is spending approximately close to two hours actually reading through these documents obtaining the proof that the person is not lying that everything is exactly as he says
and then uh you know making a So uh it takes two hours and municipality sometimes uh pro uh delays the processing up to three free three months. So we created a solution the agent that takes all these documents from various sources connects to APIs to the social security API to some other APIs collects all this information together uh reads all these PDF documents finds the names addresses
and the person's uh details and uh creates the the application uh uh in very fast like in under uh one one minute but at the very and the human is actually presented with all this processed information and have to make a final uh So this particular case worked quite well because we used uh uh several um steps. So first we have validation layer. Uh all the data
which agent found is is is double checked by the human. The process is staged. So after every step the uh results are validated. Uh one biggest problem with AI is that it's not deterministic. So this is a little bit a side note. So when you you you you have a website and then you have a button on the website and the developer actually programmed this button. So
you know when you click the button with a mouse it will press and it's like 100% true. And with AI, when you give it the same task, you're not 100% sure that for given exactly the same task, it's going to do exactly the same results. Simple example, I'm a professor at university and university students are cheating all the time. Yeah. Yeah. They take the homework, they they
put the description of the homework in charge GPT and uh uh uh we get the result and then try to present to the professor, you know, as as their own homework and they don't understand it. So uh I was working on some system which creates a multiplechoice questions uh from the homework of the student. Student submits the homework. Uh the system gives this uh homework to AI.
AI creates 10 multiple choice questions and then these 10 multiple choice questions is presented to the student. So if his student understands what he submitted, he he he should be have no problem actually checking this this this uh answering these 10 questions. And the problem with AI, I ask him to to generate 10 questions and sometimes it's out of the blue. It just generates seven. No, I
ask him 10 questions, sometimes seven. So the good mix is using traditional programming and uh AI. So you ask to generate 10 questions. It generates some sort of JSON file with with with a questions and then you have a normal old school uh uh code which just calculates how many questions are there and if it's not 10 questions if it's 11 questions then uh the system goes
back and says hey you didn't generate me 10 questions you generate me seven questions now try again you know and so you can mix things like uh uh human uh sorry not human the AI and traditional programming and obviously at the very end human oversight by the way I always check the AI generated questions at the very end when they uh generate these questions for my my
students okay so going back to my presentation uh we have staged process we have uh explicit failure modes so um inconsistencies are flagged we have uh scope authority so agent cannot not to do more than is actually restricted to to-do and we have very very clear and nice template which cannot be uh expanded. Uh another uh use case where we applied uh uh agents we created a
cargo broker. So basically the problem for the for the transportation for the manufacturing companies when they manufacture something we need to deliver this thing to the client and then they if it's especially a big thing yeah it's not simple uh small parcel where you can send it via via post or via courier uh they contact multiple uh carriers and ask them um can you what would be
the price you know to deliver this uh thing to to to to this specific city you know and so on. So um we we have a system which contacts uh some like markets like cargo LT it contacts uh the uh carriers by email. It connects it collects some uh pricing quotes and even tries to bargain with them a little bit and then puts everything ready for the
person to to make a decision. So this is a screenshot of the system. You see some you know some roads you know from here to here from here to here some of them are still being negotiated some of them are waiting for uh delivery. Um so why this system works again because we implemented all these structured and clear rules. So first of all we have rulebased controls
we have scoped access and we have a structured output not autonomous decision and everything is is is fully locked. So um uh we contact only uh approved carriers. We have a price limits. So we know then the price what would be the the range of the price. Uh so it's not like€1 million euros or something. [snorts] Uh you you compare it with the exchange price and u
the the whole process uh it typically takes about 2 minutes instead of you know 15 minutes per per per package which used to be uh before. So, uh key takeaways key takeaways of the of the examples uh from the from the from the examples we we we implemented that uh human in the loop is still necessary. Uh I heard that some um uh some companies in in
in one of the conferences I met a person from AT&T and basically he said that uh uh we use agents a lot for customer support but we have a threshold. So let's say if the amount of let's say some people are complaining about their service and asking for a refund. So if the refund is lower than specific amount then we allowed actually the agent make a decision.
If if the amount is is is big is bigger than specific threshold then the human have to be involved. So human in the loop is still necessary uh and we should make sure that the humans are doing the the right thing. I mean we're doing the decision making rather than very boring uh thing of select collecting information from various sources and putting them uh together. So let's
say instead of five unit minutes mechanical work human spends only 10 minutes just reviewing the work that was done by the agent. Yeah. So AI operates without the principle uh that uh isn't more advanced. Yeah. Just less safe. So what what we're trying to see is uh if uh someone says that they have a fully automated agent uh it's not more advanced is less safe. Okay. So
this is the end of my presentation. Uh uh I'm Rudis and we also made a QR code if you are interested in having a u a free agending AI consultation uh with myself or my my colleagues. Um you scan the core code and uh you can book a meeting and we can we can talk sometime after the uh >> [applause] >> Okay, I thanks uh for this
presentation and I think we can uh start with the questions. Uh so we already have the first one arrived. Uh how about the practice of assigning AI assistant agent a specific role like you are analytics engineer. Is it the way to go or there's a better approach? What about the practice practice of assigning AI assistant a role like you are an analytics engineer? uh yeah this is
very very typical uh uh approach then I mean it's actually in the tech textbook of uh uh chat uh how to organize the prompt correctly you give the role to the to the agent like you are a doctor or you are uh engineer and you are um you are you know some sort of uh customer support um representative and then you are given a task you know
to do and that than that. So, uh it's um it's it's very common it's very good practice when specifying the prompts uh to indicate what you actually uh what role you expect the agent to fill. It's it's it's very good uh it's it's a right thing to do. Yeah, agree. >> Okay, thank you. Uh next question. Uh do you use multiple agents powered by different LLMs to
test and validate each other? Yes, this is very actually is very useful thing because different LLMs they have a little bit different behavior and point of view. So even for example for creating a presentations I I typically have a subscription of ch and clair and uh I speak to one I speak to another you know I ask to generate some slides you know by one then review
it uh ask another one. So in the agentic type of work is very useful to to have multiple agents reviewing each other's work. So let's say if one decided to delete the database it would be better that it's not the same model reviewing the first model's work. It would be completely different model. So it's uh it's it's common practice and it's a good practice to to use
multiple uh agents. But for uh Okay, so this is the top one. >> Yeah. Uh I think we can move on to the second one to next one. Uh have you tried any offline AI? >> Yes, I tried some of the of offline AI. they typically not very uh not as good as uh as uh uh commercial AI and because they are much smaller. So typically this
small AI like four billion parameter models they are they are not suitable for agentic type of work. They are too um too simple. They they are not too experienced. So uh I tried and it uh doesn't doesn't work. Uh you have to have uh very serious uh uh uh hardware in order to to run these models locally. But uh uh it costs uh uh at least like
20,000 uh euros the the hardware the the the GPU to use it. So it's it's not very easy to try the bigger bigger models. But um for uh for uh um uh some other type of uh some other type of agentic work um I I tried using like GLM some of the Chinese-made models uh DeepS number four version four and uh they are open source but they're
too big to run locally but you can but they're still open source so if you would would the hardware you could you could use them. So GLM uh Miniax and uh Deepseek 4 Pro is u quite capable of doing agentic type of work. I use it. I think I use um GLM for my open claw uh agent uh because it's it's it's reasonably good and it's uh
many way many times cheaper than uh uh cloud or GPT agents. >> Okay. Uh uh let's go next question. Uh I would rephrase it a little bit. uh what's your vision to continuously build knowledge and decision making capabilities within Asian so over time they could be better than us. Yes, I think in the long run uh looking at the progress we made when uh the first chat
GPT version uh three uh started uh to now what is the difference between the uh between the AI uh in the forthcoming I think five years we will have very very clever agents which I think can replace quite a few few job but not now not now uh at the moment uh I don't see this is happening h I see human working human being being a supervisor
and the agent being a a worker this setup is is is working very >> okay um next question um so how do you ensure uh data quality data relevance and so forth when you uh do agents Yeah. So as I mentioned uh mixing agentic type of stuff and good old uh programming is is a very good uh uh approach. So let's say if you have some data
you can run some statistical checks you can count how many questions uh it was generated or how many questions actually it was given to the model and so on and so on. So uh whatever is is is possible to do using a software the simple IT uh programmed you know systems uh is a very good step. >> Okay and I think we have time for the last
question and I really like it because I I also face the same problem. So can you realistically monitor everything that's happening? Yeah. when especially when you run multiple agents >> uh you the good scenario is to uh uh log everything. So uh typically you don't monitor uh everything but then you have some problems you open the logs and actually check these specific places where the problem is
happening. So it's it's is is no different from the normal software development. When you have a normal software, it has logs. Nobody's looking at these logs all the time unless you looking for just keyword error, you know, something and then there is a error then you give a alert. So it's it's it's the same thing with agents. uh it's not realistic to monitor uh all the time
but it's realistic to actually monitor some um specification when things go wrong. >> Okay. So even more time uh for one more question um what do you think how long AI will still require human oversight? So will humans be always involved? U no as I said AT&T case for the small claims AI is making a decision already. So uh as uh as as the time goes and
we will see that AI is making right decisions that this threshold will increase higher and higher u my my thinking is uh within five years we will have AI that will be in comparison to Okay, thank you very much. Thanks for answering questions. Thanks for asking questions. One more round of applause.