Building the next interface - AI personas who feel as natural as interacting with a human
За тази лекция
This talk showcases Anom, a company that creates real-time AI personas designed to improve human-computer interaction through more natural communication. The speaker introduces their API that allows developers to create and customize AI personas for various applications, including customer support, virtual training, and therapy bots. The importance of these personas is highlighted, emphasizing their potential to enhance accessibility and engagement for users with varying tech literacy levels. The speaker explains the technology behind AI personas, discussing the integration of voice activity detection, transcription, and text-to-speech processes. Challenges such as latency and the need for emotive, dynamic facial expressions are also addressed, showcasing the evolution and future possibilities of this technology. A live demo demonstrates the ease of creating and interacting with an AI persona, underlining their imminent potential in diverse fields.
Пълен транскрипт
Okay, sorry about that. Uh, this is why people do rehearsals, I guess. Um, but yeah, let's get going. So, first off, uh, brief apology. Uh, you've got me, not my co-founder Quea, who is listed to speak. Um, she's quite a bit more charismatic than me, so uh, lucky you. Um, but yeah, who are we? We're Anam and Yeah, myself and Quea founded Anom about 18 months ago
with a dream of improving how we talk to computers uh through real time AI personas. Uh these AI personas are as natural as interacting with a real human. Uh they're available at scale 247. But what do I mean by that? What are personas? Uh this is definitely something that's best demonstrated with a demo. Uh so this is Queever talking to Leo. Uh but let me do do
a live one. We'll try anyway. See how it goes. Uh let me just refresh that. So this is in our customer portal where you go to create personas. Uh, and we're gonna talk to Cara. >> Hey, how's it going? >> Oh, audio is not coming out. Uh, hey Cara, >> I was looking for some advice. okay, got it. What specific tips do you need for presenting to
crowds at live events? >> She's saying, "What specific tips do you events?" Oh, you hear it? Okay, sorry. Uh, okay. Would you recommend uh giving live demos? >> Let me think about that. >> Oh, um, I would recommend focusing on your audience, maintaining eye contact, and practicing your material thoroughly. And, um, definitely avoid live demos in front of lots of people. They're just too risky. Okay, it's
a little bit too late for that. Um, what if I was to tell you that you're actually part of a live demonstration that's happening right now, uh, in front of a lot of people at Dev World. >> Oh, that's interesting. Oh, wow. So, uh, in that case, remember to stay calm and engage with the audience as much as possible. And, um, just to be clear, I really
should avoid doing any live demos. It's just so risky, you know. focus on your message and connect with >> audience. You get the idea. >> So, this is just one example of one uh but we've released an API uh that lets you as a developer uh create and customize AP uh personas in a variety of ways and deploy them in a variety of applications. Uh we've only
just launched a beta recently, but we have over 300 customers using us today uh for a variety of use cases uh including virtual training, customer support, sales enablement, uh therapy bots, uh interestingly quite a lot of use cases there. um and all manner of assist assistants, language tutors, co-pilots, uh practice your interview, practice your uh your sales pitch. Uh it's been amazing to see. Uh but I
would like to take a little step back and think a little bit about why AI personas, why why are they important? Why are they more than just a gimmick? I'd also like to then talk a little bit about the challenges of bringing them into reality uh and our solution at Anom. Um, so yeah, why personas? I'd like to first touch on accessibility. So, I don't like making
assumptions about people I don't have not met, but I'm willing to bet that you lot fall into a bucket like this on the tech literacy graph. Uh, and I'd like to think for a second about the people at the other end of the bell curve who are not as tech literate as you lot. And let's focus on Europe in particular. In Europe, one in five people have
low literacy skills, never mind tech literacy. 10 to 15% of Europeans are dyslexic. 120 million Europeans have some form of rheumatic disease uh which can cause pain typing. And as you get older, the demographic u statistically you're less likely to be comfortable with textbased interfaces. And that's why you've seen a huge surge in voice interfaces. Uh 40% year-on-year growth uh in voice search usage. Voice interfaces are
also much more convenient for multilingual speakers. In Europe or in the EU, we have 24 official languages, but many many more regional languages that voice interface uh can be a huge unlock for in switching between. So in broad brush strokes, I'd like to make the case that voice is a more natural interface than text for a lot of use cases. Uh but actually face is an improvement
still. But not just any face. We need faces which are emotive, reactive, they're dynamic, can active listen to what the user saying in real time. Uh faces that can convey a large array of uh non-verbal cues and gestures uh and also detect them in the user. I should add here that every pixel you're seeing right now is generated by our AI models. But I'd like to stress
as well that this this technology is more than just just an accessibility tool. Uh it taps into something that's innate in all of us. Uh and that is our visual cortex. We are visual beings. 30% of our brain is devoted to vision compared to just 3% which is devote devoted to hearing. So how does this manifest this disproportionate part of our brain? Well, we form impressions of
a face in just 50 milliseconds. We can detect a micro expression that flashes for just 12 25th of a second. and faces are seemingly somewhat universal. Paul Ecman, a social anthropologist, did a cross-cultural study and showed that no matter where in the world you are, people will identify these these six expressions. We even see faces where faces don't exist. And don't make me pronounce what that phenomenon
is called, but there it is. Paridolia. I think we also remember visual cues far better than textbased ones. actually research here, right right here in the University of Amsterdam found that misunderstandings in text are 50% more common uh than in with visual cues. So they're much more precise. I could go on and on, but the point I'm trying to make is that we're visual beings. This stuff
is hardwired into us and it's literally in our DNA. It's why we bother coming to Dev World talks when you could just read the transcript at home in your own leisure. It's also why people have video calls and why people turn the camera on for video calls. Some interesting research from Stanford actually showed that with with just the camera on, people tend to uh solve problems faster
by 30% than just audio only communication. It's kind of crazy. Just turning the camera on makes us solve problems faster. So perhaps if we want LLMs and AI to collaborate with us to solve problems and we need high bandwidth communication, perhaps we should be giving it a face. Now, some of you might be thinking by this point, it's a bit of an it's an obvious point. We've
had this idea for decades floating around movies and books and games. So why has nobody done it yet? turns out that reverse engineering human conversation is pretty tricky. So, this is a a high level overview of kind of what's going on under the hood when you talk to one of our personas. So, on the left, uh you have the user on their device, be it mobile, web
based, on their MacBook. uh on the right these are GPU servers managed by us in the cloud and like let's quickly talk through a life cycle of a response. So the user talk starts talking to the persona the audio stream hits our WebRTC server. WebRTC is a video streaming protocol. We then have the functionality of the ears. So we do voice activity detection to understand when the
persona is talking and when it's silent. We then transcribe that into text. We then try and estimate when whether the persona should start talking or should it wait. We then pass that to an LLM to get the text response. We then pass that to TTS text to speech service. We get our audio that then goes through our animation model. Uh then through our rendering model to finally
get RGB video frames that we stream back to the user at 24 frames a second. Now, there's a kind of actually we go back. There's a there's an an easy route to doing face generation of avatars. Uh, and that is where you take a real video of someone talking and you mouth dub just 10% or so of the face uh with the to match the new TTS
audio. But this is very limited. And we took a bit of a bet early on that to really get to expressive humanlike uh personas, we're going to have to control every pixel of every frame. And that's why our models are full generation models and they generate 100% of pixels. But it turns out this can go wrong in all sorts of ways. Uh I had quite fun quite
a lot of fun last night going through our like research channel on Slack uh going way back and finding all these uh these horrendous results. Uh I also checked we have now over a thousand failed experiments on our experiment tracker. Uh we've yeah the other thing we've focused a lot on is latency. So if we go back here, this life cycle of response that has to happen
in in less than a second. And in fact, we've somewhat arbitrarily chosen 800 milliseconds as our kind of northstar target. Any slower than this and the user starts to feel a bit of a drag in the conversation. But it's interesting real world latencies between two people who click, who kind of vibe with each other, who you can let your guard down and don't have to have a
filter on what you say. People respond in as fast as 200 milliseconds. Uh this is actually faster than human reaction speed. It's h it's faster than we can consciously think. Uh, humans do this remarkable thing of kind of pre-generating responses before the other person's even finished speaking, caching them, uh, and then firing them off at a predicted point when we think the user is going to finish.
This is pretty remarkable, but it it's pretty hard to mimic. Uh, I checked this morning actually and our uh, server latency is now at this is the average in the wild across lots of sessions. We're hitting 744 milliseconds. Uh so big shout out to the engineering team for for pulling this off. Okay, how's my time? Five minutes. I'd like to quickly show you now a demo of
like how easy it is to build a persona and then customize it into your app. Uh wish me luck. So maybe I'll start hit the website. So, this is our customer portal where you go to create personas. It's about to get an overhaul. So, uh, let's go test account. Create a persona. Okay, I'm going to have to put the mic down, I think. So, I've just uh
tapped chosen a face uh chosen Leo uh and I've just put in a very brief system prompt for the LLM or the brains. Uh I'm then going to create should pop up. So, I can go in, I can chat with them, I can edit uh how I want the persona to kind of look and feel and think. Uh, but I'd like to show the integration. So, show
you the code that's required to to actually to kind of serve one of these. Uh, so this is a basic kind of HTML snippet. Uh, so if we hang on, sorry. So this is like a hello world little little example. So it's going to start a conversation with the persona and then kind of force the persona to say hello world. So, I've just tapped in the persona
ID there. Uh, I also need an API key. Might be. We'll be rotating this straight after the talk. So, don't get any ideas. So, yeah, that should be it. So, we've saved that. And now, let's go. Just going to serve that with an npm package called serve. simple one. And then that should be it. Hosted. Oh, if I can use my Mac. So then we just hit
that and we should what have I done? Bear with me. I must have made a little mistake in HTML. Oh, Ben, your confidence is impressive. It must take a lot of self asssurance to Hello world. >> Okay. Uh, hey. >> Okay, got it. >> I'm here, Ben, but judging by your speech, it seems like you might have mistaken the audience for a mirror. So, you get the
idea. Um, yeah. So, hopefully that shows you just how easy it is to get started with one of these things. How am I doing for time? One minute. uh very quickly. Uh so what's next? What's coming up? We've got some big releases. Uh we're going to make them more natural, faster to respond, easier to customize. Uh but there's actually one feature I'd like to just highlight in
particular, and that's this uh feature called oneshot personas. And this is where you can create a new persona with just an image. Uh it's not out yet, so I can't demo it properly, but like just to give you an idea of what's possible. Uh this is me animating uh uh a crude version of the Mona Lisa. So that's my driving footage in the middle coming from my
webcam. Uh and we're animating the face Um, I'd like to also point out that what makes this really tricky from a research perspective is an image obviously doesn't include, you know, things like the side of the head or the teeth necessarily. So, the model here has had to hallucinate uh those details and in a temporally smooth way so that it comes across as a nice video. Uh,
so yeah, this is literally coming out in the next next couple of weeks even. Um, so yeah, stay tuned. So yeah, I'd like to stress if there's one thing you take away from this, it's that personas are coming and they're going to make tech more accessible and more natural to use. Uh, and I'd love to see what you guys will build with it. I reckon some of
you will already have ideas on what you could do. Um maybe some of them would even be products or companies in their own right. Uh and so if you'd like some free credits, please uh hit me up or come to our booth. Um we're by the Battlebots Arena. Uh and also we're always hiring as well. So if you really get excited by this stuff, uh yeah, please
drop me a CV. Uh thank you. I don't think we have time for questions, but uh