SCaLE

Room 106 Friday Mar. 06 - SCaLE 23x

6:59:18 · 05 Mar 2026 – 08 Mar 2026 · YouTube

About this talk

This talk focuses on the integration of AI agents in open-source telephony and discusses the implementation of the OpenFloor protocol, which enables multi-agent interactions in voice communications. The speaker, Diego, explains how AI agents, such as those utilizing Asterisk, can facilitate machine-to-machine communication and improve human-to-machine interactions. By leveraging tools like agentic AI and the shared conversational space concept, the implementation allows for more efficient handling of tasks, such as booking trips through coordinated AI agents. In addition, the speaker highlights the importance of trust and security when deploying AI agent systems in various applications, emphasizing that interoperability between agents is crucial in addressing modern challenges in voice technology. Diego also shares insights from his extensive experience with AI in communications, providing examples of real-world applications and community involvement opportunities.

Full transcript

We'll get started in a few minutes here if everyone has a chance to grab their snacks. We still got some fruit snack bars. traffic. We're starting in a few minutes. Yep. All right. >> All right. Good morning everyone. Thank you again for coming to Astrocon 2026 here as part of Scale 23, Friday, March 6th, bright and early, 9:00 a.m. Pacific time. had a great uh variety of

sessions yesterday. We have a couple great AI talks to get us started this morning. We'll talk about free PBX. Have our lunch break at noon. Uh bit of logistics. If you didn't get your cheat shirt yet, please do so and get that on for our group photo at about 2:30. We're going to try to wrap things up. That will give us some time to mosey on over

to the exhibit hall which opens at 2:00. So you can check out all the other vendors over there on the other building and wear your t-shirt and hopefully after today you've learned even more about asterisk and free PBX and open source voice AI and are ready to share it with all of our new friends here at scale. And speaking of uh new friends and might be a

new friend uh uh a regular friend uh of the project and speaker at many Astracon and a voice AI pioneer uh Diego is going to be speaking about aic AI and some multi- aent conversations with asterisk and of Diego over to you. >> Thank you Chris. Thanks everybody for being here today. It's it's a pleasure actually. uh th this was few years ago actually more than 10

years ago I think it was Nastri in Orlando Florida so I was younger but it's always a pleasure to be here because we can share insights about truly open-source telephony and communications so I'm I'm really grateful and um well before getting into the weeds I have a a simple question for you because I I figured it out that I'm going during the presentation to talk a lot

about AI agents. So, anybody can give me a definition of what's an AI agent because everybody is speaking about them. Do we have a clear understanding about it? Yeah. Ju just just uh wait for the microphone please. Yeah. If you use the traditional LLM, it does not have any any data or information about the current situation, what is happening today because it was released several years or

months back. But if you use agentic AI, you can give the ability to go and do a web search and it will tell you who won yesterday's game or you can give a ability to call a SQL query and SQL table then it will make the query everything you just have to give the prompt. So whatever the uh ability LLM lacks LLM does not have information to

your personal database or SQL query table. But this with the agentic a cap capability you can extend that the power of LLM to your own personal database your own documents everything and it will query and put that into the context window and then whatever the answer you are getting is will be less hallucination and it will be correct. So authentic AI gives you all the eyes, ears,

nose, everything you can see around the world just like a human being but it's still a language based. >> Thank you for the insight. What's your name by the way? >> Praash Quitz from UCLA. >> Thanks a lot. I appreciate. Yeah, this is very true. Uh another general definition we can we can give is that an AI agent is an entity that is capable of agency. What

does agency means? Means that is able to take some actions to perform some actions. Okay, it's not just an LLM which is going to resonate. is going to you to to provide natural language interactions with agents. We can mix the LLM capabilities with tools and other um features allowing this entity to perform actions. That could be query data, that could be sending an email, checking my calendar

and so on and so forth. Okay, so the actions are the key point of an in an AI agent. Um, let me give you just my quick understanding about uh the evolution of voice communication since here we are at Astricon speaking a lot about telefony and voice communications. in the past we were used to deploy PBX and voice system for human to human interactions, human to machine

interactions. Next, basically with simple or very complex IVR, nice drag and drop IVR or even textual IVR, whatever you like. More recently, we are able to use AI agents as well. So also in this case we can talk about human to machine interactions but on the other side we have an AI agents picking up the calls and interacting with us and this part is becoming more and

more power and also deployed I would say. Finally, we have also and I will talk a little bit about this today. You u AI agent to AI agent interactions or generally speaking machine to machine This is not just something theoretical. Uh I found out recently that actually of course machineto-achine interactions are growing but already in 2024 so more or less two years ago we reached the break

even point where machine to- machine interactions overcame the other kind of interactions human to machines and human and uh of course this is going well this is 2026. By the way, there are on the market applications like Moldbook and others. Mult book basically um a social media forum to for AI agents to have a discussions uh uploading content with humans able to observe. So this is happening

right now and u on one side it comes with a lot of opportunities because with multi- aents interacting each others we can solve complex problems. We were not able to do that before. On the other side there is of course the matter of trust especially when we talk about AI and in combination with open source the matter of trust is a key important thing that we need

to consider. So Chris already presented me just few words about my history. I currently serving a head as head of AI at Tzisquare. It's an Italian company. It's mentally an international company actually and um specializing in uh supply chain software solutions. Uh we used a lot of agentic uh infrastructure right now for supply chain. Uh previously I was in charge of the AI in a customer care

software solution. I used to work on since 2004 more or less. So I deployed my piece of PBX and especially contact center solutions and of course a lot of conversational AI solutions in the in in the latest few years. Uh well the true reason I'm here today and what I will speak about is I'm also part of a Linux Foundation AI and data team working on voice

interoperability. So today I'm going to give you a quick introduction about our multi- aent vendor independent standards. So it's called open floor protocol. Um we we talk about how to deploy the openflow protocol with asterisk. So I will give you an example of deployment. um few live demos and then I will mix this content with the previous talk last year we did in for Lauderdale uh the

other opensource project named AVR agent voice responder and uh bringing everything together to have a kind of full recipe about what we can do with multi- aent solutions asterisk and AVR and finally of course how you can get involved if you're interested. Uh my presentation is full of QR codes. I believe that sharing is very important. So I'm I'd like to share all the knowledge I'm sharing

with you today. So GitHub repositories, papers and so on. So please be ready if you're interested with your mobile phones to scan the QR code. A bit of history about the open floor protocol. What's that about? Well, everything started actually before COVID at MIT research and the the project was named open voice network. Uh the concept was to create something very interoperable to have voice behave like

the web. Um the project in 2023 joined the Linux Foundation AI and data and recently last year we renamed the project open floor protocol. I will explain what we mean by floor in a few. So this is the first piece of paper I'm going to share with you if you're interested. It was published in 2024. It uh a bit about this multi- aent concept and the standard

protocol that we designed to have multiple agents to interact each others. So feel free to scan and have a look at that. There is also an extension of this paper dealing with um the floor concept and so on and Okay, I think everybody is ready. >> How you doing? >> Uh a quick note about the This is a list of the main contributors of the project. Um,

I'd like to thank everybody there. Uh, we normally meet twice a week. It's not easy to come up with a standard protocols uh totally vendor independent. We we have our discussion. But finally, we got a version 1.0 zero last year and we are we're keeping on working on the new version with new features. Why we believe that having a standard protocol to provide standard communications between AI

agents? Well, the reason is there are tons of different AI agents in the world. When we started, we didn't call them AI agents. There were chatbot, voice bot and stuff like that. But uh they they are normally deployed with different technologies simple or complex and often they cannot speak each others or it's very difficult to do that. But there is this need because no one single agent

can usually perform all the task operation we need. For example, in the supply chain, I see that a lot in the document processing, we needed to come up with quite a complex multi- aent chain to solve complex problems and such agents needs to interoperate. So main reason is scalability and openness also. So in a nutshell we released a standard uh nothing fancy it's a set of JSON

um structure uh so JSON messages and um there are messages to send communication to manifest the capabilities of agents I will give you You can also play a little bit with the standards if you're interested. Uh we we published a sandbox. Uh you can use a playground. Um you can download that from GitHub. It will provide you a Python framework very simple to use and to customize.

uh and it will come already with our API that you can test uh a a first agent which is actually a dispatcher agent is uh recognize the user intent and route the user communications to the proper agent. So we have pit a general purpose agent. Attina which is a the smart library agent is specialized in providing book informations and zos is an agent capable to call external

tools to get weather information. Okay so very simple but something that you can use to test the openflow protocol API. Let's have a look for example with Postman if I'm able to do that now. Let's see if I need to do this. Yeah. So I hope that you can see something. Okay. So for example we have AINA here the my library agent and this is the structure

of the JSON get manifest what it does it ask to this agent to provide its capabilities okay so I'm an agent I want to know okay you are Attina what exactly you do so I can send this post message and Atheina is going to reply me here so you can see here uh author libraries capabil capabilities uh a description about what the agent is doing and so

on and so forth I can I don't know I can ask aina here this is another message um it's called utterance this event type one of the features of our JSON structure is that you can use natural language inside to message to the other AI agents. So in this case, I'm send I'm asking Attina about the book, The Heart of Darkness. And so I'm sending this Attina

has an an LLM behind and hopefully she's going to reply me. Yeah. And uh so this is the reply and then explaining me that Joseph Conrad wrote this book and so on and so forth. Okay. Uh let's have a look at ZUS. So Zus I want to know the weather in where is it now? San Francisco. Okay. And says to me okay the temperature and few clouds

in San Francisco it seems. Okay. So th that was an example of how you can play with beacon forge uh to to have a a look at at the API and now let's move on asterisk. So I was wondering okay how can I use asterisk to take some advantage and uh so in the new version of the specifications we came out with an idea about a floor

which is actually a shared conversational space where multiple AI agents can interact So we have this use case for example it's a set of multiple agent we have the convenor the convenor is an entity that you can find in our specification and is the AI agent entitled to orchestrate the communication to invite pe people and AI agents to the floor also to reject them and so on

and so forth then we have of course a human that is going to I don't know send a request in this case I I want to uh ask for suggestion about trip and travel. Then we have a travel agent. We have three agents in this case. A travel agent specializ in travel information, flights, train, event agent is going to is an expert about local event, car rental

agents and and the floor. This shared conversational space. And then in the new specification, we also thought about a sentinel agent which is an entity dedicated for security which is also as you can imagine very important in a situation like this. Um so um what I came out was okay let's use asterisk and capabilities of asterisk to implement the floor in this situation. So uh there are

many ways to do that. Okay. So my way was use using the AR asterisk AR massively with web soocket to have those three different agents inside the floor. Of course I use a speech recognition engine text to speech and so on and so forth. So I have a quick demo here. Uh it's a video. So I can also play the live demo but let's let's use this

video. So uh there is a script of I I will share with you the GitHub for this. Okay. So when I launch the script here I three agents are initiated and this is the AVR web RTC client that I'm using in this case to call asteris. >> Welcome to your AI agentic booking service. How can I help you? >> Hello, this is Diego. I'd like to organize

a trip to New York Christmas time with my family. Which options do you have? >> Please wait while we put you through to our specialized AI multi- aent conference. >> Hello, Diego. I acknowledge your request for New York. Initial check suggests a 5-day trip. >> Hold on. A music festival at the selected location takes place exactly during those dates and it requires a minimum 7-day booking. So

the first one was the travel agent and then the event agent interrupted it him or her, I don't know how to call them. Um because he found out that there was a special event that the customer could take advantage of. Okay. Uh, and then the conversation goes on and on and the three agents will find the solution. Okay. Let's play it again. Let me >> take booking

service. How can I help during those dates, and it requires a minimum 7-day booking. I suggest extending the trip by 2 days to cover the entire event. >> Copy that. I'm adjusting the rental car search for the updated 7-day period. Availability for an all-wheel drive SUV is confirmed for the adjusted dates. Our specialist agents in the conference room have identified the best solution for your New York

trip, a 7-day trip. Would you like to proceed with this booking? >> Yes, please. Let's move on. >> Thank you for your confirmation. The booking is now complete and all details have been sent to your email. Have a wonderful trip. >> Thank Okay, that was an example of floor implementation with asterisk. Um the floor is deployed with a bridge area where we mix the three agents, the

three AI agents. there is an orchestrator uh taking care of the convenor operations and of course I'm using the AR because then I deployed uh in this case a Python uh project that is taking control of the dial plan and is mixing together the agents with the human and of course is also taking care of the speech recognition and so on and so forth. so on on

asterisk I'm using the re bridge capabilities. This is an example of the asterisk CLI where you can see the bridge show number of the bridge and in white colors you can see there is the human there using the PJ SIP and then there are three channels using uniccast RTP. I'm not using PJ SIP signaling for the AI agents. Okay. Um I'm using that only for for the

human. Uh feel free to clone the project and to improve it. Uh this is the link to the GitHub. I give you few seconds to scan it. I think everything will be recorded here here. So we'll be able to do that also later. Um so we have a dual bridge design. There is the main bridge and there is also an isolated bridge. which I'm using for mitigating

echo effect especially with a speech recognition engine. In this case it's based on deep grammar but you can use the speech engine you want of course. So in a nutshell I'm using external media uh so no special signaling to inject the AI agents. I don't need that. I can control that with ARI. um an isolated snoop for the ASR and of course status because I'm using the

re uh you can find here the components massive usage of websocket for the agent handler and the gram as uh speech recognition external media is used for AI agents. So again I'm not using special signaling here and uh and status. So for people who don't know about this is very powerful. On the left side you can see an example of asterisk dial plan to route a call.

On the right side there is this application called status um status sorry that I can use to provide external API and so my Python scripts are going to intervene whenever the dial plan comes there and taking and they are taking control of of the calls and of the and everything basically. Okay. Um, so again that's a potential implementation of the floor. It's very important to say that

we are vendor and technology independent. So you can deploy the floor however you like. You just need to follow few guideline in the specification. So for example the messages specific messages about how the convenor is going to um invite other agents uh or pull them out and so on and so forth but apart from that you can deploy the technology you want. This is understanding. The final

part of my presentation get back to last year when we were in Fort Lauderdale at the Astricon event and me and Joseeppe uh Ker presented the agent voice response project. Uh it's in a nutshell is a a solution for truly realtime interactions in between normally humans and AI agents. uh this time using asterisk audio socket capability not really websocket uh the project was quite successful we didn't

expect that now we have more than 600 members there is a discord community uh those are the results on the web and this was the if you like to have a look back the presentation we did last year at Astricon And uh it was it was really good. We got uh really good feedback about that. So the project is available of course on GitHub. This is the

link I think to the main web page of the of the project. Feel free to scan that and have a look. There is also a Discord community. The main maintainer is Joseph Picareri. I like to thank him very much for the uh job that he's doing here in in his during his free time. If you like to join the discord and the community feel free to do

that. So um why we think that this conver shared conversational space floor is important and in particular also multi- aent system combined with that is very much important. Well multi- aent as I already mentioned allows us to solve complex For example, we did uh two years ago with this standard open floor at the time was called open voice network an experiment an experimental phase with the Estonian

government. They use that to provide some basic service to the citizen with chatbot. There was a first level agent providing basic information and then for security purpose in this case maybe the citizen would like to have information about personal visa or healthcare information there there was an escalation to a second level of agents entitled to perform such actions. So also for security reason it's very much important

to provide the possibility for multiple AI agent agents to conversate each others but uh not just that uh I just like to mention that uh basically me and Deborah Adal the project manager of the open floor protocol released a couple of paper they are published on archive where we did some experimental empirical experiment Sorry. Uh where u we proved that with three agents, three different agents combined

together with open floor, we could mitigate significantly problems like hallucinations and prompt injections security issue. Of course, we cannot get rid of them. It's proved scientifically that it's possible because uh AI is thinking probabilistically not deterministically. But again, multi- aents can also help you with security and hallucination mitigations. So um the full recipe of my presentation today is based on three pillars. We have a standard interoperable

way to have multiple agents interacting each others. The open floor protocol providing also the new concept of shared conversational space where we can have orchestrated securely by a convenor a sentinel agents and so on. We have asterisk a universal media gateway uh which is kind of perfect to deploy an example of floor to have as you have seen in the in the in the demo multiple agents

conversating each others to finally find the best solution for a user. And then you can combine the agent voice response project to have to to provide a truly realtime experience to a human interacting with AI agent. Contributions uh are welcome. Uh at the Linux Foundation AI and data team voice interoperability AI we normally uh gather twice a week. On Tuesday there are two kind of meetings. uh

one is more technical for the specifications update and the other one is a little bit more strategically we discuss about events and uh papers and other stuff. So feel free to join also few few meetings if you like. You can find information on our website. It's called voice interoperability the team. Feel free to have a look or contact me if you want to join. Okay. It's totally

it's totally free and any comments and contributions are welcome. All right. That brings me to the end of my presentation. I'd like to thank everybody for for being here today and I'm looking forward to answering any questions if you have. Thank you so much. >> Thanks Diego. Very edifying. Um question back here. >> So I have a question. We built the same thing. Well, very similar to

what you have. We have the technology to be able to open the channel, speech to text, text to speech, send it to an LLM. What I found though is that the problem isn't in the technology. It's in the application of it. And the biggest problem that we see is the human. How do you handle um a nobody picks up the phone and says, "Hey, I want to

book a trip here on these dates for this thing." They say, "Hi." "Oh, are you a robot? Okay, what can I do? They need that direction. So, what are you guys doing to play with that to get the human to interact? >> Yeah, I mean that that's a potential issue that in my opinion is improving because the technology is improving and now the natural interactions are are

better. Uh but it's important to say that our open floor protocol is not taking care of this. Okay, it's totally again vendor and technology independent. So we we don't take care of such technology improvements. You can build your own AI agents. It's not about our standard protocol. Okay. Our standard protocol is designed to have your technology to speak with some other technologies maybe. I don't know. I

don't want to mention any vendors that built with some logics. Okay, back to your questions. Yes, that's understandable. Um, because AI think in a way it's not just about voice. It's also generally speaking about multi- aent systems where they would like to be fed with information uh better injectable by AI. Okay. But uh yeah, it's a potential issue that is going to uh not get rid of

but I think uh we we are getting in a situation where I I can see now a lot of practical applications in productions. Now this is actually it's been one year now that I'm starting to see customers moving from proof of concept to real production system. Maybe with some issues, okay, but they are in productions. Thanks for the question by the way. >> Great question. >> I

think there is a question here. >> Yep. Right here. >> Oh, there. Sorry. >> Thank you for this presentation. Very enlightening. Um, if I have a let's say 10 of these agentic conversations going on and I want to keep an eye on one that seems to be going astray and I want to jump in. Is is that kind of uh administration or visibility built into the protocol

or how would I get to know here are the 10 people who have called in. They are going through these conversations. I want to see what they're talking about so that I can kind of jump in the conversation where I feel like okay that's not going the way I anticipated it to go >> or is that something that sort of we'd be building as a application on

top of this and then decide what to do right some sort of governance keeping an eye on these conversations >> yeah you may have a look at the specifications there are really some JSON uh messages that allows you to uh understand who is on the floor uh up to some extent. Okay. After that, it's really up to you to deploy the proper in this case probably convener

or sentinel agent depending if you are looking more on the orchestration side or on the security side to monitor what's going on there and so to understand about the content. It's not really inside the >> Just a peripheral question on that. Uh I went to the AVR website. There is a demo page. How do I log into that demo? Do you know the creds? because it's not

>> there is there there is if you look at the website there is the link to the to the GitHub okay but feel free to message us on Discord is the better way to do that thank you >> just a sec >> your hand up first my question's hopefully simple um which AVR system do you think handles uh accents in different languages best so for example you

have an accent in English and I know many international people have find that challenging. of course I I cannot provide you the name of some specific vendors now, but uh the point is there are some uh speech engine that can be trained a little bit or not trained fine-tuned in the I word I can use this language so that I don't know maybe I have a Vietnamese

ascent. Okay, what you need in that case is a bunch of Vietnamese conversation with English or English with Vietnamese as conversation, some thousands of them hopefully and then you can feed the AI models of the speech recognition engine to learn about that. That's normally how it works. Okay. So on one side we have a speech recognition engine that are constantly improving and also they are very cost

effective now. I remember 15 20 years ago the speech engine were so expensive and so closed. Now they are really accessible on one side. On the other side there is for some of them the capabilities to make a little bit of training to improve the the recognition in specific case. Thanks for the >> Thank thank you very much. Um in terms of a deployment for uh open

floor, have you uh what are there metrics? Um and if not, could you uh what would you think of uh Prometheus metrics would be appropriate for an open floor deployment? >> What do you mean by metrics? Uh exactly. Um well that's part of the qu I mean uh you've got multiple agents uh you've got uh engagement streams uh you have uh I assume you monitor the human

entry into the floor uh and their engagement pattern and departure differently than the agents themselves. Um but in terms of a deployment and I'm trying to think in uh if you were doing a um if you needed to monitor okay you you are thinking more about performance metrics >> well yeah what kind of metrics would you want to watch to make sure that the floor wasn't jammed

or in a problem or something >> yeah just a general observability uh off the top I we're probably I don't know if there's a Prometheus metric thing there but um if there were what that open floor hadn't gone bad? >> Yes. Um, of course there are many metrics we can consider. Um, if you if you look at the two papers here, for example, looking at the hallucinations

and especially the security problems, prompt injections is one of the many, but it's one of the most important one. Those could be critical in a situation where we have multiple agents not properly governed governed. And uh in this paper you will find some formula a matrix that actually uh we came out because there was no much in literature to do that and we measured with multiple agents.

Okay, this is an example. Um, other example and I don't have the the paper here. I will show that tomorrow. I will be in another talk uh at the main event uh speaking about open floor applications and another example of open floor. uh we'll share with you the more recent uh papers we released where we deployed also some caching capabilities and semantic learning inside the multi- aent

chain and that was helping us one for observability and second also for performance because of course I didn't say that but with multi- aents we can solve more complex problems on one side but the complex complexity increases and also the performances can be affected if we don't design very well the system. uh there is we we would need one day to talk about this but there is

the problem of um AI context uh how they call uh overwhelmed so too much context on these agents and everything can get wrong there are many issues that we need to consider yeah it's a good question we need more time to to get that so those are just a couple of metrics that you >> that's a yeah great question I think we maybe have time for one

more. Um, if >> thanks for the demo. I think the demo is really good. Um, by the way, my name is Sanjay Yadav. I'm the head of engineering at Parliament Corporation. We build conversational agents. Um, the question that I had for you is that in this conference kind of setup, it is completely possible for these agents to get into conflicts. If that happens, have you thought about

how to really resolve that or just wondering? >> You mean from the security point of view? No, just from because all of these agents are trying to solve something, right? And just like humans, it's possible that they can get into some kind of conflicts. If that were to happen, how does the agentic framework goes about solving them? >> Uh so the question is about say say the

question again in another way because I >> Yeah. Okay. Well, so, so basically what I'm thinking is I am thinking that there's conference room just with humans and when you are trying to when you uh when you want them to solve a problem, people have different opinions and because of that there's there's a natural tendency to have conflicts. I see that happening even in the agentic situation

as well. And if that if that were to happen, how do you govern that? How do you manage So you mean uh the possibility to get biased or something like that for a second? >> Something like that. Yeah. >> Um I I have a thought we and we've got to wrap it up. I I'd love to hear more of your talk to go. >> No, simply put

it's it's something really related to the agent implementation this kind of things. Okay. So how you can um have the agent performing better limiting the bias. But on the other side again if you use multi- aent in a proper way you can have checkers you can verify what the agent before this agent was saying. Double check the fact and improve the answer trying to limit for example

hallucinations and biases. Okay. So that's what we found out in a multi- aent capabilities. I >> I think maybe if I could just help wrap it up a bit um for our next speaker to to come on that uh I I do find fascinating some of these ethical uh moral discussions about the implications of these systems and uh what other human systems might be effective. Uh so

the first thought that bubbles in is uh Robert's rules of order newly revised. If anyone's a fan uh of uh speaking limits and times and you know maybe some of those more of those ideas you know we we can discuss as part of this I think it'd be great uh to continue this conversation with our next talk which is also about some practical applications of voice a

technology. So thank you again Diego. >> Thanks a lot. Thanks a lot. Just to adapt a very last uh QR code, you you you will find the new publications there. Okay. Also the semantic learning and cing for AI agents and everything there. Feel free to get in touch with me if you have any questions. Thank thanks a lot. >> Yeah. Thanks. Thanks. And our live stream uh

of course you can watch on YouTube. If you're in the back and can't see that uh you can pop it on your phone that is broadcasting the slides uh quite clearly as well from the scale website through their YouTube channel. So, just a minute. We'll get uh started with uh Jeff's talk on uh some voice AI. So, just uh refresh your coffee and we'll be back in

a minute. Test test test. Everybody hear me? All Okay. Just uh get started here. Thanks a lot again everyone for coming to as part of scale. My name is Chris May, open source solutions advocate. If we don't get a chance to talk today, uh I'll be around at our booth or various sessions tomorrow and Sunday as well over in the exhibit hall. So our next talk on

the voice AI track this morning is uh Jeff Lusier with Stratus talk about the local AI receptionist. Jeff, >> Thanks. Thanks. Um Jeff Lorier, is this I don't know if I can hear quite right. Is it good? >> All right. I'm the CTO of Stratus Talk. Um, we've been a hosted PBX company for going on 20 years now. Uh, we're mainly based in the US Virgin Islands,

which is where this picture came from. I like to kind of rub it in that I live on a boat in St. Thomas, and that's where I do a lot of my development. It doesn't have quite the punch in Southern California, I suppose, but uh, yeah, that's the view out out of my boat. So, um, we're a hosted PBX platform for resellers. This is kind of important.

We never really figured out sales and marketing. So, we're really reliant on our our dozens of resellers that go out and actually do the sales for us and and the frontline support. So, there's only a handful of us that actually work in the company. Uh, and most of the work is all done by our resellers. We're an AWS hosted Kubernetes cluster. Uh, we're not running EKS though.

We we uh created our cluster with Kubspray and if anybody has questions about Kubernetes, I'd be happy to discuss Kubernetes with you. I've been a Kubernetes expert now for almost 10 years. Uh and it made a huge difference in our cost structure. Um we cut our AWS bill in half by going to Kubernetes and it's really a great thing to do. So in our uh architecture pods

are free PBX instances. We did the work to break out free PBX into multiple containers. So there's a Mariah DB container, there's a Prometheus container exposing metrics, there's uh now of course an agent container, and we're fully white label. So our resellers have a dashboard, which I'll show you an example of, and we make it really easy for them to add features to the the PBXs that

they're selling to their customers. And we have DIDs in the Caribbean and in the mainland USA. We did start life down in St. Thomas and the US Virgin Islands, and more than half of our revenue still comes from down there. So the motivation behind this project uh we knew that we wanted a voice agent. We actually lost a few customers a few years ago because we didn't

have such a product and we're accosted all the time um especially in the last year by people that want to sell their AI agent to our customers. And one of the things that really stood out to me as all these people contacted us is that number one they're very shallow shallowly integrated. So, of most of the the folks that came to us, they're operating on Twilio. Um,

they're offering an agent that has a public telephone number, and I asked them, "How are you going to transfer to my internal extensions?" And they said, "Oh, it's no problem. Just give that extension a DID number, and we'll transfer to it." Well, I mean, that's that's not a great situation, especially if you've got hundreds of extensions. Uh, we're not going to be able to provide a DID

number for every one of them. Uh, and there are other things that you might want a receptionist to do inside the PBX that this kind of integration just won't be able to provide. Um, it's kind of like giving a cell phone to someone and asking them to be your receptionist from outside the company. The second thing I noticed is that because they're all based on Twilio and

other APIs, they're all permanent costs. They're paying for tokens. They're paying for minutes from from Twilio or from whatever uh provider is providing the telephone number. and the permanent costs can add up. And I don't know about your customers, but my customers do not like to pay permanent fees. U they they are very adamant about costs being predictable and affordable. So when we started running the numbers

and seeing, you know, if we're going to provide a reservation agent for a hotel, how often is this agent going to be online, how much is that going to cost? Uh sometimes that came to a very large number. And finally, compliance and privacy risks. So a lot of our customer base is healthcare and legal and there are HIPPO uh requirements involved. So we want to make sure

that none of these conversations hit public APIs and that was really the biggest motivator. So our solution um we wanted our agent to be deeply integrated and for us that meant that it should run ideally on the same machine as asterisk. So in this example there's a server we have asterisk running on the server and then there's a docker docker uh container running the agent process and

they talk to each other over loop back and that's about as deep and secure as you can get I would say. Um, so by being inside the PBX and uh actually being exposed as an extension of the PBX, it now has direct control over everything that uh a receptionist might have control over. So if you want to do transfers, even attended transfers, that shouldn't be an issue.

uh if you want to um do conference calls, you know, all all the things that a receptionist might might do for you, the agent will be able to do because it's inside the PBX. And then of course, it's flat rate pricing that we're able to uh provide because we're going to provide the GPUs So, we're offering our resellers this simple control, powerful customization. This is our dashboard,

just kind of a snapshot of the agent module. So, they can define agents. Um, they can see all the agents that they've defined. They can choose voices for them. They can choose specialties. Um, the different specialties have to do with what system prompt is going to be chosen and maybe what tools will be available. and uh they can then send sell this as a feature to their

customers and and make revenue that they weren't making before. So to get started, we had a design and it's the familiar pipeline. Everybody's talking about ASR, LLM, TTS. It's not the only solution. In fact, I did look for a little while at some of the emerging models that are speechtoech. The problem I had with the speech-to-pech models is how do you get in the middle? How do

you make decisions? U if you're just going to blindly let it talk back, then uh you don't have a whole lot of control over what it's doing or or even to to be able to get it to go in a different direction. Man, I hate this thing. It's in my face. Um so we rejected the speechtoech models. I don't think that's really the right way to go.

Or maybe we're not quite there yet. Uh maybe someday, but I don't think that's ready. So if we have to deal with this ASR, LLM TTS pipeline, then what are the pitfalls? I mean, it's not very difficult to throw these three pieces together and just connect them. But what you end up with is a whole lot of latency, uh hallucinations, uh issues with tool selection. All of

these things are pitfalls. Um, and the latency being the biggest one of them, anything over 700 milliseconds just sounds strange. Like if you're waiting and there's nothing but silence, you're tempted to hang up. Uh, and depending on how many times you're calling the LLM and and some other things, uh, that can be pretty bad. So, I'm going to talk about how we solve some of these issues.

So, I wanted to say that uh when I first put this together about eight months ago, I actually did write the initial code uh and I I wrote the code that brought all those three pieces together and I ended up with an agent that had horrible latency and all these issues and then I discovered cursor as I'm sure most of you have also discovered. And I don't

know if you've seen this meme, but this is my experience with cursor and vibe coding trying to close the blinds. That's exactly what vibe coding is all about. Um, so I think this is hilarious. it's not that using cursor and vibe coding the rest of the project was a bad choice. it was actually definitely the best choice. Um, but getting around this kind of issue uh really

required still a developer, right, that knows what they're doing and knows what they're trying to accomplish. Um, and I just wanted to point that out. Well, now how do I get past it? So, getting started, what kind of hardware? Uh, we knew that we wanted to use local GPUs. We had on hand a couple of Tesla P40s. And if you're not familiar with these cards, they're great.

They're 24 gig of VRAMm. They're slow. It's Pascal architecture from Nvidia. Um, but it can run all three pieces, uh, which is pretty amazing. And the first iteration of this actually did run on the two Tesla P40s. Um, ASR, as a matter of fact, we're using Faster Whisper, which I'll talk about in a minute. That can actually run on CPU without any degradation. Uh, and in fact,

we use faster whisper to do our voicemail transcriptions. And that does run on CPU. I didn't bother putting that on a GPU box because again, we don't really care how long that takes, right? TTS we discovered uh when when I finally settled on a TTS model and we're using the the Moshi project, which I'll talk about in a minute, that does require newer hardware. It would not

run on the Pascal architecture. So then I started playing with some other cards. Um, I bought a 5090 right when it came out and I'm regretting that a little bit. I spent way too much money. Uh, it runs great on a 5090 obviously. Uh, but we also got a 3080 and I played around with that a little bit and it runs on a 3080. So, I guess

the point of all this is you don't need really expensive hardware, right? The Tesla P40 was like $250. So, you can grab a handful of those and implement this project without any difficulty. The 3080 I think was just a little bit more expensive, maybe 350. So no big no big deal. Now in the end we ended up with L40S. Um so we're planning to provide this at

scale. Uh we want to be able to run not just one agent on these machines but multiple agents. The L40S is uh supposedly the best at inference in its class. So we've got three of them now. And I'll show you how we've split that up. it does make sense to split the load if you're going to do some kind of bulk deployment. Uh I didn't want the

LLM maxing out one of the GPU nodes while it's also trying to do ASR TTS. So we've got them running on three separate instances. And then finally uh as I mentioned we're a Kubernetes shop. My first instinct was to bring the GPU nodes into the cluster as GPU nodes. Kubernetes has some some difficulties in sharing GPUs. So out of the gate uh the problem was I could

run one pod and one pod only on one GPU node and then it was restricted. So um that wasn't going to work at all. And in the end we decided that they didn't need to be in the cluster. The APIs that they expose can be outside the cluster just as easy. So that's what we've done. connecting from asterisk. So the the most obvious thing to do is

to use websockets. Uh and when I first started the first iteration of this, I was using the websocket command uh inside the dial plan and just dialing up the uh the URI for the agent which was listening on a websocket. And the agent was then using AMI just because I'm old and have been using this for a long time and I knew AMI. Uh it was using

AMI to control the channels and that worked for the most part. Um we could have a conversation. There wasn't really a lot for AMI to do until I started implementing tools. So blind transfer worked no problem with AMI. But when I started just uh trying to figure out how to do attendant transfer where you're going to have to use multiple channels and bridges and and work out

who's going to be connected to what and at what time. Uh AMI fell over. It can't really handle multiple channels. So we ditched that and we went ARRI, which I've been avoiding for a long time, but finally I got into ARRI a couple of months ago and we were able to finish the tool for attended transfer. ARRI is amazing. Uh, one thing I will say is that

when I started playing with ARRI, uh, I managed to crash Asterisk, which is the first time I've been able to do that in gosh, at least 10 years. Asterisk to me has just been as solid as a rock until I started using ARRI. And I've I've randomly had crashes that uh that I need to get um uh I need to get diagnostics back to the team so

they can figure out what happened. But we've we've been able to work around all those and uh and ARRI is still the way to go. You got full control. You can do really anything you want to with the audio and with the channels and the bridges. Now, all that said, I think that we're not going to stick with it. Um our our next phase is to implement

SIP. And this is mainly for our resellers. Our resellers are really used to working with the SIP protocol. They're really used to handling lots and lots of SIP phones. Maybe they've got autoprovisioning already set up for their SIP phones. So, we want our agent to look like a SIP phone. And I think that's going to be our next iteration. Possibly WebRTC, but I don't think so. I

think it's going to be SIP. And then I've got I don't know if you can see it, but I've got a couple of screenshots of uh the configurations. At the top there's um the websocketclient.com, which you've got to have an entry in there for um for the URI that the agent's going to listen on. And then in from internal custom, I've got the extension that it's going

to connect as. So I've exposed extension 222 for the agent and that just answers the phone and goes into the stasis application. Uh that's how ARRI works. And the last one is uh I don't remember what the last one's supposed to be. At any rate, um you know, if you're interested in implementing this, you can definitely discuss with me how to connect it. Um but the demo

works over ARRI like this. So audio madness that we had to go the Asteris websocket is providing audio in SLIN 16 uh mainly because we asked for it to and I did that because I knew that uh faster whisper wanted SLIN 16. So it's not really a problem for the websocket to provide it in exactly the format that faster whisperer is looking for and that's kind of

important. So in the design we have u the audio coming through the websocket from asterisk and being um piped directly to the ASR over another websocket and that websocket to the ASR uh container stays up for the entire session and that's important. uh there's actually a thread that gets uh spun off by the agent process to handle ASR continuously. The idea being that you can interrupt the

agent, right? So anything that the caller says is immediately transcribed and then we can use that information to potentially interrupt the agent that's speaking at the moment and react to what was said. TTS on the other hand, so uh we've again decided to use the Moshi project which I'll talk about in a minute. Uh that sends audio in 24 kHz and that was a problem. So the

24 kHz has to be downsampled in order to be sent back to asterisk and ideally in SL16 again. So the way to handle that we uh we have another thread. This thread implements a buffer, an audio buffer. And this got really complicated, but the idea being that the the buffer is always filled with something. Maybe it's silence, maybe it's keystrokes, and maybe it's packets that are coming

from TTS, but it's tasked with constantly sending a packet back to asterisk every 20 milliseconds. So even if there's nothing in the buffer, it'll send a silence packet or a keystroke packet. In the meantime, the buffer is filled by the TTS process. So if there's nothing to send, then it gets filled with silence or keystrokes. Uh if there is something to send, then it just starts filling

the buffer with 24 kHz, but downsampled so that the buffer can then be read back out as SLIN 16. So in the end, that all that all worked fine. think I covered all that. So, ASR, all of these are running as Docker containers and I have images. Um, I didn't put it in the presentation, but if you're interested in in our images for these different pieces, please

just contact me. So, ASR is faster whisper. I did play with regular whisper. I haven't played with anything else, although um I am very interested in parakeet, the Nvidia uh ASR, and I think that's probably going to be our next piece that we we test. It's super super fast. I mean, it's almost instantaneous. So, when we're talking about latency, very little of it is coming from the

ASR process. So, when you speak, the transcription comes out almost instantly. Uh silence detection is really important. And as part of the ASR, you've you've got something that you can define that decides when the user has stopped speaking. And that needed a little bit of tuning, but not much. It pretty much worked out of the box. Um, a few things about faster whisper, though. It's not very

good with noise. In fact, for some reason, every time it hears a a loud noise or, you know, some kind of bang in the background, that somehow gets just transcribed as thank you. So, our agent says, "You're welcome a lot." Uh, which can be hilarious. Uh, multiple speakers is also confusing. So, if you've got a noisy background, you have music in the background, or if you're on

speakerphone trying to use faster whisper, it's not great. It's not production quality. So, I would say we're we're still looking for a good ASR model. Although, the accuracy of the transcription is really really good and it's fast. So the LLM we have somewhat settled on Quen 3 8B. I know that I wanted something really small. An 8B model should have uh super fast returns on a card

like an L40S and it did. So while we were testing this and playing around with it, uh the latency that we were experiencing was somewhat coming from this but not much. So the responses were coming back so fast and in fact we stream the LLM results directly to TTS uh and I'm going to talk about that in a second. But when we we stream it we're not

waiting for the LLM to completely finish whatever it's generating the TTS starts speaking immediately. Uh which is really key for minimizing latency. You can't wait for the entire LLM response to come back even with an 8B model. Uh, now that all worked great and we were impressed with the the latency from the LLM until we started adding context. And when we got up to about 4K of

context, which isn't hard to do, um, the time to first token just cratered. So, we went from having an answer in a little less than a second to the answer maybe taking multiple seconds. This is really poor. Apologize. so it took me some time actually to to track down what was going on there. Uh, I had arguments with the L40S supplier thinking that they had done something

wrong and that they' connected the card wrong. And uh, I went through all kinds of things to finally discover that the problem was KB cache. So KB cache I I didn't realize is really only looking at the beginning of your your query. So anything that changes in the front of that query will completely kill your cache retrieval. And most of our dynamic stuff was up at the

front and that was our problem. So if you're trying to put this together, keep that in mind. Uh I'm going to show you a slide in a minute that how we've arranged our queries so that we can avoid this and optimize KB cache and that pretty much solved the time to first token problem. So the next problem was auto tool calling versus uh I got another slide

on that. Auto tool calling um is something that's provided by Quen and you can just turn it on. Basically you're you're giving a JSON uh definition of all the tools that are available in the query. and then it decides whether or not one of them should be called. I'll talk about that again in a minute. Um, we had some confusion with the user supplied context maybe not

agreeing with the system context. Right? So when we have a system prompt for a skill like an AI receptionist, it's quite a long system prompt that we've come up with to try to steer the model in the right direction and having these But what if the user then supplies some context that overrides that or disagrees with it in some way? Um, and we've had some issues with

that. We've also had some issues where in the user context you might want to give rules like uh maybe you don't want to transfer calls to people that aren't working. And so if you know in the context the hours of your employees and when they're working and somebody asks to be transferred to someone that isn't working at the time, we'd like it to be able to decline

that transfer or you know have some other method take a message or something like that. Um but what I found is that when we start putting rules for tool use in the user context, sometimes the model decides for whatever reason that a tool should be called. Um so often we were having conversations that did not request a transfer and suddenly it's transferring the call. So those are

things to keep in mind. I think some of that is going to be be solved when we do our fine-tuning which I'll talk about in and then multiple LLM calls. Um, so if you want to be able to call tools, well then you're going to at least call the LLM twice and that really kills your latency. Right now we gone from one second to maybe multiple seconds

and we've got to fill that space with something. So we'll talk about that in a second. Um, but we really want to avoid multiple OM calls if if you can. And one of the issues with auto tool calling, I think I say it here. Whoops. One of the re one of the issues with auto tool calling is that by default it puts the um the tool to

be called at the end of the response. And so that's a problem, right? If we have to wait for the entire response to come back before we know if a tool call is required, then then we've already missed the chance of uh only calling the LLM once. And we got around that actually with some prompting to to force it to put the tool call at the beginning

of the And then we know that we have to stream the output to the TTS. We can't wait for the end. So this is how we're arranging uh the context to make sure that the cachable stuff shows up front. So the system prompt, the tool definitions and the user context should be pretty static across this conversation. So the KB cache will be u uh primed as soon

as the first uh response comes back and then as you continue having the conversation you should be able to to take advantage of that cache. Environmental context things like the date and time, the day of the week that gets added next, the conversation history, the current prompt and then finally any tool results. So if we do decide that a tool is is required then we throw out

whatever uh it's responded with at the time uh run the tool and then call the LM again and those tool results end up at the bottom. So TTS we've uh somewhat settled on the Moshi project the cyute.org or uh uh TTS. They have actually all three and you can use the Moshi project uh as it stands. I think they've even got a way to to just do

that as a container. Um and they're the ones that uh pointed us to an L40S. So, they claim they can get 50 agent conversations on one L40S card. So, I'm I'm counting on that myself. We haven't tested that limit yet, but I'm hoping that we can do that, too. What's really great about it is that within 200 milliseconds it starts speaking. So you send the the stream

from the LLM directly to the TTS and it'll simply start talking. And that's how we did it at first. Uh we didn't try to buffer anything. We just sent the stream directly. And for the most part when the the stream's coming back quickly, it worked fine. But it didn't work great for the emotion of the voice. So what we found is that the the voice requires some

amount of phrasing ahead of time so that it knows how how to structure the response. And if you don't give it that then the voice gets kind of flat. So odd audio effects when it can't produce fast enough. Uh it'll also slow down a little bit. Um so she'll start talking slower. So instead of getting pops and hisses and breaks in the audio, she starts talking slowly,

which is interesting um and probably preferred to audio effects, but uh but something we'd like to avoid. So in the end, we decided to buffer and what you're giving up is some latency, right? So instead of 200 milliseconds, maybe it starts talking at 500 milliseconds. Uh it turns out that's still okay. In fact, anything under 700 milliseconds is actually pretty good and it still sounds normal. So,

we decided in the end after trying buffering by word, by phrase, by sentence, we settled on phrase. So, it usually goes three or four words at a time. And instead of streaming over a websocket to TTS continuously, it does it kind of by phrase. So we make uh we make one call to the TTS, send the phrase, send the EOS token and close that websocket. And that

seems to be enough for it to prime that audio buffer that I talked about earlier so that it keeps it full and we we have no audio effects coming out to Now we don't like to hear silence. So, if you are waiting for two LLM calls because a tool call was required, the last thing you want to do is wait for two seconds for her to say

something and hear nothing but silence on the line. Um, and I'm sure what you've heard if you've talked to any agents uh out in the field, there's key clicking like they're typing or there's maybe the sounds of a call center in the background, something like that. People are doing this to get around this latency problem. uh and it really fills the silence in a way that doesn't

bother you like the silence would. So I gave a quick command here. We we found some open source keyclick audio. I converted it to Eslin 16 so that it could be loaded and and basically what happens is um I've got a uh a buffer that it reads in the beginning when it reads this keyclick file in and it takes 20 millisecond samples from it and instead of

silence it'll stuff these 20 millisecond samples of key clicks in the buffer instead and what you end up with is some key clicks instead of silence and then she speaks which works really It's important that you're able to interrupt the agent. Sometimes the agent will go off on a tangent, speak a lot, and you're not interested. Uh, you want to change its direction. You should be able

to say something and it should be able to respond immediately. And that isn't too difficult. Um, you've got to make some choices about what to do with all these async processes that are going on. So when you receive some transcription in the middle of uh of the call and the agent is speaking, we want to stop the agent from speaking. Now, you don't want to just clear

the audio buffer immediately because then you get pops and hisses. It's got to be a way that uh she kind of finishes a word, right? Or uh you don't want it to break in the middle of a of a word, for example. So there is some some uh some code in our agent that will spin down the audio coming from TTS and kind of slowly um uh

volume it out so you don't hear that pop and hiss and she just kind of uh fades away and then you can interrupt and she'll start responding to that instead. And I say she because we've just used a woman's voice and I just think of her as she. And then uh conversation flow. Do you acknowledge that you've been interrupted or we've decided no? Um but I don't

know if that's the right choice or not. Tool calling. So once we got conversations working, it's clear that we needed to be able to to call tools to do things. a receptionist needs to be able to transfer calls at the very least. So that's when we discovered that uh that auto tool calling spits out the JSON block at the end of the result and we had to

fix that which we did with some prompting. Uh if there is a tool then what do you do with the response that came back and in the end we decided to toss it like whatever it decided to say in addition to asking for a tool to be run is kind of obsolete at that point. So, we just throw that out, run the tool, and call the LM

again, basically with the same prompt, but including the tool results. Uh, tool mistakes. So, an AP model is kind of dumb. Um, and we've had quite a few instances where it decides to initiate a transfer for no reason. Uh, and some of that came from I I decided in the end the the difference between the tool definitions that are sent at the beginning of the prompt and

maybe some tool rules that we tried to mix in with user context. So by playing with that a little bit, we made those uh not happen as often. But we still had to implement something that would check to make sure the tool really should be called, especially transfer. So, there's some hard-coded stuff in there that I'm not happy about that uh that if a tool to transfer

is called, it actually goes back and looks at the prompt and and deterministically decides whether that was really a transfer request or not. And if it's a mistake, then what do you do? Well, now you've lost half a second to a second of of latency, which is unfortunate. Um, but the way we're handling these mistakes is just to call the LLM again. and then there was a

a time when I I just wanted to say that simultaneous queries like autotool calling is something that's provided by the model and I thought maybe maybe we should make a different LLM call to decide if a tool is is needed. So rather than auto tool calling maybe we could call an LMM a different LLM to decide whether or not a tool is required or not. um that

gets pretty messy. So trying to handle all these uh async calls going out and mixing them back together again, it was too too messy. So I think that the answer is really in the fine-tuning, which I'll get to in a So I'll talk about the transfer tool. Blind is super easy. In fact, we could even do it with AMI. Whereas attended was exponentially more difficult. Um, blind

I had working in a day. Attended took me several weeks. So, attended, I think I've talked about most of this already. Well, blind is is really just using a hold bridge, getting the caller out of the way, dialing the destination, and then connecting them. And that really wasn't difficult. But attended transfer on the other hand, you've got to do all this stuff. You create a hold bridge.

This is all done by ARRI, by the way. Create a hold bridge. Move the caller into that hold bridge. Create a destination bridge. Call the destination. Move the agent websocket to the destination bridge because we really want the agent now to talk to the destination. Right? So the destination picks up the phone and she says, "I've got a call for you from blah blah. Do you want

to take the call?" Now the LM has to decide whether or not the answer is yes, I want the call or no, do something else. Um, so there's some determination we have to make there. And then she either completes the transfer or she hangs up on the destination, undoes all of this stuff, brings the caller back off that hold bridge and then tells them, "I can't transfer

you to that person today. Maybe I could take a message." And all of that works, but it requires ARRI. There's no way to Next, we came up with a Slack tool. We we run our business on Slack. Maybe a lot of you guys do too. Um, so as far as taking a message, I didn't want it to just be shunted to voicemail. I would rather have it

take a message and send that message to the right Slack channel. Maybe it's an individual, maybe it's a team. Uh, and that turned out to be super easy. So um it could easily be uh extended to do email, SMS, teams. and I I've even had some thoughts. We our tool won't do this today, but I've had some thoughts about having the Slack tool interact with our employees

during the conversation. So maybe the caller would say something like, "Is Keith going to be around tomorrow?" And the agent would be able to slack Keith. Keith would respond and then the agent would be able to say, "Yeah, Keith says no problem." And all of that should be pretty easy. So that brings me to fine-tuning, which we have not done yet, but our our plan is to

have a Laura adapter for every skill set. So receptionist is going to have a Laura adapter that is trained on some number of thousand of synthetic receptionist calls. So our plan is to create this synthetic data set, fine-tune the AB model with this Laura adapter and then offer that to get rid of these problems with tool calls to get rid of odd responses. I think this is

going to be the big win and some skill sets are really going to require it. Reservation taking, pizza ordering, all these things we're working on. An 8B model is just too dumb uh to to do things correctly without some fine-tuning. and Laura adapters are great because you can we can host dozens of them on the same card conversation tips. Um here's another screenshot of our dashboard. So

in our agent module on the context tab you can specify um the specialty receptionist an initial greeting you know what she's going to say when she answers the phone and then this is the user context. So this is just an example. This could get much much longer. Of course, um we've had some context that included facts. Um you know, things that you the way that I I

tried to describe this is that if we were going to have a new employee answering our phones, what would we tell the new employee? And that's the context, right? Whatever you would tell that employee about how to answer the phone, that's what goes in this box. A couple of tips. Failsafe extension is a good idea. Uh people are going to get frustrated when they realize they're talking

to a a bot. Um so you want to make sure there's a place that you can be transferred to a real person. Uh define all the extensions that you want to be able to transfer to. So in our case, the the handful of people that work in the office and when they work, uh the hours that they work, you know, those things are important. The model knows

the date and time. So it should be able to tell if those people are available or not. Uh yeah, lower adapters. So the verdict, we were able to produce this with no public APIs. This is all running on local GPU hardware. And if you're not trying to serve 50 models at once, and you're really just doing this for your own PBX, you're quite able to do that

on a single card. Like if you had one L40S uh in the box running asterisk, that would be enough to run all three of these models, the ASR, TTS, and LLM. No permanent costs. That was really important for our resellers and for the ability just to sell the product. And then compliance and privacy risks. We don't have to tell our our HIPPA using uh clients that there's

a potential So, I did say in the in the thing that I was going to do a live demo, but I'm I'm just not that brave. So, I'll play this uh I'll play this instead, Oh, I'm not getting any audio. Oh, now it's not playing. Let me see if I can do it this way. >> Hello, Jeff. How can I assist you today? >> What time is

it? >> It's 10:08 a.m. CST. >> Can you transfer me to Jeff, please? >> I'll connect you now. Please hold. So, I don't know if you can hear it, but I I was on speaker phone, so you can kind of hear her talking as the but I'll stay on hold until u she finally completes the call, completes the transfer, and then you don't hear anything because it's

on a different channel, right? But you get the idea. So, the flow was pretty good. The the latency was pretty decent. Um, you heard the key clicks, which filled some of the silence. Uh, and it called the right tool and did the right thing. So, getting the code, I uh I meant to make the repo public and I just haven't done it yet. So, if you are

interested in the code, just get in touch with me and I'd be happy to send you the tarball uh or invite you to the repo. Um there are some database calls that we make that uh have to do with our dashboard and all of those can be overridden with uh environment variables. So if you do have the source, you'll notice that in the config.py uh there's a

ton of environment variables that would override the things that we're looking up in the database. So if you do try to implement this yourself, um that would be the way forward probably. Um, I would run the agent container on your freebx system because it normally wants to talk to loop back to talk to asterisk. Um, you can run the other models, the the ASR model and the

TTS model also in Docker. And I've got uh if you contact me, I'll show you the the containers that we've set up for those two. But you can run those containers right out of Docker and and they're ready to go as well. And of course, if you're interested in hosting all of this, contact me. And then in the future, um, this is our big plan. We're going

to start this new company, Vox Layer. Vox Layer is going to be about hosting these agents with SIP. Uh, I think I mentioned that earlier. Our resellers are really used to working with SIP and I think that providing the agent like a phone is something that they'll just get instantly and be able to quickly work into their provisioning and uh I think that's the way forward for

us. Uh yeah, so I guess that's all I've got. Um I'll take any questions. >> Oh, thanks Jeff. Yep. Here we go. Did you try making any API calls out to what would emulate customer databases? >> We haven't integrated it with anything yet, but it would be super easy. >> Yeah. Right. So, >> I know you you took care of all the latency and everything by having

it local, but that would still add a piece there that hopefully would still be better because you've isolated everything, but still could have some latency to it. >> That's true. And if we're talking about a HIPPA compliant client, then we'd have to be very careful about what API that is. I would assume that those would be internal anyway. U you know the integrations that we're expecting to

do would be consulting dollars for us. U so anybody that wants to implement one of these agents and they want their CRM to be integrated that would just be consulting work from our perspective. Uh and super easy, right? >> Okay. Thanks. >> Good question. Anybody back here? >> Thank you. So just one suggestion for your operator on that hesitation in the key clicks. One of the things

that we've done is if there's an issue or a problem or I need that little latency the fill a little um hold on uh what was that so that you can reestablish that control from the receptionist. So just a suggestion to play around >> like that. >> And then a question for you on the LLM. Have you thought about the idea of not just a Laura, but

an LLM memory across all calls? >> I've thought about it. I've certainly >> And And what' you come up with? >> Well, I don't think anybody's come up with a decent memory implementation yet. Um, and I'd love to be proven wrong. >> Have you heard of me? They're an open source. >> Mege. >> Uhuh. >> I have not. >> So, it's a open-source LLM memory across all

chats. and it's something that I'm looking into, but I was wondering if you've heard of anything like that. >> I haven't. I would bet that it's some kind of rag based thing, right? Probably. And we've played with rag. Um, we've done some implementations of rag and choosing the embedding model is key there, right? And I've had not very much luck with embedding models. Um, so yeah, that's

definitely something we're But I don't have anything yet. >> Yeah, a lot of a lot of good parts there with the VCON's discussed yesterday, I think, for for some of those thoughts, too, maybe. U any other questions? We're a little bit over time for this session, but we started a few minutes late. So, all right. Uh, see, all right. Thanks very much, Jeff. >> All right. Thank

you. started in about uh restarted in about 10 minutes here. So, if you haven't yet had a chance to go out and uh grab a t-shirt. We've also got some notepads and pens. If you're running out of uh hands and arms to ink up, you can, you know, grab that to write uh and keep up to date with uh all these fun things that we're learning here

about voice AI and integrations. And we have a session coming up here at quarter of the hour about free PBX and some updates there providing a status on that project and then one more AI related session before we break at noon for lunch. So if you can all um just refresh your cups. It looks like we're out of coffee and snacks at the moment but we do

have some tea. So, uh, with honey packets and lemon packets, delicious items. So, please do step up, grab those, and grab some swag outside. If you're a speaker, just a reminder to grab your mug, your coffee mug from the front desk as well for the speakers. So, we'll get started here in just under 10 minutes. Thank you so much, Check one too. >> Okay, check. >> All

right, you good? Whoa, I'm coming. >> Cool. A little hot. >> All right. Good morning, everyone again. just get started in a minute here. Feel free to grab a seat. Some little bit more tea left in the back. This thing does tend to move around some. >> So, good morning. Um, thank you again coming to Astrocon 2026 here at scale in sunny Pasadena. It is Friday, March

6th. We're going to talk about Free PBX now. A couple of great AI talks this morning. So, myself, I'm my name is Chris May. I'm the open source solutions advocate at Sangoma. With me, Michael White, VP of open source, and we're going to talk about Free PBX, the OG guy for the asterisk telefanany toolkit that keeps getting So, quick agenda for the day if everyone's interested. Free

PBX. Just quick show hands if you've used free PBX or regularly make use of it maybe. Okay, cool. All right, great. If you're new, uh, fantastic. We'll have a little bit of, you know, intro here and there about, you know, how good ways to go about installing free PBX. We'll talk about great place to get questions answered for free PBX is our community forums and made some

updates there. Also, some AI enhancements in our ecosystem and exciting improvements elsewhere in Free PBX, including what's coming up for our future versions. So little bit uh introduction with our who, when, where, and what. So for us right now, you know, 10:45 in the morning, right on time. Uh Mike and I Capil's over in India though, so he's probably sleeping watching our live stream maybe hopefully and

might have a, you know, question or two that we can help field online with that. But we do have people of course working on the project around the clock at Sangoma and elsewhere online. we uh you know bugs on the GitHub issue tracker uh constantly getting worked on and addressed. We'll talk more about those. Uh our current supported versions, so you're aware are 16 and 17. So

we uh had uh version 15 go EOL uh this past fall and you need to if you haven't updated yet, great time to do so. That's what we're supporting. Our got an upgrade. We'll talk about that. There's 20 years of search history there as modified mailing list originally started as migrated over to forums had a bunch of updates uh that we came through with uh running on

discourse currently open source software and we host that. Sangoma provides that as a service for the community. Our search uh has improved a little bit too. We've had some better database indexes and whatnot. So, and uh Tango, our also a participant uh you'll see here and there our beloved free PBX tree frog maset who uh just turned 17 which is legal age to drive a car. So,

if you do see a frog uh I'd get out of the way uh driving anywhere like that. So, this slide was uh not quite the hit yesterday. I thought I'd include it again just in case uh people wanted to study it more. But, uh you know, I I did hear you like Astrocon. That's why you're all here. If you've never heard of Astrocon before, great that you're

here as well. So, appreciate it. Uh myself been involved with Asterisk and open source phone systems for over 20 years. Started at Sangoma about a year and a half ago after uh 20 years being small medium business uh entrepreneurship with a number of different clients that went on from starting with phones and then how many systems can you tie phones into and turned into database, web development,

a whole variety of things. So really enjoy working at Sangoma now where I still get to do a large variety of things that we'll talk about in this presentation. And about Mike. >> Hi there everybody. Uh Mike White as Chris said I am the uh vice president of open source here at Sang. I had the fortunate luck of joining Sangoma in 2020 by way of acquisition. My

company E4 Technologies was acquired by Sang. Um, while I wear many hats here, I try to generally distill my role down to, I guess, a single point. Um, I help customers, businesses, people find their way to asterisk. Um, and FreePX, of course, if I, uh, you know, distill just a little bit further. You know, my hope and goal is just to get more people to use, uh,

Sang's open source tools. From there, you know, within that scape, I I help partners basically do what we did at my company, which is uh, you know, build a business or, you know, frankly make money around these uh projects. And, uh, while I wear many roles at uh, Sangoma, I I definitely find myself orbiting, you know, within uh, the engineering landscape, within the the PM landscape, and

just trying to, you know, help grow these projects and and bring them forward to folks like you. So, that's a big part of what I do. Dan, um, who you may have met out at the front desk, also joined Sangoma with, uh, with E4 when we acquired, and we've just been doing our best with, uh, folks like Chris to make sure that all this stuff just continues

to move forward. >> Yeah, thanks, Mike. Yeah, we've got um definitely a booth like this for the next couple of days, booth 104. So if we don't get a chance to talk now, we'll be around the exhibit hall uh this afternoon uh Dan, Todd, Mike, myself, uh Saturday and and Sunday as well. So that is a bit about Mike. At a previous uh K12 event, I believe

this was we do a lot of those uh at with schools uh IT administrators at schools installing free PBS. It's free and then needing phones and we offer those integrated as about Capil. Uh he's been with Sangoma for 14 years. VP of engineering. Work regularly with Capil. You've seen uh I guess Noah gave some props to him yesterday regards to working on some bug fixes. That's K

Gupta. Uh he's also on the forums as well. All three of us are. Uh we uh have Capil works on a number of different issues. Leads a large team, wrote the free PBX 17 shell installer and is very active on our our community forums. So you can uh ping any of us you have questions there. Talk about Tango as well, our uh official uh free PBX mascot.

Uh I I dug through the archives and apparently we we there's been some discussion that we've been at scale before. Not exactly sure all the details, but I did find this uh from from 2016. Uh a nice uh t-shirt. It could, I guess, you know, being largely a Canadian company, I love Canada, but we'll go with I love California for this graphic. So, there was not much

more metadata to go with. So, that's a great shot of a previous a couple of previous uh images of Tango. We've of course modernized that a bit with a new logo since uh but we've been around for a number of years. Not not Kermit, that is that is Tango. >> And we still do have that that uh OG graphic t-shirt floating around. So that is a current

thing that we we do share out to customers and partners. >> Nice. Nice. So we Tango in a business suit. Uh that's our new uh updated Tango. We've been using for about 5 years now I guess or so. Uh of course kit. You may have seen the asterisk mascot uh because asterisk is a tool kit for telefan applications. That's where the name kit comes from. and then

a whole bunch of other mascots to talk about some of our main products that we're using currently on free PBX 17 for example. So we're running on Debian 12. We uh use a lot of PHP uh Maria DB uh NodeJS uh and uh of course TUX and the Apache logo recently got an update. It is now uh in Oakleaf but >> so just want to add a

little bit here too. You may have noticed that over time Tango's evolved, you know, from a basic piece of clip art to, you know, this professional character that embodies, you know, the life of our open source projects. So, you know, we see Tango as or the new Tango as a fresh face. He's been around now since about uh 2021 and he was actually designed by a wonderful

woman who um went on to be a lead designer at LEGO. So she's quite the artist and we're happy to have her art represent our projects. >> I think helped design kit as well then, >> Kit and Kit and Tango as well. >> So yeah, Entangle itself first the first birthday I think for the naming of Tango I could find on our forums uh was around 2009

actually the name first got coined. So made a post about that. So you might call it I know you know lots of people familiar with the LAMP stack. I went with lamp n just so that we could add in node.js and asterisk. Those are such big components. So that is uh what we have been and and before free PBX 17 we're on CentOS7 distribution. So we're still

LANA but this is what we're at now. Um some new updates on the forums. So in case you like dark mode which has been a little bit harder to see in these visuals but online uh it it looks pretty good with the with the live stream. So we've got uh some choices now in the forums. Uh we we really updated the infrastructure, improved that uh set up

some new uh staging servers, been deeply involved with that, migrating our database and definitely faster searches, some cleaner menus. Might see like we've got more rotating banners about events that are coming up such as this one as well as some training we have in May is at with Sangoma. Um it's a lot more mobile friendly. So it is much easier to search and scroll on your phone.

just some a lot of these are benefits are just from our our self-hosting of discourse and having a lot of control over that. So there are several themes to to choose from and you can see there those are our handles for the forum so you can contact um myself, Mike and Capil um with our our various usernames there as admins so we can help out with issues

if you get locked out of your account, need a password reset, like are flagging a post. uh you know we rely on a lot of community help and feedback to help us manage the forums. >> I just wanted to add you know maybe some credit here to you too Chris. You know he mentions we as in uh you know making a lot of these changes possible but

Chris not only is the person that's in there most of the time at the forefront fielding questions in our forum but he's you know single-handedly been doing the work uh for you know bringing dark and light mode and some of this performance stuff. uh clicked some buttons I guess and some keys but uh thanks Mike. Yeah, props to yeah Josh too uh yeah who's it does a

lot managing asterisk forums and also throws in a hand on on the free PBX forums. See like we're putting a little bit of a new menu on the top you can choose now between some people really like you know light and dark mode. So looking at where you click on that down in the bottom left if you haven't done it yet of the hamburger menu and so

you may see I put a circle there on the hamburger. Why is it a hamburger menu? It's because it looks like a hamburger if you haven't seen that before. Um, then also you can choose, you know, not just your theme, but your light and dark color modes. So, you got a various options for customization, your uh forum experience. So, if you see a screenshot from the forums,

you moving forward and it doesn't look like the forums, it's just somebody else's preferred view of viewing the forums, don't don't panic. um that new header menu uh some quick links I'll just point out this one to our rag uh Sangoma's uh integrated search wiki with it. So we have a lot of documentation in the forums a lot of documentation in the modules and we also have

in GitHub and then we also have a completely separate KB knowledge base that is for all different Sangoma products. So including like hardware cards and information free PBX PB exact a number of different things that you can get one sort of integrated view. So it's a well-trained rag after decades and decades of of proper feeding and care. It is ready to go at help.synoma.com. So you can

take a look at that. It's just some basics or are a great place to start. How do I set up a zip trunk with free PBX 17 can really kind of walk you through the steps for that. So, >> and I think just a good example of what this does is, you know, the wiki search itself leaves a bit to be desired. Utilizing the help.sync search really

brings you to that core information that you're searching for. It's it's basically just a much better search function than the wiki offers. >> Yeah. So, how do you install free PBX7 if you haven't done this before? So the shell script I mentioned on Debian 12 if you've installed Freebx7 recently on a stock Demian cloud image great way to go. There's also an option with our ISO wrapper

around the shell script. So it's just a way to kind of have that as the payload to help you set up Debian quickly. We're going to look at a couple of those ways um and what that means for do-it-yourselfers. So in particular, so if you are someone who has been installing, you know, free PBX17 and running it, you know, you need to make sure that you're, you

know, taking care and responsibility to update it, maintain it, apply security updates, check your logs, go through each of these in detail. You can watch online and and flip through it as well, of course. So this the one part I want to point out in yellow here. This is a little bit different if you're coming from previous if you're coming from as a Linux cis admin. totally

familiar, understandable, but if you're coming as a free PBX 16 user or previously, we had different integration and we're working on improving that, but you were able to do system level package updates like through the guey. And so that's something that's different in free PBX 17 at the moment. So you'll need to shell in there and do that yourself. So if you are expecting for example that

my module admins free PBX PHP modules for the most part there's node in there as well and SQL you want to call that separate. So if you are only updating that interface in the web front end you're you're not getting the OS level package updates. So I hope that that makes sense for everyone. So this is a difference here that you just need to be aware of

this. A couple simple commands and in Debian stable that we're on right now. Debian 12 is very good about not breaking your your system when you update like this. So um we had some issues with it in the fall and that was honestly due to typos on our part as far as specifying repos. We fixed those, provided solutions for that, but we just need to make sure

we're being very explicit because it's now ex not even stable. It's old stable bookworm Debian 12. The latest version of uh Debian is 13. We'll talk about that in a minute about what we're doing in preparation for that. But um these are just some other more VOIPE related thoughts as well in here about uh you know making sure that when you set this up that you do

uh have some good DNS working, your network's working, you solved a lot of problems like that. uh if you haven't installed Debian before uh it's you know very friendly for having a proper loves a properly configured network DHCP server. So these are just some ideas to think about when you're installing. Um now the ISO is getting an update. We first announced it last year that we weren't

grely weren't going to have an ISO and we brought it back because we saw those wonderful looking hardware appliances that I talked about earlier with a variety of telefan cards and uh you hopefully some GPUs in there too, right? So um in the future but this uh this ISO spent some time on it uh try to make a focus on making sure that it is the e

one of the easiest ways to get a whole OS up and running. The time for that someone asked me this morning is and I I answered is about 20 minutes right now with the version we currently have in QA and we need this because we're you know selling hardware with these boxes and we need to be install it quickly and this is a full Debian install. This

isn't we're not we're not dding an image out from you know 0 to 1 million bytes uh directly. We are grafting it in bit by bit running scripts making sure it's updated and making sure that your kernel works with the current version of uh there's a couple different spice levels that you can choose from the installer. Uh so if you're looking at a fog option for the

cloud that's great. We use internal for our internal hardware and that makes sure that we stay on the kernel that is supported with our hardware drivers for our interface cards. Um something been striving for for a while now has been reproducible builds. So, there's a few uh thoughts on that about making sure that you have uh correct date stamps that are synchronized, usernames, permissions, some subtleties that

you wouldn't necessarily see if you looked at MD5s of all the files, but when you start combining those in the ISO, that metadata is affected by those timestamps. So, you need to take that in account. So made sure now that this allows us to as part of it being open source the the build script SNG FD12 you can actually take the input ISOs that we use from

Debian and upstream and with those same ISOs produce the same ISO output as us verify the hashes as long as you're running like a modern Debian 12 or 13 with modern Zerizo tools and whatnot. So, that's been a focus. Um, and I did want to have a bit of discussion and love to hear thoughts on this afterwards and at the booth about uh folks feedback on some

of the partitioning choices that we're making in regards to how much size we're allocating for your call recordings versus your temporary log files. Uh, I see, you know, a few uh questions come up on our forums about this. I I want you to know that one of the thoughts, you know, we're moving with, you know, isolating some of these partitions a little bit more is that we

can get to ideally more readonly situations for, you know, critical files. First, first, our biggest thing to solve is not filling up the disc with logs or call recordings. So, you crash the OS. That's a big issue. We'd like to avoid that. And then moving on in the future, where does that go? Um you know some people it's not you see just some maybe things that are

different right now. Uh for example when you restore a backup in free PBX which is our recommended way from going from 16 or earlier to 17 is to back up that previous system install a new 17 and restore the backup. You need to you can just point you need typically before the temp directory was used to store all this data and that is rather small in this

new setup because it should be not that enormous. It's temporary but you can point PHP to look at a different tempter. So that's kind of the the pro tip on the bottom there about how to do that and and pass it in the environment. Um and it is going to be a lot bigger. Uh so it is including a lot more packages with you know hopefully on

the path towards having it completely contained. So it could be an offline installation once you download the ISO again. So that is uh some new wallpaper to go along with it which you know due to our colors you can't really tell but uh it's got a little bit more purple in there instead of just black. So anyhow to look at that later. Um now let's talk about

some AI enhancements. of course. Uh so this was a whole morning of AI talks. Then I made sure that we included a separate section on that. Uh this is a cool one that came out of some open- source contribution to the systems recording module. So if you haven't used this before, it's a fully integrated way to once you do configure your your API key uh elsewhere. Then

you can go and just use a drop down to request a number of different providers from 11 Labs, Open AI, third parties uh or Scribe, Sangoma Scribe that we offer. And that lets you type in the text that you want and have it turn into a pretty wave file that you can use in your IVR. So that with all without leaving the guey. Uh Mike I think

has >> been working on that. >> A little bit into that. So this is built directly into system recordings. You may see in the list here that uh one of the options is Scribe. So, if you're familiar with our commercial product, Scribe, Scribe gives you the ability to um transcribe your voicemails, do sentiment analysis, summaries uh to the actual call recordings or Q calls. And uh we

brought this into the system recording manager to provide the same functionality. Again, you can bring your own solution to the table as well, but um this gives you, like Chris said, the ability to type in your IVR uh messages with the Scribe module. It gives you different English types, you know, UK English, American English or various dialects uh within that scope and you even have the ability

to modify them later on. So if you make a change or uh need to make a change to that recording or IVR later on, you can jump back into recording manager, select that recording and make the change accordingly. So um that's some of the additions that are brought in with the Scribe module itself. >> Yeah, more more credit than that. Mike's gone through and uh tested and

and worked on this thoroughly updated most of our IVRs. Uh so we're regularly getting, you know, updated messages, which is great. You know, just a little easier way than having to get your voice talent back online to re-record the IVR. and and that's been a popular, >> you know, and kind of one other point to the IVR that I'd just like to mention too is that if

you're using free PBX, you know, part of the the groups that I manage at the at the company, if you've ever logged into Sangoma portal, you'll look up in the top right um you'll have a dedicated account manager. So, we have real people, not AI, that are here to answer questions about anything that you might have uh free PBX or asterisk related. We call ourselves the uh

NaOS or the open source open source group. And uh if you log into portal or call into any of those numbers, obviously this is going to be the the recorded message that you hear, we're using the the voice called Andromeda, but you can hear that by calling any of the numbers that are listed on freepbx.org or or any of the open source numbers listed at sangoma.com. But

you'll reach a live person. Those people are in my group. You can even reach Chris >> occasionally. Yeah. Uh um so I just want to talk about a few things on on if you are integrating AI just some other thoughts and ideas. I mean there's already been lots of great thoughts and and discussion here. Um, we do have a couple of things to think about when it

comes to uh, you know, what your main ways to integrate and I think some people have talked about this with audio sockets and websockets uh, between asterisk and there's a lot you can do of course directly with asterisk and if you need to expose some of that and we don't have module integration of free PBX yet uh, the extensions_custom.com is a great way to do that. that's

a separate but a dial plan that gets included by free PBX when it reloads and you can do whatever you like and and native dial plan generate JSON by hand for those you know or or integrate with other AIs. Um now I'll say that you know there are some struggles but uh of course you know how is this AI trained and you know what's the copyright and

trademark issues some of the ethical stuff I was hinting at earlier and love to have more talks about you know AI not being a substitute for good judgment and what what is how do we program that in and you know what sort of guardrails do we put on it and where do we want the AI's input and where do we absolutely want it not to declare a

thermonuclear war you know 95% of the time I saw the other day. So, you know, there's some some things that can go off the rails on. So, we'll we'll we'll rein that in. Um, we do have a a request as a part of our forum policies uh both asterisk and community free PBX community forums to disclose your use of AI if you are including that. uh there

was recently uh some updates to the asterisk uh added a AI contribution policy and this is something that you know we're looking at elsewhere but you know you have to ask this question about you know did you did you just ask someone like hey uh you know I have a friend who plays guitar I say make me a song and and I can't I can't copyright that

song right like he he wrote this beautiful love song or whatever no it's mine because I said make me a make me a love song like that's not that's not fair. So that that's the kind of thing that like you know how much did you put into it? Like were you the prompt engineer or did you did you vibe it so loose that you're just going to

declare everything under the sun is yours? That's I that's more God's domain, right? So anyhow, this part about like a FreeBX AI voice agent uh labeling of your project as bad and another AI voice agent for free PBX being good is something that an AI is not going to tell you. Maybe it will now because we fed it a little bit more of like here's how to

do things AI, please read this and respect what we've wrote here. Um but you know, we want to make sure where we put our nouns uh is is important. And so if you are interested in in having a new vibecoded project, fantastic. Just please don't put people in a spot where they're confused about like, okay, is this asterisk now? Like is this free PBX? Like what what

is this that I'm running that I just downloaded? I' I've never heard of asterisk before. Oh, it's that cool AI tool. Like no, it's a toolkit for telefan that does all these other things. So just you know that's that's one of the reasons we want to talk about that just lightly. So um exciting improvements otherwise besides the again uh we've got a lot going on. Uh we

had a K12 focused event uh in Denver over the summer at at Free PBX World Summit. I gave some of uh parts of that uh emergency calling talk. Mike Perine gave a talk about planning out what a school network might look like. We had a couple other speakers uh education space talk about how they're using this in the classroom and how they are working with you know

students on demos. Um one fella had some kids who went on to like the state final uh uh meeting for uh business class students to talk about how they could you know reinstall free PBX on the school's phone system and save a ton of money which is a great opportunity if you're just you know getting out of school or interested in going there and doing that. we

have lots of interest u I'll say grassroots because it percolates up and and uh people you know jump in front of Tango and and are really happy to see because that's their school phone system. Um >> I just want to say you know K12 is an interesting environment as well but you know we have a number of universities throughout the country that actually teach classes on free

PBX. So in terms of finding new users and you know bringing it upstream like Chris said it's just been a great opportunity for us from a a business perspective you know K12 is still one of those environments has large user counts small budgets so free PVX is a perfect fit for K12 and that's why a lot of our discussions you know past and present are about K12.

>> Yeah. So I had my tie on there I realized wow >> fancy. >> Yeah. So, I want to talk about another really cool open-source module that's uh part of a lot of free PBX installs that we see now. And this came out of some work um from a developer in um back east uh Adam Volko. He we did a webinar. So, I'm just going to show

a couple of slides uh earlier in the this guess it was the spring last year about this great way to like look at um your call flow. So, you know, they kind of developed a lot of this in-house uh so that as a as a hoster and reseller of free PBX, they could see what the heck was going on in their systems that they were getting on

board. So, it's a fullon analysis in the guey of automatically of your dial plan. And I mean, not the raw asteris dial plan. This isn't going to show you, you know, each line, but how the free PBX parts of it interconnect. So you could consider each one of those bubbles as like you know a wrapper for a whole lot of other raw context. >> Just a maybe

a quick question. Is anybody here using DPV today? Show hands. Maybe >> got a couple. Got a few. Great. Awesome. Yeah. So that's just if you haven't seen it, uh just to zoom in there. Um you know just showing along these arrows in this this design. You can export these as PGs. You know it's a great place to print it out. um you know as far as

identifying like you know uh situations where maybe there isn't a call flow happening there I didn't show like some of the error conditions when things like that occur but like it's continuously getting updated uh these are all clickable now in a lot of ways you can actually go into the right spot in the guey so sometimes I see a lot of people go to the guey the

free PBX guey it's it's like looking at the space shuttle and and everyone goes to the search bar in the upper right corner to type and what they're looking for. I was like, "Yeah, that's needle." You know, that's the great way to go about it. But if you have this dial plan visualizer, you can click on these individual items and jump right to it. So, you don't

even have to think like what necessarily what was the name of that module. Again, we do try to, you know, keep some pretty uh sane names of modules, but occasionally it's it doesn't doesn't come to top of mind. So, this is a really fun one. There's a webinar online on our YouTube channel about that if you haven't seen it. and and I know it's getting uh you

know a lot more updates uh all the time. So it's fun project to watch for. We have a lot of other posts and discussions about on the community forums too. So >> and this is receiving a ton of updates from Adam. You know if you get a chance or an opportunity to check it out, you know, run out and uh show some love, give it a star

on GitHub if you could. >> Yeah, totally. Yeah. Yeah. So I just looked at uh some of our code stats for where most of the work is happening um framework and core still uh those modules it. So we have a about a hundred repos right now in free PBX GitHub that are part of free PBX. So each one has its own you know individual change logs and

whatnotss but they all do you know a lot of work together. Um you'll notice a few on the bottom however are not actually modules. So uh the the free PBXW install on the very bottom is the shell installer. Uh these are updates over the past year like number of commits. So you may be able to see at the top there's 86 and at the bottom there's 18.

If that gives you an idea on the scale uh free PBX CI actions is a new one I want to talk about in the next couple of minutes. Um backup uh that's a that's a free PBX module uh and SNG FD12 we talked about that's the ISO generation script. So you can uh download the main branch of that and build your own ISOs and hashes are discussed

in the forum thread on that topic as well. So about um oh security man. So these bots have just been going wild with our stuff. I mean it is nearly another full-time job. Lots of fun. Love it. Uh but we have definitely had an uptick. Uh I think Noah gave a great talk about that yesterday and some things to look into. Um we have you know there's

articles out there about this that the other. I I just want to talk about like I think like what the most the critical ones here in the middle. We had 15 last year out of something like you know a thousand issues and most of those were um like a quarter of those were on uh uh our GitHub and then we have you know parallel issues internally for

things as well and these couple of critical ones were um what you know we obviously critical we should pay attention. So some of this stuff and I have some thoughts on the scoring lots of thoughts on the scoring but we we had one that was a remote you know unauthenticated bad news right like that that's probably the worst one that came in it got reported by a

victim on August 21st so you know appreciate Noah's reporting of that web server off one that he talked about yesterday uh but we we think that in most cases that is very unlikely scenario that someone has opened up web server O which is a non-default mode of operation and then that they failed to secure it with uh more more tools in front like a proper Apache config

right so not as concerned about that for most installs but if you're a reseller and you're running that like you know just walk out quietly and fix it afterwards I guess but it's a lot to fix. Um, this particular though, uh, the remote one that we got from a victim, uh, we we worked with some of our partners to set up some honeypotss, fresh honey pots, and

and caught it quickly within a day or so. Um, and then analyzed the network packets. It, you know, this isn't something you could pull out of the Apache logs because they came in on this Ajax path and into the in an old unused part of the endpoint module that was still floating around. So I wouldn't say it was um we we had very limited use for it

and it was easily migrated away from that part. So that's our our immediate fix was to correct that problem initially of you know catching wind of it. That's the one that went on like a few weeks later and became a larger problem for some users, but we already had security updates in place. We didn't publish the PC, so uh we're we're not doing that in bold at

the bottom. Um when we do put these out on our GitHub, we clean up the descriptions and we provide a lot of details and information, but that is not the approach we've been taking. So also you may notice in our GHSAs that there's some alternate scoring information at the bottom as well as sometimes some historical analysis and that uh alternate scoring I came up with a little

bit of an idea about you know maybe some ways to put this in more perspective. I mean it is very hard to score some of these less than an eight a lot of times when like the real impact is probably closer to a four out of a 10 scale on the score. So we can talk about that more later, but if you haven't seen it, that that

code for this alternate scoring is available and then we are including a lot more details and information on the GHSAs. So if you want to follow along, a great place to do that is in our GitHub in that separate security repository and repo. That's where you should report these. And we're also, you know, making some updates into our forums that when we do discuss these issues, if

you subscribe in the forums to the security topics, you can actually just get email notifications when there's new posts in those areas. So, it's a good way to keep up to date on that stuff. Um, >> once again, more of Chris's really good work there. Credit where credit is due. >> Yeah, we just put out four more this week. So, that was, you know, 15 last year

and, you know, we're up to, I think, almost a half dozen now so far this year. So, that's going to keep on going. Thanks, Mike. Um, uh, so as far as, uh, new, uh, free PBX stuff, that's other stuff that's cool in the in the free PBXCI actions repo. This is a bunch of GitHub actions for trying to automate our Weblate. So, big announcement here. We're bringing

back Weblate for language translations. So, you can go to weblate.freepbx.org. Uh, we'll have a webinar in a couple of months. Oh, next month about that. Um, you know, we're we're thinking that uh, oh, scratch that weblate beta in the first option. We we just took that out of beta. So, it's it's uh you still need to get your account approved. So, uh shoot me a DM on

the forums or open up an issue in the issue tractor if you'd like to help. Uh we you know, we already support dozens of languages, but we when we migrated the project to GitHub a couple of years ago, this got left behind. And so, I've just, you know, we've been working together on bringing that back recently. And you can go in and see uh all the strings

that are there now at weblate.freebx.org. you can start signing up for an account and get a you know through the approval you have to click on the CLA same thing you do when you submit patches into the GitHub repos very similar idea so that's available so exciting that free PBX will be in your language or of of choice in the future um so you know we we're

also kicking around the idea of you know putting pirate language in there as well as Cllingon so if you feel the need to to jump in there and contribute to that as well or you speak those fluently Please join us. >> Yeah. Yeah. Yeah. That's that's on the list to do as well. Yeah. Um so how's it work? I there's a whole diagram like this in the

GitHub repo. Spare the details. Running out of time. Um but suffice to say we're trying to automate a lot of it and in separate branches so we're not pushing directly to master from the this place and or or main as the case may be in certain modules and we're trying to you know adapt that those terms. um we're trying to incorporate, you know, an opportunity for translators

to review things in the weblate area and then have bots take care of a lot of the rest of the work. Uh but we do want, you know, some human touch and intervention, you know, on these things along the way. And we also want the opportunity for you, separate weblate branches, then separate, uh release branches, and so we're not uh you know, cross polluting these and we

can have faster translations. And then as part of this work uh by bringing more GitHub actions in in a separate repo we're going to you know we're already doing some more linting and some more QA automated QA processes and we're going to continue to hopefully expand those. Uh I mean I don't think we have a dedicated talk about it but you may have seen there's a continuous

amount of work in the asterisk GitHub actions which are an inspiration for a a lot of this. I I mean there are scores of tests that are run every time you submit a patch that are making sure it compiles etc. And we are trying to you know work and cross-pollinate a little bit so we can do more of that on the free PBX side of the house.

So there's that. Um but we do need help from editors. So we want to make sure as well that we have good source strings to start with. And so we want to focus on not just the developer who gets to come up with the string, but that it's spelled correctly from the beginning and it makes grammatical sense in English and then go on to translate it into

others. So this is sort of a flow uh humans on the top and what the bots are going to do on the bottom. So you can see we're looking for, you know, developers to make sure they use uh th those uh expressions to make sure their string does get translated. Pretty common if you look at the PHP code. And then the editor is going to go through

and make sure that these new strings are appropriate and then a translator will go through once those strings are ready to translate. So multi-step process, humans involved uh and also some bots. So keep that in mind. And uh of course weblade has dark mode available as well. Uh that's a exercise for the reader how to enable it but is available if you like to translate it in

the dark. So V18, what do we have coming up? Got about five minutes left. It's in version 18. Well, this is our full version calendar which we put together after last year's Astrocon if you haven't seen it. Now this is awful a lot of forwardlooking statements of course. So you know when we talk about what's planned and for 19 and 20 we're looking pretty far out. So

pretty optimistic. None of this is set in stone beyond but this is our goal. Um you some significant ones I mentioned 15 was end of life uh last last October. Um we look to be supporting free PBX16 until the beginning of next year and then we also are going to start on free PBX17 development uh very very soon. So we're hoping that uh you know as part

of this process sometime in um you know sometime next month that we're going to look at all right what PHP versions we know you know we've got to get to 8.4 in Debian 13 but we are going to continue on that as our base OS uh for and and make sure that inline upgrades you know are available for for Debian uh from in in Debian 13 with

free PBX 18. So, if you do try to run or accidentally try to install FreeB PBX on Debian 12, Free PBX um 18 on Debian 12, I'm not sure about that, but if you most what we've seen so far is folks trying to run uh free PBX 17 on Debian 13. And some people have some success with, you know, changing up their PHP libraries, but it's not

supported. So if you need help with it, community forums are there, but it's not something Sang is going to officially support and we we definitely have some, you know, code updates. There's been a lot of changes in PHP. So that's our our big calendar. It's just look for free PBX versions on the wiki and that's all there and you can see our target asterisk versions as well.

I know that's really helpful for a lot of people to think about, you know, what features are going to be available in the future. So that's that. Um, in this uh next month we should be starting on free PBX18 and that beta is planned for September. Hopefully by January 2027 we'll have it out which means that version 16 goes EOL at that time. Um, I mentioned the

large PHP update and there is a milestone. And it looks like a small uh arrow with a line like a a a guidepost if you will like you know this city 500 miles away or whatnot for the milestones and you can see some things that we've been trying to move into 18. So you have ideas on that. Uh you know it' be great to talk about that

what you like to see in 18. Uh I've just you know been taking some of the issues that have been raised on GitHub and trying to say yes this is something we can you know we want to look at for version 18. So appreciate all the feedback and help there. And um if there's any questions, we've got uh about two minutes left for that on free >>

I'm the question guy today here for moral support. >> Anyone? >> Everyone's excited about Okay, >> I guess we did a pretty good job there. Thank you very much everyone for uh joining us. >> Yeah, great. Thanks so much. Yeah. And uh Mike Berdine is coming up for his talk on some websockets work here in just a a couple of minutes. And we'll have a that's a

30 minute presentation, not our usual 40 45. And then we'll break for lunch for an hour. So if you uh can make it back, then we'll have a couple of talks this afternoon and a wrap up a summary of the Asteris project as well and what new developments have been there. So thanks. Thanks again. How's that sound? >> Sounds all right. >> All righty. >> All right,

everybody. I think we're going to get started with our last talk of the morning, some more AI discussion. Thanks again for coming to Astrocon 2026. And with us today is a one of the core senior software engineers on the asterisk project. Uh number of sang for many years in a number of capacities. Uh Mike is you know an expert on websockets work. He's been doing a lot

on uh internal AI projects and development making sure that asterisk is smoother and cleaner when it comes to talking to your AIS. fewer pops, cracks, whistles, slow talking things. So, if you appreciate uh all that work, you know, we've definitely had a focus. It's always been a great platform, right, for for AI integration and uh it's just getting better. So, uh Mike is going to talk about

I think some of that the way of websockets. Mike Perdine, do you? >> That's right. Awesome. Thank you, Chris. So, that was a quite an introduction. I don't know if I deserve that, but I'll take it. Um, so before we get started, can you guys hear me? All right, I sound up. I talk quietly and I talk quickly. So hopefully that's a little bit better. Is that

okay? You got to turn me up. Chris got turn me up. Hop over there. How's that? There we go. Alrighty. Great. Thank you, Chris. Um, so before we get started, I'm I'm gonna start with a poll because that's everybody's favorite, right? How many people here I'm assuming most of the people here but how many people here have written some sort of agent or real time service on

top of asterisk right at this point most people some 5050 okay of you guys how many of you have used audio socket okay how many of you have used external media with uh raw RTP one it's still it's my it's what I did last year it's I understand um how many people have used Chan websocket. Excellent. That's good. That's what we want to hear. Hopefully by the

end will be everybody's hand. All right. So, you know, as we've learned, right, asterisk has to be more than just a PBX. It's got to be a a home for your uh your applications, right? Everybody wants their applications to live alongside their PBX. Um and while dial plan is versatile, it can only get you so far as we found. Um and that is where ARRI comes in.

So ARRI was uh originally designed to be the foundation for building your own applications on top of asterisk um and not necessarily having to rely on dial plan to do everything. Um ARRI for a while now has had websocket support but that was only for asynchronous messages coming from asterisk. So you would send a HTTP request to asterisk for for ARRI and you receive uh the asynchronous

events back via websocket. Um external media has been around for a while. Uh that allows you for sending raw RTP streams to and from asterisk. Um but that requires your application to actually build an RTP stream. And building real time RTP streams is not the most fun thing in the world, right? So that means the tools were there but to build your own application you had to

have awareness and use of at least three separate protocols right you had HTTP you had to use websockets and you had to be RTP or audio socket or something else um RTP in particular you know if you're can you know you're trying to service uh an agent on one side you're trying to service asterisk on the other side can be really wonky in terms of it's very

very sensitive in terms of uh in terms of timing for for those RTP streams So if you were uh at Astrocon last year um we had built something called the Astros Voice bridge which has since been released open source. The link is here on the page. Um this was a research project that we did in order to learn sort of what an application built on Astros would

really need and what it would really look like in the real world. uh that used ARRI for call setup and control that used an external media uh RTP stream for sending audio to and from asterisk and then we used websocket connections to deepgram and open AAI. So those were our three pieces. Um again more information can be found at this GitHub link although at the end of

the presentation I'll have a page with just all the reference links because this is kind of a dense gets a little dense towards the end. So what did we learn from from building the voice bridge? We learned that processing voice streams uh requires more than just passing audio to asterisk to and from asterisk. There's more a lot more to it. Generally, the audio that you're getting back

from either your AI or your text to uh speech service is faster than real time, right? You're not just getting a real-time stream. It's usually coming faster than real time. So, you've got to be able to buffer it. You've be able to process it before you h hand it to asterisk. Um so, you know, I know Open AI, I think, sent it. It was roughly almost a

year ago when I was working on it, but it sent it at about one and a half to two times real-time speed. Uh, deep Graham would just send you these big blobs of audio that you have to deal with yourself. There was not really any kind of timing with that. So, you've got to be able to buffer and play back at the appropriate speed. Applications need to

be able to pause and resume playback so you're not speaking over the user. Nobody likes to have their AI talking over them, right? If you're it's even with a regular IVR, if there's a long introduction and you already know what you want to do and you want to know what button you want to press, the last thing you want to do is listen to a 30 second

pre-recorded announcement when you just you just want to talk to Bob. So, you need to be able to at bare minimum pause and resume playback. Um, and then again, the buffer may need to be able to purge. So on top of pause, on top of resume, if let's say, you know, sometimes someone just goes and clears their throat or something, you want to keep resuming. Sometimes it's

actually you want to stop talking and move on to the next bit. Uh the real time uh AI services will tell you when to do that generally. So, for instance, with OpenAI, if it if you're streaming your phrases back to it, it's going to know um hey, stop talking, but or or you know, don't send the rest of the audio. I sent you, but it already sent

you the audio. So, you can't just rely on its stream to do that for you. You have to be able to do that. Um, and then the last bit here is that the stream that you are sending to the real-time service generally is going to include silence even if it's to a speechto text service because it needs to know when you stop talking so that it can

cluster that and send you end. You know, if you use deep grammar, if you're familiar with deepram endpointing, that's one of the things it does is it waits until there's not just silence, but that it hasn't generated a a word in a while. And that's when it it knows, you know, that you're done and it's time to move on to the next stop. Um, so yeah, the

services will mark the beginning and end of usually your audio service and segments and their audio segments that you need to be able to process that. Um, we also learned that websockets are used across the industry uh to interact with those real-time services. You know, like I said, deep, open, the other services generally for real time are using websockets, which makes sense. um they almost all use

JSON uh formatted text inside there for their metadata and you know sometimes some of the services will encapsulate their binary data within a JSON message sometimes they'll send you binary separately but generally you're using JSON for that exchange so this all led us to the question uh why not just you know jump on the websocket bandwagon for asterisk and use websockets ourselves this is the way this

to thank uh Alvida for this slide. So why use websockets? Um all the services we found were already using websockets. Like I said uh it sort of has become a de facto standard for any kind of application voice application that you're running. Um they can handle binary data. So you've got your audio stream set up. They've got uh your commands and your metadata via JSON. So you're

set up there. Uh ARRI already was using websockets for for their outbound events. So we we had that there. Um and there are like I said there are just wellestablished methods at this point for doing voice applications over So the first piece of this puzzle uh was to implement ARRI rest over websocket which from presentations we've heard over the last couple days it sounds like people already

take advantage of this which is fantastic. All right. So, historically, um, like I mentioned a little bit earlier, uh, using ARY required using two different communications channels. So, you make an HTTP request, your post or your your put to, uh, asterisk. Um, but then you need to maintain an open websocket the entire time to receive these synchronous events back, right? You can do a rest request and

a response for the synchronous events but you need that additional uh websocket to be open all the time for the asynchronous events. Um but now we are letting you do everything over the web websocket. We use just a simple JSON wrapper inspired by swagger socket inside of the theRI request or inside the websocket request. You send a rest request and you receive a rest response. So this

looks exactly like it did over HTTP. you're just doing it over websocket. Now, um you can there's both the concept of a transaction ID and a request ID that we've added that helps you correlate the request messages with the responses because you might have multiple requests that are within the same transaction but have different requests. Let's you maintain all of that together. Um and then within that

there are four different options for uh specifying your parameters within the request. But we generally select say, you know, once you've chose one of these, stick with it. Right? So, we've got a query string and the URI. Again, that's to be sort of like mostly compatible with existing HTTP requests. Uh, query strings array. Um, your your form URL encoded message body or a JSON message body. Um,

I personally tend to lean more towards just doing the JSON message body if you're going to be incorporating some kind of uh thirdparty, you know, API or library to do the JSON for you. um you're already using that for everything else, so you might as well use it for this too. So that's the first part of it. The second part of it would be ARRI uh outbound

websockets um which particularly valuable if you're running inside a container. You don't want to have to you know making outbound requests out of a container is a lot easier than you know keeping ports open all the time. So yeah, historically we've uh accepted websocket connections from external ARI application. Um but since a stasis application can only be handled by one uh websocket, um this limited opportunities for

really building a a scalable uh application set on top of it. So we wanted to add the appbound websockets so that you could uh more easily do load sharing and redundancy. Um there are two types of outbound web sockets that you can make to your external applications. Uh persistent connections which basically open up when asterisk starts and maintain uh stay up for the entire time asterisk is

running. Um all the activity for that configured stasis application will go over that one connection or you can establish uh protocol connections which can be a little bit uh even nicer in terms of of load balancing or sharing. um you you essentially set up a template connection ahead of time and then when you invoke your stasis dial plan application it creates that uh outbound websocket connection that

stays up for the life of the call. Uh in both cases asterisk will attempt to reconnect if the connection is broken. Um for the persistent one it's we'll just keep trying as long as asterisk is open. The per call you can set a number of times and then it will fail out but we'll we'll get more into Uh configuration happens in two different config files. ARRI.com and

websocket.com. Um in ARRI.com, you uh build your application. You give it a name. The type is always going to be outbound websocket. You give it a uh a client ID that is references the config section at websocket uhclient.com. Uh you specify a list of apps that are going to be servicing this connection. And then there's a subscribe all um which is you know the equivalent of a

subscribe all that we had before. So basically you can say am I going to subscribe to all the events the typical ARRI events you get before. And then lastly you specify a local ARI confuser uh an ARRI user um for uh the handshake for establishing the connection. The other part of it is configured in websocketclient.com. Uh the type there again is a websocket client. Uh the URI,

you specify the URI you're going to connect to. You can specify WS or WSS if you want to use uh TLS or not. Um the protocol is just going to be ARRI. You have your O username, your O password, and then you're going to specify is it going to be persistent or is it going to be a per call configuration. Um and then lastly for uh for

the per call configurations, you're going to specify reconnection attempts. So how many times you know do you keep trying to make this connection during the call if it fails? Um there are sample ari.com there are sample websocket.com files that are up in GitHub. Recommend taking a look at those. They will go into more detail about how these all work. So persistent connections uh when the connection is

initialized this automatically creates the uh stasis application for each app listed in your apps list. Um, essentially the same thing that happens if you were to connect and do a get like uh with the previous HTTP methods. You can see here specifying app one, app two. Same thing. You're just going to create one of each of those. Um, and then for your convenience, a dial plan context

uh is created for each of those with the stasis and the app name. Um, just you know so that you can send stuff to it. Uh, incoming channels can be sent directly to that context or create your own. Um, in either case, the uh once you know you call the dial plan app with one of the the stasis names works, you know, very similar if you've built

stasis apps before. Uh, once the apps are registered, uh, it'll attempt to connect to your application. Uh, using connection timeout and the reconnect interval parameters that you specify. So, you can specify, you know, how long do we keep connecting and how often do we keep connect? And then each time you connect or reconnect, you're going to get an application registers. uh event for each of the stasis

applications that you have listed. The per call configs are a little bit different. Uh connections are a little bit different when Astra starts. Uh you're only going to create the dial plan context. Um and then nothing happens until you actually call Stasis from your dial plan. Um then what that happens, Stasis is going to go and look at the inter the internal registry, see if there's any

uh persistent configs connections. If they aren't, then they're going to see if there's an outbound websocket with a pro per call type configured and if there is one, use that. Um, it'll then create an ephemeral uh AI stasis app with that with the app name and the channel name and that's what shows up in your application registered event that your application that you'll see over the connection.

So, when it first comes up, if it's one of these protocol, you'll see app name channel name. That'll live as long as that connection stays up. um if and if again if it fails you'll get uh you you'll retry based on the number of times you you configure. Um and then if you uh if it doesn't connect your application goes down or for whatever reason the stasis

will return to dial plan with a uh the failed and you'll go and go back to your dial plan. Um the big thing that this allows for the the big improvements is that you now have the o opportunity to have several different types of load balancing and failure recovery that you didn't have before. Um the two sort of methods we suggest uh you could use a common

DNS host name that routes between the different applications so that if you know you you basically just rely on your DNS server to to choose between in case of a failover or have a roundroin process for for the different applications that get invoked each time. Um in either case if a connection drops uh you know you go on to the next one. The other option is to

have a websocket aware uh load balancer that which then you know you're configuring it at that point and you're distributing it you're you're in control of how you're distributing that amongst your different applications. Um, you know, not just, you know, if you think about that, there's more to it than even that. You know, you could essentially have the different application, you know, everything go to one load

balancer and then the load balancer looks at the stasis application and chooses between, you know, which application's actually going to service it based on the, you know, channel name or the actual app that you're using. So, you can distribute not just based on load, but on this method or or app if you want to call it that. Um George has written a bunch of sample code uh

that show these different methods in in use. This is they're written in Python. They can be found here. Um and then again I will put up a list of references at the end with uh with all the different samples. So the last piece of these three things is the new channel driver Chan uh Chan websocket was designed to ease the burden of uh ARRI application developers getting

media in and out of asterisk. Uh the external media rest endpoint already has uh existed and had two other channel drivers which were audio socket and the RTB that we we both mentioned earlier. Um but they both require uh manipulating the binary audio the packets themselves. Uh not the most fun if you've ever had to do it. you you know you've got to time it and then

in particularly with RTP you're also framing it right so you're taking your your whatever is it u law linear whatever the packets are you've got to wrap them in an actual RTP packet and then send that over the external media port um and then just a note here that uh when we first uh we first put this in uh the the the control method the control uh

packets were just in plain text but we realized pretty quickly that uh JSON was the method to be using and not just plain text. Um, so we've then since uh added to JSON, we recommend sticking with the JSON. Um, we had already done releases at that point, so we're, you know, it's in there. We can't just yank it out. Um, so the support is still there, but

going forward for new features, we're going to be using JSON. Um, and then the last bit is that this new uh driver is available starting with 23 uh.0, 22.6, 211, and 20.16. going to take a little bit deeper dive into Chan websocket. Um so you can send and receive media using most codecs. Uh TLS is again is supported because you can do WSS. You can send arbitrary

packet lengths to asterisk. The the channel driver is responsible for breaking those up into actual uh suitable byte sizes um to convert to RTP. You don't have to time it. Um, obviously if you fall behind and you're not sending audio, you know, asterisk isn't going to magically make up the audio and put it in there, but as long as you keep that buffer full, it'll take care

of that for you. Um, the channel driver can accept incoming websocket connections, although you do still have to dial them when we'll get into that. And it can also make outbound connections to your application, which again for containers can be very useful. Uh, just having outbound websockets. Um, and then again, while it's targeted at ARRI, external media users, it's not specifically tied to ARRI. So, in theory,

if you had um a service that um, you know, simple service that was running outside of asterisk that you just wanted to pump audio to and it knew what to do with and you didn't have to do any, you know, ARRI with it, you could open up that up on websocket directly and start sending and receiving Um again it's it's implemented using both text and uh binary.

So you can answer um you can hang up uh you receive certain information about the the audio which I'll get into I'll cover in a familiar slide next slide. Um but it's all over the same connection. Uh so for outgoing connections you have to preconfigure a websocket client and a um that you're going to reference into your dial string and then so here's an example that we'll

see with this. So provided we had already defined connection one in our uh you're just doing a dial websocket you give it the connection name and then in this particular instance we're saying with the codec uh with a mu law codec fairly straightforward right once the application accepts that connection you're going to get a media start uh text event saying you know we're here we're going to

start sending that and it's going to have the basic channel info and the channel variables um which again you still configure in the comp files which channel variables are going to come through. So you can like I said you could because you can't I mean you could do this without ARRI you can really kind of put everything you want in inside the channel that you send to

your to your app. Uh incoming connections are a little bit different. you still have to um use the global HTTP server uh to um create that connection like you would with any other you know websocket connection to asterisk. Um but and then you also still have to dial it as well. So you would dial it, it sets up the path, you call in through your websocket. And

in this case, you can see we're specifying in, you know, looks pretty much exactly the same, but instead of specifying a pre-used con uh configuration or connection that was defined previously, we're just going to use the incoming flag and that tells asterisk to that that's coming in. Um the channel will be created immediately. uh the media websocket connection ID channel variable will be sent to that ephemeral

connection ID and that what has to that's what has to be used in your incoming URI. So when you come back in that's how asterisk is able to connect the dial with your incoming request. Um and then again once it starts you'll see a media start as the same as going the outbound. Um, and then the last note on here is that whether it's inbound or outbound,

um, by default, we're going to answer the channel. But if there's any reason that you don't want to answer the channel, uh, you can specify, uh, an N, which we'll see some examples later, um, and then either answer it via ARRI or as we'll see on the next slide or two, you can actually answer it directly uh, from by sending the text uh, over the the same

websocket you're using for the just a little bit. Like I said, this is pretty packet dense, so I'm going have to pay it pay it back a little bit slower later. And I talk fast and slowly or and quietly, which is perfect. Uh so the message size is basic is based on what the internal as frame size. So if your SIP side is u your packets are

U law, RTP packets but without the RTP wrapper around it. So if you're p times 20 milliseconds, you've got a 660 p byte packet. Those are the chunks that asterisk is going to send to your app through the binary part of the uh the websocket connection. So it's almost exactly the same as an RTP except there's no RTP. There's no header around it. It's just the raw

data. You can pipe these directly to a file or you can send them off to your text to speech or your AI or your conference app or whatever it is. um you generally can just take those and send those directly over the pipe as is. And that will include, you know, there's no uh silence detection in there. So if your person stops talking, you're just going to

get silence packets that are going to forwarded on to you. Um media sent back to a from your app is a bit trickier because, you know, for the most part, you know, your SIP and RTP on the other side. So if you're sending short or too long or mistime packets, you know, if again anyone who's written one of these apps before, it doesn't not a great experience

for the user on the other side. Uh you get pops, you get clicks, you get gaps, you get uh slow talking, you get choppiness, all that kind of stuff. So to help you with that, um Chan websocket will do the framing and timing automatically for Um, but there are a few exceptions and rules you have to follow. And some of this is still in flux because um,

even as of a couple of weeks ago, we were padding short packets and inserting silence. And I'm not sure if that's actually gotten merged yet or not. Um, but we found that as long as the uh, the audio that you're throwing in doesn't have any extraneous header information or anything like that, uh, you generally don't need to do that. So, um, George had put out an email

about it on the the mailing list. If you guys haven't seen that, we encourage you to try out his his his uh his pull request. If it's not in already, give it a shot. See if that works for you to know if we're 100% going to adopt it or not. Um, so when it's created, you're going to get a media websocket optimal frame size channel variable that

tells you uh, you know, how much audio is needed to match your p time packet size, right? So again, you'll get that basically for Uline, you're going to get that 160 back. And so you'll know, okay, a general I need to do this. Well, you can, but you don't actually have to. Um, if you send a websocket message with the length that's a whole multiple of that

size, right? So n* 160, uh, the channel driver is very happy to break that up into 160 byt pieces um, and send them one at a time to the core. And Asteros is very happy to take that audio and pass it along. However, um depending on which your audio source is like so for instance uh you know uh OpenAI in my experience the audio streams they send

back are pretty re they're consistent frame sizes that they send you back. So it's pretty easy to break that up but again like deepgram will not deepgram is just going to send you audio blobs. They can be odd sizes they can be even sizes they can be anything. Um so to deal with that there's a new start and stop uh buffering commands that are added to Chan

websockets. So if you've got a blob and you don't want to deal with breaking it up into N160 bytes, you can just say start media buffering, send your chunks. When you're done, you send stop meter buffering and asterisk knows that this is the set of this is the audio I have to deal with to break up into into the right size fix uh pieces. Um and this

is done again, you know, because Asterus needs to know the beginning and the end of that. These are the commands that you can send to the websocket channel driver. I'll just go through them quickly. Start and stop. Mentioned those before. Flush media and pause media and continue media. We alluded to that a little bit earlier. So if you detect somebody talking, you can say pause the media.

And then maybe they're done talking, you can resume it. Maybe they keep talking, you want to flush it. Um, again, you know, if you even if your AI has said, hey, you know, you need to stop, you know, I'm or the AI knows to not keep sending you audio because it detected that the person has t has started to talk. If it's already sent you audio because

it's sending to you faster than real time, you need to be able to flush what's already in the pipe. Uh, get status, which will trigger a status message. Report Q drained. So you can tell it uh tell asterisk that you want to know when the queue is drained or not. Mark media which is um quite useful. So like let's say you send a series of phrases in

in in a set of blobs to asterisk and you want to know when you get to each different part of those those segments, you can mark it and then basically when when asterisk processes through the audio up to that mark, it'll send you back an event to let you know it got through that part of the phrase. Um and then you can answer and you can hang

up. Uh and this is what it generates back to you. You get a media start which we mentioned before uh you get at the beginning when you connect. There's a DTMF end event which is very useful if you're building a hybrid IVR for instance or if you want be able to take you know a voice command or you know you've got an alternate you know zero press

zero press one you know whatever you need to do. It's very useful for your application to be able to get DTMF events as well. Uh uh media X off and X on. So you know Asteris has a finite buffer that it's going to to to hold on to audio for for you. Um obviously we can't you know otherwise you could send in theory right you could send

like a gigabyte of audio to asterisk and it would crash immediately and no one would have a good time. So we have a fixed buffer size with a high watermark on and off. The status which will be in response to your get status. The Q chain which again is in response to you requesting one. Uh the media buffering which is different from media buffering completed which is

generated in response to if you do a start and stop you'll get the the the completed when the stop uh the the the audio demarked by the stop has been proceeded. Uh, and then the media mark processed, which is just your echo back of that mark for your your spot inside uh your different phrases. Uh, here's a couple examples. Um, create and dial the channel to connect

your application using connection one. Very straightforward. You know, you post and then again, this isn't necessarily tied directly to uh ARRI uh over websocket, although it makes sense to use them together. But this is what that would look like. you specify you know again websocket your connection uh and then your codec avail the next line is essentially the same only for an uh and you you can

choose and if you see you'll notice too that these aren't direct these aren't using uh external media there's some slight differences which even I'm a little bit murky on what the difference is between when you use external media for this or not uh but it sort of changes some of the the uh channel variables you get I believe uh and And then the other ones are essentially

the same thing except for you're specifying external media. You can see you're specifying the transfer data uh in JSON format. Um the external host in this case it's specifying a media connection just it's just a different name you know that you're you it's named media connection so that you know this isn't your application when you defined it in your websocket client. Um and then the last one

again wait for an incoming connection very similar. So putting it all together there's outbound uh websocket connection for application control and complete right you have outbound websockets that can create the automatic connection to your application they can be persistent or per call websocket or websocket secure and you can load balance via DNS or a websocketware load balancer um ARRI over rest uh or yeah AI rest over

websocket you can make and accept calls without having to use HTTP it's same as using you before only you now use the request requ uh rest request and get the rest response which is encapsulated in JSON. Um there's a websocket connection for audio. It can be incoming or outcoming. Uh you can answer any calls and you can control when audio starts, stops, pauses, resumes. Um and then

I want to throw a thank you to George for his hard work on implementing all this. Although I'm going to shout out a little Ben worked on this recently. So he just added the ability to choose uh the direction. So you can have in only, out only or both. Something Ben just added and I think that went in what a week or two ago. Um and then

I'll leave this up. This is a list of references um where all the slides are essentially pulled from different things on here. So if you're interested in doing that, I would uh recommend taking a look at these. >> Thanks, Mike. Thank you. >> Do we have any questions? Oh, Alex. >> Hi there. Uh, thank you for all this amazing work. Um, I might be just narrowly confined

in my imagination to a very particular kind of use case for Chan websocket, but I'm having a little bit of a hard time. And this is going to be a moronic question, but what is the kind of use case or call flow that is you have in mind for the incoming the dial to incoming connection? How does that work? What what would you use that for? Essentially,

if like let's say that was your other you wanted to open up a second channel uh on the same call like let's say you're building a conference or something like that and you're actually doing the conferencing offbox or you wanted to merge in a secondary stream you're connecting multiple agents or something like that your call initially would come in it would get dialed and then you could

then invoke the dial and created the incoming there anyway. Um, it's also just for, you know, depending on what your network topology is or, you know, there might be a case where it's easier for you to make a websocket connection into Asteris than out of it. Just leave it there. We wanted, you know, this is our first stab at this. We wanted to really make sure all

the options are open for people to do different things. >> Maybe is it helpful to ask an immediate follow-up on the manual part of the dial invocation? You could do that through the ARRI as well. So, you could that would be the next line maybe. Is that how how to actually initiate that dial? It's not uh >> right. It wouldn't necessarily be another spot in dial plan.

The ARI application that you're connected to could say, "Hey, expect an incoming connection from this other websocket app or something living in another container or something like that." >> Okay. I don't know if I helped it, but maybe I made it worse. >> No, I just don't I'm not personally familiar with a lot of real world services and applications that would actually want to connect back in.

>> Um, >> again, it was We're mostly trying to cover our bases, but I think you know particularly if that other you know like said talked about you know websockets are generally easier to make out of a container than into it. So like maybe asterisk is running on bare metal but your app is is running in a container somewhere and it needs to stream from >> Good

question. Anyone else have a question for Mike? >> Okay, thanks so much Mike Perdine. >> Hey we all get to go to lunch. about an hour for lunch. Um, if you can make it back here at 1:00, so little under an hour right now. Uh, if you haven't grabbed a t-shirt, please do so on your way out so we can come back for the picture. At 2:30,

we uh have two talks. I did check um we uh Conrad is speaking out about uh soft phones afterwards. There is some AI uh involvement there as well, I think. So, it is a full voice AI day today. every speaker. Congratulations. Grab your mug if you're a speaker just because you're a speaker, not because you included AI. Uh but we'll be back here at 1:00 with that

talk and then we'll wrap it up after that with a state of the Asteros project and then have a group picture. So if you want to wear your t-shirt for that, that'd be great. Thanks again. See you in a little under an hour. Byebye. on. Yeah, sounds good. We'll get started in a couple of minutes. Uh I think there's still some bags of chips in the back.

Looks like there's fresh coffee, tea. So, one grabs a snack before we sit down and get >> All right, I think we're going to get started now. Everyone, welcome back. Our last couple of sessions the afternoon, Astrocon 2026 at scale. My name is Chris May, open source solutions advocate. Bit of housekeeping again. Start out. We are uh looking at our after this talk a uh asterisk overview

of what's new and what's changed in the past year uh via telesatellite connectivity from Josh and answer some live Q&A. we have that and then after that we'll try to take a group photo quick if we can and then we're looking at uh migrating over to the exhibit hall for the rest of the afternoon saying asterisk booth so and you didn't get a chance to say hi

yet um you can walk in to the left uh it's in the south uh east corner of the building so we're alphabetical under asterisk uh you know if we're I guess we're kind of near the Facebook or the Meta as it is now um booth. So, there's a huge number of sponsors, a variety of uh swag. I just picked up a small penguin myself um on my

my walk over there. We have more stickers and t-shirts if you want those. If you haven't gotten one already, please do. Uh pens, books, all that good stuff. So, and we'll be there throughout the rest of the weekend. So, lots of other, you know, talks and events going on. So please do share your asterisk knowledge with the wider scale community. And speaking of sharing some knowledge, a

great talk on building a soft foam for scale by Conrad. Thank you. Yeah. So today I'm going to be talking to you about building um a softphone. This is our little journey so far. Uh my name is Conrad Devet. Company is called Superb. And I'm going to open with a quote. Uh this is from Sam Alman. Um he said it a few times. Essentially he's just saying

that AI is going to make it possible for one person to build a billion dollar company. Um so show of hands. This is a bet. Does anybody think it's possible? Billion dollar company. One person. The question really to me was, do we really want to make a one person billion dollar business? And I don't think it's a very good idea. In fact, I think it's really quite

a bad idea. Um, some of the reasons, well, I think you have an incredibly high reliance on a single individual, not a very good uh for a succession plan. I think you tend to believe your own arguments quite a lot. um easy to agree with yourself. It is of course easy to give up. Um and it's probably quite boring. So there's Sam at this table. Um but

it did get me thinking, you know, if I wanted to build a business for the um you know, what what in fact would it look like? um taking into account, you know, AI and what it could do for a for a business now and in the future. So, I like to structure it as such. Um I would make it trust owned, certainly not owner operated, although that's

something that often takes a little while to get there. Um I had the number eight in my head, eight humans sort of as a panel with any number of artificial intelligence behind that. Um I I think uh a business that we should businesses that we make today should be globally scalable. So that's sort of some of my key points there. What would we make? Well, for me

obviously it was quite simple, but I wanted to build and provide the service of a softphone with uh the built-in proxy and uh I believe in the web RTC principles. Um I believe that calls should be encrypted and the signaling should be over TCP link. So that was something that was important for us. Um, how would we do it? Rely heavily on AI. Uh, I wanted to

build a mobile first but not mobile only business. Um, and I wanted to build an application which had a single user experience. Uh, that single user experience I'm going to go into a bit of detail now in a sec. Um, and I want to build a business for scale. Goes without saying for today. um, back in 2024, uh, after finding myself a little bit stagnant in the

browser project, I saw browser phone early on today. Um, I think it was Diego. So, it's a popular repo. I find it popping up in a lot of places, but for me personally, I think I'd got into a little bit of the Sam Elman problem. Um, and in March 2024, I hired some developers, sat down, we formulated a plan. We wanted to build this as a product,

as a and as a I'm going to take you through today three sections of the business. the development, the tech stack, and the so you know, being a business that's designed to have a small number of people and sort of working quite largely with AI, it meant that we took different decisions in our development stack and our tech stack and our support environment. So one of the

things that we wanted to achieve was that this idea of a single user experience. So it meant that we wanted not just a similar user interface across all the devices. We want the actually the same user interface. We wanted to present the same UI on all the devices across all the platforms. And we did this using one technology which is actually better than anything else we came

across and it's actually quite simple. HTML and JavaScript. It by far was uh a better or by far a more similar experience across all the different flavors of Android, iOS, browsers, desktops, applications. So like in the example of those three screens below, it's actually the same HTML being rendered across all these devices. So I think it's something like nine different platforms in the end of the day.

what we used to achieve that, Tailwind. Tailwind is a fantastic toolkit that allows us to develop the styles and patterns for the UI. There's a bit of an example underneath. You won't be able to see that, but uh, essentially the way we did it is we built a single HTML document that populates all the user interface. Okay? separating out UI from code. Uh I think this was

a very old principle from the C and C++ days where you had your UI completely separate to your sort of business logic. And over the years it's sort of mixed in. You get sort of more React style looking interfaces or code nowadays. And for us we we wanted to separate them out again. So HTML is a separate document. That's all the user interface elements in one space.

Then we have our JavaScript core. So our JavaScript core ended up looking like this. And it allows us to have one team, one focused skill set for the entire product. Uh starting in the middle there, we've got the browser phone SDK, which is the sort of evolution of the browser phone. That's the GitHub project. Um wrapped in a thin layer. On top of that is what we

call the web phone SDK which kind of gives our core arms and legs and various appendages that kind of makes a phone but it's not a full uh product. So then the what we call the superb layer on top of that is now the branded layer. That's the product that we sell and that's what people interact with. um anybody could sort of enter in at any one

of these layers if you really wanted to. Um and then the sort of the product that we sell and everybody sees mainly is the superb layer. And then just to illustrate that's the layer that's deployed down onto browsers, onto desktop applications and mobile phones. um and sort of how that kind of looks if I illustrate it on a mobile phone for example. You know, we layer down

a web view which renders the first layer and then the superb layer and then of course the actual phone integration integration layer which would be something like call kit um for push not uh voice calls, push notifications, things like that. That's the other layer. So it's a it's a kind of neat layered approach but it does mean that we rely on web view. Uh web view is

an interesting one because we end up in this web view land. Uh operating systems don't like us. They tend to unload you very easily. Um we can achieve the same kind of performance. I mean it's not a game. We're not doing 3D rendering or anything like that. So, uh, the user performance is fine and the scrolling is just as good. So, we don't really have any problems

like that. Um, but it seems that operating systems just end up unloading you for um, memory conservation. So, we've just got to handle that a little bit more carefully. It also means that sometimes it's another step in to communicate with your JavaScripts on that layer from the phone layer, but you don't have a lot of integration, a lot of interaction between the two. So it uh it's

generally fine. Uh we find though that essentially the performance and the um the ability for us to deploy code is uh it you know sort of makes up for any of these um disadvantages. Um but I just want to sort of pause a little bit on the development side. Is it you know a couple of months ago it was tab tab tab you know we work in

our code editors and produce code magically. Uh now we're just given a prompt right make me an app. So you know it it sort of made me sort of think a bit building this future business to say but you know what does it what does it mean to write code now you know what is what is the value of code when somebody else can just come along

and make me a soft phone right easy um so what it what it meant right from the beginning in fact even in 2024 um we had this idea that it's not so much the code that we write. It's actually about the service and the infrastructure that we provide. That is what makes us a business. Um so just in sort of like summary, I feel like in the

future, you know, companies would be sort of judged or companies are going to be successful based on, you know, how well they integrate. It was API, now it's MCP. Um kind of just imagining a you know, you just ask your agent, make a call to mom, you know, it's it's it's not an app anymore really. Um, you know, and the AI might find an MCP that's a

good route and a reasonable rate and send your call out through there. So, you know, that's probably a little further down the line, but the idea is that, you know, the the value of the code drops with AI and the value of your service and your infrastructure increases and for us that was something quite important. So, we really did need to create a tech stack. So, that

means going onto the tech stack, this is what we we did. We realized that obviously, you know, this is almost brings in a whole new department when you're putting down equipment. Okay. Um, if you the the usual story, if you put down one server, you need to put down two, right? Um, failovers, all that kind of stuff. We speak in arrays now. We don't have a server,

we have a server So in our hosting objectives, uh, one of the things we wanted to do because we could now because we're providing it as a service, we can get the passwords out of the browser and that was something for WebRTC that's always been a bit of an issue. Uh, we can hide your IP address, right? It's the proxy's IP address, not yours. So these are

two things that we could do. And then of course we could grow the server side features um endlessly right we add on AI integrations all this kind of stuff recordings transcribing all these kind of things can happen on the server side um how we did it we use Amazon Web Services where I've pretty much been part of the Amazon Web Services um for years and It's uh

I've just stuck with it. We use serverless as much as possible. We have a handful of EC2 instance in instances but mainly use um Fargate. this is the tech stack. This is what it looks like. Um we use open SIPs as a proxy. Um it's it works with your asterisk box. So there you can see in orange on the left your asteris box either in your office

or in the cloud hosted etc. We sit in the middle between the client the softphone client and there's two registers. the one register which is the register between our SBC and your asteris box that's a permanent registration and then when a call comes in then of course we can alert your your instance say it's the mobile app and it's in your pocket we send the vo push

there's the call or if it's in the browser or if it's in the desktop app it's all the same another thing that we do in that is that we obviously provision down your details. So you just launch the app, all the details about who you are and all your credentials, that kind of things all pulls down for you. So you don't type it in into the All

right? So you know, when you make these systems, you think this is how it should work ideally, but then of course everybody has their own sort of ideas. Um but yes there are a lot of local asterisk in in instances installed in people's local lands. So one of the things that we needed to be able to accommodate for is people's asterisk boxes in the on their local

lands. And so we brought in something called the websocket registration mode which now just bypasses our proxy. Okay? It doesn't go through us. It unfortunately means your passwords are now in the browser. They provision down the passwords, but they're your passwords. So, they get go down to the device. Um, but the call is made, the the websocket connection is made directly with your asterisk box. And as

of fairly recently, we can now do that on UDP, uh, TCP and TLS if you're using the desktop application. Um, and we accept self-signed certificates on the websocket connection as well. This doesn't obviously work on mobile devices. There's just uh mobile devices are just not, let's say, online enough. They'll be easy to switch over to 3G and then of course your asteris box in your office is

just as unreachable as it was just now. So uh it's not an option for mobile devices but for the desktop and for the web application that is an an option for you. but every business is going to have this problem eventually right you need support. You have customers and as I described earlier on we're a small team less than eight people and we need to now support

customers. So how do we support customers? We not we're not going to have call centers. We're not going to have people around to take your call and hold your hand through these things. So we looked to AI um but we had some fairly strict uh of what we were going to do to solve our support problem with AI. Now one of the things that I learned in

business many years ago was that it's very important to focus your time and your efforts on the things that make you or your business unique. Um, you know, if if you sell coffee, you're not going to spend your time developing an invoicing application, right? Just use one. They exist. You focus on selling coffee. And the truth is that, you know, even with Microsoft Word coming out many

years ago, probably people were probably asking the same question. Oh, do I need to now become a Microsoft Word export e expert in order to work in your company? Well, yeah, you probably do. You probably need to be able to use Word. If you pitched up today and said, "Yeah, my typewriting skills are great." You're not going to get a job. So, we applied this to our

support and said, "Do we need to take on support? Is it something that we can outsource to AI? Is it something that we need to own? The truth is that we felt like with quite a bit of work in AI, we could get it to be better than an actual support person. So our objectives were this that we wanted to be able to answer customers in real

time, right? No cues, no waiting, no nothing. get your answer straight away. We want to be able to answer them on multiple channels, email, text, chat, and audio. All right. And we wanted the same agent, right? We want the a this agent with with all the information at its disposal across any of the channels. So, if I emailed in and said, "Hey, I just sent a WhatsApp

message. Did you get it?" or whatever the case may be. Um the AI can cross over the channels. Um and we felt like over time this would build up what we call the customer digital clone. So armed with all of this information across your different the AI agent becomes this digital clone. but of the customer residing in our environment. It has access to the knowledge base and

the website obviously it can access the customer's admin control panel. It can access transcripts, text, call history and support tickets, over time, this customer digital clone actually becomes far more useful at answering questions with this customer than we ever could be. We just don't have that length and breadth to know every text message that you sent in. But AI can know every email that you ever emailed.

AI can. So, what it turns out we'd be able to do is in fact ask the AI agent, not the customer, you know, give me a customer engagement score. How's it going with customer X or Y? They'd know, right? We'd ask the AI, is there any outstanding items with any of the customers? Do we need to get back to anybody on that? AI can give you that.

Anyway, as things can progress, we can even get it to send emails back to the customer periodically, follow up, ask them how the kids are, all that kind of stuff. Although, we've put the handbrake on that one. Think that might be jumping the gun. So, what are the channels that we build? Well, well, first of all, we're doing this all through OpenAI. Um, it provides everything that

we need and out the box, we use the chatkit SDK. That was a fairly straightforward implementation. Uh, you can see in the middle there that just drops into your Uh, it's tied in with a vector store, which is your sort of knowledge base. um and the sort of text prompt and response system. Um it keeps the history all by itself because if you use the agent builder

in open AI, you've got to remember to use your conversation history, your conversation ID because otherwise you don't actually have history. So we store the only real thing we actually store is what channel you're you're communicating on. um a reference to that channel. So if you called in, your caller ID is your reference. The channel audio and the conversation ID. So each channel is its own conversation

ID, but the agent can go across channels, but you'd have to ask it. You would say, you know, I called in yesterday. And then it could say, let me go and have a look at your calls from So the other channel we have is uh Telegram. That was a pretty easy one. Uh the the just add a bot and tie in the callback URL to a lambda

function. And the same with uh WhatsApp. WhatsApp's got a couple more hurdles you have to go through in order to achieve that. But they also have very similar sort of end product where the callback URL is initialized on every message. This lands up in our case in lambda and from lambda it goes through to the open AI prompt and response system where we keep the conversation ID

going. So uh endlessly the conversation just keeps building on itself. even if you come back later, you just continue on the same conversation on the open AI side. And um and then of course the last little piece in the in the puzzle with that, I don't know if you can see it, but uh there's a small little call button on the top right of the WhatsApp uh

message. So we can call in using WhatsApp. All right. So it's text and audio. And as it turns out, last year, probably about the middle of last year, WhatsApp opened themselves to SIP termination. Um, and so did Open AI. Go figure. So all we do is take the SIP call from uh Meta and send it in to Open AI. I'll show you what that looks like in

a sec. But in order to achieve that, we have to set up audio calling. And you use a special agent inside OpenAI called the real time audio agent. Configures slightly differently, but it's very similar. Um, we use asterisk as a sort of a stop gap in between that that allows us to utilize the backto back nature of WhatsApp of of Asteris. So the call terminates on our

asteris box and a new leg is then onwards dialed into open AI. This is what the config this little snippet of what the configs got look like. Not sure if you can see the the dark blue there, but uh so your pjip.com just sets up a fairly straightforward trunk really. Uh only thing unique is that we have to use TLS. They want you to use TLS to

to open AI. Um you don't use WebRTC. just web RTC iss uh I have allowed opus I can't remember if it has to be opus but I have it in the list there I generally use direct media is no force rport and RTP uh symmetric is yes it's pretty much a default for me I didn't really ever try anything else and there's the contact where we use

sipiopai.com So that's the configuration in order to reach open AI also fairly simple just dial you dial the project ID so the project ID is given to you inside the console of of open AAI uh you dial the project it's project underscore and then the ID of your project and you can see I'm referencing the open AI trunk from above with a timeout of 30 seconds. how

that all fits together and how we build it for scale. I mean, we at the scale conference, so let's talk scale. Uh AI agents are great, but what happens when you have many calls, right? Billion dollar business going to take a couple of customers. Okay, so um for this we turned to Fargate. I had been using Fargate for a couple of years before that. Uh Fargate being

a Docker runner. Um it's a native AWS product. Um I like it. Uh and it does what we need. So again, we didn't really think too far out. Um and I like working with Docker with asterisk. I know this is a bit of a funny topic for people. There's a lot of people think no, you can't run asterisks in a Docker or there's a lot of problems

with it. Uh ports often come up, but it doesn't actually work that way in Fargate. Uh they give you an IP pair that's a local and live IP address. And that IP address gets attached to your Docker instance. And it's up to you with your security groups to essentially secure that. All the ports are open. So UDP ports for media are not blocked. All right. So your

security group is the thing that secures it. But essentially it's a thing that also keeps the ports open for media. All right. So we also like the CI/CD flow using Docker because we're not logging into any consoles asterisk minus R. Right. Okay. If your instance is running, it's running. Leave it alone. If it's not doing what you want, rewrite the Docker and press deploy. And that's it.

So, we tied in the deploy into the GitHub repo actions. And that gives you a workflow of edit Docker or or your configs, run test locally if you want, push to GitHub, pull to release, and your actions perform your deploy. You just sit back to relax. Deploy takes about 8 minutes and you got a new machine. We build each time from um from from source. So uh

we do the entire installation process um from essentially from the GitHub repro. So that's our sort of CI/CD process with this. Um and then of course the big question then is what about IP addresses? As I said uh you are working in Fargate the IP address has to be known to asterisk. So this is how it all ties in. So earlier on I showed you that there

is a there was a TLS transport configuration applied to that trunk. So here it is. um you have you would have to use a substitution in your uh in your run script in your docker installation and you set the docker IP address to your bind IP. There you can see I've used port 5060 on the docker IP address and we use the external media address entry. So

using external media address with your live a uh your live docker instance not the local the live docker instance IP address allows media to flow over you can see the orange arrow at the bottom there the media flows on its live IP address and the way we've done this which I can highly recommend is that your local IP address is the one that you're using for your

signaling. So in this example, I have the asterisk box with its local IP address for signaling. The invite arrives there and the live IP address for the media means that the ports can be open with relatively little risk. So a call flow starting from the left through any channel maybe a cell phone PSDN WhatsApp as we said or a normal VoIP call comes in and it lands

on our SBC that's our live IP address exposed to the world wellprotected server uh it's a an array uh of the invite arrives on that box. Now we have some rooting logic obviously not every single call goes into uh an AI agent. So it'll follow the routting logic and if it is for the asterisk array it finds an available asterisk server. So you will have a lookup

process. Remember of course if you want to replace Docker instances and things like that you'll have to drain them of existing calls and connections. So you'd have to find an available asterisk box that I believe up to you to work out the me mechanics of that. But the invite arrives on the asterisk box on its local That is then the second leg, the B leg. Remember the

dial uh instruction in the dial plan dials through to open AI referring to the project. Um, the call establish establishes and media flows from OpenAI onto your asterisk box on its live IP address thanks to the external media address and then back on to wherever you came from. Okay? Now if you did that you would land up in a situation where you wouldn't actually get a call

because there's couple of things that come into this. this call flow is a sort of like a madeup version of what you saw earlier on rather than using anything new in Asteris. No web sockets, no chan audio, no what was it? Snoop channels. Yeah, nothing. This is standard vanilla asteris. In fact, it could be whichever asteris you wanted so long as it supports the TLS connection because

what we did is we created a node service that runs on the same box as the And the call flow would go as follows. It would arrive on the asteris box. You create the new leg and the invite goes through as I mentioned earlier into open AI. In return, open AAI then uh call your registered web hook. That registered web hook is assigned in our case to

a lambda function. Now that's quite important. Our Lambda functions are always available, highly available and they can perform any number of tasks. Get into a database, find customer history, all that kind of stuff. So that's exactly what we do there at that lambda function. Uh we can collect call history, we can collect all the customer information based on the caller ID or whichever identity you want to

use. uh we store in Dynamo DB and we return that back to the Lambda function. But Lambda has got a hard cap 20 minutes that's it. You're not running any scripts longer than that. So the truth is you have to be answering the call and be on your way. So we get the call caller information and we then do a post to our own service. So that's

the node agent service because the one of the the parameters that we pass ourselves is which instance of asterisk is running this call because I obviously want to give this information back to the instance of asteris that's actually running the call because then I can perform functions on on the on the call. I can use any number of interactions uh rest APIs all that kind of stuff

in order to control the call. So the call information is then sent back to this node service. I called it agent servicejs. You can call it whatever you like but essentially it's a small little express service that's running waiting for a post. It receives this post. The next step that you have to do because of course up until this time the call is not in fact answered

yet right you have to post back to open AI's rest API you have to post back with a call ID and slashanser right that's the signal to the user that your call is answered you get back the 200 okay at that point all so that's our answer. But at this point in the call, it's answered, but the agent says nothing. Okay. So, what you need to do

now is immediately after answering the call, you need to establish a web socket. And that websocket is your call control web socket. All right? I think we heard something similar earlier on, but essentially you can reference your call ID on their system and feed it transcripts, feed it information and get back things like transcriptions, trans transcription delta, which is each individual word. And all that information passes

between open AI and your own service. So you have everything in control. You can control the call and you have the transcriptions. You also run inside that service your tools. So you define your tools inside open AI's console environment where you would say for example I want to get previous WhatsApp messages right and you run your tool from inside your agent service so the agent service would

be initiated this user wants to see his uh previous call history or something like that and that's where you would go and look up the call industry. One of the things that I did for this agent was that I primed the agent with information. But timing is quite key here because it is a real time call. So when the web hook posts the information to us, it

doesn't contain too much information so that the timing is quite quick. And what happens is then while the agent is saying hello, how can I help you? All that kind of stuff, we use the web socket to pump in the call history and the context of the call. So it has the what's called the instructions. That's just how the agent behaves. But I want to give the

agent the previous interactions, but I don't want to give it to the agent before the call. That just slows things down. I want to give it at a time that is convenient. And we found the best time is when it's busy saying, "Hello, how can I help you?" and all that. So that's a fantastic time to really pump in a whole lot of information. You can get

quite a lot of information in there. Um, so that is essentially how your agent works. That's what we built as a support agent. We're currently using this system now. It's live and people are texting the agents, some people are calling um, and it's got enough of the knowledge base uh, to be able to answer the questions. The the big key in all of this is actually to

give it enough information. That's when you don't get mistakes. If you give it enough information, enough context, essentially that's when you you get the best So, bit nervous about this one because I don't know about the audio quality, but anyway, I recorded it, recorded it in interaction, uh it's well worth uh listening to. Um, so I hope it's going to come out and hope it's going to

be loud enough. Um, it it it it uh it was pretty close to this, but I did cut it for different uh purposes. Yeah, if you just ready to kind of tweak the audio a bit there. Um, what you're going to see is I call into my own agent. Obviously, I mean, that's unrealistic because it's really for uh for WhatsApp, but I have a route into my

own agent and it's going to demonstrate its knowledge in the in the knowledge base. That's the first thing. Then we're going to get back to a couple of things like the call drops and the call back and it demonstrates that it can uh listen, it can pick up where we left off. I mean, calls drop, right? And you don't have to explain yourself again to the the

agent. Um, what else does it do? Uh, oh yes, I use a tool call to open a a support ticket. Um, and then I call back again and I inquire about the support ticket, but then I need to log in. Okay. And that login process, uh, it takes you through the login process, which it's well trained how to do. And then uh after the login process it

can read and feed back the information that it had off the ticketing system which is simply Zoho. So here we go. >> Hi this is the superb AI assistant. How can I help you today? >> Yes. Hi. Does uh Superb support asterisk >> Absolutely. Cyperb fully supports asterisk PBX. You can connect your Cyperb phone to an asterisk system using the Cyperb proxy or via websocket registration mode

depending on your setup. Whether your asterisk PBX is publicly reachable or on a local LAN, Cyper can bridge the connection and give you a seamless experience. Can you tell me how I connect my PBX to um the softphone? >> Of course. Connecting your PBX to the Cyperbone involves creating what we call a connection. This connection acts as a bridge between your PBX and the softphone. There are

two main options. If your PBX is publicly reachable, the Cyper proxy can sit in front of it. You'd set up an inbound or outbound connection in the admin control panel found in the softphone settings menu. If your PBX is on a local LAN behind a firewall, you can use websocket registration mode. In that mode, the softphone registers directly to the PBX over the LAN. Would you like

to walk through the steps for your specific setup? >> Uh, no, it's fine. Just tell me something. Um, using connections, does my Asteris PBX need to support WebRTC? Good question. Your asterisk PBX doesn't need to support WebRTC natively. That's exactly what the Cyper proxy is for. The proxy handles the WebRTC side, translating it to SIP for your PBX. So, your asterisk just needs to support standard SIP.

The Cyper proxy bridges the gap, allowing your saw phone to connect via WebRTC while your PBX Hi, this is the superb AI assistant. Sorry, it sound it sounded like we got cut off there. Um, what were you saying? >> Hi again. Sure. Let me recap. We were discussing how the Sipperb proxy works as a bridge. So, your asterisk PBX doesn't need to support WebRTC itself. The proxy

translates between WebRTC from the softphone and SIP for the PBX. >> All right. Uh, you know what? My Asteris PBX needs to have a TLS connection. Does uh, Superb support TLS? >> Yes, absolutely. Superb supports TLS connections with your PBX. >> You know what? I would really want somebody in your support department to confirm this. Could somebody do that for >> Absolutely. I can open a support

ticket for you and our support team can confirm all the details about TLS support on your asterisk PBX setup. To open a support ticket, I'll just need your email address so we can pass this on to the support team. Could you please confirm your email address for me? >> It's comrade at superb.com. >> Could you please tell me the subject and a brief description of the issue

or question you'd like to include in the ticket? >> The subject is whether Superb offers support for TLS connections. >> Got it. The subject we have is confirm if cyperb connections support TLS transport. Now, could you give me a bit more detail about your setup or what you've tried so far so we can include that in the description for the support team? >> Well, I would just

want a human support person to confirm that TLS is in fact supported um with the connection system between Superb and my Asteris PBX. >> I've opened the support ticket for you. The support team will follow up shortly to confirm the details about TLS support with the connection system and your asterisk PBX. If you have any more questions in the meantime, feel free to ask. Your support ticket

has been opened successfully. The ticket number is 333. The support team will get back to you soon with the confirmation about TLS support. If anything else comes up, let me know. That's the Zoho system. Hi, this is the Cyberb AI assistant. How >> Yeah, I'm following up about the support ticket that I just raised with you. Is there a response to it? >> Thanks for following up.

At the moment, you're not verified, so I can't check the status of your support ticket directly unless we verify your identity. If you'd like to proceed, all you need to do is go into the settings of your Superb softphone, enable audio call support, and generate a PIN. Then you can share your email and pin with me, and we'll get the latest update on your >> Okay, I've

done that. My email is conrad >> Thanks, Conrad. Now, please provide the PIN that was generated in your Superb soft phone so we can complete the verification. >> It's 6888. >> I'm validating your email and PIN now. This will just take a moment. Your identity has been verified successfully. Now, I'm retrieving your support tickets and we'll have the latest status in just Your most recent support ticket,

ticket number 333, is currently open. The ticket requests confirmation from the support team on whether TLS is supported between the superb connection system, the superb proxy and your asterisk PBX. >> Yes, that's the one. Um, you know what? Is there a reply to it? >> I've checked the latest details on ticket number 333. As of now, there hasn't been a response from the support team yet. Actually,

there has been a reply added now. It looks like the support agent responded and confirmed that TLS is indeed supported. and they even mentioned hoping you're enjoying the asterisk conference. The ticket is still open, but the confirmation is there. If you need any further assistance or have follow-up questions, just let me know. All right, so there you go. That's uh essentially the Oh, wrong one. yeah, so

that's my uh my support agent for me working hard for me there. If you're wondering, um, there's a couple of breaks or mispronounced words or something like that. I still think it's quite a good feat. That call was made from the bottom of Africa in Cape Town through to the UK and from the UK. That was where the asteris box is. And from the UK through to

open AI, wherever Open AI is. And so for that kind of distance to get that kind of low latency I think that's something uh really worth considering. A lot of people are doing or a lot of people are battling with latency. But when you sip dial directly into the open AI structure, it does the transcoding, does everything all for you there and you still maintain all the

call control with the web soocket and you can get your transcription fed to you live in Delta. Um, and you essentially get a very similar experience. Okay, that's uh essentially that's it. That's uh information about Superb on the on the the link there and that's the Telegram bot over there. The QR code if you really want it. Thank you, >> Conrad. >> For a one or two

questions before we move on. Yep. Conrad, thank you very much for that presentation. I like the um the way that you interrupted the bot. I thought that was really good, but I had a question on how to handle the ums and the a's and the pauses and the the human element. What did you go through or what have you gone through when somebody says, "Oh, hold on

a minute. Let me go." And then comes back and all those human things. Um >> yes, so we we started um we started the turn detection with with a simple voice activation. I think that was the the default in the open AI infrastructure and then it does actually have another option which is sentiment which is then where it's listening out for you ending your conversation or your

sentence. Um, I tested it a bit. Obviously didn't want to put it in that video, but things like clapping your hands while you talk or while it's talking, moving, squeaking your chair, things like that, it doesn't interrupt it. Um, the only tricky one is if you actually if while it's talking, you say, "Oh, okay." And you know, it's quite natural to encourage people when they talk. You

you agree with them or you say, "Mhm, ah, okay, cool." Things like that. and and it doesn't really handle that very well. But I think these things are getting better. So it's really the kind of thing where I don't really want to take that architecture on, I'm very happy to go with what is available. And by changing from voice activation, which was terrible, anytime you squeak or

make a single noise, it just jumps in. And so going to sentiment inside the settings inside Open AI was absolutely wonderful. It made it much smoother, made the conversation much more uh consistent and more humanlike. Great question. Yeah. Anyone else have a question for Conrad before we move on to our next session? Oh, Mark. >> Hey. Um, yeah. So obviously you guys are a softphone company but

have you thought about marketing this AI that you use for tech support? >> It was quite funny because I was um I was do I' I've done a lot of this in clawed um code. I mean I I vibe coded half of it. So the whole time it's prompting me the whole time coming back and says in qualifying questions in the beginning and saying is this a

product? Are you making this as a product? Are you sure it's not a product? No, it's not a product. We just want to make our own business more efficient. >> It's a that's a great origin story for asterisk itself, right? I mean, wanted to make own PBX efficient for Linux support services back in the day and that's why we got >> one thing one thing I would

say is that the style of architecture of that system does mean you take a lot of the technical burden off yourselves. There's no websocket audio packet uh stuff. You simply dial straight in. So, it simplifies things a hell of a lot and you get an audio agent in in minutes. >> Thank Thanks for sharing, Conrad. It's great to see. Hopefully, everybody can benefit from that. Thank you

again. Our next session will start in just a couple of minutes while Mike gets prepared for Josh's remote in. I think he'll be taking some Q&A as well. So, uh Josh Culcor's been the asterisk lead engineer for a number of years. Uh leads the the core team, has been a uh at least a dozen years or more. Started some uh many code contributions for voice recognition back

in the in when we called it uh speech to text. And so uh he has, I'm sure, some thoughts on AI and improvements that we've had in Asterisk over the past year. and he'll be sharing those in just a couple of minutes. And that's our last session. So, if you need another quick break, um looks like there's still some coffee left and tea and t-shirts outside. So,

if you are just visiting and stopping in, this wasn't your main event, please do feel free. I think everyone should have gotten their t-shirt who uh paid for it so far. So, we've got those available. Um reminder then will be over at the booth. If the t-shirts did just make their way over, that's where they are if you missed that. And at the end of this presentation,

I think we're going to all try to cram in over around the uh general sparkle uh stage and take a group photo at the end of the conference. So hopefully stick around for that right afterwards. All right, get started in The volume goes up. Yay. going to get started with our last uh session here. Um looks like we've got uh 30 minutes or so of of Josh

speaking to us remotely. Maybe he's even able to speak to us in between. uh we'll we'll bring him online at the end for some Q&A about uh the Astros project. So uh if um in case you're not familiar uh you know it's Astrocon 2026, our last session here at scale on Friday, March 6. So thanks again everybody for for coming. Um, we'll be over at the booth

in the uh southeast corner of the exhibit hall uh for throughout the weekend to talk over at the booth says asterisk at the top. You'll see our banners and things if you want to get some more pictures. So, speaking of that, we'll have a picture at the end here by the stage. We should have some t-shirts left over hopefully if you want to grab one of those

and throw it on before. Great. And we're going to uh hit play here and then answer take some Q&A from Josh Culpa, director of engineering on this state of Mike Berdine with the controls. Thank >> Covering approximately the past year or so. Uh I am your humble Asteris project lead Joshua Culp and I wish I could have been there to spend time at AstroCon with you all

but unfortunately circumstances did not allow. So I hope this past version of myself in recording form is good enough to give you some useful information for Astr from the past year. With that let's get started. So covering from a version perspective Astros 20 we did 11 releases with the latest being 20.18.2. Asteris 21 had nine releases latest being 21.12.1. Asteris 22 having 11 releases, latest being 22.8.2

and then Asteris 23 which was released in October um having the latest release of 23.2.2. Something to note is that 21 is now security fix only. So it goes end of life this year on October 18. So if you're on asterisk 21, I would suggest actually moving to asterisk 23 instead. or if you want a long-term supported release, asterisk 22 instead. We had 1,800 commits across all

branches and 1.7 million downloads. Uh, which in my opinion is thanks to whether it's good or bad AI specifically um the application of using asterisk for voice AI sorts of things. So from an asterisk forum community interaction perspective um the forum's accessible at community.aststerisk.org I'm on there others are on there so if you're having issues it's a great resource. Please also help out other people there if

you have useful insider Uh we had over 8,700 new posts and over 630 new contributors. So, Asteris 23 is available. This is a standard supported release. So, it will see a year of bug fix and a year of um security fixes. It was released on October 15, 2025. It had 360 commits and 67 individual contributors. Our top contributors by commits, uh George Joseph, who is a Sangome

employee, works on the Astros team, was number one. Naveen Albert uh was number two. Sean Bright was number three. Uh I was number four for once. I actually wrote code. Yay. Um and then Ben Ford was number five and so on. So thank you to all the contributors um who helped make Astro 23 what it is Now I always like to go over the difference between a

standard Asteris release and long-term supported. So standard releases get one year of bug fixes and one additional year of security fixes. That would be asterisk 23 currently. Long-term supported releases, you get four years of bug fixes and then one additional year of security fixes. Those being asterisk 20 and 22 currently. Um realistically these days you can if you wish just jump from version to version whether it's

LTS or standard release. We're fairly good on keeping things stable um and not breaking things if we can help it. So there are totally people who are doing that and seeing great success with it. Now the past year, what's new? Uh I'm going to talk about geoloccation and E911 which from a code perspective is not new but kind of from an industry perspective is becoming more relevant.

I'm going to talk about ARRI outbound connections, ARRI rest over websocket, media over websocket, performance, performance, performance. Um, note there if you did not see my prior talk on performance when it comes to asterisk, then stay tuned to see the recording of that online. Good information in there. Uh, we also had some new open source code called VoiceBridge. Uh, and as always, some miscellaneous fun. uh as

always, animation is brought to you by Chris May, who I'm sure is in the room now. Chris, if you're there, raise your hand. If you raised your hand, good. If you didn't, shrug. I can't do anything. I'm in the past. Uh let's talk location, location, location, location. So, geoloccation. Geoloccation for those who may not know is just the ability to provide location number or location information when

you place a phone call. This is not new. It's been around for quite a while. Uh but it is becoming more relevant within North America. It's known as E 911 or NextGen 911. And it allows communication of more specific location information for emergency calls. so they can um better better get emergency personnel to the specific location that they need to. The way this works is that in

the beginning address information existed on a perid basis. So if you wanted to be able to have a a location tied specifically to a phone, then that user or endpoint would need a specific unique DID. This address information would be statically configured upstream with the provider for the DID. Call to DID would have to use that configured or call to 911 would have to use that specific

DID. 911 would get that uh statically configured address information and they'd know where you are. Now, if you did not want to do that, uh you could assign a general one, but that general one is not in all cases specific enough to get emergency personnel to the correct location. So, this inherently has problems. Complete information needing configuration ahead of time and it's not easily changeable. It couldn't

include more specific location information like floor, room, suite, that kind of thing. And from a mobile perspective, mobile could only present a preconfigured home address or a course physical location. So, geoloccation aims to help with that. And like I said before, geoloccation is not new. It's actually made up of existing RFC's and standards that have just kind of been put together to allow the transport of geoloccation

information to an upstream carrier. Um, it can sometimes be confusing because specs may say one thing, another may say another. So, it's kind of a complicated thing. Fundamentally, it boils down to an XML content type and zip headers. What this allows is specifying location, and that's not just an address. It can also include um GPS coordinates, information in the call that you place to emergency services. However,

it does allow location information to be statically configured with the provider ahead of time, but not on a DID basis. Instead, you get a unique identifier that you can attach to the call. So this removes the need for a DID per user or endpoint. From an asterisk perspective, we've actually supported it for a few years now. And the reason I bring it up during this state of

the project is becoming is because it's becoming more and more Specifically, we've seen it's starting to see um at least testing from a Europe perspective. So it's gone beyond North America. Um so we have received a few reports of issues and resolve those. Uh so for an asterisk perspective configuration files exist to support configuring it alongside some PDIP options and we also have some dial functions. And

one of the fundamental things is that you can be as static or dynamic as you want. You could technically do it completely dynamically within dial plan for example. There are some things to be aware of. SIP packet sizes grow ever larger as the days go on, it seems. So, the addition of XML information within the SIP invite can actually cause the SIP invite to become quite large.

This can exceed UDP sizes. So, TCP or TLS may be required for signaling of the 911 call. as well. Since the SIP invite can contain location information, TLS fundamentally may be preferred just because that can be sensitive And finally, provider support can vary. You have to talk to them. You have to test with them um because they may require formatting in a specific way. Asterisk itself does

not validate it. So, you need to test and confirm with the provider. Yay, another Chris animation. Uh, ARRI gets some great improvements. A focus over the past, I would probably say year, six months, something like that, has been um on ARRI, making it better for deploying ARRI applications, making it easier to scale ARRI So, one of those ways is ARRI outbound The way it's worked for now

is that an ARRI application has had to establish a websocket connection with asterisk. On an application basis, you could only ever have a single inbound websocket connection. So, all traffic would flow over that connection. An approach some people took is they wrote load balancers that just connected the once and then load balanced and distributed out in an ARRI aware fashion. We've kind of changed that a bit.

Uh you can still do that. The inbound still works. However, we have also added outbound support. So the way it works is that Asterisk will establish a single outgoing connection for multiple channels or you can do it on a per channel connection basis as well. This has some cool things because it makes it easier to deploy and scale ARRI You can stick a websocket aware load balancer

in front of one or more ARRI and then load balance that way. You can also do it from a serverless approach and spin up your ARI application on a per call basis if needed. You can also have it so instances can be drained as needed during upgrades. And really this just ultimately allows easier scaling of ARRI applications. In combination with that to really make it easy we

added rest over websocket support. So up until now what has had to happen is that ARRI applications do HTTP requests outside of the websocket back to Asteris. So those travel two different paths. What we've done is add the ability to do those HTTP requests over the same websocket as the events. When coupled with the previous outbound connection support, this makes it even nicer for scaling because you

can just use that single connection. You don't need to know where asterisk is. You can just use that single connection and just go with it. So underneath what is it using? Um it's actually using kind of an inspiration from an abandoned draft protocol from Swagger called Swagger socket but fundamentally it's just a simple JSON based protocol with a route identifiers parameters and type. Um it's documented on

docs.astris.org. You can search for ARI rest to actually find it. Um but it's not difficult to implement. So how does it work? You send the JSON payload over the existing websocket internally. We handle the HTTP request just like normal and then the response is sent back as a JSON payload. Now there is a caveat. It does not support binary file transfer such as the downloading of recordings.

So if that's something you need to do, this is not something you can currently Um, if that is something you would like, please file a feature request on the Asterisk feature request repository on GitHub. Um, because I would be interested in knowing if there are people who would actually um, use that if it were available. Now, media over websocket, this is something that um, has caught a

lot of attention. Uh I've seen or I've gotten people telling me that this is why they switched from other open source solutions back to asterisk. Uh so let's talk about it. So media website until now asterisk has supported external media. However, it's only worked over RTP and audio socket. Uh there's nothing wrong with either. Audio is its own protocol. TCP based um RTP UDP based but it

does have headers. There is just complexity with both of So we've now added a channel driver called Chan websocket which will do external media over a web soocket. It works normally using dial or an ARRI. I'm seeing it when it comes to dial what I've seen a lot of people doing is just making a an application an external application which for example connects to open AI on

one side and then accepts a Chan websocket websocket on the other and then they can just dial open AI that way instead of using um any uh open AI SIP mechanism or anything plus that way you can put in additional additional AI providers if you want. The nice thing is that it does not use RTP. There's no headers aside from the normal websocket stuff which is generally

handled by whatever library you're using and it just flows over as websocket binary packets. Now, this does provide some nice benefits over the other options. A question I've seen a lot is why would I use this over audio socket for example. So one of the reasons is you can do connections inbound to asterisk or outbound from asterisk just like you can do with ARRI now. But the

really nice thing is that it supports buffering of audio in proper timing. So you don't need to provide a real-time stream to asterisk if you want to provide us audio. You can provide us somewhat large chunks and we will properly time it out. Um which makes it easier on the application side since you don't need to provide that real-time stream. So that's great for AI in integration.

It greatly simplifies it. Uh and something else is it also supports DTMF events over the websocket if that's something you're interested in So what can it be used for? Uh AI integration being the number one thing and what tons of people are doing. So you can do that birectionally or unidirectionally. uh we actually have a patch up for a unidirectional perspective that will reduce the traffic going

back and forth because you'll be able to say as an application I only want to receive audio I never want to send you audio or vice versa you want to say um I only ever want to send asterisk audio I never want uh so in comparison to audio socket and RTP um less code involved overall a lot You can also use it for recordings of calls to

an off instance application or service. Uh and then you can also do live call trans uh transcription. From the perspective of ARRI that would be using snoop channels and external media or you can do that without ARRI using Chan Spy and Chan websocket or app Chan Spy and Chan websocket. Wee. Now, everything comes with a cost. Performance. Uh, like I said at the beginning, I did a

performance talk when it comes to Astroscan performance. If you weren't able to see it, uh, check it out on YouTube when that goes up because I did talk um, more uh, specifically state of the project is going to go more lower level a bit. Um, but the actual performance talk was a bit higher level So the channels container we all use this um whether you realize it

or not because asterisk has to keep track of the active channels on the system. This is stored in a container called the channels container because it contains channels. A sangoma developer George had a hunch that this could be an area for improvement. So he did prototyping and testing of various replacements to see what would actually work best. And in the end, the one that worked best was

a C++ map. So this is really the first introduction of C++ code into the core, I would say. Um, kind of as a test to see how well it worked and if there's other avenues um that we could use to improve things. So in the future there may be more C++ code, but um we'll see as things evolve. So it's completely optional. You have to go and

enable it in menu select. And there's been an uh an option added to use a C++ map for the channels container. If you don't enable it, then it will continue to use the existing support called um AO2 AOBS 2. This can reduce CPU usage. Uh it's most noticeable on heavily loaded systems. So, if you're just probably an everyday user, uh it wouldn't have much of an impact,

but it could be something for you to try out. Uh something to note though, make sure you're on the latest releases. There have been issues that have been uncovered just because C++ map is uncharted territory and the intricacies and expectations of the channels container um were not there were just edge cases that we didn't uncover until later. So, make sure you're on the latest version. Uh, something

else, AMI event filtering. AMI itself can produce a lot of events which may not be needed by the AMI user. Now, we have had event filtering for quite a long time. However, yet again, we did some investigation and determined that the filtering itself could be performance hindrance just based on the way that it's actually written. So George also came up with advanced AMI Uh so filters are

more specific about what to filter which allows it to be more optimized in how the filtering actually happens and it covers the common filtering patterns such as event names and headers in a more efficient manner. Uh you can check out manager.com.sample for more details. Uh, something to note is that legacy event filtering still available, still works if you're using it. Upgrading won't break anything. Don't worry. Now,

time for my personal favorite thing because this is the um this is something I actually got to code and sadly coded it mostly on my vacation. Uh, so task processors, a bit of background. Task processors are a mechanism to have a queue of work. It's worked through in the background. And now if someone in the room is yelling out, "Ah, uh, you may say that task processes

are problems. I've seen warning messages. They overload." Yes, but task processors themselves are not the problem. test processors are actually extremely efficient to the degree that I've tried to optimize them. And by trying to optimize them, I made performance worse because they're already so simple and so optimized. What is fundamentally the issue when it comes to task processors is the users of task processors themselves. That can

be stasis, that can be PJ SIP. It's really how you use task processors that influences their performance. So I've created a new API based on uh some theories I had and experimentation called task pool. Task pool is a pool of task processors. Basically instead of a single queue of work, it acts as a pool of multiple cues. uh internally a selector chooses the best queue to put

the work in based on how loaded that queue is. It supports both asynchronous and synchronous tasks. It has the capability to do guaranteed ordering of specific related tasks and it can grow and shrink as needed. Uh or for best efficiency, you can really have a fixed pool of cues instead. So why did I create task pool? Uh we already kind of had a similar thing called thread

pool which as the name might in suggest uh is a pool of threads. So it's arbitrary what work those threads would do. Um however I had a hunch that thread pool was expensive to use and through prototyping experimentation profiling this proved correct. Thread poolool is great for medium to long life tasks. Um, but it's very heavy weight for quickfire things which most things are. So a long

lived thing might be a channel executing dial plan while most everything else quickfire things like handling an incoming SIP request or publishing a stasis message um are actually fairly quick on how they actually execute. And the reason that threadpool is so expensive is that it's not queuing a single task to execute. It's actually queuing two. So the first is the task that we actually want to execute

and the second is a management task for the thread pool itself. Another thing is that the task that we actually want to work goes into one queue at the beginning. Now once the thread is actually chosen that task comes out of that queue but the result is that single que can actually grow depending upon how fast the threads can actually wake up operate on uh cued tasks

and move on. Task pool on the other hand doesn't suffer from this. When you cue a task to a task pool, it's immediately allocated to an individual queue that has a thread that then works on that task. So the first major user of task pool and really the reason that drove me to do the investigation was stasis. So for those who don't know, stasis is our internal

message bus. Uh it's the foundation for a lot of things. So AMI, CEL, CDR, device state, hints, just tons of stuff uses stasis internally. And I saw forum posts and posts elsewhere about a stasis/pool task processor growing quite large. So I uh so I created Taskful. So, Stasis became the first major user of task pool. I switched it away from thread pool over to task pool and

this actually resulted in about 20 to 30% CPU savings. Uh that depends on the system itself and it's even dependent upon the underlying processor. So, um an Intel CPU might see less of an impact there while an AMD CPU might see a greater impact. The nice thing is that there's nothing to really notice aside from some new configuration options. I did make it so thread pull options

do continue to work. Um and as well there's some different output in core show task Instead of there being a stasispool set of um task processors, there's now a task poolstasis um set of task processors instead. So realistically, what this has resulted in, it's using less threads to do the same amount of work because things are just so much more efficient versus using thread pool. So um

if you haven't upgraded to latest versions, highly suggest doing so. You may see some automatic CPU savings. So PJ SIP was also a user of the thread pool. Uh and that's because of its dispatch module. So it's now also been moved over to task pool as of the latest releases. This was not as substantial as a CPU decrease as stasis. So realistically I've only seen about 5

to 10% CPU savings depending on the system. Um just like stasis uh just configuration changes the old options still supported you'll see different output and core show task processors and I also took the opportunity to switch the default options to better fit usage uh because of the needs of pjet previously with thread pool it was extremely I would say aggressive on creating threads and using threads Um,

under task pool, it's a lot less. Uh, during some testing of 400 calls per second with a one second call length, just throwing SIP traffic and PDIP um tasks at it. I primarily uses a single thread to do all of that work. There's a little bit of spillover on my system to a second and third thread, but it's really not that much at all. Um, that's just

goes back to how efficient task processors themselves are. So, some real world results. Uh, I actually had a user who was testing these changes before they were actually merged who said that on their system their actual CPU usage went from 120% to 50% for their workload after they incorporated them. So that was extremely pleasing to see. So, but as always, like I said before, very dependent on

usage patterns and also CPU as well as generation. Uh if you're using a really old CPU, it seems to really help there for some reason. Um so update to the latest version of things, see if you see a CPU decrease. Hopefully you do. So something else we open sourced is the Asteris voice bridge. This is under the GitHub uh Asteris repo. It's a go-based ARRI application. It's

under an AGPL3 license. It's basically a demonstration of some of the work that we did over the past for doing a voice bot using AI and ARI. So contains examples for all the pieces needed to build such a thing. Text to speech with endpointing and speech detect with pause and intera interruption support using deepgram. Dynamic AI prompting and function mapping. real-time AI interaction using It is modular

so things can be replaced and adjust as needed. It comes with an example config and docker file. Uh and it uses external media. It does not currently use Chan websocket. It was the inspiration and provided guidance on Chan websocket, but it has not been updated to it. Um and as always, it's not meant for production use. It's sample code. Um it's just something we thought we'd put

out there, see if people um find it useful or interesting. So take a look. Uh and all of the voice bridge and bot work was done by um Mike Berdine, who should be in the audience. Raise your hand, Mike. If you raised your hand, I don't know. Uh and now as always, on to some So Chandi now has call waiting deluxe and last number redial support on

analog for those who are still using analog on Chandi. Uh the next one is something George Joseph did and actually I think we're one of the few who now support this. Uh so we have Shaw 256, Shaw 512 now supported for authentication digest algorithms in PJ SIP. Um I believe we did also have to upstream some support there to PJ SIP for supporting such a thing. So

um that's an that's an area where we upstream some more stuff. Uh we now have a new option for PJ SIP called suppress mo on send only and this is really useful for providers. Uh some providers will pass through music on hold requests all through their network to um to their users. Uh I've had this happen on a provider uh up here in Canada actually. And essentially

it means that if uh if you are calling your provider and the remote side puts you on hold, you will hear your own your own hold music and that can be extremely confusing. So we now have an option to basically suppress that. We won't start local music on hold. Um you'll just get silence instead. That's what some people want. Uh PJIP show contact now shows more information

than before. Uh, stir shaken identity headers can now be suppressed on a callby by call basis and they can now be set for unknown telephone numbers as well. Uh, Reso DBC now has support for storing its connection pool as a queue instead of a Uh a big one for people who are using ARRI, you can now control, you can now delegate essentially handling of incoming SIP transfer

requests to the ARRI application instead of having asterisk handle it itself. And we also have a fun fix for PJ SIP to prevent overwriting of caller ID during a DTMF intended transfer. This was a fun race condition that Mike Berdine yet again um fixed up. But essentially, sometimes caller ID would get overwritten, sometimes it wouldn't. It was very interesting and odd. Astros CLI shell access can now

be disabled. Um did everyone actually know that you can prefix a something with the exclamation point in the Astro CLI and it'll run it as a shell command? Um, probably not everyone. Now you do. Uh, that can also be disabled. CEL can now provide events for DTMF and playback starting and stopping. CDR now supports the cancel disposition. There's now an ARRI operation to uh indicate progress to

a channel. So that'll send out a 183 session progress on SIP. Uh, menu select now has options to enable AES1 192, 256, and GCM for SRTP support. Wait tone now supports a custom tone. Hangup cause can now access reason headers. Uh, and then this one is something that George did. ARI recordings and playbacks on bridges now support a sample rate higher than 8 kHz. So this got

overlooked during the initial implementation but essentially those were all locked to 8 kHz and uh it is now more intelligent. It will examine the channel um or rather channels on the bridge to determine the best um sample rate to uh dial timeouts now support fractional seconds and there's now a digit sum dial plan function. So some upcoming stuff in asterisk more performance improvements as well as some

quality of life ARI improvements. So performance we're continuing to work to improve performance. Uh George worked on some usage scenarios and profiling to determine some hot areas. And so based on those, we've been working on things here and there. Uh our current focus is CEL and CDR as well as hints by the time this is actually shown because I'm in the past. The past um CE and

CDR is hopefully up for review and may have been merged for next releases. Hints will take a while. I'm working on that. Uh prototyping shows good stuff, but hints themselves are very old and organically coded and understanding all of the quirks and intricacies of them um will take time. So uh look forward to that. ARRI early bridging. So this is the create and dial operations uh in

ARRI which allow you to essentially create a channel stick it in a bridge with someone and then dial it. So you can have uh early media passing back and forth. Uh the implementation doesn't really handle offnominal scenarios that well. It can cause weird events or behavior. Um so we're looking at changing this behavior or rather changing the implementation so it works better for users. Another one is

something I've heard for quite a while. Uh doing multiple uh to get channel variables is annoying. So you do one request for one variable, one request for another over and over over. Uh we'd like to add the ability to be able to do that as one request and one response so that you don't have to um don't have to do the multiple requests to get the actual

information you want. just to simplify things down. So, we'll be looking at adding a route Another one I've heard is attachable state. Uh people implement this right now on channels using channel variables, which is fine. Um but we'd like to add the ability to do this as a first class operation. Um so that would be adding attachable state for different objects. So, bridges, channels, etc. Um, allowing

you to attach, delete, request they be in events or just go and get it as needed. Uh, something new I'm doing for this is, uh, showing off or touching on a community project. So, the first one I'm going to do is T140 real-time text. So, T140 real-time text is the ability to send text over RTP in real time as you Now, this isn't just a message that

you send and then forget. It also allows you to, for example, backspace and delete on the receiver side. So, it's much more of a real-time interaction. This is becoming more prevalent in certain countries due to regulation, specifically over in Europe. Um, so there is a community poll request up on GitHub for it. um it could use input, further testing, help seeing it get through. So, if this

is something you're interested in, uh please check that out Um as always, some good stuff to know. Uh if you ever want to know when an API or dial plan application was added to asterisk, check the documentation on docs.aststerisk.org. uh it should show when individual things are introduced. So if we ship something uh in the latest releases for example then that version would show up there. This

also does work in the CLI for dial plan applications, dial plan functions, that kind of thing too. Some general reminders, asterisk 21 is As always, keep track of what's happening in newer nonLTS major releases of Asterisk. If you don't, you can potentially experience big surprises when you move forward. We do try to minimize the surprises, but um we're not perfect. We're human, so they can still occur.

Keep apprised of the Astros versions wiki page on docs.asters.org as well as the module deprecation wiki page. Uh we don't currently have anything slated to deprecate. So you should be good on that front. And now we're going to try to do a live Q&A if we can. If we can't then if you have any questions feel free to post them on community.asters.org for example. Um and I'll

happily answer there so others can see them. Um and if the live Q&A doesn't work then I'll say thanks everyone for coming. Um I hope you enjoyed this state of the project. I hope you enjoyed Astrocon and uh I look forward to seeing you next time. >> Thanks, Josh. Virtual Josh. >> It's not virtual Josh, it's past Josh. >> There's virtual Josh. All right. Excellent. Um thanks

for staying awake for us here. Uh I guess, um we've got a couple of questions, I'm sure, for Josh to answer remotely. Can you hear us, Josh? >> Yeah, I can hear you. Uh, I also wanted to, uh, add something else about documentation, um, which just happened over like the last week or two or whatever. Uh, it's currently in pull request status. Um, but much like we've

added version information to plan, applications, functions, and etc., we're also adding the module that they belong to. um which should help those who want to run a um a trimmed down asterisk so they can more selectively um disable modules for example. Anyway, thought I'd throw that out there. >> Some quality documentation as always at docs.astros.org and it's all you pull out and make revisions if you like

yourself to it, make improvements. Anyone have any particular questions? Here we go. Yes, thank you for your presentation. I have a question regarding RFC 5031 uh regarding uniform resource names like source service and so on. Do you know when or if ever asterisk will support them? Thank you. >> Uh there's no current I I know of no one currently working on such functionality. Um we do have

a feature request repository um that if you'd like to see support then that would be a place to create an issue in so it can be um organized and um interest gauged and that's just um asterisk or asterisk-feature- requests in uh the asterisk organization >> It's per personally one of my favorite repos to just drop things in. So um anyone else have a question for Josh about

asterisk? >> I mean ideally asteris you can ask other stuff doesn't need. >> Yeah sure. >> Hi Josh. Thanks for uh doing that virtually. Uh I'm going to embarrass myself and say that we have been using Chan SIP up until just this past year. And one of the big complications of moving to PJ SIP, I would say a good dozen of our customers were using ring groupoups

that had more than say eight or nine extensions in the ring group and those won't work over PJ SIP. We've discovered that we can only get about five uh extensions in a ring group with PJ SIP. And it amazed me that this hasn't been resolved and and I wondered if you knew about this. uh no current issue and I am aware of people doing larger ring groups

than that. Um so I'd suggest file an issue. I think in my testing I've done upwards of 10 to 20 um just for my performance >> When when we dug into it I I noticed that what it was doing is truncating the dial string. Uh the pjip dial string that gets created by a large number of extensions is much larger than in Chansip. And I think it's

the the syntax. >> Uh if if you are using PJIP dial contacts, then yes, >> I'm not sure I understand what that means. >> Um so PJIP dial contacts, uh when you dial a PJ SIP endpoint, um if you just do like in Chan SIP, um SIP slashalice for example, um you can do the same thing in um in PJ SIP PJsalice. Um the other option is

to use uh PJIP dial contacts which is a dial plan function that will return a dial string but it embeds the contact URI which can be quite long. So if you do that then in comparison to chance SIP the dial string can be quite a bit longer. Um, if your endpoint will only ever have one thing registered to it, then you don't actually need to use the

dial plan function and you can just dial PJIP/ALIS instead. That may be why the dial string was so long. A good question. There definitely some changes in Chan Sip, the Champ SIP, that's for sure. Um, anyone else uh have maybe something tangentially related to >> as past Josh said, if something does pop into your head, I am readily available on community.aststerisk.org. I probably spend too much time

there to be quite honest. >> Oh, great place to bump into you. Uh so yeah uh Josh uh also on the free PBX forums occasionally when someone asks a Astra's question that bubbles up. It's very helpful. So thanks for that. I think without much more um Josh you would you like to close out Astrocon 2026 virtually for us here. >> Uh I mean you're really putting me

on the spot and you're not my boss. You'd have to ask my boss over there. It >> it it was a friendly uh you know request. So yeah, but um up to you. >> Uh sure, why not? Uh thanks everyone for Um I hope you had an enjoyable time. You had lots of good information downloaded into your brains and had some good conversations and contact. Um, a

reminder, if you did miss any of these presentations, they will be available online probably fairly quickly. Um, we'll see how long they take. Um, as well, uh, after this, unless it's happened already, you should be doing a Yes, Chris, >> you for the reminder. Yes, a picture, everyone. Uh, we have a couple t-shirts left. If you haven't yet grabbed one and would like one for the picture,

I think we're going to mosy on over here to the uh sparkly booth that Mike's at and try to crowd around there and smile pretty. I'm going to turn up the lights and um we'll uh we'll have to have uh Josh in there later on. Notice I didn't say Photoshop. We're at the open source event. So, thanks so much. >> Give consent for that. we'll we'll clear

that with legal later. Yeah. All right. Appreciate it. Thanks y'all. Uh, Astrocon 2026.

From event

SCaLE

05 Mar 2026 – 08 Mar 2026

All event videos
Back to Watch