SCaLE

Room 103 Friday Mar. 06 - SCaLE 23x

7:59:26 · 05 Mar 2026 – 08 Mar 2026 · YouTube

About this talk

This talk focuses on the importance of open data in promoting transparency within local governments, specifically through an open-source model used by LA County's analytic teams. The speakers discuss various initiatives tied to the justice system, highlighting the challenges of data standardization, particularly around criminal charges. They elaborate on a standardized approach to categorizing charges, including using authoritative sources for clarity and accuracy. The session emphasizes the role of collaboration with community stakeholders and justice partners to refine data methodologies and enhance the usability of public data. By leveraging open-source technologies, the analytics team aims to create meaningful insights and foster public trust in government operations.

Full transcript

because they are upset by corruption, they are upset by violence or something. Okay, sometime it just start by you learn that your government is cheating on you and taking some money outside. Okay, I'm not doing politics but I know it happens in some country. Um I don't think in your country it happens in my country. No, no, no, no, no. Believe me, I don't want to start

a never ending discussion. Believe me, just go to Africa, just just go somewhere. It's terrible. Okay, of course, for sure. Okay, it can happen at some basement. But you know, there is a beautiful joke about this. Maybe here or in other democratic country, you take 10 or 15% back. In other countries, they take 100% of the budget and the project never happens. Okay? So it's good at

least to make it public then the citizen can decide to vote or not. We have our voice. We still believe and this is why also we do open source. Okay. We still believe about sharing about having a community of people who just want to do something together for the good of the community. Whatever happened outside, we know whatever happened outside, but it's not a reason to stop.

So this one I will go fast because I'm little Oops. Open data should make your life uh better. So some example for example, it's your data you are paying your tax for this. Okay. So in in in return of your tax, you need to have access to the data. That's the minimum you can claim to your government or city. Sorry, I will confuse the word government, region,

city. It's a different level of government, but it's all the same. Okay. Uh some nice data portal. There is this one. And you will be surprised there is not that many. Another one. So based on uh the I don't know who is feeding this one. This one is uh all the data are coming from the people who decide to load. So and many of them are dead

now. When you click on this one, you click on the website and you say 404 no more website. Okay. Because this one they load it for years and sometime after three or four years the website will stop for some politics reason. Welcome back Dorian. this one about French just because two weeks ago before this uh presentation there was uh a report from uh OICD and the French

was number one at least we are number one in something so I need to promote what in term of uh opening the data to the citizen okay in my country we have this culture believe me okay so we have a lot and lot of website and open data website. Okay. >> I'm sorry you just >> uh I this one like this I don't you can you can

access it later. >> It's there. We are number one after Korea, Poland. Okay. Okay. I maybe maybe that's that's mean that means this uh my my law in my country if your city is more than 4,500 people and you have more than 50 employee in your city you need to open your data only one data set is enough only one you need to open at least This

one data set can be and you will see it. The data set is a map where you can have your dog bring your dogs. So at voo really we see it. It's a map or the cemetery or the school. The law in my country we go step by step after 15 years is you need to publish one data set minimum. If you decide to put 100, it's

your own decision. Okay. So maybe this means that 95% of the city in my country who should open their data did it. Okay. But it's not all the city. In my country, we have 36,000 cities, but more than 34 are less than 1,000 people. Okay? I don't know how it's divided in your but in my city we have like 50 or 100 large city and many many

with uh 200 people. Okay. Next one. Sorry. So some standard I I talk about the word catalog. Okay. Catalog means the metadata and the data. So it's like a catalog of uh the food you can buy. It's like a menu. Okay. But this menu is very sensitive. You have three different kind of catalog standard. This one is more like uh data grid. Okay. And this one is

a European directive with JO special uh metadata. Okay. So that's how it works. And this one is just generic one. So when you publish your metadata, you need if you follow one of this one uh standard it means that another website can automatically connect to your website get your data catalog and make some comparison. That's the power of the standard. Okay, you can have a website that

does not follow any standard. You just put an Excel spreadsheet or a PDF, it's enough. Okay. But it means you don't want people to reuse your data. Okay. Another word O data. Okay. O data is not open data. Okay. O data is a protocol to uh issue a restful API. Okay. So standard uh platform like Microsoft PowerBI can connect to O data database. Okay. My platform is

using in back end posgress. So you can use it you can access using GDBC whatever and others there's no problem. Another one is API. Is your platform providing API or not? is the old way of MCP in a way because before if your platform whatever it's open data or another kind of platform uh like an accounting platform HR platform if you have API a program can connect

collect the data and interact with you using uh uh Java or any languages now with MCP you use natural languages it's it's more easier. So that that's some of the standard that good open data platform should follow and the one we meet in tender or or those things they all follow this one. There's no problem. In US you have something called fair. So maybe it's in comparison

with the standard metadata and catalog. Okay. In US there is a way uh a project make your data [snorts] findable, accessible, interoperable and reusable. Remember the eight principle of open data in us. This so I I just put in the slide some uh website reference if you want to look further about this and is your open data platform fair or not. Okay, reusable for example means or

this one means you need to access it with a an API a standard API take it and reuse it. Okay. And there was something I saw recently in New York City. They they have an open data platform for the city and they want to do one for the In in my country sometime you have uh collaboration data meanings for example you open a map and people can

go on the map and indicate for example in this street there is a problem with a traffic light in district there is a problem with of security or blah blah blah. Okay so but problems is that if you have this you need somebody to monitor the data. Okay. So sometime they want to separate the community data and the official city data that's so okay and because we

are here in Los Angeles and very happy for me very happy to be here I hope for some of you also uh there is a Los Angeles uh website open data which is a very nice open data website okay I mean you can bruise it there is a lot of category of data there is plent plenty and plenty of data and very I put also the one

of San Francisco uh there is something interesting I think there is less data on it but sometime too much data kills the data okay that's it in France just by opposition we have an open data association work on what should be a data and they put nine n data set that are recommended to be published. The nine are birth of the children, the death. Um, and for

example, birth of dead, they say mandatory you can put the name. Okay, and they they built a structure of each data set. It's month and month of work. Okay, we call it a taxonomy. I don't know the maybe it's the same. Okay, and we have also in France schema.datag.fr FR which list which has a long list of to describe everything for example to describe a street to

describe a traffic light with all the metadata on it. Okay. And there is nine metadata recommended for open data. One is the bid the tender one is the subvention the grant I learned two word today. Uh and hovers that you can read when you connect on this website. Okay. And something very nice we will see it we will see it my government has open it's open data

um website with an MCP server which means if you have cloud if you have wind surf if you have cursor or even libro chat I will demo you with libro chat on it Okay, you can connect and talk with the French uh open data website. This French open data website has something uh something different than other website is just collecting the website from everywhere. All the city

that started an open data can push back or allow to be collected their data to this central database. The database is not perfect. Okay, it happens but the datab exist. It collect a lot of data and we have also for the cities who don't have open data website because they have no money for this. they can still connect with their uh official uh name and password. If

you are a member of a region, you can connect to this website and publish your your data set. Even your city has no um open data portal. You can maybe you will find some data about your city in this portal. Okay. So it's just like a central national uh open data portal with thousand of data set. So too many information kill information. So it's very difficult to

to understand. And now it's good because you just configure the MCP server. Okay. You have your own API uh K and you run it. It's very easy. And you >> Yes. It will it will it will allow this because I know you like when I'm zooming. So sorry this one is in French but I will do it also in English. Uh you just write this now before

you have to write a SQL sentences. Okay. So you have to know SQL. My mother don't know SQL. Okay believe me. Okay. So It means now MCP allow more people to access to some information because they can talk with the machine. Okay, that's something which is very nice with MCP protocol and our government now support MCP protocol. I will demo you. Okay, so I was I'm very

late but it's okay. Second part I will go fast. It's a lots of very nice slide. Okay. Uh some example of open data uh platform different level region city uh a uh [snorts] some arguation of city when we group together. Okay. So this one for example is one for Lar Roelle. Lar Rochelle is very well known. It's uh on the west coast. Okay. And you have when

you reach the data set they put in front you can download your catalog. Okay you can search you have even a uh map search integration and you see the website is uh is very nice. There should be the reference I forgot. I will add the link to the website. I think it's after. Okay. So different interfaces for um a you have the metadata. Okay, so it's all

data about the data set and there is standard about this. For example, if it's map data, geo data, you need to indicate longitude, latitude, altitude. If it's data about a lake, uh the surface of the lake, blah blah blah, many things. Okay? Then you have a data grid if it's available. Okay? So these people will understand. You can have also the map visualization. You can have some

graph and all this is inside the platform. You don't need ex additional tools like PowerBI. If you have you can access no problem. It's open. But inside the platform you have all these things. Okay. Uh this one it's another kind of data visualization. This one is the export format. For example, if you go on this one, maybe it's nice because you have the uh flat file format

uh geographical format. So in this you have gojson shape file KML flatbuff geo package. So if you have a QGis you just download the geo package and all the data is available for you in an instant. I mean in two minutes. So it's very easy. And the last one is API. Sorry this one is API interfaces. So it will generate you an AP HTTP request that you

can use in your own programs, okay, to do data extraction every day. For example, I will give you a real example of this which is nice. So some visualization with here you can uh you have two or three different data set on this one. Okay. And if you go to this website, you can you can build your own maps yourself. It's free. I I did it myself

and I can extract the data here. I put some different data set. One is blue, the other one in red. Okay. Because I want maybe to compare some information. And uh I can decide to do a show uh color graduation many things on it. It's it's very easy. You have KPI also. So you have a KPI studio and you can see the KPI uh information. So when

you connect to the website you can directly the KPI for your city or something the things that uh are pushed in front of you. This one is dashboard. Okay I will just show you this one because it's lots of side but I I want to show you one example on uh this one. I think this is this one. You see it? No you don't see it. Sorry.

Uh how to stop like this? It's stop. Okay, that's a website. Okay, so on this one is just some data. You can put this one like this. You can change your data like I don't want a grid, but I want uh for this one I want like this. You know, you can do whatever you want. You can have the PowerPoint mode so you can see the one

after the Okay, it's very easy. You can have also uh something which is nice on this one. You can have an instant snapshot because at the end people don't want a dashboard. They want the print screen of the filter and get the information back on their uh computer. Okay, that that's a big difference between report and dashboard. Everybody wants dashboard but everybody is saving report. Okay. And

something which is nice also we have this if if you don't understand for example this data you just collect this there is an analysis which will build real time a statistical description of your data. Okay. So I'm French I'm connecting with a French uh uh interfaces. So it will report in France in French. Sorry. If I was connected with an English profile it will write this in

English. So it gives you a graphic analysis data summary inside directly on this one and on this one I have also over some information the data set used if I want to go back to the information. So it gives you confidence okay because every time you can go back to your information and to to make a dashboard like this make it publish this one need uh 10

minutes to do this ju just for somebody non-technical okay very easy to do this this URL it's public for example if you want to type this one you access it Now ah the name of this one it's just a website example okay it's uh but uh if if you want this one or another one I can share with you is no problem this one is an example

of a dashboard I I can I will show you after another example for over city open source open data platform oops I come back on This one I'm sorry. So this one. So that that's how you can use with dashboard. You have also a map studio. So you don't have to take another tools to build your maps. Okay. So I give you a map studio with some

information. Uh we have an arrest module. arrest module is very nice to connect to another open data that follows the same standard data catalog and so and collect some of those data because sometime you want to compare your data with another city. Okay. So there is a harvest modules, there is a scan modules. Sometime when we start a new project, we arrive in a in a city

and they have a lots of u uh cuis project file. Okay, they have a lots of uh Excel data sheet. So we have a scan module with AI that will read all the Excel spreadsheet and extract the data and publish it at data set. that that's uh something very useful. We have also data chat modules. So the this modules is there. I I can this one I

can demo it also. It's much better. If I go there, okay, here for example, I'm connected to a website. I can set to this one. Okay. So I don't know if you will saw it, but I will write create a map on something. to create a map on data set So it's just thinking and and I have it in a smartphone. So if I talking to my

smartphone it's much better because I don't have to write. So he's just finding it and he will okay now build a map. It's finished. Okay. If I want for example uh another information like a graph, it's uh 5G antenna. Okay. So I want a graph of 5G antenna by provider and by status. Some antennas are broken We have to create a graph of number. And what is

nice is that if I misspell some word, the LLM will correct. Okay. And still create the graph. Here I've got some suggestion. It's like CH GPT or others. It's just reading the the data set and telling you with this data set you can do this, this or this. Okay. So now he's telling you that the the colon is not called statues with a S but statues with

a T. So he's correcting me. Okay. Bad luck for me. And uh I think in English it's with a S. He said I made a mistake. So I'm just talking to the machine and he's trying to find the good information and he said I apologize for the other side. Let me correct the code just behind is [snorts] there is a percent miss. Sorry it's a demo. So

bad luck. I apologize. Okay. Repeated error. Bad luck really. So I hope you will succeed to do something. Okay. He's not happy with me today. Yes, finally. Just for my my heart. Thank you. So that's it. Okay. And this one I can save it if I want, share it, do whatever I want. Okay. Create a dashboard on it. It's just so easy. If I want a 3D

graph with interactivity, I can have it. It's a Python library. it's plotly behind or another library. It's very easy to to do. Okay. So that that's uh how you can at least as long as you have access to the data you can do it. So that's something and that's uh for example the my my visualization library which is here. I can view all my library of data

here. And now it's just after it's an internet connection. Okay. So it's all my visualization. It's like my in chat GPT all the image I created and I can create dashboard about on it and it's very easy. Okay. Oops. That's my last session. So I can come back and reopen a session, share the session with somebody. Uh as long as you have access to the data, you

can do it. It can be on open data but it can be also on SQL data. So even only this uh modules you can deploy it on your SQL data open data website or not at least we don't care at one time we just want to make the data available to anybody with a limited cost license okay we are not fighting against Microsoft BI it's a nice

product but sometime you pay too much for something you you don't use only 20%. Okay. So just be realistic and sometime it's better to use an opensource dashboard system. Oops. I come back on this one. So this one is this modules QG integration. We have this also because we saw a lot of QGIS project. I don't know in US but in France we have a lot of

QGIS project. Okay. And so people for example when we have this kudis project automatically we transform the kudis project into maps it's done in one instant I mean 10 seconds believe me you take your kis project you use extension you publish on some website example so if you reach after you you can go this one this one is nice because it has a lots of KPIs it's

has lots of temps of study of data okay with visualization and other I will go fast I'm sorry uh this one is a metropolitan so they have a few data data set only 200 it's only 200 okay but they have more than 30 city and they publish for example this one is all the trees they just record all the trees because it's a green city and where

you have the trees this one is an election uh result maps. This one is also one of the nine data set recommended. You need to publish your election results in the data set This one is another one. It's a region. So regions it's the equivalent of state for you. Okay. This one is one of our oldest one. It has now eight or nine years. It's uh French

uh antenna. So every Thursday, every Thursday at 7 p.m. we publish the update of the data set. BIV during the night our four major provider of uh telco I don't know in US the name but in France we have four. Okay. they just take out the data because they want to compare with with the other where the other competitor deploys their antenna. Okay. Is this a 4G

antenna? 5G antenna. Now it's all 5G. But in 5G you have different kind of 5G antenna. So they can because it's it's open, it's low, they need to declare what they install and so your competitor can have access to this. Okay. Another one, it's a Caribbean one. Well, friend, it's a data about land development. So there's lots of maps in this one. Uh this one is data

chat. I saw it before. And in data chat you can have this kind of uh very nice uh visualization. Uh sorry I mean uh the four the third part MCP interfaces. Okay. So MCP interfaces I will show you not in data in data chat but in other one I will show you an an easy way to integrate data into any kind of external application. For example, we

have easy rag. It's one of our development. It's a chatbot or virtual assistant. Okay. So, if I go in this one, I I will better show you the result. It's much better than a slide. Okay. If I go there in easy, I just declare in my MCP server or maybe some of you know what is a webbook tools. I just promise at the end of this presentation

it's my 10 last minutes but it will be little more technical. So we have MCP server. I just declare my MCP server either the one data go which is there. Okay. I can test it showed me it's okay or my data for citizen one. Is this a SQL database or a standard database? I just go there. Okay. And I can test it also somewhere And it will

be it will tell me it's okay. Now that I have it, I will I will do it with you. I can create a chatbot. Create a new chatbot. I will call it uh uh Scalix scale 23. Okay, that's it. Uh my system prompt is sorry. Uh so I am uh an open data manager. And because I'm not good on this, I will create this. Sorry, I just

take this and I have a prompt assistant and it will recreate the system prompt. Okay, that's much better. Now that I have this, I will just choose my MCP server like I will take the I will choose uh create data set. It's all the the data. So I will take everything because I'm lost. Get data set should be okay. It's solve the API that are embedded into

the MCP protocol for those who knows. Okay. And now it's done. Okay. It's varied. Scalix 23. I need to I will do something because you like to have open uh to URL. So I will call it this one slash scale scale 23. I will not put the X. Sorry. Okay, now it's done. I can either view the embedded code and put it in a website to have

a chatbot. Okay. Or I that's it. And here I will type something in English or in French. Okay. what? Sorry, I need to take Okay, perfect. This one. So, it's there. Uh, what are my data set? Because I don't know I connect to a website. So I just need to ask a question and it will tell you you have data set about this this this and after

if you are happy with the result of the data set you can simply here now it's just telling uh some information so I need to uh enter in a dialogue with him because maybe my uh question is too large so he asked me additional information and I need to communicate with him like a onetoone discussion between two people okay to make precise question to get good answer

okay so that that's it and when you have this one you can build automatically a graph or grids I will take another one which is uh much easier I think this one and I will take uh I will put my So I just ask him in this one to list all the data set of one terms. He give me all the data set because this one I

know it better. You can view the data set if you want. So directly access to the data set, view the maps on the data. It works like a charm. Okay. It just needs some some second to display the data. Okay. And you can come back on this one and set for example for this Okay. For this one create a graph. And it works. Just give need some

uh some thinking. And again we have it like a smartphone application on Android. It's a flutter app. So it works on smartphone Android and uh iPhone. And um he asked me some more information precise information because he said it's too large. So I need to to do it but at the end you can get something like this Another integration for with Slack. Okay. If you have Slack

or Teams, you just uh build a channel and immediately you can ask information and you get it. I'm not talking about security access on the data. We have it. But we're talking about the open part. I'm somebody like a citizen. I've got a Libra chat. Maybe you know Libro chat or Slack or something. I have the channel available. I can connect to the data. I get the

information. Okay. I will do it with Libra chat also because it's a nice implementation. So for Libro chat is there. Okay. I will do this like this. I need to choose my model. This one is an is a good one I think. And for this one I will uh connect here not it's just uh MCP here I will connect to data so the French one and not

data for citizen not mine I will connect to this one so he's trying to connect to data he said me it's okay now on this one data go is okay and I will say for example I'm living in Leon So I will say give me from the city and everybody has this kind of uh chatbot or virtual assistant now it's deploying very fast in enterprise so everybody

can use it just to get some information and some data. So now he's thinking and uh he connect to the MCP server of my uh French national website and uh we'll return some data set you saw. I'm sorry I could not find could you please I don't know it should find um okay it not find it. So maybe I will write in French because it's a beta

implementation it if it was this it's really uh h yes it does not understand English. My god. And it's recorded. I'm sorry for the French man. So here it's just understand French. I don't know why. Okay. Should understand any kind of languages. And here it gives you all the data set for the city of Leon recorded on a French national website and you can go to one.

It's a catalog of all the data or this one the the grant the grants. You click on this one and you get access to all the grants of the city of Yong. Just a question of internet speed. Sorry. Uh I access a website and here. Okay. And that's uh that's a PDF. Okay. So I need to I can I can access the data. I can get the

PDF and [snorts] then work uh on the data by myself. Okay. You have also in this you have in re reusage usage and API. It give you who has take the data when you can take the data you can build a graph or a map and declare it on this website for this data set. So it means as a citizen you have done the job you have

taken the PDF you have built a visualization and now you put it here so that everybody will see it. Okay. So that's how we do it. It's not perfect. Okay. Like here it does not understand English, but you you you get a list of all data. You can collaborate with it. You don't need to learn and know SQL anymore. You can get access to the data. It's

MCP. If you are a developer, you can do some vibe coding and with the MCP server, it will build you an interface very easily. Okay. It's really, really easy to do. I think I'm done because I'm just on time, but you give me five minutes more. Okay, that that's it. So that's that's one. For example, this one is very nice because I ask him to create a

map. So he create a map. There is longitude and latitude, but there is not open source, open street map behind. Okay, but that's really a map. If you go there, it's a sea. Okay, but at least it does not understand. You just create this and and it works. I was so amazed. Okay, same in Windsor. Me I'm using windsurf. Okay, I go in my windsurf. I just

connect to the to the data. Okay, I declare the MCP in it. I just open it and I can insert ask a question and get the answer. Okay, so all this again just to finish it's open data. is based on standard O data standard catalog standard subject is to create KPI or standard data set so you can compare between region between city and take some decision for

example if you are looking for a new job or for relocation in another city data should be available let's be real it's not always available okay but we have to start it's only 10 to 15 years. We are all getting older. We are all getting now with the way and the the facilities to get information on this and we want it. So maybe it's not for today,

maybe it's for two years, in five years, maybe in 10 years, we don't know, but it will happen. Okay? In my country, we had a lots of uh green party who win some big cities last election, next they will lose because uh they don't know how to manage a city. It's another subject. But uh when they win the election five or six years ago, they really push

the open data movement. really really really push it because it's part of their roots and what they want to do to embed the city then into the city management when people from the far right party and far left party whatever they are talking they they don't want people to have access to the data okay that's it I'm just on time uh do you have some question about

this does it bring use some information. Yes. Question. >> You've set up some of these. So if let's say a town around here would want to set one up, how much work like how many man weeks does it take to make one of to set this up? Not the cost of the software, just the work of setting it up. >> It can can last one week, it

could last one month. It depends if you have the data or not. I'm not talking about data quality. When you publish your data just to set up, it's one day. One day it's a docker image. Okay? You download the image. You just uh put two or three parameters your API key to open AI or in France we have Mistral. I don't know if you know, okay, but

we need to be compliant with Mistral otherwise we don't work. Okay. But we are compliant also in OpenAI. You just download the image. You you put a docker compose. It's done. >> No. No. They always in my city they go through tender. So they lose six months to write the tender, issue the tender. >> Depends if we do it or they do it. One week. Okay. >>

Yes. Yes, it's recording. I mean, if we >> Okay. [clears throat] >> When you set it up for a city, you spent some amount of time setting it up. So, did you say that it took you about a week? Okay, that was the question. Yeah, >> I will tell you why. If they have the data available and they know what they want to publish, the only thing

we have to do is to create the catalog, scan the data and where we spend the most of the time is what are your terms, what are your keywords, what is your web chart, which logo you want, uh which layout you want. If you came back on some presentation sometime we have to to spend I would not say waste but we have to spend time with people

just to accommodate with the logo the color all this can be done in two or three days but loading the data itself it's one day okay really it's one day if you have the data you load it it's done and then after you can practice learn with it all is open so for Sure, I did it like my president. Uh you need to lock some data set

because you don't want those data set to be published for sensitive data or something. But it's not a one month or two month project. Okay. Then after we go in cities and we train the people if they want to do it themselves. Okay. Because it's a tender you have some rules to follow. Okay. If we go to a private to the private private company, it's more easy.

Okay? Because you discuss direct with the people and uh some some company they use it just to share internal data between uh department just not to send data using Excel spreadsheet or blah blah blah. They said if one department is in need of one data set maybe this need will become regular so let's put it in a process it's no cost okay it's just an extraction for

example from a SQL database transform it into a data set refresh it every month it's done it's just that you need to know where to click in the interfaces I mean uh that's it other question. Thank you very much. Thank you a lot. Thank you very much. I hope you understand my French English. Okay. With some real time translation on some word. Don't hesitate. There is my

name somewhere. Uh send me a mail. Some people can tell you I answer to the mail. Okay. Uh I'm here. I can share with you about all those things. Uh even some that you know for example the easy rag which is uh chatbot but connected to data. Usually chatbot are only connected to document. So our is connected to data which allow you to deploy a chatbot about

data which is which can be good in some enterprise when you want to have access direct to a data on something and don't want to take a call to somebody. Okay. And I'm here I stay in this room until noon. Okay. Because uh they invite me to stay. So I'm happy. Thank you very much again. Thank you for coming. Thank you for staying also until the end.

and uh maybe see you later. Thank you. >> Thank you very much. test. Hi. Hey everyone. I guess we're just going to go ahead and get started. It's about 12:30. Thanks for coming everyone. Welcome. All right. Today we are going to be talking about the transparency stack which is LA County's open source model to public facing analytics and building public trust. So my name is George Miranda.

I'm a senior data scientist with the county of LA working within the CEO's office within that the chief information office within that the chief data office. So um kind of like nesting dolls in there. I've been working in the public sector within the data science and analytics space for about 15 years now or a little over that. Um I'm gonna go ahead and introduce my partner in

crime here to Fay. >> Hi. Uh my name is Fay Woo. Um I come from the same office uh like George. We actually have a um even smaller unit called analytic center of excellence. I'm the uh lead research scientist there. Um I've been working in government um closer to 10 years. I came from a um academic background. All right. Thank you. Okay. So just at a high

level, this is what we're going to be discussing today. Uh we really just want to provide context for what we do. Uh so what what we do why we do it and why transparency is so important to us you know I think um it's very easy to understand like you know the government does a lot of things and as the public we want to be able to

know what it's doing and how it's doing it especially when it comes to data. So, we'll talk about why it's important to us as an organization and how open-source technology um enhances that transparency, enhances and also we'll walk through in more detail what does that actually look like um at the county. So, go ahead and pass this over back to Fay to talk more about what we

do. So, um, how many of us here are LA County residents? Yay. Then I don't do I need to um introduce that we are the largest county in the country, one of the largest. And uh um our local government serves more than 10 million uh residents. And uh um LA County has more than 35 no 37 39 um departments um within our county uh including more than

a 100,000 um uh employees. Um we are a small team but we have mighty um responsibilities. We have the privilege to um um you know integrate and use uh data from multiple departmental sources within LA County to um serve uh you know um uh high level program evaluation and policy decision. Um just as a background, we have data from uh uh you know um public hospital systems

within the county and uh mental health and public health uh service networks, um social services, child welfare um justice systems like sheriff's department and superior courts and uh probation and all those uh data we can integrate them in at a individual level to see um individuals um service utilization trajectory within our county. Um yeah uh that's oh we also have some uh happier data like park and

rags and the arts and culture department data. All these um can be found on our open data portal. Um so our function we are a multifunctional team very small but um we support our data engineering team for um platform building um but our main focus is on making data meaningful to inform uh um high priority board motions and I'm going to give a few example projects which

cover a significant chunk of our car current um efforts. Um so the first one you're seeing is our justice data hub. Um I'd like to demonstrate our data structure um and analytics uh workflow with this chart. So um we use a medallion data structure within data bricks. We transform raw data from uh sheriffs and superior court in the bronze layer into silver layer through cleansing and standardization

processes making it analytical friendly. In the gold layer, we built out project or research question specific tables that allow us to um do like analysis with rearrest or failure to appear in court. Um that's pretty much a longitudinal pattern of individuals trajectory with uh within the criminal justice system within LA County. Um so uh I listed on the left hand side three major programs or functions we

support with the justice data hub. Um reduce jail population. That's number one. And number two, map out path from incarceration to community based treatment. and uh also number three that's um the focus of today u which George will um do a deep dive um the standardization of uh data definitions and uh interpretations. Um a lot of the people when we talk about criminal justice data um people

will say yeah the California penal code is the s sing single source of truth but when in practice there's multiple um individuals personas with different interpretations and categorizations. It's a very complicated um process um which um George will show you that we um this the um our attempt to standardize this process of charges. Um before I pass it on to him, I'll also touch on some other

areas that we um like is our current uh focus in the county in our data team. We also support county programs ser serving uh various homeless populations. Um for example, the military and veteran uh affairs office and the um transitional age youth. Um not sure you're familiar with that um term. That's youth um aging out of the child welfare system. Um we're supporting um projects to strengthen

home uh housing services for these for these uh youth and also um the poverty alleviation in uh initiative as a whole to identify needs and service gaps to make sure we're utilizing our limited resources in an efficient and equitable way. we also um like we also serve as an engine um to enable data modernization uh at our department with our partner departments and a research collaboration uh

uh hub for these departments too. We help them uh upgrade from their legacy systems by showcasing our PLA platform so they can envision how much they can achieve with a um more advanced technology. We also connect departments with external evaluators like universities and think tanks and serve as the department statistical or methodological reviewers to ensure quality and ensure quality uh evaluations and translate difficult um tech terms.

Um so basically um this this role is kind of interesting. I don't I don't know um how many of you have ever tried to talk to departments no like people who are not that techsavvy and uh people who focus on um very specific research questions and technologies and methodologies. Um this closing this gap is uh part of our um is a big part of our uh daily

job. And lastly but not least, we also um help departments and initiatives tell their stories on our open data portal using uh uh multiple tools to um you know uh synthesize their uh data stories at one uh website or um at our uh open um data portal. They have a little section so they can uh go there and one click all their data stories are there. is

this still mine? Yes, sorry. Um transparency is one of our core values in providing all these services in the county's new technological ecosystem. We tell departments how we keep data securely and uh um reiterate our row and um data sharing rules within our county and with external partners. Um for example, recently we just updated our um DUA that's governing all these services with um um we're going

to communicate that with uh departments and uh um we also showcase what innovative technology can do for them and created uh departmental workspaces in our envi environment for them to try out. Um that's um what George is going to show you. um the part of um um standardizing uh charges. We also solicit for feedback with department partners to improve our work um to inform us on how

to best ingest their data and use them. we inform them about limitations of our data and our studies and methods and so our results are defensible in um in terms of accuracy and ethics. Um, for example, the reunest analysis that we've been working on, um, it's again one of those terms that seems seemingly simple like recidivism, but in practice that has so many different uh um interpretations.

It can be a rearrest. It can be a rebooking, which is a processed arrest. It can be a prosecuted booking which show shows up on the court case records. It can be a convicted uh crime. So it really depends on how you define these terms and uh um based on the research question. So with that, I'm gonna toss it back to uh George to talk about his

amazing uh work with charges charges categorization. >> Thank you. Yeah. Um I think and what FA is kind of uh getting to or touching on is you know these different silos and we've got different interpretations because we have different reasons uh that we use the data different purposes and so we are being a central department the CEO's office we are trying to be that glue. We want

to be able to have everyone be on the same page. Even though we might have a different slightly different definition of something, we want to still be able to build like a common language. Um, so I'll just talk a little bit about how the open-source um both technology and ethos really enhances the transparency that we're going for. Um so we invite stakeholders to review the methodology that

we use um by sharing our code. So we I I have this image here because I really feel like it's unraveling the sacred scroll for everyone to see. You know, it's not like a secret. It's not a black box. How do we determine rearrest? How do we determine what's a serious conviction or violent? These are these are not just simple words I think to the everyday person

like the public they have an idea of what violent means but um and that's that's maybe great for a public facing dashboard but for specific use cases violent really um has like a very specific meaning a very legal meaning. Um so we want to be able to communicate that nuance um by kind of sharing the the process and the method that we go through. Um, so our

justice partners, we work with a lot of different folks from superior court, uh, from the sheriff, um, and then not just those agencies, but we also work with community- based organizations. Um, people that, you know, are advocates for communities in, you know, that are impacted by the justice u the justice system. So they they want to see how we do things, too. Um so we're able to

share with them uh you know the our processes for enriching the data and our analytics and our reporting. So we get a lot of guidance from them along the way and they really do help us interpret what we're seeing. So essentially it's kind of like we do all of our work and we do it out in the open because we want to be able to show the

math. we want to be able to um but our partnership with justice departments and community organizations is really noteworthy and so I'm going to talk about um a specific use case in charge categorization and I'll walk you through the different layers of of our transparency stack. Uh so essentially starting with the open data portal that's that's the most visible and publicly visible u component of it. That's

where we share the statistics. We'll share dashboards, the different methods um that we use to get to those numbers. Uh we also share source code through GitHub. And even with data bicks, data bicks is the most like private and most secure, but there are tools within data bicks that allow us to share data with researchers. Um so, you know, through each one of these um tools, we're

able to promote transparency. So this is what our open data portal looks like. Um it is getting a makeover so it might change in the next week or two. Um but essentially it will have all these parts to it. Our team maintains this um to serve as our primary communication tool with the public. So you can actually scan that QR code and it should take you there.

Um so you can see here there's you know there's um a couple of dashboards on there. There's some things that uh you know there's a glossery so you can see what some of our definitions are. So there is an analysis of charge levels across demographics which I'll show in the next slide. So here, so we really get u into showing the research, showing the methodology that was

taken um to understand uh charges and uh pre-trial reform that was happening and how it relates to homelessness status and demographics. So in this one, this is this is where sharing the methodology for analyzing complex metrics uh really comes into play. So, like it was mentioned before, a rearrest sounds very simple enough to understand and to analyze, but you ask different stakeholders and they will give you

their own opinion on what a rearrest means. So, this is a this is a report that's in response to a board motion, the board of supervisors, and we're laying out exactly what we did and how we did it. All right. Um I'm going to highlight right now the charge code cleanup process. So this is here so that we can share you know at a high level the

methodology for cleaning up those charge codes and I'll get into some more specific um examples. Here we go. We collect data from a lot of places, but I'm showing you just three of them right now. The LA Sheriff's Department has booking information. The LA Superior Court has court case information. The LA County Probation Department has their probationer case information. So, each one of these has information about

the charge related to their population. So, it could be robbery, burglary, you name it. um they're all recorded slightly differently which makes analysis a little more challenging, makes it very time consuming, especially with the amount of data we have. We have millions of of records. We want to be able to to label these charges uh with more meaning aside from a 211 PC or a 459. We

want to be able to tell you exactly what that means in layman's terms um so that we can do analysis. So, just a couple things to notice right off the bat. Hopefully, you can see it. I realize you probably can't, but um there are but one thing is that all these charges are entered manually in each source system. There's no descriptions in these systems. So, you're looking

at a 211. That's a 211. I hope you know what a 211 is if you're trying to analyze it. Even if there were descriptions in here, they would probably have a lot of variation because they are manually entered. You might see something like, you know, assault on a police officer and then you might see assault on an officer or officer assaulted. You might see so much variation,

but it's really pointing to the same charge. So, we needed an authoritative source to tell us what do these mean? And of course, the authoritative source is the law and it's there on a website, legislative info. Uh you can pull it down, but even then it doesn't it doesn't really mirror what goes on in these systems. And I'll give you an example. There's something called a literal

ID. So within within these systems, a literal ID basically helps identify a very specific sentence within the law maybe or a part of a sentence. So a 459 PC, that's the the numerical part is a statute. The PC is the code. The literal ID might have a few letters on the end like res, which means residential, uh coom, which means commercial. That matters actually because when we're

trying to figure out what's violent, a commercial burglary is not violent compared to a residential burglary. So when we're trying to categorize these things and trying to create meaning, those little identifiers really help and they're not always there in the data. So we we really need to try to find a way to to um use all the data that is provided to us but standardize it and

uh standardize it across the systems and link it to some labels. So this is this is the authoritative source. So, I'm putting authoritative in quotation marks really because, you know, it earned that that title by being the cleanest and being compatible with our data sources. So, it's it still has a few flaws, but it's so far the best that we have. Um, so you can kind of

see here that there's the statute, the code, these are all in separate fields, the offense level, which is really important. So, a misdemeanor or a felony and then the literal ID. So, you know, examples like AU would be auto or vehicle. G A R is garage. So, 459 probably has the most that's why I use it as an example. It has the most variation in literal IDs.

Okay. So, what we did is we looked at this authoritative source. We took it to our partners and said, you know, help us work out some business rules here. We want to be able to use this as, you know, this being our the best source that we have. And they all agreed. You know, we had several meetings to kind of see, well, how does this actually play

out? Let's actually start using this and refine it over time. So, we it was an iterative process with our justice community, uh, where we were able to clean up the data in a way that still represented the original record because we want to be able always to say, you know, like this is what actually happened u why the person was booked. This is what they were actually

convicted of. We want to try to maintain that kind of fidelity to the original source and that's where we get to this. So that process of working with our community um it was really facilitated by well number one by GitHub. you know, it's a lot of the analysts that work in these different departments and work for, you know, different um community organizations, they're, you know, familiar with

Python. And so, we let's write things in Python to clean up all this data. So, we made a a a Python package. And we developed really advanced cleansing and matching rules. And so, when all was said and done, we really needed a repeatable way to implement that code, to implement those rules. And so having this package made the most sense. Although originally we had things in a

datab bricks notebook um for people to kind of share and use. It involved a lot of copy and pasting and copying and pasting is not the way to do it. Um so we really had to create this package. Um right now it lives in our enterprise GitHub account. Um but it is available to many of our justice partners for them to pull down. They can use it

on you know their um their data sets if they're no matter how big or small. This this is compatible with both Python pandas and Spark. You know, we use it mostly in Spark because we're building pipelines with this. We are um using this on millions of rows of data, but it works with millions of rows or just a single row. So, here's just kind of an overview

of what that package looks like. So, these are just some of the main functions in it. So there's a lot of components to it. Uh it to again it took a lot of uh collaboration to to see exactly what we needed to do. But essentially a lot of the functions here extract components of these different um statutes. So the numerical portion even the subsection is important. So

the subsection is usually that uh part in parentheses where it's like a 459A that points to a very specific subsection in the law or A1 or a D2. So those kinds of things those are all very important. So we capture all that. This parses through it. And let's see here what this is is yeah this is an example of one of the extraction functions. So this will

go through and we actually had several meetings where we were explaining you know this is what we're seeing in the bookings data. The bookings data is the messiest. It's the the least tidy. And so we used it because if it could work on the bookings data it could work on all the other data sets. Um so we wanted to have uh these these sessions where we actually

worked with other analysts and some of the subject matter experts and said you know implementing these rules we get these results. Uh there was a lot of feedback. There was a lot of you know well that's great but if you do it this way you might you might um you know overlook some of the the nuance that they were going for. So things like attempted attempted crimes

they're coded slightly you know differently with like the letter A in front. There is in the law there is no letter A in front of the the number. It's just that's how we do it. So, we we've tagged the letter A and that seems simple enough. You okay, we'll just anytime you see the letter A, that means it's an attempted crime, attempted murder, attempted burglary. Um, but

then sometimes that A is not in the same spot or sometimes there's like a A and then like a funny character after it. So, you know, all of these little things, we wanted to really vet them. We implemented them and got their feedback along the way. Uh this is essentially a wrapper function that really takes all those extractions um and puts them into one function so that

we can uh pull it all out in one in one fell swoop. So this actually puts together the it extracts the statute, the code, the subsection and the literal ID. So how does it how does it actually work? um after it's cleansed there's a matching process that happens. So the standardization was like the first step. Now the actual matching needs to happen. So we're matching to our

source which is that that list I was referring to earlier. And we start with the most restrictive set of rules down to the least restrictive. And so the most restrictive is basically match it on every single piece of information we have. So, will it match on the statute, the subsection, code, level, and literal ID? So, an example of that at the top is like a misdemeanor 12702AH

HS WPR. So, WPR is that literal ID. If it matches on that, we have a very specific description of what that is, the most specific. If we don't have the literal ID, which in is often the case, then we get the next best thing. we get, you know, a slightly less specific but still accurate description. And so we just really work down our way to the bottom

here, which is the least restrictive rule, which I think is the most interesting. Essentially, it's a tossup at this point. So the statute if that's all they've put in their database if they just put let's say I think I saw one was like 11 8 it was like 11187 or something like that. And so I showed this to our justice partners and I was like I don't

know what this is and they're like oh that's the health and safety thing. They knew exactly what it was. Like it doesn't exist anywhere in the code. Mind you, if you go to the website, the California legal website, I don't know, there's maybe like 30 different codes, health and safety, penal, vehicle, business, administrative, tons of codes, election codes. So, she knew exactly which one that was. And

so, I looked I looked at, you know, the the data we had and realized, okay, we could use this, but it's kind of iffy, right? So, that's why I say it's a tossup. Basically, if if the code exists in one other place, then we can safely say it's probably that. But if it's in two or more different places in the code, if I find let's say 187

somewhere else, then I probably don't want to use it because I really the algorithm won't know and we want to kind of air on the side of safety. But it this does really help when we see typos. If there's like a 187 HS, I mean, most people would be like, well, I know 187 means that means murder, but HS means health and safety. It's not found in

health and safety. There was a typo there. So there, that means it's a, you know, they really meant to put 187 PC, which is a penal code. So it helps in those situations when we see typos. So how is this actually implemented with code? So these are the general steps that we go through. Um essentially we'll load a reference data set. Um the package does come preloaded

with a cleaned up reference data set. We prepare the raw data for matching and so the I think I showed you earlier the function that's a wrapper function essentially that prepares the data for that matching. So prepare charges prepares that. Um and then the charge matcher class actually handles that very complex matching process. So all of this really makes for a very simplified interface that can work

for you know any number of of um charge So how well does it work? So here's a few uh numbers here. So before we implemented this this is these numbers are based on the booking and so before this package before we had a process to clean these up and label them uh if we just did like a very simple join you would have gotten about 68% sorry

uh we would have got about 68% match rate uh which is very poor I would say and if you did this manually you would we had, you know, like 300 lines of code um to to even just get it to a place that was I I would think is subpar, right? You wouldn't you wouldn't want to spend like hundreds of lines of code and still get anything

less than 80% match rate. So with this package, we're able to provide like consistent results and essentially like the bookings data now has a 90% match rate. Um, basically we're able to say, you know, we can tag all of these almost all of these u charges with an actual description and category. And the table on the right breaks down how many were were matched um for each

step in this process. And this is another component of that transparency part. So the output of the of the package and this code will also tell you how it was found. So, it's not just here's a label for you. This is what we're going to call it. It's telling you how it found it. So, it found it based on the statute code and the subsection. Um, it

gives it a score. So, that way the analysts, you know, that are outside of our team, the analysts that are using this at different departments that are using this um, you know, at any other organization, they can use this and have more confidence in the the labels that they're seeing. Right. So where do we go from here? Basically we want to improve this categorization work. Uh we're

continuing to do that by um very soon we'll be leveraging um a couple of methods. Hopefully uh leveraging BERT and weak supervision models uh to classify our statutes and to continue to work with our partners to refine those results. uh one day it would be really great to make this taxonomy available um to more jurisdictions throughout the state. Uh for now we're just going to continue, you

know, strengthening those partnerships and building new relationships by openly sharing what we do and how we do it. With that, I'd love to take any questions you might have. >> Thanks. Yeah. I'm part of a fire safety committee for my neighborhood which is a high very high severity fire risk and it would be super useful to have some sort of a risk score for every house in

the neighborhood based on whether they've done the home hardening. But that's kind of sensitive information and block by block the neighbors that I talked to would be okay with sharing that within the neighborhood but a little nervous about sharing that just publicly. So to your point about having like aggregated data being something you want to share and sensitive data being So questions like what guidance do you

give for how to manage data like that? And then if you did want to both host it and then share it with the city or make it public, how do you think through stuff like that? It's a big question, but I just wonder if you have any examples of community data that plugs into actually government county owned uh All right, thanks for the question. Yeah, that is

a big question. Um, you know, we don't have at least I can't unless F has a different experience. I can't think of a data set that comes from the community uh that would be public like on our platform or a government platform. Uh but that's a really interesting idea. I I mean the way we handle sensitive data between agencies is you know through contracts essentially you know

so that would be that would be the only thing I can think of is you know you have to develop some kind of enforcable contract between people. Um, yeah, cleansing the data I think is is smart, but you know, even that can can be challenging because you can always deidentify things or not always, but often we think we're deidentifying when we don't do it enough. Someone can

always outsmart you and deidentify identify someone again. Does that Okay. Yes, sir. relating to his question here. Um, last year when we had the fires, we actually used uh the watchduty app. Have you used it? Yeah. So, and that aggregated uh public data along with a lot of other sources on the ground and it worked better than the public sources because it's hard to get to the

actual data. I'm wondering beyond what you presented, what else is LA County putting out there for the public to use that would help us in situations like this? because it sounds like what you're doing is obviously it's a very special use case. There's also a lot of other data that should be public. I'm sure GI you have all that stuff. What else? Can you talk a little

bit to what else LA County has that is made accessible through APIs and other sources that we can feed instead of doing a batch download? We want to get to the live data and build an app with it. For example, a little because it's so cold in here. [laughter] Um we do have a um LA County open data portal. Um there's um public health data like community

health survey data and uh um you know what I call happier data like arts events and the parks and uh recreation and um also um from other departments even some of the um uh criminal justice data. Um I think the difficulties for us is what this gentleman just mentioned. We cannot share any identifying information especially with locations. Even if when it's a address when you zoom in

close enough you can identify the individual house or uh apartment there which is strictly prohibited on our website. So on our open data portal, it's very hard for you to link to any um underlying data, but um there's also other um resources you can find at L County GIS portal. We created a lot of boundaries and a lot of um um our county offices service service areas

responding uh response catchment areas or those data on top of your individual level data or community level data. Hopefully that can provide some uh context or or for you. But um I know the problem is like nobody knows we have these resources. We gota um we got more outreach uh tasks to do. I think actually I have a comment and a question. I used to be a

uh correctional counselor within the department of corrections and it was really difficult handling the data especially um with people that came from other states or that were uh prosecuted at the level uh federal level. So is there anything being done in regards to the different states how you know those how the codes change state to state and then federal and all that? Would it be awesome to

create a system that actually standardizes data throughout the whole country? I think that would be the most useful way to do it. else. >> We I mean that's a good idea. I would I'm I'm just trying to get my head around what we're doing [laughter] about LA County, you know, but um you know, I I mentioned at the end, you know, I really would like to expand

what we're doing throughout the state and create some standardization. That would be that's like that would be amazing. That would be a goal for me. Um if we can go beyond that, even better. Yeah, >> This the impact is pretty big. just it sounds like just changes charges categorization, but if you think about it um people who are in treatment, people who are diverted from courts to

treatment, people who are going through the whole re-entry route, um this categorization decides whether they qualify their charges qualify for treatment and um the evaluation of these critical programs. Whether through these treatment, if they were rearrested, if they were rebooked and the charge levels, whether it was still violent, whether it was still uh strike eligible, that makes a huge difference in one individual's life and also like

whether we're spending resources effectively. Uh, I have two questions. Um, how does like the LA County government like detect web crawlers trying to like access information and then um for bugs in the code? Do you guys do contracts for like a third party that tests for like bugs in it or you have like an in-house team who looks for like bugs in like the Python code? Um,

so to answer your first question, I I really don't know that that whole uh detection of, you know, web crawlers and things that's that's going to be outside of my purview. I can't really speak to that. Um, I do know that we have a very robust like uh security team, you know, in the CIO's Um, but then you're in regards to the other question. Sorry, remind me

of the test for bugs. Yeah. Oh, thank you. Uh with our, you know, largely with our our subject matter experts, you know, there end users. I mean, we have a lot of obviously like best practices when developing software, you know, developing tests and that kind of thing. Um, so we have, you know, our test procedures, but it's it's internal, you know, we're we're developing them internally and

then our end users give us that feedback. We iterate through that way. I was wondering if you ever had feedback to the to the legal code itself. So if you had some analysis is not possible, could you actually propose like legislative changes that would make it possible? Like would you talk to a you know a state rep and say like you know my customers are saying they

need to do this and they cannot do it. Maybe this law, this law should be changed in a way that makes it easier. >> That's a really good question. I've definitely thought about it. In working on this, I've absolutely thought like I want to go back to whoever wrote this. I've I've I've gone down the rabbit hole. Yeah. Um but never done never actually gone through with

that. I don't I don't know if that's the right move for me to do. Likewise, likewise, it would be nice if they maybe uh had those literal codes uh listed somewhere in a way that accessible so you don't have to hardcode it into your reax that you showed on this screen. That would have that would be nice if you know as the as the codes change and

the law new laws get passed that could get updated automatically without you having to go in and update the reax. U but that's not my question. Uh my Uh my question is uh I maybe I missed something from your talk and if you could uh maybe repeat something that I didn't catch. Uh but I've seen the word transparency mentioned a lot but the transparency seemed to be

contained to transparency within or and or in between the agencies. Um, how does the public benefit and what parts of your system are transparent to the public? Great question. Really great question. Um, so the the public, yeah, it's not it's not like everything is fully transparent. Um, but what we're what we're doing is is making, like I've mentioned earlier, the open data portal that contains a lot

of the data. Um it does contain sort of like a highle summary of what I've just kind of shown you like this is how we arrive at this process of charge Um so a lot of the ways that it reaches the the like everyday person is is either going to be through like those community organizations that are doing like advocacy and doing like that kind of research

on their own or other external like researchers. So that's like one way. Um, and I don't want to but I also don't want to like diminish like the the transparency transparency between that is something that didn't always exist either. So it is it is like a really big step especially between um you know the the fact that we have government or I'm sorry LA Ali County government

agencies. So they all belong to the county government. Then we have superior court which is not a county government you know it's it's a state entity. So just interfacing with them with this level of transparency is really um facilitated a lot of the work that we've done um the results that we're that we're seeing. Uh but yeah, I mean I I think one of the things I've

really thought about and really wanted to do is to for me one next step would be to make the the GitHub repo more public. Um yeah, it is enterprise right now. Uh which is still very limited, but yeah, that would be I think the logical next step. Uh it is still something we're working on. So I mentioned like we're still working on those categorizations. So, we're hoping

to refine it. And basically, when I get it to the point where it's like 1.0, then like the public can can really see it. Like, it's going to be good enough for anyone to kind of like take a part. Right now, it's like still it's still internal. Even though it's like a big big internal project with different agencies and organizations, it's still still kind of small. So,

yeah, I'll acknowledge that. Um I'd like to add to that answer in addition to our internal like process to break the silo the data silo and um um we do have mechanisms um like every quarterly we have like working group um presentations where um CBOS will be involved in the process. So the transparency would be for us as a data unit would be a methodology transparency that

here's the results. The aggregated table and this is our technical report of how we arrived at this. If you have any um other ideas or better suggestions on how to do things, we are like on the receiving end. We are always there to take um suggestions. And also in terms legislative change, we take our supportive role very seriously. Um all these county departments sometimes make data requests,

make analytical requests to us. We support by providing data by lead leading by um letting them lead the data projects to a position that will support their argument through their representative with state or federal um agencies. My question is, we talked a lot about like the specifics of things like the definitions of violence definitions. So, and it it occurs to me it's um there's going to be

differences in local. So, like county, city, state, federal, they all are going to have different definitions of what violent is. And then they're going to all be different definitions of all these individual little codes. Do you have like is you part of your process to have somebody on staff who's a legal expert who helps you to define these codes and then the sort of secondary question is

it sounds like you're doing a lot of hard coding of business rules and then you talked about maybe going into some sort of like weekly supervised learning or something like that. Um I know that there were like there's another part of this process which is like um like when you're talking about people who are recidivists how do you know that you have the right person when you're

joining across the data so if you have a John Smith like that could be any number of people locally I know that there are my name is Mark my name is Mark Rhden and I know that there at least one other person named Mark Rhoden in the LA area who is a programmer who is about my age. And so we get we like how do you do

that dismbiguation? I also work with somebody whose name is Chris Smith and he Yeah. Well, he gets all kinds of stuff. So yeah, there you go. Maybe probably not the same one. So >> CWMD part. So we have another um process in the um that I earlier mentioned. We integrate data from different county sources. Um that's a separate process from our analytical process. So every individual from

each source department comes into that identity solution system. Um we take in um all their PIIs into consideration. uh names, SSNs, uh addresses, phone numbers, gender, race, age, yes, um all those information and we have a um algorith algorithm to um identify whether individual A is individual A showing in uh sheriff's department is the same individual A showing in court is the same individual A showing in

uh um like welfare um system or the same individual showing from history in child welfare system. All all those um things can be linked but it's a separate um process that um analysts cannot see. We have a electronic ID assigned to each individual. Using that number we can uh pull all their uh service administrative records from other departments. This way um we can do the link um

we're not um for any project we're not exposing individual even um us as analysts we don't have access to the personal identifying information. Of course, if you think about it, if you have addresses or um phone numbers for each individual, um you you kind of can identify individuals, but um that's another process with being doing all the demographic information is also separate and all those address are

break broken into sections and we're allowed to see um only certain sections to pro um protect privacy. But we do have a certain um accurate um accuracy rate of the match that we are confident 85% confident that this person is uh um this person we see across system and do you want to take the coding part the coding sorry remind me again I have a voice memory

right you know, um, we're really, yeah, we don't really have, um, insight into city specific codes, municipal codes essentially is what you're asking. Yeah, we we have them in our data and those are actually that makes up the 10% or so when I was pointing at the number 90% match rate. Um it's because the other 10% are usually other jurisdictions. So municipal or federal typically. So yeah,

we don't really have that level of detail integrated yet. I don't know if we ever really will. Yeah, that that is I think there is going to be a natural like ceiling to to what we can do. Another question. Thank you for your question. I got a couple of questions. Uh how difficult is it to get the data from agencies because dealing with like the sheriff's department

uh social services or probation and is the data that they're giving you uh I I'm guessing it's not clean or anything. You guys have to do some cleanings and and uh transformation. inter agency. So, like in the Orange County, if there's a person that has committed a crime there, how do you uh kind of like follow that path to there uh to LA County if if he's

been arrested and all that stuff? Uh are you able to see that? Um just uh data within within counties is is difficult. How do you deal with that? How do you get the buy in from the sheriffs to be like, "Hey, give me that data." Uh yeah, intercount data. We we don't we are not lucky enough to get data from Orange County. So really the only data

that we can ever see is if it was a transfer between counties. If a if a body a body person was transferred from a c, you know, Orange County Jail to LA County Jail, that can show up in our data with the charge code. Um, but we're not tracking people necessarily that committed a crime in San Jose. Uh, we're not able to see that. We're not able

to track their criminal history across, you know, the county or the state. So, that is, yeah, that is a pretty big limitation for us. about tracking. LA City actually has a crime map that's public that shows pretty much in real time something happened with an address. It obviously doesn't put personal information. So technically from a technology point of view, it's pretty straightforward if you have a schema

and definition of what's personal, what can be shared, what can't be shared. And I would I would think that if we're talking about transparency and you've spent all this great work doing all this analysis, a lot of it can be put out there, especially the the business rules that say, what is the code? What is the rule? And not putting it out there is actually a detriment

to the value that you actually already put into it. So it could be done. You don't have to wait for it to be perfect. In other words, if you wait for it to be perfect, it'll never get out there. So, I would I know it's not your choice, but I would suggest uh putting that out there as it's not only a I think it's a duty of

an agency to put it out there, but it's also a risk reduction for things like fire and crime and whatever to to actually put it out there so the community knows what's going on and could help build applications like watchd duty or whatever that we can respond to emergencies because when the next emergency happens, we don't have this information. it's going to be on the county that

we don't, you know, the city or whoever that didn't put it out there. So, I I would strongly suggest getting that message up the ladder. Anyway, >> yeah, thank you. Thank you very much. >> Thank you all. Thank you all. >> [applause] >> Check. Check. Yeah, sure. Okay. Sorry. Yeah. Uh so yeah, so this is like a source code example where you know th this is a

really interesting one because you know creating contracts and documenting agreements is something that um is already like a barrier to entry in the legal world. But having a dispute is like a whole another level of barrier to entry. Nobody wants to like get get into a dispute. even small claims court is like a huge uh pain. Um, but you know, if somebody if you have some issue

with like your open source license, you could use harness uh and you could basically have it like create a standardized uh complaint um and and you can you know people could work together to build that out uh like you know systems of lawyers you know technically the way it works is these types of tools are supposed to be used by licensed lawyers and that's how uh legal

licensing works. Um, but you never know like the there are already laws uh like that people are trying to pass to open up and allow more uh access to to this type of tool. But even just lawyers can use these types of things to create standardized templating um for documents. And I think that that that can make access a lot better. And so here's the example where

it basically created the document harness then validates it makes sure it passes all the rules um you know creates the output uh that that then needs to be filed and can actually create like multiple documents from a single set of the templates. Um so that's another example. Um here's a here's another example we created. This is like a litigation example. So something that you do and this

is a huge cost right like why do lawsuits cost so much money a big big one is you have to go through all the prior documents all the prior testimony of the witnesses and some you know group of people have to sit in a room and review that and write up the takeaways and then prepare a seven-hour outline of questions for that person uh to try to

get them to you know trip up or make admissions. And so you can actually standardize that. And and I can't tell you how many times over my career um moving between cases, between projects, working with different people, everybody has a million different preferences uh on these documents and how they should look. And um it's like, you know, you you you learn the hard way oftentimes or you

don't get what you need. And so if you could just provide like these types of things then it it would it would make life a lot easier. Um another example is like one of the big drivers of cost in litigation is there will be documents that aren't marked privileged and confidential. And so if you talk to a lawyer, your conversations are privileged with the lawyer when they

give you legal advice. So no one should be able to get that and use it against you in your case. But so many times uh people are creating documents and they just forget to put that stamp on it and then it can cause huge battles around did you wave the privilege? Do you still get to claim it? And so this is something you could just standardize um

using harness, you know, across uh any any document and then it could be tied later to like a potential harness uh document standard for reviewing documents in a case- which which could help. And so basically, you know, there just many examples. uh I'd say a huge cost reason for the high cost of legal services is everything is bespoke and the uh lawyers just are constantly reinventing the

wheel uh in terms of document creation and you need you need something more than just like you know a template document in word uh to to be able to do it. So, so yeah. So, here's the just to recap like the pipeline um in terms of how harness works. Um you can create your templates. It can validate the rules. Um you know can generate output documents. You

can have all sorts of you know kind of like simulated processes that that you would do. You can get it signed or served uh and delivered and and things like that. And so, um, so yeah, so that's that's kind of, uh, where we're at. It's built in Swift, as I mentioned, uh, and so, you know, they but as I also mentioned, it doesn't have to be. Um,

you could do whatever you want with it. Um, but there are some benefits to that. And, and so, you know, you can look on on the foundation. And I think there's um there's some other things that they're building uh with it. And so so yeah, so this is you know out there. I think that the uh legal kind of open-source uh tools for legal AI adoption are

really light and they're just just starting to kind of spring up. And so this is a really interesting one I thought and that I have used and again like I did not create um at all. Uh so so yeah. So that's where we're at. And so I I see we have time. I don't know. I'll open it up if there are questions. You know, I do use

like AI a ton uh in my law practice and just happy to answer any questions or whatever would be helpful. Yeah. I think it's pretty nent. It well it's a it's affiliated with the UNLV Law School. So it has some institutional support, the Neon Foundation. Um and there are active multiple active developers, lawyers like that have you know like I have a big team that we do

you know, 10,000 hours plus a year of legal work. Uh, you know, and but but it's I don't think it's like, you know, in the in the lawyer software kind of AI world, it's very, uh, early. Uh, most are not using like clawed code as lawyers. They're more using like co-work. So, yeah, it's not um, you don't have as many like on like a open source project,

let's say for like a DevOps tool. It's not at that scale at Yeah. Uh, sure. I've seen a lot of those like either drag and drop document assembly tools or other types of of them. I think just like making it a CLI is really interesting to me because it allows you to go faster and and there's this sort counter incentive I guess for a lot of lawyers

because they charge by the hour and I do too a lot of times but so you don't want to you you know there's this idea that like well if you do everything faster you make less money but what I'm seeing now is like if you can dramatically up your speed and then just take on more cases then you can actually net expand and that's what I've been

doing and then um you can also change the the economic model maybe and start charging flat fees which is more predictable and so um so yeah does that answer the question I guess yeah um on the GitHub it had you know I think there's at least one example and we can put up more but um it'll really be every lawyer or law office is going to want

to customize it but but you could start from somewhere. Yeah, >> it's probably Yeah. I think you're right. We and we actually created like a bunch of these last night. So, and we could probably send up some templates. Um, yeah, it's going to be continued to be developed in Swift for sure. >> Uh, yeah. Well, it's interesting like a lot of lawyers will interview clients to get

information just live or on on a Zoom, but um I have started talking to more lawyers that are designing essentially like avatars of themselves that will interact with their clients And so they still train them and um will go in and review the conversation and follow up on it. Um but I think you could use something like this to to interact with the clients. And a lot

of these days, every client is feeding everything you tell them into AI anyway. And so like what you get a question, it comes from the AI that the client used. or if you give them an answer to the question, they'll go and ask what the AI thinks. Um, and so yeah, this could be a way that maybe you could have a little more uh control over that

process in the attorney client relationship. That might be good. Yeah. >> Sure. Yeah. I would want it to like autopop populate templates based on prior examples of things I did, which I think it probably could do with with a little bit of massaging. Um, and then I think the idea of having like more a template library sounds interesting. Um, I guess if you guys have ideas of

like what you, you know, any any ideas, feel free to tell me and I'll pass them on. Um, because it is under active development. >> Yeah. and to create work like work to simulate workflows to also have like review and validation workflows to let's say check uh all the client file materials and make sure what we're saying is accurate or check the court rules about how the

document's supposed to look like and customize it for that or you know make sure the client signs off on it. Sometimes clients have to so just whatever the which is right now kind of this like it's split all the information in these different rules in people's heads. Uh just kind of standardize the workflows that you would have to provide it or like create the rules to go

and and check it. Um, so that is still kind of like a collection process and every law firm would have a different kind of set of those that matter to it. Um, but you can kind you know I would know let's say if I spent with a bit of a boot up cost I could feed everything I I needed it to know into like a certain place.

Yeah. That's a that's a hard question. That happened in in New York. There was a judge a few weeks ago that he made this order where the guy um it was a criminal case. So, the guy was arrested and he had been talking to either GPT or Claude before he hired a lawyer and the judge said uh it's all discoverable there. It was a little unique because

the government was suing him, right? It's a criminal case and Claude technically says like you you can reveal what we we can reveal what you say to us to the government in if and so it was like kind of a tortured scenario that was very I think distinguishable from most but and the judge did say like if it was his lawyers using it or if he had

already hired lawyers and it was like a more of like a let's call it like a team project or something that might maintain the privilege. So I think the answer is your lawyer should use it. You should be able to try to like communicate in a sandbox that you and your lawyer have access to and that over I I believe and I I taught a like a

class at a law school about this. If you go back to when the cloud was developed initially, uh it was the same thing where people were like, okay, does it wave privilege to upload and then eventually the cases kind of came out and yeah, if you upload it with like a publicly accessible Google Drive link and then post that on Twitter, you might be waving privilege, but

if you if it's like between you and your lawyer, like probably not. Um, so yeah, that but that's very early and the decision so far there's barely any uh on cool. I think it's wills because wills and trust is like pretty expensive to talk to a lawyer about it. It could co go it could cost the way I think about it as a lawyer is like it

could cost a lawyer $3,000 basically of their time and their staff to like talk to a client. And so as it's not the type of law I do, but it's something I need. And so and I do I've seen just over over time and like when I was doing startups like it's an area that startups are always trying to tackle is that wills and trust use case

and I think this can actually do it and it it does it in an interesting way where it's like just an open source tool that any lawyer can use and so that's how it ties into the like the foundation that is helping develop it trying to make uh legal services more accessible. All right, cool. Well, thanks everyone. Uh yeah, thanks for coming to the talk and hopefully

it made sense uh in some sense and uh yeah, have a good one. Anyone that has questions, please. I'll hand you the mic. Hi. Um entrepreneurship support not guided at all on and he was like dreaming of having an open source project but and he's been coding for And then another couple of students who were asked to participate in a campus and I read the uh contract

and it was from my department. Um, and so I'm just curious like do are you running into issues with uh the people on campus who are trying to produce as much private as possible? And what's your strategy around that source for me to send students to who need somebody to look at a document and say, "Hey, maybe you have another option. And you could open source this

or this is what would happen if you did. There's like Um, hi, I'm Carla Padilla. I'm over at the UC San Diego um, open source program office. Um, so I've been working a little bit with the TTO's on some of these problems or issues. And what I want to say is that the um, UC regions, which kind of, you know, sits above all the UC's, I consider

it like the federal government and then each campus is its own little state. Um, so they have policies about copyright and who owns what and there is a website um, you know, just Google you U you um you UK cop um um copyright and it'll tell bring you to a page where it kind of outlines who owns what and under what conditions. Um, you know, like generally

like Stephanie was saying, if you're doing something as a staff member, if you do something as a research project, because a lot of research grants will specify who owns what, right? Sometimes, you know, it's not owned by the students. It might not even be owned by the university. It could be owned by, you know, whoever's um um providing the grant. um sometimes with like what students do

on their own. Yeah. Generally, you know, undergrads do own it except that if they use significant um significant um university resources now, how that's defined and you know what is that line can be a little bit blurry and that's really where tech transfer office or your office of innovation or commercialization whatever it's called on your campus would um be the ones to weigh in on it. once

you know so we can help guide you to kind of those resources we can also talk to you about the different licenses you know what licenses are kind of favored by the UC if UC has some of the ownership but they're going allowing open source or if like you know as a student and you completely own it we can help you say like these are the different

licenses here's what they do but we have to be careful about not providing actual um legal advice but you know there there is um opportunity for us to help provide that guidance that I think you're looking for, your students might be looking for. We do have still time for questions. open source versus what should what what is defined as IP. I think like from a different angle

of the same story bridging the gap. Uh everybody talking I will be talking about bridging the gap between open source and academia as well. So what I am interested in uh so you talked about like empowering uh facilitating like students researchers and uh basically like uh all you are doing for the student body so what I'm curious about uh what are your relationships with the faculty uh

do you have any influence on curriculum uh do they come because from for me as a founder of nonforprofit one of the goals is to bridge this sad gap I find actually faculty being like the most inertious part of any educational system. Uh and uh I'm still something they need to pay attention this regard. The greatest faculty care about their students. So by connecting a open source

project to the Check. Check. One, two, one, two. Hey. Hello. Hello. Okay, seems that it's working. I don't I'll just do it like this. This This works, right? >> Oh, okay. No, I don't want to hold it. >> No, this is good. Thank you. Okay. Um, I'd like to invite people to move up to the front. Uh, we have a lot of space here, so if you'd

like to No, no pressure. No pressure though. >> Okay. Welcome. >> Okay. Yes. Got it. Okay. All right. Thanks everybody. Thank you for coming to our presentation. Uh I'm going to begin with breaking governance capture how sort can transform organizations about me. My name is Leonora Camner. I'm executive director of sort USA. Until very recently our name was democracy without elections and we've just underwent a name

change. We wanted to be more focused on what we're for and not what we're against. So now we're sort USA. Um, I have experience in movement building and organizing. Uh, I was previously the executive director of Abundant Housing LA, the proousousing organization. So, the main idea of this talk is that randomness is powerful. So, just engage in this exercise, this thought exercise with me for a minute.

Let's imagine these things without randomness. What would happen? Like if we took randomness out of clinical trials, what would happen if we took it out of jury trials, cryptography, statistical research, risk modeling? Um, I imagine, you know, if I propose taking randomness out of these things, you might say I'm crazy. You know, you might say, "How would clinical trials work?" You know, maybe I made the argument,

no, the problem with clinical trials is we're just not getting the right people into clinical trials. We should just get better people, the right people to be in the clinical trials. You'd say, "No, that that doesn't work. That we'd have an issue. They would be manipulated." Okay, that's one issue. You wouldn't get a representative sample. Okay, there would be skewing. There would just be so many issues.

You can't take randomness out of clinical trials. Okay, if so you might wonder why don't we have randomness with important processes like policym and governance. Okay. So, if you agree with me that randomness is powerful and it's crazy to take it out of things and that um you know maybe things are not perfect but randomness is an essential important quality in keeping things free of bias, making

things more reliable, um preventing capture of those things like preventing manipulation. Randomness is so powerful doing that. uh it allows real representative sampling like we can't do that without randomness and it provides more system resilience. These are all things that we need in governance and policym too. So take so the fact that we don't have randomness in it you know if you're if you're following me so

far I feel like you've gotten the main point of this talk like I could just end here. Okay, randomness is so powerful and so important. So, uh let's think about the governance capture problem in organizations. You know, we don't use randomness. So, we have these problems. We have the issue of dominance and influence by certain interests. We have um board entrenchment, you know, like certain ideas and

personalities staking out a place in leadership. you know, we're not getting turnover, we're not getting new ideas. Uh, low member participation. You know, people feel checked out, disenfranchised, disillusioned. You know, sounds like the general public about politics. You we'll also see this in organizational governance also. Uh, low transparency, you know, people not having a direct way to get involved, like not having u their own place in

the organization. you know the there's an issue of the the board or the governance being far away. Uh and then internal politics over mission you know so when you have a situation where people um might benefit from these leadership positions they might benefit in terms of you know power or influence there's that incentive to keep that okay and and to prioritize that over the core duty of

fulfilling the mission. Um so this this is really about organizations. Um but we can also think about these problems you know in in governance you know uh public governance we can think about it internally in like organizational governance also. Um so you know another comparison that one of our members brought up is if you take cyber security you know randomness and cyber security uh makes those makes

it it more difficult to attack and capture by outside influences and manipulation. So similarly we could see like randomness in organizational governance can maybe do that security role too of uh of protecting against capture from lobbyists or special interests you know or like particular uh power interests like manipulation. Um and you know we we see this both in the organizational setting and in public governance you know

you have to follow the money right. So we have an issue where like donor donor interests donor preferences can take over in an organization. So enter the solution sort what is sort of government by lottery. Um so here's a quote from Aristotle. It is accepted as democratic when public offices are allocated by lot and as oligarchic when they are filled by election. Um so sort has a

deep history. Uh it goes back to ancient p practices like ancient Athens uh where you know you saw that quote before where people considered sort to be actual democracy. So something to think about and that tradition of sort has continued in many forms um and we know one of the most important forms um that we appreciate today is jury trials. Uh today we also see sort in

practice with citizens assemblies sometimes called civic assemblies. These are lotterybased panels of regular people who are participating in policym. So uh one really big example is the Irish citizens assemblies that's pictured here. Uh and so it has um become a almost codified practice in Ireland to use these citizens assemblies to tackle really difficult national issues like really polarizing ones that uh the nation was stuck on for

many generations such as abortion and same-sex marriage. So, it was the citizens assemblies that really broke through on these issues. Uh, and then another example I'm really excited to talk about today is the LA Charter Review Assembly. So, we're actually right now doing a Los Angeles civic assembly uh you know a true cross-section of Angelinos um an impartial group of people are learning about and deliberating on

some key issues with the charter review the charter reform and the charter is LA's uh document of governance. So, it's basically our constitution. So this is a really big process. We need to have, you know, do you want to see the current politicians who have power deciding that process? Do you want to see them, you know, at the front seat? You know, I don't. I would like

an impartial group of people to listen to the facts and deliberate and make a recommendation on that. So that's why I'm really excited about the LA Charter Review Assembly. And if you want to learn more about that, you can go to um um Rewrite LA, Rewrite LA, and our local group is called Public Democracy LA. Okay. Um, so why does randomness work? You know, we've talked about

this already quite a bit. Uh, randomness protects against politics and campaigning. It protects against factional blocks. Uh, it helps us deal with dominating influences and interests. And of course, it protects against this issue of a lack of representation. um and it produces a statistically valid cross-section of the community. This serves to connect the constituency's actual opinions and lived experiences with with policym. You know, I think we

often have this issue both with public governance and internal organizations where the the real opinions uh are not being represented because um we don't have this type of like uh like statistically valid process to assess you what people's opinions are, you know. So like the the decisions and the policym is there's like a gap between that and what the actual constituency wants to see happening. On top

of that, civic assemblies have a really important component which is a deliberative and transparent process. In civic assemblies, people are paid for deliberation over days or months. This allows deep understanding and dialogue to inform decisions. You know, unlike many types of voting that we might do in in the general public where, you know, we we are not able to have a deep understanding. we're not able to

find the time in our own lives to do the deep research and understanding necessary to make an informed choice on on a particular vote. So, you know, we're paying people so they can take that time out of their day and study and deliberate on a particular issue. Um so one concern that people sometimes have about certues experts. Um this is not what happens. Expertise is actually elevated

when it is presented in front of impartial people. Okay. So you know I've seen this time and time again from my past doing housing advocacy. You know we have great experts in urban planning and housing in Los Angeles. you know, people paid like we the general public pay these people to serve in government and provide their their expertise in terms of like how much housing we actually

need. Like what are the actual numbers of housing? But then, you know, I've been told by some of them they're not able to present the actual numbers to to the politicians that, you know, serve on certain decision-making boards about housing. numbers because they're too big and scary. Okay, so this has happened time and time again. And then when they do present watered down numbers, those the politicians

choose to ignore it and they say these numbers are too high. Um I think our we are devaluing experts currently that that is throwing expertise away that we pay for. we pay for them to come up with these uh studies and then we just ignore it in the political process. Uh so expertise gets a fair chance when it's in front of impartial people who don't they're not

worried about where their campaign money is going to come from. They're not worried about, you know, their what future office they're going to run for. You know, they're not they're not worried about like what favors they promised for their endorsements. So at certition USA we put certition into practice for for our organization. We we live cert. We have a lotterybased board. Um this is a snapshot of

some people who served on that. And um we do open meetings. So they're very transparent open board meetings. Uh and an important part of that is we need a strong onboarding process for when people are lottery selected onto the um you know instead of like the the usual way where you look for people who have experience you know we're getting people with no experience straight onto our

board. So we give them good onboarding. We do staggered terms so that when we get a new cohort on we're not starting over again from zero. You know some of that knowledge is carried on. Um we also have an advisor's committee of experts and we do training and agreement to legal duties. We do screening for basic disqualifications and we have strong rules about conduct. And so by

living it, you know, we've discovered through direct experience the strengths and challenges of of this of doing lotterybased governance and we've discovered that we need a system that's resilient to strong or difficult personalities. It's just that and it's a totally solvable problem. you know that's why we have a strong code of conduct process and we have dispute resolution processes. Um we've also discovered that we need that

advisor's committee like we need that expertise and off officers and others who can who can provide that guidance to people who really like don't have that experience. You know they um uh they're bringing that you know representation and impartiality and then we also need to pair that with expertise. Um it one awesome thing that I've seen happen I you know I didn't even really expect it is

that we get people involved and engaged who might otherwise not be you know so sometimes people just kind of they're they're sort of members or they're on our email list. They're they're like oh this is an interesting idea or you know they're just like not motivated to be that involved. Okay. But once they get lottery selected onto our board, it's amazing to see they they transform into

real committed champions for the cause and they they bring all these ideas and projects. Um, and we would have no idea, you know, if we just like sent out a general call, oh, who would like to be on our board? Or, you know, if we we tried to hand select the people on our board, we'd be missing out on all of that. And so we get that

through this lottery selection. And also, we've discovered that regular you know, non-experts, they're more capable and bring more fresh ideas than might be expected. And we're like shuffling those new ideas onto the board. For example, you know, we we've had one board member recently who's a physicist and he he has no experience with organizations, nonprofits, or movement building. Um, and you know, you think like how could

physics, you know, what could be the benefit of phys physics on a nonprofit board? But it's amazing to see like his perspectives, his ideas, they're fresh and new. We would otherwise never have his ideas on the board if we were just only looking for nonprofit experts. So here's some ways to implement sort in an organization you might be involved with. So you could go all the way

and do it like we do with a lotterybased board. That's one option. And you know we are here to support that. um we can pro provide consulting services and guidance and materials. If that's what you're interested in doing, please reach out. Um you can also implement it in committees. So maybe you have like a particular project committee. Um you know, you could explore doing a lottery selection

for that. Another option is dispute resolution. You know, it's almost like a a mini jury. Um, if you need internal dispute resolution, doing a lottery based process is a great way to do that. Um, also if you're doing internal reviews, you need some like internal oversight, that's maybe where this is most beneficial because, you know, you want to get that um that unbiased impartial perspective. And again,

if you are interested in any of these, like implementing them, you have some ideas for how to do these, um, we'd love to work with you. Please reach out. So, how can you learn more and get involved? Join us at certitionusa.org. And again, reach out to us. We have models, training, and resources that can help you implement certition in your Okay. Thank you so much. Okay. Now,

I'm gonna Thank you. >> Okay. Oh, would you like to take a question? Um the question is is bike shedding. Yeah, that's that's a great question. So the question is what do you do about this problem where you have a group of non-experts on you know considering a very technical issue. How how do you avoid the possibility of that group kind of gravitating more towards um things

that they are familiar with? Um oh I think maybe George has some thoughts on >> here. Okay, for general purpose, you know, uh, governance like boards of directors, those are not specialists. If specialists are needed, they are elected by the boards. In classical Athens, for example, the position of admiral was elected by the board. You didn't randomly select an admiral or a treasurer. So, if it's a

special committee on leakage in the steam generator, those would be steam generator people picked by the Um, also I think you know sort of the the science and art of civic assembly facilitation is is very deep and and there's a lot of different organizations out there with with different approaches and different theories on this. I think you know they all do an amazing job. Um, and if

if you'd like to connect with that network or learn about it, I'd love I'd love to help you do that so you can see what they say. But I you know I for example I think one thing I I've seen written about um by the facilitators is the need for um for a good remmit. Okay. So it's like they often say the civic assembly we have to

be very clear on what the output needs to be from the civic like where we all need to go. Um and like it should be specific and clear but you know we also need to balance that with with some flexibility. if you know the civic assembly needs um to kind of think outside the box. So, I think that's that's sort of like an interesting challenge, but I've

seen, you know, currently the LA uh charter reform assembly is very focused on the question of um the size of the city council and um I know that's like a new topic for for a lot of the people on the civic assembly, but I think like because they're getting a lot of expert information and a a lot of that factual information, I feel I think you know

the the discussion has been very focused on that. Um, do you do you have a question also? Yeah. >> No, please go. Yeah. >> Uh, so I'd like to try to kick the tire from the other side and focus on this uh this idea of a committee of experts. uh what mechanisms I guess a sort of like Ulisses packed type functionality can be built in to the

system to prevent for that committee of experts to become an entrenched uh sort of a cabal of of people who push through their ideas and push aside those good ideas that you said lay people may come in with and put on the table. Does that make sense? >> Yeah, it's a great question. Um I think it's a really good and important question and um something we think

about Um but uh so um when we talk about large public civic assemblies, one idea I'm a really strong believer in is having what is called nested assemblies. And so the idea there is that you could have sort of like an oversight assembly. Um and you know they they do this in what is called the OB Belgian model, the East Belgian um permanent civic assembly that is

a part of um the city's governance there. Um so there's a there's a lot there's a lottery selected um body that there one job is to decide what issues should be the the subject of like the next civic assembly. Okay. And then you know we could go a little further I think and have you know oversight assemblies also that oversee the process and facilitation. Um and you

know often when the civic assemblies are done they um the the gold standard is to um allow that assembly maybe like seating them with some initial experts but giving them control over like what experts to call you know. So especially you know with long-term ones like they they find their own experts and so I think that that helps reduce the issue that you're talking about. Um but

I think it's it is a challenge I think in smaller organizations where you can't benefit from these from nested assemblies or larger processes. Um so I just I think the solution is maybe uh you know like a smaller version of what we do in like the larger civic assemblies where we need to make sure that um the the board the lottery selected board and like these lottery

selected committees have a lot of latitude to go out and find like new information new experts and that they're sort of like it's not the experts controlling them it's the other way around. But I think it's like a really great point and more challenging in the smaller settings. Um Okay. Yes. And then Okay. Yeah. Um maybe like I'll just take one more question and I can take

more at the end after Liz goes because I know Liz has a lot of Yeah. Yeah. Okay. Let me just hand it over to Liz next and then Okay, for everyone who had questions for Leonora, please make a note to self. Remember your question, you can ask them at the end. Um, I'll be quick so that we can get right back into it. Um, this is the

segue slide between me and Leonora. We both care about ensuring that community governance can become a form of scalable and sustainable decision-making that's resistant to capture by billionaires, authoritarians, and corporations. So, my name is Liz Barry. I'm the somewhat recent executive director of Metagv, which is a nonprofit originally stood up by Larry Leig at Harvard, the creator of creative com uh yeah, creative comments. Um, Metagv is

a both a community and that nonprofit organization. We kind of think of ourselves as a research to infrastructure laboratory because we start with a lot of academics and experimentation um around uh research topics and we frequently move into advocating for um technology to become infrastructure that serves everybody for self-governance in a digital age. some of the things we've done are analyze a billion dollars of investment into

digital public goods like how did that investment go? So, so um we created standards that are that currently 19 billion in onchain value are running through um with the uh like web 3 um like crypto there is a corner of web 3 that is not scammy and that is where this organization came from before I got there. Um I'm terrestrial which you'll see my work in democracy

that we come up with but anyways what is up with Metagv? We worked literally with Switzerland to stand a public aai.co. So we built the inference utility so you can chat with a sovereign LLM that was transparently chained instead of um uh using OpenAI or Claude. Um me personally, why I highlighted it, I'm co-directing around $2 million of investment into democratic infrastructure and advising on $10 million

more. Um what do I mean democratic infrastructure? Well, I'm going to start by just telling you about one tool. Now, in some crowds, this tool, Polus, is widely known. Can I see hands if you've heard of Polus? All right. Um, for you all, thanks for your patience. Um, I'll take it pretty quick. Um, I have been working with Polus. I co-founded the before Metagv, I co-founded the

computational democracy project with the creators of Polus. So, I'm kind of OG in the Polus community. Um, so here's my quick talk about it. Um, it it helps large groups of people figure out what everyone thinks about what everyone else says. Um, enabled by advanced statistics and machine learning. And I'll explain Voxit shortly. Um, in one slide, you go from statements on cards where you see what

other people wrote or write your own. You vote up, down, or pass. And out of all that sparse data um in highdimensional space we do some projection and C means um so we cluster people who share similar opinions and we figure out what people who think differently may yet all agree on the so-called bridging statements or groupinformed consensus. Um and this little technology which was a knowledge

transfer from computational biology, some genetics research that was also familiar with very sparse matrices and how do we detect pattern in that um has come over to help out democracy a little bit. And so already this open source software has been connected to a lot of different types of processes including what Leonora was just talking about. Um this is before generative AI. So I'm happy to get

into the use of AI and democracy at the end of my talk, but for now we're going to stick with machine learning. And in this particular system, all the content is written by humans and all the meaning is made by humans. The machines just help us see Um, so these are seen. So, we're now back in 2014 in Taiwan where the legislators um under the influence of

the mainland made a trade deal that was far more favorable to China than to the Taiwanese who elected their own representative representatives and people flooded out into the street. Now, luckily that country had invested in facilitator training for 10 years. So instead of looking like what most of the occupies looked like in the west, they organized themselves into small groups, passed a talking stick, took careful notes,

carefully help find the people who had the most divergent opinions talk to each other and figure out what their common ground might be, and then passed those notes around to all the other small groups that were doing the same thing. Yeah. Unimaginable. But I was actually I had just been flown in to keynote because of my previous group um called public lab which I helped create um

a global open science resource uh research community to collect evidence on polluters and bring them to court. Anyways, that's why I was in Taipei. So I saw it for myself and then that turned into um honestly like the global darling of digital democracy right now. In the green shirt, we see a very young um Audrey Tang, who's now Taiwan's cyber ambassador. I'm not sure what other country

has a cyber ambassador. And they used this u they reached for a bunch of open- source tools to scale up this crowd decision-making um to their nation. And the first thing they fought was um colonization of their transportation sector by Uber uh companies headquartered outside of their um their political boundaries but who which were extracting a lot of wealth and they brought together multiple stakeholders. >> Hello

All right. Polus has been used around the world and in Austria where only 84 members were be were able to be sat in a citizen assembly like what Leonora just talked about but they use this kind of software very lightweight you know look at what someone else wrote wrote your own agree disagree or pass to consult 6,000 more Austrians. Um, Urguay used it. Um, the uni in

the there was a moment where a corrupt political party, go figure, decided to put their entire party agenda, 150 points into a single yes no national referendum. They it had been in the legislature but they exploited the rules so that um the the legislature gridlocked and triggered a provision which is after 30 days discussion had to stop and it had to go to national referendum. And so

um folks from the university which ranged from computer scientists to communications experts um and designers stood up one of these polless systems. Oh, I guess my GIF is not animated. All right, that's fine. Um, and 16,000 people participated to realize that it was just one single provision about public safety that swung the vote. But otherwise, um, people would have the relatively poor people taken nationally would have

voted differently on issues around education and social welfare. Um, but they were afraid and that political party knew that. um bundled it all together. Um in Slovakia, you know, under um uh information assault from Russia, um a few years ago, their society reached a tipping point. 50% of people wanted to leave the European Union and the young people, in this case, a lot of young lawyers. So

there's a lot of different professions who find interest in this. They partnered with a national newspaper to hold a different kind of comment section and they just went issue by issue, you know, by labor, by cultural heritage to say, can we talk about these things without going into flame wars and they reached for polus and started actually building modifications on polus so that they could find um

what people held in common faster than they could find what pissed everybody off. Um, this has also happened in the states in Kentucky. This is in 2018. And recently there was a lot of press from Google Jigsaw. Google Jigsaw was so impressed by what happened in Bowling Green, Kentucky in 2018 um because of the newspaper running a pace ahead of a town hall meeting, by the time

the politicians got on stage to start whipping the crowd into a frenzy of national culture wars, the crowd itself booed them and said, "We don't want to talk about that. We know we all want to talk about fixing the dangerous traffic pattern on Main Street in Allen and also we want to stop doing business with Comcast and we want municipal broadband. That's what we're here to talk

about. And [clears throat] the group knew that there's a kind of a group self-portrait aspect to this. Okay, the UK is doing it. Um fish are very important. I have to go fast here. uh the UNDP, Pakistan, Bhutan, Timor Lean 30,000 youth conversation across national boundaries and at this point even the parliamentarians were so into it that they got into this sort of promotional free ride situation

along with the facilitators who were who were helping the people who jumped in these tuk tuks uh to get a free ride learn how to participate in the conversation. And once in a while those people would actually meet their um parliamentarians. Wow, this was funny. So those used to be flags. These uh I think that says Taiwan, Great Britain, Finland, Netherlands, and Canada. Thank you. Twoletter codes.

Um so those are all nations whose staff have officially contributed to the repo. and then Taiwan, Great Britain, Finland, and Netherlands have been running it on government servers to do their Um, and Finland and Netherlands just got together. Finland made a new front end and Netherlands, no, sorry, Netherlands made a front end. Finland had a backend. It's not totally together yet, but there's a new repo called

Boxit that they're going to do governance like European governance over because it may end up heading into Euro stack um as part of the European Commission. Okay. Um thank you for listening to that. That was all my old job, which is just one tool as head of partnerships around the world working with peaceuilders, journalists, governments. Um, now at Metagv, I walked in and there was over a

quarter million dollars to convene a bunch of different tool makers. So now I just met 30 tool makers and counting to work on interoperability. And we did a bunch of granting. We got together in Berkeley. We built a bunch of stuff starting with flat file exports. And we um Oh, that was a video. Um we put we started organizing like what are the functions all these tools

are providing and similar to the UDA loop if you've run into that sort of air force um you know observe orient decide and act there's a sense of a self- steering system going through a loop cybernetics will be very familiar here I'm abbreviating this to a dole learn and decide loop and saying that these kind of collective intelligence softwares are mostly working between the learn and decide

phases And so figuring out what they all were variously possible to do and then we started decomposing them into um let's skip that into what supports the health of individuals as they are learning starting to contribute and encountering other people who then are learning from them and are encountering their ideas. And what parts of these tools support relational health, which Audrey Tang, Taiwan cyber ambassador, and I

define as the information that is between us. Um, and here you'll probably recognize some um process, the combinatorics powers a lot. Um, finding the paro front um features in some very interesting things. Um, once in a while you just run into good good oldfashioned deli. And then the bridging algorithms that I mentioned of course now AI is showing up everywhere and in another talk I will tell

you how I think about that. Um so we had to get all these different tools to talk to each other which I refer to as ontological spaghetti. But then we got together in Denmark and then um we got advice from the W3C and now we've licensed an ontology that you can find in Metagav's repo. And these are just the very first switchboards that we're now able to

make. Um so the government of Scotland showed up with a million great British pounds and um wants to have this kind of deliberative switchboard where people can take what processes they need and into making um the collective intelligence workflow to make policy. And this is my next to the last slide. I want to say underneath all of this, we're doing governance, but what's actually happening is that

we're learning. We're strengthening our ability to rule ourselves. We're getting better at this. So, whatever the policy output is from one process, the thing is is that we went through it as humans and we gained capabilities. And that's when I design these systems. Again, I'm co-directing $2 million of investment into systems like this in the states and $10 million internationally. Um the impact ultimately is on increasing

our ability um to collectively determine our own fates with assistance from AI um not replacement. Uh, keep your eyes out for a new speculative fiction contest we have going called the Protopian Prize. This is not yet launched, but that'll be kind of a fun thing coming out from Metagv. We have Kevin Kelly on the judging panel, the founder of Wired. We have Gideon Lichfield, the former chief

editor of Wired. We also have um Annaise Nuititz, Ruth Anna Emmeris, and Carl Schroeder. Um, probably some of you all know some of them. Okay, thank you. >> So, let's do Q&A now. How much time do we have remaining? >> 20 minutes. Okay. So, um let's do Q&A for both u me and Liz. Okay. Um and let me bring the mic over because I know it's hard

to hear people's questions. Bill Buckley once said that he'd rather be governed by the first 2,000 names in the Boston phone book than the city council. But um and that might fit in what you're talking about. But what bothers me is randomization is not random when you have to be self- selected to be on the group from which the random people are selected. So therefore, it's not

totally random. There's a group of people who won't get in because they won't volunteer. And unless you make it mandatory, you can't make it random. >> Yeah, I no disagreement for me. I would love to see a situation where it's mandatory, more like jury service. Um, you know, on our our board, um, you know, we we talked about doing our lottery selection just from our membership. Um,

but we actually opened it up to anybody, any email that we have. So, it's like actually any email that we have. I wish we could do a lottery selection of just like anybody. I don't I don't think we're quite able to do that. I think that's like something we would all love to do at Certician USA. Um just like have a true random selection. Um but yeah,

I totally totally agree and it shouldn't be self- selecting. Another thing about that is um we have a really interesting lotterybased process here in Los Angeles um called the civil grand jury and they do an amazing job. I actually I mean I would much rather they govern than what we have and their recommendation to just like go to um the LA County supervisors who then just promptly

ignore all of it. Um, so I mean I would prefer that body be in charge, but then that's similar to what you're saying because they actually they have to go through a process of selection to even be in the lottery pool. So it's it's not a true random selection. I think what's happening with the with our local um civic assembly with with a true cross-section, I mean

like like a real cross-section of Angelinos of like all different income levels and backgrounds. I think you know Um, okay. Next. Oh. Um, let me George, since we have an opportunity to talk so much, let let me go to someone else. This um Yes, you had a question. >> Yes, it it was just previously about um the charter reform commission. Um maybe you could talk a little

bit about what you guys implemented for that because I actually went to a city council meeting recently and that as far as I know I don't I don't know I'm pretty stupid admittedly so I don't know a lot about uh uh how our government works um locally um but as far as I know the charter reform commission was selected by council members and I'd be interested in

in hearing how your process just maybe inform their decisions or I I'm just a little confused. Sorry. >> Yeah, you know, I'm confused, too. Um well, we got support from from the charter reform, the appointed charter reform commission. We had their support to do um the civic assembly um and they endorsed it. Uh but you know I I would I you know I would love a a

bigger commitment from them to follow the recommendations. Um, but I think they struggle with the same problem because it's like they don't have a commitment from the city council to even follow their recommendations and I think they are very concerned about that and that's something like I think that's why they see value in the civic assembly because they they want more attention and accountability in the process.

Um, but I think it's a really good point and I think everybody's a little confused about that and you know I wish we had a better charter so that we would have better systems in the first place. Um, and I and I wish there was the political will, you know, and I think, you know, the more we spread the word about this, the the closer we can

to get get there, but it's like I think there needs to be more political will to turn these from just like recommendations like the civil grand jury that go nowhere to to real accountability and and real impact. Um, and I also want to uh encourage people to ask questions for Liz. are also doing questions for her. Did you Did you have any comments on those questions already?

By the way, I I should have checked with >> Just I think it's great that you asked that. I mean, we're here in LA. There is this major process going on on revising the charter of the city right now. So, I mean, it's like it's a constitutional process. So whether you want to get involved um you know through Leonora and Rewrite LA or just start independently attending

some meetings um something's going to end up on the ballot in November. So there's a lot of room, you know, for a for a person who has some skills to lend, there's room for you to help now in some way um to make sure that ballot referendum is as good as it can be. Great comment. Let me go to this question next. >> One quick idea on

the um actual randomness is maybe getting together with the UBI people. Uh democracy dividend where the requirement to receive the dividend is you're eligible for sortion. Um that's a quick fix. uh on um Polus uh there are aspects of it that look similar to Lumio. Uh are you familiar? Have you talked to them? >> Yeah, they're the same age. Like so much like like Lumio, Ethus all

kind of like sprung out of the the social movement uh social media movement age. um where there was just so much turnout and everyone got in the street were like we're all mad. Okay, but what do we want? We don't know and we don't have enough mechanisms to figure it out. So there were a lot of there were a lot of like collective intelligence tools that kind

of sprung up at that time. And so yeah, the the case stud the water that those early tools carried is what's coming to help in this moment as everyone's realizing that social media is tearing us apart and that the wrong algorithms are running between us. Um yes. So yes yes >> great >> and and also a great idea. Thank you for the UBI idea. So I was

thinking in particular a lot of the cases you were describing where there there was significant amounts of money that was getting allocated uh based on what some of these groups were deciding right is that fair to say >> it grew to be that yeah fair fairness. So the metaphor that is most immediately available to me to understand what you guys are talking about is the jury selection

process. And one of the things about juries uh is that they have very strict rules about jury tampering to protect the integrity of the jury. How how did they protect these effectively juries that were making these decisions when there's that much money at stake? There's always going to be someone trying to manipulate uh the the group that's actually making the decision. I would imagine online. >> The

hybrid ones you're talking about. So yeah, the technology I'm talking about right now is um really um wrestling with trying to figure out how to move from being a product essentially to being infrastructure and having um identity and data storage solutions that are commensurate with the risk that vulnerable vulnerable people take to exercise political speech. when they exercise political speech that could get them targeted, attacked and

in some places killed. So there these are why the investments just the very earliest investments are coming online for that and it's bringing us into contact with a lot of digital human rights people and decentralized identity folks and like at proto and first person network and um uh MOSIP identity solution that the open source version after India success so there's a lot there's a lot of exploration

going on right now um to make sure that these system because the that want to assert that their results are legitimate need proof of humanity. They want to know that it wasn't bots participating at scale like wow wow so many bots. No humans with so many humans. On the flip side, the participants uh protection and probably need data ownership. No more consent. No more after the data

is out in the system. We probably need to move to doing the stuff on the other side of wallets and Taiwan's already doing that. So the need for proof of humanity and data ownership are wrestling with each other and it's thrashing the entities that are willing to commit half a billion dollars to the result. So that's why there's some philanthropic investment coming in right now um to

try to get to something that's like more like engineered like a system appropriately and is more like getting towards infrastructure than than the sort of the products and the proofs of concept that we have now. Um, also I just wanted to talk a little bit about juries and the rules around juries because I I'm actually a lawyer and um you Yes. So I I remember in law

school we used to talk about the problems with juries a lot. Um and uh you know I definitely wouldn't say I that I agree with you know all of the ancient traditions that we do around juries. you know, if we were designing a new process today, I think it would be a bit different, which is why civic assemblies are a bit different. Um, because, you know, we

often would talk about the problem where um because it's it's so strict about what outside information and knowledge can be brought into the courtroom, you get an unintentional consequence, which is very like a selection of like very ignorant people onto the jury, which is, you know, not the intention. Um and um but uh you know I what I'm seeing with civic assemblies is that the facilitators and

the practitioner organizations are taking a different approach where they're uh they're encouraging the delegates to um to seek out outside information. They you know at the LA Charter Reform Assembly um all the delegates have access to a laptop and they're like looking up. They're like, "Oh, what happened with, you know, that corruption scandal?" Okay, they're looking it up. Like, that would never be allowed in a courtroom

at all. Um, you know, and I think the idea is that they they are in charge of like seeking out that information and deciding what information gaps they have and like getting that information, like getting those experts, which is different from a a jury. Um uh however I think it's it's an interesting idea like is that does that create vulnerability to outside manipulation? Um you know I

think there's like that's an interesting question and you know it could you know I think what's exciting about this is that we're charting a better path for democracy right now and you know I think there's a lot of uh ideas about facilitation and the process. So I you know I hope people become be a part of this and help us think through you know like that balance

about do we want to have no outside information or like do we want to encourage that and like how do we you know structure that so that it's resilient against outside manipulation. I think these are like really important questions in the space and I hope you all join and be a part of that. Um, another question. Um, I'm gonna go to the people. Sorry. >> Just a

quick question about like uh the marketing or visibility of something like this or like the broad systems you were describing. like you can get a lot of data and maybe have a perfect infrastructure system, but how do you inform like the average person? What's there? Do you have a vision for that? >> Um, considering the state of things right now, there's a lot of experiments going on

and and generally the the attitude of the CNI run in is that we need more governance experiments in general. So, some of the big ones going on right now, um, there's the Bloom project that is seeking to organize like local civic who are who themsel across themselves have diverse memberships and they're saying, "All right, let's support the existing durable local organizations with local leadership who are known

quantities. have them use their convening power um so that all their members who are on different sides of issues can talk to each other. So that's one model. There's another model called the forum or A1. so I'm direct I should say I'm directly advising Bloom like directing I'm advising the forum A1. Um their model is um to go at a state level. They're working in South Carolina,

Nevada and New Hampshire which has a long tradition. New Hampshire is very participatory, but their idea here is to use um use their ability to reach um electoral candidates from all parties to have them co-convene the voters to talk amongst themselves. Do some agenda setting like form um an agenda ahead of the election so that no matter who ends up winning the election, they're already pre-bought in

to the people's agenda. So the people get what they want no matter who actually gets seated. Um and so that yeah, that's a kind of a big experiment going on. Um, and then there's a lot of different experiments um, with media partnerships in general. Like a couple I talked about like in Kentucky, what really made the difference was the newspaper decided who naturally convenes the audience. You

know, democracy doesn't work without journalists. So, I think there's a lot, you know, we're talking about, all right, civic infrastructure, where does it really land? You know, you heard me say, hey, some governments are building this up. You know, I can feel ambivalently about that. Um, sometimes it's building journalistic capacity. Okay. Sometimes it's building kind of civil society capacity. So, there's a lot of experiments going on

right now. >> George, are you okay with me getting some new new voices first before going to you because I know you've had your hand up a long time. So, Okay. I was just going to say it's interesting that you mentioned New Hampshire, South Carolina, Nevada, early primary states, but what you describe with getting the voters out to talk amongst themselves sounds a whole lot like the

Iowa caucuses, which I mean, same time frame, I guess. >> I bet. Yeah. So, Kristen Hansen of the Civic Health Project, I'm sure that was not lost on her. I think that did go into the design criteria. Um but their goal is to um reduce polariza. So I I can't I can neither confirm or nor deny, but some of the impacts they're looking for is that um

it reduces polarization among the electorate as well as gets more bills passed in our legislative bodies that have ground to a halt. They're not functioning. It can't get business done. And we're still while we are still a republic, we need to try to get these institutions working. Um, meanwhile, there are lots of other experiments, you know, like very like rogue people's political parties that also imagine constituting

themselves very like Chris Kely recursive public recursive public style where they're making their own means of assembly and grassroots funding for a party for which they are setting the agenda and running candidates on completely autonomously. Um, yes, it's a little it can get a little wild and then the bots. >> Next question. >> Thank you. Uh, I'm Arand. I represent an organization called Equal Vote and we're

really focused on uh voting method reforms, but um most people in Armenia I'd say are also very enthusiastic about sort and how to use that more. I'm curious uh your thoughts long term um what the ideal blend would be between voting um method or voting voting reforms and uh civic assemblies. Are there certain elections that you think are particularly ideal for civic assemblies and others that are

better for voting reform or is it more of what Liz was saying where you like just experiment with all reforms and all elections to kind of figure out the the ideal approach from there? Yeah, I'm curious what your thoughts are there. So, it's kind of a tough question because um you know I think there's different opinions on this in our movement and so you know I don't

think I'm I I don't necessarily represent that with like my opinion. So I just want to say that but I I'm personally I'm not you know I'm I'm happy to see uh voter reforms but it's like to me elections and voting um are just so problematic. Uh I you know I I am very concerned about voting systems. Um I think uh you know there's like this topic

of conversation comes up a lot. um where I think like a lot of practitioners in the field they talk about like okay we have these um many publiclix in the form of civic assemblies but there has to be some involvement from the macro public you know whether that's like the recommendations and go to a referendum vote something like that or you know I think like Liz is

doing a really great job kind of using pollless um you know to to make more people a part of the conversation like use that data in the civic assemblies. That being said, I just like anytime there's voting, you know, I, you know, I've just, you know, I have some experience doing like political organizing, like I just know that it's, you know, it's so ripe for manipulation. Um,

you know, so anytime that is involved, like anytime, you know, money and organizing can influence the outcome, I'm just I become very concerned. Um, I'd rather see a setting where facts and expertise are presented in front of an impartial representative group like every single time. So, okay, sorry, [laughter] I'm not giving up my vote. Um, I put it like this. So, the bullets to ballots transformation is

really, really critical. When I talk to peaceuilders internationally, the vote is not about getting equality policy outcomes. The vote is a conflict management approach, right? We it reduces violence when we have the vote. So, I don't want to give up the vote. Also, in terms of a decentralized system, the the voting in the United States is like a great example of a decentralized system. Um, far more

decentralized than anything we're doing with software. So what I would say um is realize when are when governance is doing conflict management when governance is doing information processing and when governance is actually developing human capacities for self-ruule so that we can live in democracy so that we are the people who can have democratic governance. Great. Okay. Um, do we Oh, um, George, did you want have a

question? Because Okay. Do you want to do a question? Okay, let me just He's been waiting this whole time. So, um, if you had, let's say, if somebody asked you for advice, let's say there's a professional organization where parts of it help set standards for industry. Uh, what would you propose? like you give me like a skeletal structure that you'd propose for something like that. >> You're

saying uh like a standard model for civic assemblies. >> Let's say there's a dentist society and they're proposing standards for uh tooth drills. You know, they're setting standards for a whole lot of stuff. There's a lot of influence coming in from manufacturers. The people who are in the society don't know who to vote for to direct the society. It's influenced by manufacturers. How would you change the

way that professional society would operate using your methods for both of you? >> Oh, yeah. Well, um I think you know where we're going with this is definitely I would want to see professional societies implement the reforms that we're talking about the lottery selection certition open governance you know um to address a lot of the problems we talked about you know um uh donor influence you know

sort of like lobbying that's happening you know backroom dealing that's going on like you know personal uh connections like power, you know, people using their power and influence to kind of control and stay on in those positions. So, um definitely I think lottery selection is the way to break through that. Is that kind of like what you're asking, George, or am I misunderstanding? Do you want to

say like do you want to say what you think? >> Um selection from members. Let's say if it's a let's say if it was a software community uh let's say for a major uh open source project people who have committed a certain number of lines of code uh would be selected at random to be on a board of directors rotated uh and they would for example be

able to pick the lead designer or something like that. uh and if it is a standards committee they would pick people from the specialties that apply to those so that the status committees represent the developers rather than the manufacturers who are funding the >> Great. Okay. Thanks. Um are we at time or did you want to >> Okay, I think we're at time. So, okay. >> testing.

>> Hello everyone. Thank you all so much for coming today. Uh my name is Amy Iris Parker. I'm from UC Irvine and today I'm here to tell you all about uh the history, the current state of the art and the future of censorship evasion. Before we get started, uh, two quick little disclaimers. A lot of this talk is going to be, um, primarily US focused and using

US references. Uh, most of the information in this talk is globally applicable. There are plenty of examples of everything mentioned in this talk occurring in other countries as well. However, uh, to make it more known to the current audience here, most references are to US things. Uh, second, uh, if you have any questions, I would love to hear them. just please save them until the end of

the talk so we can get through this quickly. Before we begin, a little bit about me. Um, uh, as I said, my name is Amy Iris Parker. I am a first year PhD student at the University of California, Irvine. Uh, my main research interests are in distributed embedded systems, computer architecture, uh, networking and, uh, the social impacts of computing. I'm currently in the Dut research group working

on kernel memory proactive compaction. I was formerly at CC Fullerton where I studied evasive VPNs uh and Kimu's intermediate representation pipeline which if you would like to know more about that come to my talk tomorrow. Um also the secretary of Orange County DSA. I do a variety of different tech things for that chapter down in Orange County. Um and I have been a decade long Linux user.

Uh I currently use Nyxos for my laptop and my home stuff. I use Proxmox in my home lab. do a lot of task automation and use Linux everywhere in my day-to-day life. Uh in addition, uh portions of this presentation have also been at uh US6 security 2025 in Seattle and Latincom 2025 in Antila. Um the latter had a paper published from it. I would encourage you if

you want to know uh a lot more about the current state of the art with regards to evasive VPN protocols in particular to go read that paper. references will be included at the end of the presentation. So, an overview of what we'll be going over today. Uh we're going to uh talk a little bit about early censorship in the pre- internet area. Uh how different media were

censored by governments. We'll go on to basics of uh commonplace censorship, IP blocking, DNS blocking, protocol blocking, uh as well as the effects of laws such as the Communications Decency Act. We'll then move on to how censorship grew from the institutional level at universities or libraries to the nation state level with entire governments censoring the internet for entire countries. Then we'll look at the rise and fall

of VPNs and onion routing, how those were used as a primary censorship evasion tool and how they have fallen out of favor. And we'll look at the current state-of-the-art which is evasive protocols and also recent research on how sensors are learning to identify evasive protocols and block them. And lastly, we're going to go over the future new encryption methods uh new protocols, alternative media for distribution, and

as I'm sure you all have heard about uh the rise in age verification laws. So, a little bit on early censorship. Uh, before the internet, um, there were lots of different ways that humans communicated. Uh, ideally, a lot of communication was done in person. Um, and in countries that had free speech protections, inerson communication was by and large protected. However, as people found new ways to communicate

such as telegraphs, telephoneony, books, radio, television, all of those, uh, government started looking for ways to censor that information. if you were not in person, they wanted to find ways to limit the mass spread of information. Um uh yeah, and so uh a lot of the methods that were developed early on for censoring these are still fundamentally present in what guide censorship today. Um uh f some

of the first censorship that happened here in the US was male censorship with the comtock act. Um uh mail was common place distributing lots of information, societies, groups would all send people uh information through the mail. It was cheap. It was for the time relatively fast and it was government run. Um that is one the perfect opportunity for distributing information and two the perfect opportunity for the

government to censor information. Uh the comtock act was passed in the US to prohibit obscenities. Um the initial target was supposed to be pornography sent over the mail. However, it would also be used to crack down on descent information about uh abortion and a whole variety of other things. The Comtock Act is still actually on the books today. Um and it continues to cause problems if you

ever want to try and send information via the uh other media that have that were censored in the past. Uh books such as Ulisses um uh uh telegraphs were censored during the Civil War to uh prohibit the rapid distribution of news about what was going on in the war uh to protect troop movements. Uh Telephony continues to be censored and has been censored for as long as

it's been around. A recent example is uh Oman's sweeping um regulation on telephon. uh SMS and MMS have been censored for as long as they've existed as well. Movies and television were previously self-censored uh by the industry, but have also been censored directly by uh governments both national and local. And the radio is also fully um censored and regulated uh here in the US by the FCC

and by the national radio authorities of other countries. Um libraries were also how a lot of the early ideas about censorship started. Uh the very first uh book ban to happen in the US was in 1637 in Massachusetts. Uh there was a book that went around that was criticizing the Puritans government at the time that was banned. Um it was the first known case of book manning

in the US during the 1950s. McCarthyism and the red skier led to a lot of libraries being forced to pull books about communism as well as other things like queer issues um uh under threat of being reported to the House on American Activities Committee which would investigate people's lives and make them a living hell. Um, in 1982, uh, another case was Island Trees versus Pico, which, uh,

the Supreme Court precedent in that set the foundation for what we now know as the modern bookman ban movement with plenty of states having varieties of methods, uh, to ban books in school libraries. Um, now, how do people evade these? There's three general categories of how censorship evasion was done pre- internet. There are ciphers. Ciphers make it so that a sensor cannot tell what information is being

transmitted. These work well when you're trying to conceal information, much like encryption today. However, um a sensor can easily just say, "We're not going to let any ciphered messages get through." Uh a more robust way is through using code works, trying to slip in information that you normally wouldn't be able to send into otherwise completely permissible information. Um, we'll see more about the idea of code words

when we get to masquerade protocols. Uh, and lastly, there's alternative media. If you can't go through the channels that have been sanctioned and regulated by the government, you go through other ways. Uh, we'll learn more about this when we get to synchroness. So, basic blocking ideas for IP, DNS, and protocols. Um if any of you have interacted locally with internet censorship at the local level, you have

probably seen it in institutions, libraries, schools, maybe your workplaces. Uh these often block websites that they deem inappropriate, obscene, um irrelevant to the activities of that institution. Uh two main ways that this is done right now are through uh DNS blocks where you can block either full domains or you can block um individual words or phrases in domains. So for instance you might want to um a

sensor might want to block an individual website but they also might want to block say anything containing the word porn, sex, drugs etc. That's a common block that we see in schools for instance. uh another method uh and if DNS queries are routed through um that sensor's DNS server, they can easily just refuse to return information or they can return fake information like what we see here.

This is Turkeykey's national sensor uh returning fake IP information for Discord. Um however, you can get around this by using other DNS servers having uh manual tables of routes or several other methods. Another way is through IPbased censorship. Um uh this works similarly to DNS censorship. You can have a deny list of IP addresses. And in the case where you have multiple domain names pointing to the

same IP address, whenever you block something based on a domain name, you can add that to your IP block list as well. Um the tools for getting around IP censorship are pretty much just having multiple different IPs through which you can access a service as well as more advanced techniques that we'll get to Uh there have also been attacks directly on the nature of the web itself

for uh censoring web content specifically. Um uh HTTP readers and web crawlers are very common. Um before the widespread adoption of HTTPS, web traffic was sent in plain text. So sensors could very easily look through what was returned in HTTP responses and see oh hey this contains blocked words content we don't like. we're going to terminate this connection. Web crawlers uh actively search the web for that

information and preemptively block um websites and other content uh that is not uh permitted for the sensors. Um uh there's also uh extending the idea of um the DNS censorship now that HTTPS is more common. While you can't censor the text directly and you can often try and you know work your way through DNS, a big problem has been the adoption of SNI TLS has a field

when you make uh connections to a server called the SNI field. Initially it was largely unused because it wasn't really necessary. However, we live in a world where we have lots of proxies um uh like Cloudflare. We have lots of load balancers and those rely on the SNI field to tell what service something is trying to reach. By default, all of your browsers likely send SNI whenever

you make a web request. While TLS payloads are transmitted um encrypted, the SNI field is transmitted in plain text. Uh so sensors can very easily block based on the SNI field. And for many modern web services, if you change the SNI field to try and evade censorship, uh you will simply be unable to access the service as the server you're making your initial request to has no

idea where to route it. Um, another option that some sensors use is trying to block protocols instead. Um, one of the earliest ways to get around people using alternative DNS servers was just to simply block DNS connections on port 53DP to all other servers. Um, but other more specific um, methods can be targeted. For instance, uh, some protocols like SSH or other VPN protocols are commonly blocked.

And in the most extreme example, you see protocol whitelisters that only allow certain protocols. Um uh most common example of this we'll get to later is Iron's protocol uh allow list that only allows DNS, HTTP and HTTPS during critical situations. Um now for a quick timeline of protocol based censorship. Um 1991 for uh many here in the US AOL became their first ISP, the first commercially available

residential ISP. Uh and in the early days of AOL, most services were heavily restricted to AOL managed content. Um that was heavily heavily scrutinized. Any user generated content was, you know, heavily filtered and monitored. And for any external content, they had a lot of direct filters on accessing that and restrictions on how you could access it. Uh in 1995, we saw the CompuServe and Prodigy cases. Um,

Prodigy did a similar thing to AOL of restricting access to external services and limiting what could run on its own network, while CompuServe did not do that. Um, Prodigy ended up getting sued by Stratton Oakmont. Um, if that name sounds familiar to you, you probably heard it from uh, The Wolf of Wall Street uh, for having users expose information about Stratton Oakmont. Um, CompuServe had a similar

case. Prodigy lost. CompuServe one because Prodigy was seen um because they were actively censoring content. They were seen as a publisher and not a platform. Uh this distinction is what led to uh what we now know today as section 230 of the communications decency act which has all of its own issues that could be its own entire talk here at scale. And then recently in 2023, we've

seen a revival of uh protocol-based targeting with WMS 2.0. Um uh this is used in China's great firewall. Um while most countries have had widespread adoption of HTTPS, um even some of the long holdouts like HMPG.NET for anyone who's familiar uh have moved over to HTTPS. China still only has about 50% adoption of HTTPS. Uh and so the great firewall has taken recent measures to do more

active deep inspection of HTTP connections for content that Chinese sensors uh do not feel uh should be allowed to reach citizens. So from institutions to nation states um the first the uh the first country to do national level internet censorship was contrary to popular belief not actually China. It was South Korea. uh as South Korea started to modernize their business access uh to the internet, they also

realized, hey, we should uh probably, you know, make it so we can get control of what's on, you know, what our citizens are accessing on the internet. Um this was during a time of heavy tensions as always with North Korea and they wanted to make sure that the internet was not going to be used for spreading information from North Korea. Uh to this day, South Korean internet

censorship laws are still on the books. They are not as strict as what we see in other countries like China, Iran, uh and Turkey, but they do still exist. And they were some of the first to deploy primarily again through DNS and IPbased blocking just at the uh BGP exit level instead of on local networks. China quickly followed suit. Um they created a whole variety of different

uh electronic national security programs called the golden shield. But the one that is most famous is the great firewall. Um the great firewall uh was initially implemented in the mid in the mid90s um to create a method for censoring the internet. Um uh you all can see this picture of jello stuck to the wall with a nail. Um that is because back in the 90s um US

President Bill Clinton uh thought that there was no way that internet censorship could ever be done at such a scale and so the US did not really push anything on it. Uh and he had specifically said that trying to censor the internet is like trying to nail jello to a wall. As we can see in this picture you can nail jello to a wall and as we

have learned from China uh you can pretty effectively censor the internet. Now, of course, some jello is eventually going to slip off this and some information does slip out from inside the GFW, but most of it is still there and most uh internet traffic in China is still censored to this day. Now, the actual structure of the GFW itself at its initial founding is not what made

it novel. It was uh the idea of having through local boxes at ISPs censorship that can span across a nation of 1.5 billion people. That was an idea that had never even been pondered at that point. Um, another technique is what's used in North Korea. Uh, North Korea's Kongyong is, uh, their internet. Instead of having connection to the broader internet, which only a certain few people have

through one specific BGP uplink, almost all internet traffic in North Korea is actually on the North Korean internet. Now, it still uses the same internet protocols that we all use. Um, if you were to bring your laptop to North Korea, um, plug it into an Ethernet jack, you would still be able to access the North Korean internet. You just not you would not have a link out.

And today, if you are a very strict sensor and you need absolute control over what information is accessible from your network, running an internet is currently the only way to guarantee that. Um, North Koreans of course have found other ways to get around the lack of internet access, mostly by smuggling information physically into the country, which we'll get to, but that is one of the main tactics

used at the nation state level is just never connect to the internet in the first place. Now, as we talk about nation state sensors, we need to understand the idea of collateral damage. Um, anytime you try to censor information or censor a protocol, you are inevitably going to censor some things that you actually want to keep open and available. This is generally unavoidable. Um, it is a

part of doing censorship. You are always going to end up cutting some stuff out. The question at the nation state level is how much can they do? can uh what modes of censorship can they use that keep collateral damage to a level where it is tolerable and it can be dealt with. Um uh even even packet loss rates of 0.6% 6% um as seen in one iteration

of the GFW from 2021 are so destabilizing to national internet that they were considered unacceptable by China which is why they rolled back some of their more aggressive models in the GFW. So on to VPNs and onion routing. Um, this is what most people are familiar with here in the US when they talk about basic level censorship evasion. Um, I don't want to put anyone on the

spot here, but raise your hand if you ever got the chance to use a VPN in school. I know a lot of you are on the younger or on the older side, but some of you some of you have. So, you you've had experience with this. And this is one of the the more basic ways nowadays of doing censorship evasion. VPNs uh create a tunnel from one

physical network to either a virtual network or that virtual network can be linked to a physical network. When you send packets over a VPN, um it's as if you are sending them out on that virtual network, which is often someone else's physical network. If that physical network is outside of your sensor, then you have now tunnneled through your sensor. Uh VPNs were not initially designed for censorship

evasion. They were designed to connect different networks and create virtual networks. However, because of their ability to effectively tunnel through they became an incredibly viable option. And uh in the early stages, nation states had to consider, okay, we have a lot of people who use VPNs for other purposes, research, connecting to offices, connecting to labs. Do we want to block do we really want to block all

VPN protocols? No. because that had too high collateral damage at the time. Uh, onion routing takes the principles of VPNs and ups it and increases it up a notch by routing traffic through multiple different nodes. Uh, in theory in an onion routed system, no one can know where uh the no one can pair end traffic sent out on a physical network with the user who originally sent

it. Um there is a class of attacks against onion routing called civil attacks if you control a critical mass of nodes that you can use. However, onion routing was uh commonly used as another way to evade censorship by creating even more distance within the tunnel between the output and who it came from. I said at the beginning that we had the rise and fall of VPNs. So

what happened in 2011? China uh started looking over the rates of VPN usage in their country and they realized there's not really anyone using this for employment purposes. It's mostly just people using it to evade censorship. And so they directed all of their ISPs to block all of the well-known VPN protocols at the time. Uh this would continue to be an ongoing battle, but China was the

first to do this at the national scale. Um, of course, this did not block all methods of censorship evasion, but it eliminated the viability of most common VPNs for doing censorship evasion. Um, going through 2012, 2013, sorry, you get to 2015, 2017, we see a lot more widespread adoption by other nation states of VPN blocks. Um, examples include Iran, Syria, um, uh, at some points. Uh, a

lot of different countries started doing it and eventually institutions started catching on and they started blocking these protocols as well. Uh, to this day, if you go to a lot of, uh, community colleges here in California, uh, you will not be able to use SSH or WireGuard or OpenVPN. Um I've gotten into many fights with administrators over this uh because they have decided that the uh benefits

of those protocols are not worth their inability to censor traffic on their Uh and then finally what was ultimately you know starting to show the death nail of VPNs for censorship of Asian use was in 2020 um during the early stages of COVID uh Iran deployed a full protocol allow list as I mentioned before and they only allowed uh DNS HTTP and HTTPS through um any other

protocol whether it was on a different port or whether it was on the same ports but running a clearly different protocol was outright blocked. Uh, and that got rid of most common VPN protocols. There were some uh that had been built by that time called masquerade protocols which did survive Iran's blocking, but that was um an early warning sign that regular VPNs were no longer enough. And

so that is why researchers and developers started moving to evasive protocols. Um, evasive protocols come today in two main forms. There's obuscation through what's called full packet encryption. The main difference between say a full packet encryption based VPN and a normal one is that the header information is also encrypted. Um, uh, any proper encryption should look like random noise. And so when you have the entire packet

including the header information being encrypted and servers just having to guess what the type of packet it is and decrypt based on their known list of users that can be really really effective um because you can't identify that it is a VPN protocol over anything else that you don't know of or that is indistinguishable at least in theory. We're getting to that. The other type is masquerade

protocols. Um because a lot of services like HTTP, HTTPS, DNS are very commonly used, people have the idea of well what if we just run our VPN through one of those protocols. Uh there are a couple different options that route uh VPN traffic over HTTPS. It has a lot more overhead. Um, and there's a lot of things that you can do that accidentally leak that you're running

a uh VPN, such as through the SNI being released, but it has been a viable method and to date there is still no general purpose way of defeating Masquerade VPNs. Um, there are threats to these protocols in general. Uh, the big six are SNI leaking. Again, this is one of the big things that has haunted a lot of evasive protocols is if they run over something that

uses TLS, oftent times they leak the SNI and at that point it's trivally easy to block VPN servers. Uh the second is protocol handshakes. Uh the earliest full packet encryption VPNs uh had handshake systems that would establish uh that a user was going to be connecting. Um, and the reason for this was simple because you can't decrypt the protocol information. You can't know in advance which user

is going to be connecting. So you would have to try essentially every encryption key that you know of uh in order to uh figure in order to decrypt a packet if you don't know which user to associate the key But these handshakes were often detectable. Um, and that has been the downfall of many different full packet encryption based protocols. Another one is TLS's three-way handshake. This is

not at the VPN protocol itself, but the payload. If you are connecting to say um on HTTP 1 or two, you are still using regular TLS. You're not at HTTP3 with quick yet. And one of the parts of a TLS connection is the initial three-way handshake for establishing connections. Well, if you have three really small packets um at the start of each connection and then you go

to a bunch of larger ones, but it's still some indistinguishable packet. Depending on the timings of those packets, you can actually determine that that was the start of a TLS connection in a VPN. Um there is a dissertation on this that was done uh that was able to pretty reliably about 98% of the time identify um any arbitrary um evasive VPN solely based on its encryption of

the TLS handshake. Um the fourth is through pop counts and bite sequences. Um the 0.6% 6% collateral damage number I mentioned was from when the Chinese great firewall decided to implement um statistical measures based on pop counts. Pop counts are the ratio of zeros to ones in any given bite or sequence. Doesn't sound like you should be able to do much with that. But what uh Chinese

sensors found out was that because encry uh because full packet encryption appears random, it has a pop count that is pretty close to 50%. Regular network traffic is a lot more regular and as such is going to be further away from 50%. So they create statistical models that would with pretty high certainty be able to detect after a stream of however many packets they were too random.

Um, patches for this have been implemented and ultimately because regular network traffic would occasionally be caught up in this, the collateral damage was just too high and China backed off. However, um, they have shown a willingness such as in 2024 to reactivate those systems in times where they would prefer greater censorship um, with some additional collateral damage. Um, number five is active probing. This has been used

by a lot of different nation state sensors. Um, essentially if after some statistical guesswork, you think a server might be an evasive VPN server, try sending packets to that server. Try starting a connection with it. If it responds at all, trying to send an error message or asking for more information, you've now confirmed what protocol that's running and you can block the IP of that server. Um

this was incredibly effective in China uh for combating shadow socks, one of the um more ubiquitous evasive VPN protocols. Um protocols have solved this by patching things that if anything is not a valid request, you do not respond to it. However, for protocols that still have this, it is still a major issue. And six is timing attacks. Um oftentimes if a server's response is too quick or

takes a certain amount of additional time that can be identified as um being characteristic to a given implementation of an evasive VPN protocol. And even though the data may seem random, the timing of when that data is sent can be used as a side channel for determining uh when an evasive VPN protocol is being Now I said earlier that we would get to one of the big

issues with full packet encryption VPNs and here we are. Uh as of last year research found that pretty much all full packet encryption based protocols are inherently compromised. The idea of trying to make your traffic look random seems great at first and in fact all current full packet encryption protocols are indistinguishable from random. However, regular network traffic does not look that random. Um, the idea was, well,

a network wouldn't want to block all protocols it hasn't seen before, but turns out there's a pretty standard tradition of how network protocols are developed that uh, machine learning models can pick up on. So with nearly 100% certainty and with almost zero collateral damage um it is possible for some machine learning classifiers to basically perfectly detect when something is an ease of VPN protocol or just random

noise uh versus when it is actual legitimate network traffic even if it's an unknown protocol. This is very very bad. Um and researchers are currently exploring new ways to try and do um censorship evasion that moves beyond pill pack and encryption. The only reason that uh these classifiers have not taken hold yet in China is because machine learning takes a lot of computation power. And while we

have plenty of AI data centers today, when you consider the model of the great firewall where you have millions of local boxes at ISPs, all of which need to be able to do all of the computations necessary for all of the internet traffic in a country. It's currently not computationally feasible to run these models on widespread internet traffic yet, but we're not that far off from it

being possible. Um, right now if you have uh the testing system we used for this was a Ryzen 7 3700X. If you use all available threads on one of those, you can process about 4 million packets a second under our initial test models. Um, more optimizations can be done to that and of course you can use um, uh, AS6 and other techniques to uh, get you know

much higher performance. But we are probably looking at a timeline of about 3 to five years before these models could be fully implemented at the level of say China's GFW or other count's national systems. And that's why it is very important um that researchers work on getting new solutions out as fast as possible which we'll get to later. So the future, what are we looking at moving

forward? Well, first thing is masquerade protocols. To date, there is still no no known general method for identifying when a masquerade protocol is being used. If you can avoid leaking things like TLS three-way handshake information, uh a masquerade protocol can be theoretically perfectly undetectable if robustly developed. And so more and more development continues to be put in today to using masquerade protocols. There's also some very interesting

new concepts in masquerade protocols. Um that diagram is from a another paper at USNIX last year uh where researchers explore developing new protocols on the fly. Um developing new ways to masquerade yourself into other protocols on the fly before sensors can keep up. Um it's yet to be seen whether that actually works well in the real world. um they do kind of rely on uh side channel

availabilities that are not necessarily guaranteed as we're going to get to later, but it does provide an interesting starting point for looking at further development of masquerade Uh next thing is community networks. Um for those of you who were here yesterday, uh I hope you got to see the meshtastic workshop. It was brilliant. Um uh a historically uh a lot of information before the internet and in

uh places with very little internet was distributed using sneaker nets essentially picking up large quantities of data and running it around to different people. This still works. And in fact, uh, if you went out into the expo hall, you might have seen one of the Wikipedia projects, which is putting Wikipedia and a bunch of other information from the internet on drives to distribute around the world. That

still works, but we can move further than that. Things like meshtastic allow new uh new networks outside of traditionally government regulated ones to be formed ones that are distributed and not as subject to uh government crackdowns. Now, of course, meshtastic is not a perfect protocol. Um, uh, in addition to any of its own issues, as all protocols for this have, uh, it is also subject to, um,

uh, radio interference and attempts to jam signals. Um that's a very high level of collateral damage during regular times. But we have already seen for instance uh right now in Iran with the ongoing war uh signal jamming is being used as a way to shut down communications. Um so it is something to consider. However, community based networks are another way of getting around censorship. Um another promising

thing is server side evasion. Almost all censorship evasion has been focused on what CL um what clients can do to evade sensors. uh because you know say say you are Google okay your primary audience is US European customers you don't particularly care necessarily if you can add China as a market immediately especially if you're going to have a very limited few who manage to use things like

VPNs to get access to your services. So Google um and other major companies don't really have an immediate interest in trying to use things like serverside evasion. However, tools like Geneva from the University of Maryland have shown that it is actually possible to implement many censorship evasion strategies completely on the server side without clients having to do anything. uh this is this provides a lot of potential

if developers of web services um I'm speaking to all of you on this um uh actually take the effort to start implementing some of these best practices implement things like Geneva into your applications we can make it so that uh users don't have to worry about doing the censorship evasion anymore um especially for nontechnical users this is a major benefit and something I ask you all to

consider as you're developing is how can you make um your uh how can you make your services more accessible um especially to those under censorship regimes. Um another thing that's been coming up recently is age verification laws. Um y'all have probably heard about this from multiple different states, the UK, Germany, uh Australia, several other countries. uh age verification laws are becoming a major thing and ultimately a

large part of that is as a way to restrict access to information and provide additional grants for censorship as well as mass personal data collection for destroying privacy. Um and I I like to think, you know, here in California that we have things going good for us and that we'll be protected. Unfortunately, this is a California law that will come into effect um January 1st of 2027

uh that requires age verification at the operating system level. Um yeah, that that was my reaction too. Uh raise your hands if you think that uh operating system level age verification is something any of us are ever going to implement. That's what I thought. Um, obviously none of us in this room ever intend on complying with this law, but it does give us an indication of the

types of threats that we need to expect and be ready to face because if they're trying to come all the way down to the operating system level, this is more than just initial political reaction. It is an attempt to create a structure where people don't even have the technology um accessible to access to combat censorship. Uh, and so we need to keep these things in mind and

fight against them and look for ways to circumvent these types of issues. Um, by the way, if you want to know more about how to deal with this law, look at the fedora forms. There's a great thread on there about it. Okay, one last principle through all of this that I want you all to keep in mind. We as developers, engineers and especially if you work on

censorship evasion technologies need to be constantly in contact with those who are actually using your technologies. And what I mean by this, I mentioned we have about three to five years before full uh become essentially unusable. That means we have three to five years to develop and distribute solutions. And as you know, new methods of censorship come up, it's paramount to keep pushing out solutions as fast

as possible, these sensors cannot be allowed to get a leg up because if we fail collectively as developers even once sensors manage to make the first move and completely block access to something before new technologies can be developed, doing censorship evasion from the user side will become become fundamentally impossible for most. Um, t take China as an example. Um, if you are in China and you you

get updates to your VPNs or to your censorship evasion software regularly and that works because you're already connected to the outside world. Say that stops. You're no longer connected to the outside world due to a new censorship system. How are you going to get the tools to circumvent that? You're not connected to the internet anymore. you can't or at least sought to the services who could provide

that to you. Sure, maybe we can try and sneak it into different websites, but the number of people who are actually going to reach that is pretty limited. You can try and sneak it in in person putting yourself at physical danger. Um, and even then you have to hope that you can get the um, new content distributed widely before sensors catch on and manage to block that

new protocol. Remaining in constant contact and staying vigilant and ahead of sensors is the number one thing that needs to be kept in mind for censorship evasion. We have to constantly win and come up with new protocols to fight sensors in the race between sensors and censorship evasion. The sensors really only ever have to win once. That uh if you look at the slides online, you can

find all of the references mentioned in this talk. Other than that, thank you all so much for coming. Thank you for listening to me. all right. Thank you. Uh, if you if you have any questions, uh, please feel free to raise your hand. Uh, I do need to get going very soon, so we'll only have a limit of time for questions. If you want more information on

how to contact me, uh, there's my phone, signal, email address, SoCal mesh, um, call sign, and matrix, uh, yellow shirt. >> yes. Uh CDN sensor CDN. Yes. Uh the question was previous attempts to evade censorship use CDN's um as a proxy. Why is that not being done anymore? Uh there was a study I believe came out in uh 2021 that showed that China in particular was able

to uh they had China has started it yet. They've used it now that China was able to uh detect when irregular content was being spread with CDNs and block access to that as well as just start limiting access to uh non-Chinese CDNs in general. Um so CDNs are no longer really a viable bypass method. People are still looking into it but there's no current viable solution at

this time. >> Uh the slide uh the question was where are the slides online? The slides are on the scale website. You can Google uh my name and scale 23x or the title of the talk uh or go through the schedule and you can find the slides there. Yeah. Uh the question was um on full packet encryption HTTPS is still encrypted. So how is full packet encryption

any different? The difference is is that um HTTPS which runs with TLS as its base still has identifiable header information. If you open Wireshark, you can identify that a packet is TLS. If you can identify that, so can the sensors. Um and TLS indicates what the next protocol is. So that identifies that it is HTTPS. And it is very trivial for a sensor to then determine okay

this is an HTTPS packet whereas this is some random noise that maybe we should consider blocking. Yes. Huh? >> Yes. Uh there's lots of different um uh masquerade projects currently going on. Uh one of the big ones that has the most protocol support is XVPN. Um I don't usually recommend it because it is not open source and the and the company that runs it. It's kind of

weird, but if you want a starting spot on looking at what currently is being done in the industry, they are good to look at because um their main mode is HTTP and HTTPS, but they also do over DNS. they've also done over uh SMTP. Um I forget if I mentioned what the question was. The question was um uh uh what uh can masquerade be uh how much

work is there on doing masquerade over other protocols? Um, the question was, are there examples of developers getting in trouble for doing server side evasion? Uh, to my knowledge, no. Um, if you are an active developer of serverside evasion technologies, I would not suggest that you fly to China at the current moment. Um but within the US um and within Europe as far as I know there's

not been a case of anyone gone after for uh for doing serverside censorship evasion. In fact uh the US government themselves is actually working on serverside um evasion at the Naval Research Club. That's a great question. And the question was, how do we um given the law's intention is to deal with social media companies intentionally targeting youth that are otherwise protected by laws by COPA, how do

we balance the interests of uh privacy um and avoiding censorship with needing to protect children? Um there's a couple points on this. Uh, one is that we already have laws in place that uh, when it can be demonstrated that a that a social media company has been intentionally targeting children, they can be sued or even shut down. California has some of the strongest laws on this and

the California Attorney General has been successful in those lawsuits. Um, >> yes, but the the issue then is stren strengthening those strengthening um minimum liability for that even more. Um as for um you know as for age verification itself um personally I am strongly against doing age verification in general because um at the end of the day it is the responsibility of uh parents to decide what

their children have access to in consultation with their children. Um we already have existing systems for doing this effectively even on Linux. Um if you have I I know at least it's on Fedora. I believe it's also on iuntu. Um there is a library called libmeal content that is very well integrated into gnome and many current browsers. Um that allows for setting parental controls. Um ultimately my

my proposed solution would be go directly after the companies on liability issues without imposing age verification requirements that require us all to give up our privacy and open the door to more censorship. Uh are there any projects related to evading timing analysis? >> Uh yes. Uh after the uh three-way handshake dissertation was released, um the author of that dissertation um uh ended up working with a lot

of the evasive VPN protocols to manually add in um additional randomized timings to make it so that timing attacks would not be reasonably feasible. Uh the UMD team from Geneva has also done similar uh webdeave or >> uh the question was uh are sensors blocking webdav turn or stun? Um I've not seen any personal evidence of these but I wouldn't be surprised if um projects already exist

for that. Um, blocking webdev is in a similar vein as uh blocking other censorship that goes >> Huh? Oh, that is uh on Longfast. >> Yeah. Uh I mean my my note is uh pretty weak. I have pretty awful line of sight. So, uh, if you can't hop to someone in the Irvine area, good luck trying to reach me. I suggest one of the other Uh, that

was XVPN. X-VPN. Uh there um is I don't believe I included a link in the references to XCPN because I don't want to promote them. But um uh if you go to uh the last link in the references, a lot of the uh individual things that I didn't put in the references but that I mentioned today are included in this. It's about 42 pages of the current

state of Uh the question was um uh would the California age verification law impact open source operating systems? Um we're going to have to see how it plays out in court. Currently the state legislative research office as well as individual lawyers who have been consulted on it say yes it does apply to open source operating systems. Um so yes Debian could be sued for not including it.

Um there is discussion on the Fedora forums because Fedora is a major uh open source operating system distributor about how to handle this if they are forced to come into compliance. Uh the two prevailing ideas right now are to create a DBA service that you just enter your agent to because California's law does not require ID based verification. Um uh you just have to enter you just

have to enter the information. Um and the other option is uh mal lib mal content which is already included in distributions like Debian and Fedora uh already provides the age ranges required by the California law. So, you can just say, "Okay, well, whenever we go through the GUI to create an account in Gnome or your other favorite DE, uh, just add a add a malcontent um age

option." That if if the law does require compliance, that's probably the best way to do it. And right now, most lawyers have said that would be enough to cover you under California law. Well, thank you all so much for coming. Uh, I got to bounce over to Delmare Station. Thank you all so much. Have a wonderful day. And if you like uh if you like cross architecture

emulation or you want to know how to run uh binaries for other systems on your system, come to my

From event

SCaLE

05 Mar 2026 – 08 Mar 2026

All event videos
Back to Watch