About this talk
In this talk, Michael Carduch discusses the critical importance of the semantic layer in AI. He explains how AI has recently fallen into a 'trough of disillusionment' and highlights that the root causes of many AI challenges stem from insufficient data quality rather than model size or prompt engineering. Carduch emphasizes the philosophical underpinnings of AI, suggesting that knowledge requires justified true belief, which current models often fail to provide. He introduces the significance of methodologies like JSON-LD, RDFS, and OWL for creating a robust semantic layer that enables knowledge graphs to facilitate reasoning and context. By aligning data with shared meaning and context, organizations can enhance AI capabilities, remove barriers to integration, and drive innovation while addressing AI's limitations.
Full transcript
I have been excited to tell you about this not just this year but for the last 10 years to solve a problem that the industry is just waking up to. Uh, as Kate introduced, I am Michael Carduch if we have not met in person. As she said, I am a lot of things. But let's dive right into the topic because I have so much to give you.
I can't paint the complete picture, but I can give you the road map so you can go back to your organizations and follow it. But let's start with this. There's an old joke in the technology industry. Maybe you've heard it. There are two hard problems in computer science. Yes, cash invalidation and naming things and off by one errors. You've heard that joke? Well, fun fact, only one
of those is a computer science problem. Anyway, let's dive in. Gartner, you know the hype cycle people, they just made another one of their famous prognostications. They told the entire technology industry that, and I'm paraphrasing here, the semantic layer is a non-negotiable for AI. And when they said this, most of the world heard those words and just said, "What? What does that mean?" In fact, we even
had that in this morning's keynote. The gentleman from Target was talking about what we need to succeed with AI, and he listed a bunch of things, and he quickly said semantic layer. And almost nobody noticed. It's the non-negotiable. But what does that mean? Well, I'll tell you, have noticed that in the last year or so, AI has gone from here all the way down to here. The
trough of disillusionment where we realize that the promises are just a little bit overblown. But you see, there's this whole other set of technologies over here that are already way up there. And that's the maybe one of the most important things you can take away from an an event like this because these technologies were created to solve the very problems that currently have AI trapped all the
way down here at the bottom of the trough of disillusionment. The reality is there are whole classes of problems in AI that will not be solved with a bit bigger model or a better prompt only better data. And this is why semantics matter. I know. What does that mean? Don't worry. Bring the audio gain down just a little bit. It's a little hot. Thank you. But don't
worry. The reality is this stuff is new to almost everybody. Historically, this has been a very niche pursuit. And there's a reason I'm going to get to it in a moment. I'm going to teach you what it means. More importantly, I'm going to teach you how to do it. But to get there, we might have to upgrade our mental models just a little bit because AI is
unique in the tech in the history of technology. It barges head first into philosophy because now we're dealing with so much more than than predictable deterministic logic. This is where philosophy starts to enter the picture because everything else we've built has been built on abstractions. If you look at the latest and greatest in the computer science world, not the generative AI stuff, but the real hard logic
that we create, every abstraction that we work on has a foundation. Every foundation has a foundation. You can trace that all the way down. But with AI, there's nothing underneath that abstraction. And that's the problem. That's why this stuff actually starts to matter. It's a very different world from our simple world of ones and zeros and boolean logic. And that's one of the reasons that these technologies
didn't spread too widely when they were first introduced. Uh practitioners called it semophobia. You talk about semantics, every like that's a big word. I don't know. I don't know if I can handle that. But y'all are smart. We got this. And plus, you got me. I'm your semantic sherpa. Semantics is the question of whether or not we are talking about the same thing. Now I got to
give you an example of this because we might use different words but we can still mean the same thing. So what do we mean when we use those words? I'll give you a quick example. When I was very very very young I was just barely able to walk. I was just barely able to talk. And in fact as a child I invented my own language. Kids do
that sometimes because it was too slow to learn from the adults. You could tell I was born to be a software engineer because I was born with not invented here syndrome. So we were out camping in Yellowstone National Park in Wyoming in the United States. And I'm sitting there, my sister's uh toasting a marshmallow over the fire. And I light up my I get a big smile
on my face. My eyes light up and I say, "Hi, Titi." Now, my mother had spent enough time around me to even though I used different words than she used, she knew what I meant. And that's why she didn't look at me. She didn't say what. She didn't follow my gaze and try to figure out what I was looking at. She scooped me and my sister up
and we threw us in the camper and locked the door because she knew that word meant danger. Hatiti was understood. It was a different word that means the same thing as a word that she had. Although I had the word hatiti and my mom had a different word, we still had something in common, a shared concept. That concept was a bear. This thing that we call a
bear. By the way, fun fact, I don't know if you know this, the word bear is not the original word that we had to refer to that animal sort of. If you trace it all the way back to the Germanic roots, bear means the brown one. We never dared utter that word and then we forgot what that word was because nobody dared say it. And so we
have these concepts, we have these things that we use words to describe. They are separate from the words. They exist independently of the words. So what things are there? Well, that's a whole different branch of philosophy. That is ontology. Now, ontology is the question of what exists and whether or not we agree on it. So, communication require of any idea requires both shared meaning and shared concepts.
And this is actually way harder than it sounds. And that's why That's why we are here. Let's look at just one major limitation of AI today. One of the things that we want from AI is the ability to ask a question and get an answer. And we can ask a question. We will get an answer. But that's not good enough. If we're going to make a major
decision based on that answer, we probably want to know if we if it's accurate. Can we trust it? Can we trust the answer that the LLM gives us? Nope. Because there's no mechanism in the architecture of that system for truth, only probability. That's all we have. And that's a major problem because it could be your butt on the line. It could be you getting that phone call
at 2:00 in the morning. But why? Let's actually look at what's happening under the hood. We ask AI a question. Right there it is. There's the question. It's going to tokenize that. It's going to turn that into some vectors. It's going to apply position embeddings. It's going to feed them through a bunch of transformer layers. What comes out the other side is a probability distribution. That's it.
Now, it's a very clever probability distribution. It's very sophisticated. And the emerging things that show up in the AR in the AI are impressive sometimes. Sometimes it goes completely off the rails. This is why all models will hallucinate. And more importantly, once you realize the mechanisms underneath, we realize that all responses are hallucinations. Every single thing that comes out of an AI is a hallucination. It just
happens that sometimes the facts and the response overlap. But that's probability. That's not a guarantee because there's no mechanism in the architecture for This is why we turn to the philosophers because the philosophers gave us a little bit of a hack, something that we can use. You see, the thing is, your model doesn't actually know anything. The smartest models in the world smart, don't know anything. Now,
you probably agree with that. You might not. But how about this? What does it mean to know something? What is knowledge? What is it? Is it just something you know? Something that's stored in your in your noggin in the steel trap? Well, you could have something in there that's wrong. So, if you know something, if you have an idea, if you have a fact stored away, but
it's incorrect, do you really know that thing? Or what if you are correct, but you know for the wrong reasons? Do you really know that thing or is that just happen stance? Well, it turns out people have been asking that question not since chat GPT dropped, not since computers became a thing. They've actually been asking this question for thousands of years and they gave us really useful
tools that we could build on. Now, one of those people who are asking that question was Plato and Plato did some important work. He's best known for the Republic and his theory of metaphysical forms which led him to try to define knowledge on the individual level and he came up with an answer. This is what Plato decided that knowledge was justified true belief. In order to know
something, you have to believe it. It has to be true and you have to be able to justify why you believe it. And now we look in the internals of the LLM. The response might be true. The model believes it to be the case in the context of your prompt or at least believes it's possible. But there's no way to justify it. It cannot be knowledge without
all three parts according to Plato and the tripartite test. So let's justify it. Ask a question. I don't just get an I get sources. AI done brought receipts. So there we go. We just need to connect the AI to what we already know and trust and use that to justify the responses and we get some kind of rag framework. Now there's three hard problems here. Finding the
right data, making sure it's relevant, and then finding just the relevant bits. Those are all hard problems. This is why Rag demos well, but it's enormously difficult to build. But there's another problem. Most of the state-of-the-art of Rag pretends that your enterprise landscape looks like this. We have documents, we have media files, whatever. But our enterprise landscape doesn't look like It looks like this. We have all
this structured information as well. Now, years ago, this is an area that I've been in for some time now. I started a little boutique consulting company, a semantic consulting company. And this is the problem statement that I have on my website. Modern enterprises face pressure to swiftly adapt to dynamic markets. However, crucial knowledge to power such evolution is typically spread across different sectors of the organization managed
by fragmented information systems that are difficult to integrate. This situation obstructs a clear understanding of the business's health, its core challenges, and potential opportunities for innovation. I wrote this before Geni Gen AI. I wrote this years ago. This has always been a problem, but like so many things that AI has done, it has amplified it. So, can we use the same trick? If everything's in these information
silos and we have rag, can we just use the same trick in addition to our vectorzed stuff? We could just connect it to all of our existing information silos. That's easy, right? Not as easy as you think. And in fact, if you actually go down this road, these things are a nightmare. all these different systems from different vendors representing different domains with different interaction modes and different
data models and different syntax and different semantics. But you know what? We've solved this before completely globally. In fact, this was the same problem that the worldwide web itself, the internet as we know it, had to overcome to work. And it works. It works amazingly. Never mind the fact that we put a bunch of garbage on it and we put a bunch of like flaky apps on
it and everything breaks on your phone. You know, there's a lot of issues with the web. But the web itself at the architectural level is amazing. Imagine we've got these two JSON docs from two different systems. Now beyond the question of how the agent knows where to query and how to query if this is going to be in our agentic rag whatever given these two fragments we
can guess but we can't provably reliably form a cohesive single model of a book but we can borrow the ideas of the web. We can treat these as information resources on the worldwide web. Information resources have identity, strong, stable, globally unique, dreferenceable identity, right? We already have the example.com/bushbook441. And then in our JSON doc, we do things That 441 does not mean anything outside of that database
or outside of that system that's in front of the database. But the internet, the web gave us uniform resource identifiers, network friendly, self-contained, resolvable foreign key that even prescribes an interaction model. So now when we look at this instead of just some magic number that means some database key in the context of one system now that identifier has everything we need to know to look up the
other thing we have author ID example.com/authorid72 huh I don't know what that is how do I find out the same way we do everything on the web click on it dreference it do a get request and now we're connecting ing these JSON docs. So now you don't have little islands of JSON, you're building a web of data that can be explored, that can be navigated, that can
be discovered. And that starts to change the economics of integration. But we can do more because we've got a web, but it's not enough because it's a web of data. And your AI is really bad at data. Now, wait a minute. Are we sure we're talking about the same thing? Let's align on semantics. I'm going give you an example. When I was young, I went to college.
I actually studied physics. I wanted to study philosophy, but that's a different story for a different day. I studied physics, but I took an IT class because I'd been programming since I was 8 years old. I'd been using computers since before the web. I was on the web. I was on the internet before the worldwide web. I had a little bit of an advantage over my cohort.
And so I took that IT class. I figured I was an outlier. It would be an easy A. And the first day of class, we got tasked to fill out a test just to kind of see where everybody was. One of the questions was this. What's the difference between data and information? What? That was my reaction. What? Like I didn't know the answer and it angered me.
Aren't these the same thing? Well, it depends on who you ask. It depends on your idea of semantics. Cuz in my world, the technology was just an electronic frontier to explore. But to much of the rest of the world, computers were a tool to solve business problems and information was central to all of that. So, I left it. I'm like, whatever. I'll come back to this. I'll
go to the next question. What's the difference between information and knowledge? And this was all beginning to feel like a very naval gazy exploration of how many angels can dance on the head of a pin. I hated this philosophical crap. Well, I'm going to give you the answer because I finally figured it out. Took me years. Finally figured out. This is data. I'm a believer in datadriven
decisions. I have data. What should I do? What decision should I make based on this data? You cannot answer that because all you have is data. Now if I gave you information that will change the conversation entirely. Information is data with context. So look at this. Is this data or information? Hands up if you think it's Hands up if you think it's information. Okay. information is winning.
I might change your mind. So, here's another JSON doc. Okay, it's information, right? We have we have labels for these magic strings that are the values of the data. Labels like title. Oh, yeah. The title of the book, Elizabeth the Queen, the life of a modern monarch. Okay, title means the title of a book. Except we actually use this three different times and it means three different
things every time we see it. There's the title of the book, a title of nobility, and a job title. We use the same word to mean different things. And your model can only guess and it will only guess right some percentage of the time. There's a great slide that I stole from AWS. I don't know what the original picture was or the talk or anything else, but
I like the slide, so I stole it. But at least I kept the logo there. Hopefully, this counts as fair use. JSON is not easy to understand. And the challenges we have with agentic rag in structured data systems are proving it. JSON is not a data bottle. JSON has no semantics. They are not there. It's just magic strings and magic strings. The only time it has meaning
is when we either directly push the buttons or we prompt the code with some document and other context to write the code to interpret it and give it context. Your data is never just JSON. You always impose external semantics. And this is where your probabilistic inference starts to fall apart. There's a lot of uh $5 words here. I don't know if we have an equivalent idiom here.
a lot of uh 2,000 rupee words to 20,000 rupee words in there. But the thing is now we're starting to realize why this is beginning to feel hard because unstructured data has too much context and we have to do a lot of work and usually guesswork to strip it down to reduce that. But our structured data has too little context, too little to guess from. And both
tend to fail us for the same reason. The semantics are only implied. Now, we're used to meaning being contextual. Like, nobody here skipped the beat. Nobody had a segmentation fault when we use the same word three times to mean something different because you intuitively understood that the context was changing with those embedded objects within that JSON doc. But the thing is, we stand on abstractions that are
not yet loadbearing for AI. AI sits on top of an abstraction with nothing underneath it. None of the the philosophical grounding and we cannot ignore it anymore now that AI is amplifying the problem. So the reality is all these labels, they're just magic strings. They don't mean anything. And we can change these and change our integration and it will all still work because the meaning is not
here. The meaning is in what's interpreting it. Can we make our JSON more expressive? There you go. Can we make the JSON more expressive? No. JSON can only ever be decontextualized name value pairs. It's as simple as that. Now, people say to me, well, MCP has schema. Schemas are not semantics. Oh, that's oh, title is a string. Oh, well, that that tells me everything I need to
know. It's a string, people. What are we going to do with that? Well, let's look at our JSON here. We're using URIs as identity that allows the data to connect. But ID doesn't mean anything. It's just a convention that we implicitly assume means what we tend to conventionally use it to mean. It's still a magic string, but we can extend our JSON. We can create a single
reserved keyword that means not just the concept of identity but means formally defined in a standard this is the identity and the identity is a URI strong identity that everybody understands and because nobody's really doing this with their JSON the people who came up with this superset of JSON said okay well at ID means it's an identifier and it means it's something that we can usually follow
to learn something new. It's just one centrally defined standard term that means a URI lives here. And unfortunately, that's uh the only thing in this JSON doc that we can define globally. We're not going to get the entire world to agree on what we mean when we say name, what we mean when we say title, what we mean when we say But that's okay because the meaning
of title is contextual. And in my context, I can name things whatever I want to name them. And in your context, you can name things whatever you want to name them. So if the meaning of everything is contextual, every object already has an implicit context that you work with all the time. Like this isn't new. You work with implicit context all the time. But the thing is
a language model can only understand a word in context. That's that was the breakthrough. That was the breakthrough that Ilia whatever and OpenAI came up with to build language models as we know them today. So let's add the context. There it is. That's it. And in the same way that I can follow one of those ID links to learn more about the Queen or Sally Bedwell Smith,
I can follow the link to the context to understand more about the data. So just like your JSON doc is just an information resource on the web, your context is just another information resource on the web. So, a resource, if you're not familiar with all the the Restaparian rantings, is anything that has identity. Anything that is important enough that we give it strong name. Now, that's a
pretty broad bucket. There's a lot of leeway in there. That means we can name a lot of things with a URI, including our terms. Really easy. It's really easy. Instead of a magic string, we back that magic string with strong identity that means this thing in that context. That starts to change everything. So instead of a magic string title, what if title had a pointer to a
definition? Now we don't want to go back and rewrite all of our JSON. We've got clients that rely on certain syntax and everything else. And if we just go rewrite all this stuff, we're going to break it. So let's not do that. In fact, you don't have to do That's what the context is for. So in the context, we just define that that term in this context
is backed by that URI. That means that term. Okay, that probably seems a little trivial. You're like, okay, so we're taking a magic string and turning into a URI, which is basically another kind of magic string. But it's got some massive in implications that we're going to get to as we go through because it's not just a magic string. Your AI might not know what that URI
means, but it also might. But the important thing is that URI is a pointer. You can work with the pointer if you already know what it means or if you don't, you can dreference it. You can look it up. You can essentially your AI can essentially click on the link. So let's say we follow that. What's on the other side? Because this is key. This is where
it all starts to change. Like any web resource, you can link to it, you can look it up, you can save it, you can store it, you connect it to other resources, you can combine it with other resources. Because by creating standardized links, we're not just creating documents. We're building webs of data and we're enriching them so they become webs of information that are machine navigable. So
this is what's on the other side. This is what a machine gets when we dreference that term. Let's unpack this a chunk at the time. So you remember what I said about the flexibility of resources. Well, this is kind of what we're doing here. Title is a term in our vocabulary that we're giving formal meaning to. What we're doing is we're defining the semantics. Now other people
have defined vocabularies with well-defined meanings that we can build on. that we can actually we can extend as part of that chain that gets down to the bedrock capital T truths of reality. But at a high level, we're basically defining namespaces. And everybody gets tired of typing fool your eyes. So we're using these to essentially shorthand a lot of this stuff. So instead of having to say
httpwww3.org/20001dfs /2000/1 RDF-S schema pound property down in the body we just say RDFS property it's a shortand it's compact URI uh now we have another things uh so at base is my agreement this is my namespace this is a URI I control where my definitions live so everything I define in this vocabulary lives under that address remember what we said about ontologies are just an agreement about
what exists This is where I publish mine. Uh, so what else do we have here? Then we have the term title. Uh, we don't have to spell the whole thing out because of that at base. Its full identity is https/acample.com/nshtitle. It's not a magic string anymore. It has an address. You can follow it. It lives on the web next to every other defined thing. By the way,
I want to point this out. This doesn't mean that you have to make all of your data public to the web. You can do this internally, you can do this externally, you can make you can make your vocabulary public, but keep your data private. There's a lot of different ways to do this. This is the same set of ideas that we're already building our APIs around. So
that's kind of an important thing. And then what is this? Uh this is what is this? We're saying it's a property. It's not a thing. It's a relationship. So a property describes something about a thing. RDFS property is a globally understood way of saying just that. Any system that understands RDFS, and there are a lot of them, knows exactly what this means without me or you having
to explain it. And so at type, that's another one of those centrally defined reserved words in our JSON superset. It identifies its class. And we'll get into why that's super useful and super flexible here in a moment. What else do we have? Uh uh yeah. Okay. Uh we have RDFS label and RDFS comment. This is where the term is actually defining itself. Label is one of many
available. This is localizable but this is a way of saying hey when I use this term and I have the URI this is what we show the humans comment this is the human readable definition and this is all stuff that you don't have to hardcode into a UI now and now this line RDFS subproperty of this is doing something profound Dublin core which I've I've uh namespaced
up on the third line off the is one of the oldest and most widely used metadata vocabularies on the web and in the world. Libraries use it, publishers use it, archives use it. By saying that my t my property is a subpropy of DC title, the Dublin core title, I'm not just documenting a relationship. I'm making a machine readable claim. Any system that understands Dublin Core can
now infer things about my term. Well, I'm able to build on the knowledge that is already out there. I'm standing I'm not teaching at my vocabulary from scratch. I'm standing on the shoulders of giants. Remember Plato justified true I'm not just asserting that my title means something. I'm pointing to a chain of agreements that grounds it. And then in here as well, this is just a way
of saying that the property is always a string. This is where I'm giving its type. Now, you might say XSD, we don't use XSD. That's fine. You don't have to. But XSD is a vocabulary that is standard and widely understood. You don't have to use XML to to say, "Hey, you have a concept of a string. I have a concept of a string. My concept of a
string is the same as the XSD string and your concept of a string is the same as an XSD string. Therefore, my concept of the string and your concept of a string are the same things. No guessworks, no hard coding. It's just facts that now can be inferred. This is how you don't just build webs of data, you build webs of understanding. Now, I handle handwaved over
some of these. I skipped some of those. We'll come out of that because we started with a JSON document full of magic strings. Title meant three different things and in three different objects. Your AI had no way of knowing which was which. It could just guess. Probabilistic semantic inter inference. But remember, it all falls apart when the word means different things in the same document. So, can
JSON be more expressive? No. JSON can only be decontextualized name value pairs. So, we extended it into a formal standard that works with all of your tools and can turn your existing JSON into JSON LD with as few as zero edits to the payload. You don't have to rewrite your systems. You don't have to swap out databases. You don't have to completely break all your contracts with
all of your clients. You're just enriching what you're already doing. Every term has strong identity. Every term is self-escribing. We have label, comment, type, range. These are human and machine readable. And every definition stands on the shoulders of existing agreements. We're building a loadbearing foundation under the data and therefore under our AI. My vocabulary has identity too. It's not just some private schema sitting in a config
file somewhere. It is a first class citizen of the web of data that you can build. I can share it if I want. I can keep it internal if I want. And so remember I said this in the beginning, communicating any idea requires shared meaning and shared context. JSONLDLD gives your data both. It gives your data both. You've already got this. They're no longer magic strings. They're
defined. They're grounded. They're justified. Your data has superpowers because it's no longer data. It's information. And we're not done. Because your JSON documents don't have to be documents, little islands of data. They could be graph fragments. People are starting to wake up to the fact, hey, you know, it's really cool if we connect information. Yeah. Yeah, you're right. People talk about graphs, talk about knowledge graphs, but
knowledge graph without the rigor, without the substrate, without the foundation. I'm talking about knowledge graphs with the foundation. One of the cool things about knowledge graphs is you can just keep connecting more stuff. And we're using these strong identities, URIs. You don't even have to do any work to connect them anymore. Everything just snaps together like Lego. Oh, that's what my next slide says. They just snap
together like Lego. This is what URIs give our JSON LD docs. The the JSON LD docs grow in understanding not just in nodes. And guess what? This is the exciting Uh I already said this actually zero changes to syntax. I feel like I let let you down. I had this big guest. What? And then you I Well, you mostly just sat still. We just had lunch. I
get it, right? We're all a little lethargic. But I like in my mind I like to imagine that you moved just a little bit closer to the edge of your seat and then I said something I already said. I let you down. I'm sorry. Oh, but you know what? There's another thing that's cool that I haven't told you yet. All this JSON LD stuff that I'm talking
about, your model is already fluent in JSON LD. Fluent. Not familiar. Not able to guess. fluent. So, what does this really look like in practice? Well, about a month after the public beta of chat GPT dropped, a knowledge graph engineer who's been tilting at the same windmills that I've been tilting at for years said, "Aha, JSON LD's been around for a decade by this point. We just
found its killer app." and he whipped together a little demo before vibe coding, before anything else, and he showed us just how powerful this begins to be. Now, this is not the full vision. This is just the first piece. And look how powerful this is. I'm going to show you a little video. GPT is a super powerful tool, but to really make it useful, we need to
find a way of connecting it back into our own What most people haven't really figured out yet is that while GPT has been learning by looking at all of the text on all of the different pages on the at the same time 40% of those web pages have got little islands of JSON LD embedded within them. [clears throat] These JSONLD islands form two things. Firstly, a graph.
So we can ask a question of GPT and then we can get the data back in a graph format. But secondly, they're like little anchor points that we can link our own data back into. So GPT will come back with But if we adopt the same model internally and use JSON LD within our own organization, then we can blend our data in with the data that GPT
is holding, linking it into information that for obvious reasons we would want to be keep private to our own organization. So we can ask general questions using GPT's general intelligence and link them back to specific data that we hold internally. This pretty much applies I think to any [sighs] and sometimes you will need uh to extend the model in schema.org. So yes have all your data available
in JSON LD yes uh use schema.org as your basis but have your own internal version of schema.org. or hosted within your company that'll let you build your own models. But when you build these models, make them extend the ones that GPT is being trained on out of the web. I think the potential of this technology is actually quite vast. Imagine that not like this toy example, but
in a real example, you have all of the data available from all of your various different systems. They're all linked together and they're all linking back to GPT's general intelligence. Perhaps stop and think for a second about the disruptive types of this to your organization and others might be so mind-blowing you might have missed it. He talks about little anchor points that connect your data to what
comes back from from GPT. Those are your URIs. Understanding it, being able to connect it, interpret it, do all of that. This is graph rag that completely solves the problems that we're still trying to figure out how to solve before a soul, a single person in the world ever uttered the words graph rang. And this wasn't some researcher at a semantic web conference. This was somebody who
got early access and within days understood something that most of the industry still hasn't figured out. He asked you to imagine it, but I'm here to tell you how to build it. We're only beginning to see the power of this approach by making explicit what is only ever implicit using a standard a serialization standard that you're already familiar with that you can already use to express your
data in context JSON LD using the and then using the same standard to formally express your semantics. When you do this, the economics of integration completely change. Integration goes away as a problem. This is one of the capabilities that link data gives us the LD and JSON LD to connect and understand any data set wherever it is effectively for free. No more ETL, no more EL. Your
data just connects and the meaning goes with it. Well, now we look at this and it starts to feel a little more feasible. And you're early on this. I want to point that out. Everybody here in this room and you watching this video, I see you are early on this. We used one standard to connect meaning and we use the same standard to express meaning. How far
can we take it? As far as you want. As far as you care to. further than you can imagine today. I guarantee it. It's a simple model that is infinitely extensible. So, JSONLDD connects terms to meaning. The terms form vocabulary or ontology if you're feeling fancy. And we use another vocabulary to define our vocabulary, but this time it's built on top of a loadbearing abstraction. Instead of
just probabilistic guesswork, there is a foundation. You can build your data model on a rock instead of the stand. That's what we're getting with RDFS. It is a formal standard. It has meaning that goes all the way down. It's the RDF schema language. And you've already been reading RDFS. So, let me show you what it actually is. But it starts with RDF. If it's the RDF schema
language, well, then what is RDF? RDF is the underlying data model. Plain JSON has a two-part structure. Name, value. RDF extends that into three parts. Uh, a three-part data subject data data structure. Subject, predicate, object. The thing that we're talking about, the property that we're describing, and the value of that property. It's a complete sentence, not just a label. So, JSON LD serializes triples instead of pairs.
JSONLB serializes sentences about your data and those composed to describe it in ways that we're only beginning to imagine. So RDF is the entry point into our world, our ontology. But remember, and whether we agree on it. Different people can just can agree disagree on things that exist. Within a domain in an organization, there is a known set of things that exist. We've already got the tools
to model this. And in a different domain, there are different things that exist. And that's okay. We don't all have to agree because guess what the philosophers figured out a thousand years ago? We can't agree. There is no consensus on what exists globally. But that's okay. We don't have to figure it out globally. We just have to figure out the context of one system. And then once
that system is understandable, the next system understands as well. And so it starts with enumerating the things in our world. And RDFS gives us tools to do this. So with RDFS, we have classes. These are just the things that exist. They're similar to classes in OOP, but they're a little more foundational. They're a little more they're a little less uh domain specific. OOP has a very domain
specific definition of class. This is the broader definition. Classes in the RDF world and the RDFS world are really like sets. Like Kate said, I am a skydiver, but I'm also a magician, but I'm also an author, but I'm also a speaker, but I'm also a husband and a human and everything else. I'm a member of many sets. And now we're starting to express our data as
it truly is, not in one-dimensional representations that work well in code, but in polydimensional representations that work well anywhere. Then we have subasses. These are relationships between kinds of things. A magician is a performer. A juggler is a performer. Ari Kaplan can juggle. I am a magician. We're both performers. So, a magician is a subclass of performer. We have properties. These are the relationships between things. These
are how we describe things. There are subpropies. These are constraints on those relationships. And that's it. That's the vocabulary. There's a little more in there, but that's the the heart of it. And the thing about a vocabulary, once you have one, anybody can write in it and anybody can read in it and anybody can automatically translate their vocabulary to your vocabulary. It's already built into these standards.
Now, some of this stuff sounds hard and historically people have said, "Oh, I don't want to do this. It's too hard." Well, now that AI is in the mix, we don't have a choice. And it's not as hard as you think. 20 years ago, a bunch of volunteers decided to do exactly that, to start describing the world. And they they they realized that Wikipedia had useful information,
but it'd be cool if we could ask questions of that information instead of just do searches. And so they started using these standards to make the data available to machines. They decompose the documents into a data set called Wiki data or DBPedia. Sorry, DBPedia is this one. There's another one called Wiki data. That's the official one. And currently there's about 9 and a half billion facts in
that data set and they're all connected and they all connect to other things. Now in 2007, other people said, "Hey, this is really cool what the DBPedia people are doing. Let's do this as well." Volunteers on their own time started doing the same thing, just formalizing what was already there. It's already there. You've already named things. Now you're just making it robust. And so they published their
data sets and those data sets automatically connected together. Nobody had to coordinate. Nobody wrote ETL. The standards allowed the data just to connect. And so as other people said, "Hey, this is cool. We're going to participate, too." The graph got bigger and bigger and bigger and bigger. This is what it looks like today. And the data sources are integrated as fast as they come online. And that
means that all the information that you have with these data sets, you can augment immediately. And this is just public data. Your private data, you can do the same thing privately within the walls of your organization. This is more than a web of distributed data. This is distributed understanding. understanding at web scale without coordination without everybody having to cooperate because the vocabularies are published and shared and
interconnected we can reuse them we can extend them we don't always have to roll our own think about integrating systems this is Google if I go to Google I was in St. Louis a couple weeks ago St. Louis Missouri and I just typed in hotels tonight in St. Louis, Missouri, and I got a map of all the hotels and how much the prices are and what's what's
available. But think about the effort involved here. Every hotel chain has a different backend system with different JSON payloads, different syntax, different This feature could not exist today the way we currently approach integration. Google said, "Hey, if you want to show up, give us the data with meaning and we'll figure out the rest. Use JSON LD." It's as simple as that. If you give us decontextualized name
value pairs, we're not going to do the effort to figure out what the hell your stuff means. But if you give it with meaning and you express your vocabulary and talk about how it connects to my vocabulary in the same way we did that with Dublin Core, now they can understand it. Not only that, anybody who understands Google's vocabulary will understand their data. It's transitive. It accelerates.
It compounds exponentially. And so the hotels didn't integrate with Google. Google didn't integrate with them. They didn't rewrite everything. They didn't change all their APIs. They didn't write a single line of of ETL. They just agreed on what hotel means and Google could read all of it. This goes beyond our concept of integration. In fact, this isn't integration. Integration implies effort at the boundary. But this is
something different. This is interoperability as a consequence of shared meaning. This is the difference between mapping and semantics. Mapping is bilateral. System A learns system B's language. Shared semantics is multilateral. Everyone learns the same languages and suddenly everyone can talk to everyone. The network effect is quadratic. Every new participant connects to every other participant for free. This is what the purple thing is in this graphic saying,
"Hey, if you focus on the data instead of the applications because the data is always more important. Integration, you get integration for free." That's what this picture is showing. It starts with the vocabulary built on RDFS classes and properties and subasses, everything we just learned. And the key thing is those vendors didn't have to build their vocabularies from scratch. Remember ontology is what we what exists and
whether we agree on it because in the domain of hotels there is a broad agreement on the kinds of things that exist. You can use your you can say bear I can say hot tit but we know we're talking about the same thing now for free and as as we're mapping the world we can collaborate rather than working on an island. Tony Seal in that video talked
about j schema.org That was a project that started in 2011 by Google and Microsoft and Yahoo and Yandex. These are companies that are not friends. They don't like to collaborate, but they realize this would be valuable not just for them, but for the entire Let's agree on some of the common concepts of the web. The result was schema.org. 800 types, thousands of properties, all kinds of different
things. And they're all built on standards instead of magic strings. We didn't invent a new concept of a hotel. We just stood on the shoulders of giants and Google could read all of them. So as you connect your data, this is the most powerful thing here. As you connect your data, you're amplifying meaning, you're amplifying understanding, and ultimately you are amplifying possibilities. So we have a mechanism
to look things Can AI do more than just look things up? Can it know things? Now remember, we just proved that AI doesn't know anything. And now here I am asking, can AI know something? Well, the model I introduced with JSONLD, link data, shared vocabularies, this is the foundation of our missing semantic layer. And semantics are incredibly powerful. Our model can describe more than meaning. It can
describe knowledge itself. And that's a different thing Remember this guy? He defined knowledge as justified true belief. Well, let me I am married to Kate. Maybe that's all you know about Kate. It happens to be true. If we're going to apply the tripartite test, it happens to be true. I can justify this fact with a marriage certificate. Here it is. And if you believe me, you also
now know this fact. [music] dramatic. All right, let me tell you another fact. My best friend is Draco. That statement is true and justified since I happen to be the ultimate arbiter of who my best friend is. Consequently, if I make this statement and you believe me, you meet the three-part criteria. This is something that you now know. But I've only told you one fact about Kate,
one fact about Draco. Let's see what else you know. Hands up if you think the answer is yes. Is Kate married to me? Okay. Okay. All right. I I wonder if how many people are just like, I don't want to raise my hand or I'm too tired to raise my hand or you're not sure. Uh which is fine, but that fact is also true and you probably
believe it, but if you raised your hand, how did you justify it? And you're saying, "Well, that's just how it works." Okay, if that's just how it works, here's another question. Am I Draco's best friend? Raise your hand if you think that's true. Okay, a couple people. I maybe I don't know. That's that's for Draco to decide. See, we can we can apply reasoning and infer the
answer to the first question because we know that a married a married to relationship behaves differently to a best friend relationship. One is what we call symmetric. So if I'm married to Kate, Kate's married to me. So I could say, who is Kate married to? Even if that's not in the data set, we can figure it out provably, explainably, reasoning that is actual reasoning and not the
statistical approximation of reasoning. We're using logic. We're going all the way down to the bedrock capital T truth of reality. Inference is based on logic. And maybe you've never formally expressed the semantics of the married to relationship or the best friend relationship. These are things that we intuitively know. But for the philosophers and knowledge systems, intuitively know is not always enough. If you apply the tripart type
test and you say, well, my justification is I just know. That's circular reasoning. It doesn't compute. So knowledge is therefore more than just storage and retrieval of facts. It combines facts with logic and reasoning. Reasoning allows us to infer new knowledge from existing knowledge. And logic gives us a framework for this new knowledge to be This is the nuance of semantics. Some relationships are symmetric. Some are
inverse. Some are transitive. If A is larger than B and B is larger than C, then A is larger than C. Some of them aren't. Your data is full of these relationships that your AI currently is just going to guess at and it's going to guess wrong 10 to 50% of the time. But Owl is another standard built on top of everything we've talked about. That is
how you stop guessing. The web ontology language. Owl isn't built on top of nothing. It's built on top of RDFS. It starts with the world of classes and properties, but imagines a bigger world. Same grammar, more expressive. You can describe the nature of relationships like a symmetric property. If I'm married to Kate, Kate is married to me. It's declared once ever and inferred everywhere. Uh inverse of
if I am the parent of someone that they are a child of me expressed once inferred forever. Transitive property. I assert one relationship and the reasoner can derive what follows automatically, explainably, provably, justifiably across your entire graph. And this is why it matters for AI because right now your model retrieves information and generates a respon response. It can't show its work. It can't trace the chain of
reasoning that led to an answer. It can't distinguish between what it retrieved and what it inferred and what it hallucinated. But with Al, your knowledge graph declares the nature of the relationships. Your AI can reason explainably. It can say I know Kate is married to Michael because Michael declared that he is married to Kate and marriage is a symmetric relationship in this ontology. Remember what I said
in the beginning. There are will not be solved with a better model or better prompt only better data. And this is not a bigger model. This is not a better prompt. This is a better substrate. This is our justifiable reasoning. So it started with Plato and his justified true beliefs and rag gave us retrieval which is something to believe. JSONLDD gave us context something to justify the
terms. RDFS gave us the vocabulary something to justify the structure. AL gave us reasoning something to justify the inference. The semantic layer isn't one thing. It's a stack of justifications T truths of your reality. And we can ask direct questions. Vector lookups are cool. They're unreliable. But this model has a graph query language called Sparkle. Sparkle protocol and RDF query language. The P in Sparkle is very
powerful because again it means you can leave your data where it is using the technologies that you already use that you already like that you already trust that you're already familiar with but you can express that Sparkle endpoint. And the protocol allows you to ask a question and federate that query across your entire information ecosystem and get things back and get things back that connect to your
LLM in the exact same way that Tony Seal showed us in that video. And it's all based on a simple idea that if we make our data self-describing and universally understandable, then any we can make any data self-describing and universally understandable. In other words, if we if we can describe our world, we can describe the world. But that's a big thing. You don't have to. You can
just carve out a little slice, one system, maybe one API endpoint to begin, and other people can describe their version of it if they want. And the most powerful thing is we don't have to agree on what we call things. We already know this is a big problem. Domain driven design taught us this 20 plus years ago because underneath all of our disagreements, we still have shared
concepts. We have more uh and as the more data and systems participate, the smarter everything gets and the easier everything gets. And we've accomplished all of this with selfdescribing data, the APIs and the endpoints that we already have. Uh what other kind of integration inferencing problems just disappear? Well, agent protocols are more chaotic than we like to admit. But there is another vocabulary built on JSON LD
called Hydra that describes your API provides a standard uniform universal interface to all of your existing systems. And you can layer this into what you have right now without changing or breaking a single client. And then I mentioned some of these. We talked about schema. What about shackle? If Al describes what your data means, Shackle describes what your data must look like. These are agreements with teeth.
Validation becomes a first class citizen of your semantic layer. And that means with some with another vocabulary like shackle which is another formal mature ready to go standard that you can adopt Validation becomes a first class uh first class citizen. Your data can be verifiably wrong now, which means it can be verifiably right. So long story short, Gartner, the hype cycle people just made another prognostication that
the semantic layer is a non-negoti non-negotiable foundation for AI. If you really want it to do what your the businesses and the world is hoping it will do, most of the world said what? But you know what it means. It's not a product. It's not a platform. It's not something you buy. It's a discipline. The discipline of making meaning explicit. Turning magic strings into defined terms. Turning
islands of JSON into a graph of understanding. Started by people who saw this moment coming 30 years ago. The people who spent the last few decades creating this, the standards, the toolings, the vocabulary. They gave us the entire stack. And now you know what it means and more importantly where to start. AI is just the pain that motivates the gain. And you don't have to do this
all at once. You don't have to fix it all at once. So start with you where you are right now. You have JSON, magic strings, local keys, islands of data. We add at context. Now your data begins to know what it means. You can start here this week one API. And then you could start pointing at schema.org and extending schema.org for your domain. Now your data starts
talking to the world. And then when you add your ontology, your data is formally described. Reasoning becomes possible. And then we can bring in things like Hydra for the agentic layer. Agents navigate. Everything gets smarter as we get more participants. My advice to you, start simple. You don't need an enterprise triple store. You don't need an ontology engineer. You don't need to rewrite everything. You don't need
to boil the ocean. You need one JSON doc that knows what it means. It will grow over time. And if you adopt these ideas, you start moving these things forward, you can help move your organization out of here. But just remember, there are whole classes of problems in AI that won't be solved with a bigger model, only better data. And while the rest of the industry shrugs
at this idea of semantics and finds it impenetrable, you you understand. Heat up >> [music] >> here.
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59