Submind YouTube summaries
Thumbnail for Leading IAGOV: The Right to Explanation—Meeting Regulatory Demands for Interpretable AI

Leading IAGOV: The Right to Explanation—Meeting Regulatory Demands for Interpretable AI

Watch on YouTube

Video summary

The webinar "Leading IAGOV: The Right to Explanation—Meeting Regulatory Demands for Interpretable AI," hosted by Mark Horseman with Lisa Winrich and Stephanie Parody, explores the critical distinction between technical explainability and stakeholder interpretability within the context of growing regulatory pressures. A central argument presented is that organizations cannot fully govern or inspect probabilistic AI models because requesting logic from them yields only a reconstructed "best guess" rather than a deterministic log, making full explainability unattainable. Consequently, interpretability remains inherently subjective to the specific stakeholder, necessitating a shift in focus from the internal mechanics of the model to the environment surrounding it. To address these limitations and mitigate associated risks, the speakers introduce the concept of a "decision record," a structured artifact that serves as a design specification, execution guide, and accountability mechanism. This record captures essential business reasoning, enforces approved vocabulary through governed glossaries, and clearly defines knowledge limits to ensure consistent interpretation across all stakeholders. To operationalize this decision record effectively, organizations are encouraged to utilize "competency questions" that define exactly what stakeholders expect the system to answer, thereby scoping use cases and establishing boundaries for AI responses. This approach aligns with the NIST Risk Management Framework, which outlines four core principles: providing evidence and reason, ensuring outputs are meaningful and understandable, defining the scope of knowledge limits, and acknowledging the inherent inability to verify internal model mechanics. The discussion also highlights ongoing legal precedents, such as litigation involving Workday and Navy Federal, where courts are currently deliberating whether proprietary internal variables must be exposed to applicants. In response to these challenges, a scalable implementation strategy involves identifying critical decisions, assigning dedicated decision stewards, and establishing system guardrails that translate human-based controls into deterministic code or metadata constraints, moving away from inefficient spot-checking of individual outputs. The conversation further examines the nature of Large Language Models (LLMs) and whether they rely on data or inference to connect language with physical reality. The consensus indicates that while semantic layers can create deterministic ties between language and data, base models lacking such configuration rely entirely on inference, resulting in outputs not anchored in authoritative reality. This reliance raises philosophical questions about AI mimicking human thought processes, exemplified by the concept of "persuasion bombing," where LLMs simulate human reasoning to convince users they are wrong. Regarding accountability, the speakers agree that humans must remain responsible for deploying AI and managing outcomes, even when the technology generates specific outputs like meeting summaries. This stance mirrors legal precedents establishing that claiming "the AI did it" is not an acceptable defense, similar to plagiarism in academic settings, though a nuanced view allows for scenarios where AI bears responsibility for generation while humans retain accountability for application. In conclusion, the session acknowledges the complex shades of gray surrounding these emerging issues and emphasizes the need for a balanced approach that respects both technical limitations and regulatory demands. By integrating decision records, competency questions, and robust governance frameworks, organizations can navigate the evolving landscape of interpretable AI without compromising on safety or accountability. The discussion ultimately thanks the audience for their engagement, leaving them with a clear understanding that while full transparency into model mechanics may be impossible, structured human oversight and defined boundaries provide a viable path forward for meeting regulatory expectations in an increasingly automated world.
Read the full video transcript
Hello and welcome. My name is Mark Horseman and I am the data evangelist for data. We would like to thank you for joining today's data webinar, the right to explanation, meeting regulatory demands for interpretable AI. It is the latest in the monthly webinar series leading AI governance with first San Francisco partners. Just a couple of points to get us started. Due to the large number of people that attend these sessions, you will be muted during the webinar. For questions, we will be collecting them by the Q&A section. If you would like to chat with us or chat with each other, we certainly encourage you to do so. To open the Q&A or the chat panel, you'll see the icons for those features in the bottom middle of your screen to answer the most commonly asked question. As always, we will send a follow-up email to all registrants within a couple of business days containing links to the slides. And yes, we're recording and we'll likewise send a link to the recording as well as any additional information requested throughout the webinar. Joining us today is Lisa Winrich and Stephanie Parody. Uh Lisa is the executive adviser at First San Francisco Partners where she advises Sea Level's executives on data and AI governance and leads the firm's innovation and service department practice focused on AI governance and semantics. Her work centers on the decisions and accountability structures that determine whether data and AI investments hold up, who owns which decisions, what data means across the business, what data is fit for a given AI use, and how governance operates as a working practice rather than a policy document. She specializes in building capabilities where none exist. Before FSFP, Alisa spent over two decades at Sherman Williams building data management, analytics, and AI capabilities. She made analytics and data science directly usable by business teams without IT intervention. Backed by the data literacy and governance programs that made it safe at scale, she built and led teams behind those capabilities with a sustained focus on career pathing, upskilling and coaching. She also held technology leadership roles across data network application development and point of sale. Her work spanned manufacturing, finance, marketing, sales, and procurement. Stephanie is a principal consultant specializing in building and maturing data governance programs through the effective unification of business and technology through data. She brings hands-on experience from previous marketing and data stewardship roles in addition to CDMP certification. Stephanie is an accomplished data governance leader highly adept at helping clients transition between strategic vision and tactical execution. and her experience spans across the various subd discciplines of data governance and data management with a focus on how to activate right-sized culturally relevant data enablement programs. With the uptick in focus on AI governance, Stephanie is actively responding to the required paradigm shifts in traditional government and management frameworks and approaches. And with that, let me hand it over to Lisa and Stephanie. Welcome, my friends. >> Welcome. Thank you. Welcome to everyone here. Good afternoon, good morning, and good evening to some of you. Um, we're so excited to be here. Uh, as you can see, our fearless leader is not here today, but she's here in spirit. She is taking some well-deserved time off. So, um, she has left you in what we hope are capable hands uh, of Stephanie and I. So, um, today we are going to dive into a number of things. Um, but as I was looking back, uh, stuff, if we could go, we're seeing the, um, the Could you put up the, uh, >> that could be my fault. >> Yeah, >> because I had to reboot Zoom earlier. >> Messing up. All right, it is mean minute two and Mark has already messed up our presentation. Kelly is going to be furious. >> It's It's all my fault though. Try again on the share. My apologies. Okay, let me >> Okay, here we go again. >> Awesome. I see I see the deck. >> Looks wonderful. >> Okay, perfect. >> Now it looks perfect. Thank you, Stephanie, for cleaning up for Mark's mistake there. Um, so if we go to the next slide, uh, Steph please, and we talk about where we've been. This is session nine, which I can't believe. Um, thank you to those of you who have been with us through nine sessions, and those of us those of you that are just here for the first time, don't worry, we'll catch you up. Um, as I was looking through some of these prior titles, one of the things that caught my eye was the very first session when we talked about cutting through the noise and thinking about the um words, the the vocabulary words, the bingo card we created in that session and how far we've come just in lexicon across the industry over the last nine months. Now we have words like slap economy and persuasion bombing and botsitting. All very interesting new terms. So maybe we'll do another bingo card soon because the the field keeps changing. Um so today we're going to talk about the right to explanation and we're going to talk about the difference between interpretability and explainability which Stephanie will explain because I can't interpret it. Um if we go to the next slide we can talk let's talk about where this fits. So if you recall in July in July we talked about decision clarity. We talked about owning tracing and naming as a way for the internal audience to understand what they're signing up for. And this is where the decision steward comes in um and is the kind of coaleser and owner of this process. When we went to August, we talked a little bit about decision interrogation and we also had a stellar guest speaker in Amy Bojack from Baker Tilly uh who is the CISO there and she took us down this path as well as several others um just in a really great flea free flowing conversation. If you haven't seen that episode, it's one that I would highly recommend. She really broke things down in a way that was accessible for those of us outside of her realm. So today we're going to talk about regulating decisions who for the people that are affected by it, the people on the other side. Um what do we owe to them? And we're going to look at some case law that is setting precedent right now for this very topic. Next month we're going to talk about the mechanics. So demonstrating how things work. So I just want to give you that little preview and know that we will talk about that next time. All right. So getting into it. Here we are back in our boardroom where we've been kind of at the beginning of the last couple of sessions. But here's where we have a couple of different questions um being asked the same question being asked two different ways. Right? So, if we have a regulator ask asking us why we declined this applicant, can we answer it for that specific case? But then the board wants to know this was one applicant. How do we know that we can answer this question every time? Um, and that's really what we're going to break down here and be able to understand if we can say yes with confidence. >> Right. So today we're going to talk about two sides of the same coin, interpretability and explanability. And there are many similar words around these topics that are all used interchangeably. But today we're going to dig into the nuances and why that matters. We're going to look at how we can improve and defend explanability using a method that defines its requirements once so that we can reference that same output for multiple stakeholders and we're not going back and reinventing the wheel or trying to reexplain something multiple times. And we're also going to look at real technical limitations that exist so that we're not chasing pipe dreams, but we're also not shifting organizational responsibility onto the vendors. we know exactly what we are responsible and accountable for and we're able to defend that and then we're able to push back on the industry where it makes sense because that's where the real limitations lie. So the whole premise of the session today rests on what is really an uncomfortable reality in the industry today and that is you cannot see into the model itself to know how it is operating. You can't open it up and explore hundreds of thousands or millions of layers of neural networks and compounding that uncomfortableness. If you ask a model how it got to its decision or its output, you're actually getting its best guess at reconstructing its logic, you're not getting an actual real time decision log, which a lot of us assume that we are. Um, and so therefore, we can never really achieve full explanability of a model because we can never fully explain the mechanics of how it's operating. And because of that then in consequence to that the ability for this to then interpret that output is impeded. We can't say with a 100% certainty that we've interpreted something correctly or we've interpreted all the possible facets and inputs or if we don't know how that output was created and how it led to that particular output or decision. So essentially what we're saying is that we can't govern the model and because of that there are real risks that organizations are forced to accept when they adopt AI particularly anything that is probabilistic in nature. So to counterbalance this is where governance steps in and really shines and really has its value moment. governance is going to help govern everything outside of and around the model in the absence of being able to actually govern the model itself. And we do this through something called the decision record. And the decision record is an incredibly powerful tool that is not only going to help us in explanability and interpretability of the solution, but also it's going to serve as a design and configuration specification, which is going to be super useful when it comes to helping to close that gap between business need and any assumptions that the technical teams might be making or may be forced to make based on technical limitations. So, we'll get into the decision record more in a little bit, but first let's hone in on how and why explainability and interpretability aren't the same thing, but are related. So, before I said that they were two sides of the same coin, but probably more accurately is that they're really a sequence of events that go hand in hand. For this session, we're going to anchor on NIST and we're going to use that as our authoritative framework throughout the entire session. So all of the definitions and all of the concepts that we reference are going to be rooted in that framework. And in the riskmanagement framework, NIST presents seven characteristics of trustworthy AI systems. And explainable is the first in that sequence of bucketed, explainable, and interpretable because it forms the foundation for that entire characteristic. Explainable refers to how a decision was made in the system. It's really referring to the mechanics of how the AI got to that particular output or why it made that decision that resulted in the output is how we can think of that. Interpretable is then the consequence of explainable and it's going to answer why a decision was made by the system. Um it's really going to be on the stakeholder of how they interpret the output and how they interpret the supporting information in order to be able to make a determination on whether or not that output or the decision that was made is acceptable for their use. So explainable is really a capability of the model itself. It's something that's inherent to it. So it's a little bit more objective in nature. Whereas on the other hand, interpretable since it rests on the stakeholder, it's really more subjective in nature and so is therefore a user capability. And since we can't control how a stakeholder is going to interpret the model's explanation, we can't get into their mind and force that explanation. But what we can do is we can influence it. And that's what we're going to use the decision record for. So let's talk through an example of this to sort of um solidify these these two related terms. So explanability is going to answer the question, how was this decision made? And so the answer is going to be something along the lines of debt to income ratio came out at 30% was checked against the approval threshold which is 35% and was flagged as being acceptable. Interpretable is then going to answer the question of why does that matter and what does it mean to me? And so interpretability is going to be what someone is able to interpret or make sense of based off of that explanation and the supporting information. So you're they're going to be able to interpret that it was the ratio and not someone's credit history, for example, for why the loan was approved. So, a little bit of a illustration there, but those are in explanability and interpretability. If they're characteristics of an AI system, their qualities or their properties, they don't explicitly say what needs to happen in order to achieve them, which can make creating a roadmap to achieving those difficult for organizations. And this is where NIST comes in again to help. So an accompanying document to that riskmanagement framework is a document called four principles of um explanability. And what it does is it takes that characteristic and breaks it down into four specific requirements and what we can do as an organization is we can actually execute against each of those requirements. So we first need to baseline requirement is we need to be able to provide evidence or reason for how and why AI came with its output and what this is going to do is this is going to mirror the human process of how we would work through solving this particular problem. It then has to be meaningful so understandable to those intended consumers. And this is where we are can be very careful on the language that we choose when we are writing these explanations. Um, this is where a governed glossery comes in to really make sure everyone is interpreting this language the same so it's consistent across the organization or across stakeholders and we're not coming to different conclusions or making different assumptions on what particular terms in an explanation could mean that could deviate. Um then we have knowledge limits which are going to be really the scope and the boundaries for the operation of the system. So how do we know that a particular model is only answering questions that it has been built or sanctioned to answer? And so providing knowledge limits to that um to those teams is going to help when they're building the solution to make sure that they are flagging for in scope keywords and out of scope keywords that will help filter out any sort of input that is coming into the model that it may inference off of. And then fourth is explanation accuracy. And this can be further split into two two parts. The first is what we is that uncomfortable truth of today that we we can't actually see into the model. So we can't really take a look at for at a forensic level whether or not the explanation that for how the model came up with the output actually matches what we intended it to be. Um so instead what we need to be able to do is to provide something that is consistent that is governed and approved and is defensible record for all of the guidelines that we've put around the model in that system in order to influence its outcome. Hey Steph, looking at this for the uninitiated like myself, it looks like we can get pretty far down the to the principles without ever really looking at AI. It looks like the first three things here at least are governance activities that could happen perhaps well ahead of the actual build. Is that correct? It is which is leading into how the decision record becomes part of the configuration requirements and the design of the system itself not just an explanation for the output. Um so as we said we can't govern the model itself but we govern everything that is surrounding the model and NIST is giving us um a framework with which to do that. Okay. So, like I said at the beginning, let's see what the courts are saying about this because we all know that um AI regulation is coming fast and furious and some of these court cases are really directly setting the precedent for the kinds of conversations we're having today. We've talked about Moy versus Workday um a couple of times in this series and it kind of where it stands today, it kind of closes the loop on the it wasn't us defense that workday was trying to employ. Um and the courts have said it most certainly is. Uh and that litigation is ongoing. Um, but what we want to look at today is the to Stephanie's point about the other side of the coin, what's happening with Navy Federal because what's what is different about this is that we are asking to see those things that Stephanie just outlined. Particularly the fourth thing, what were the what is the reason that you decided and how you decided it? And the courts are now deliberating on whether that um internal proprietary variable set is um ex should be and can be exposed to the applicants. So taking this one step further from we can protect ourselves to how are we protecting our clients. So, if you'll remember in July, we talked about the credit decline. Um, Stephanie rephrased it as a credit approval, which is probably better. Uh, but we said that there were three things that needed to happen. And this is for internal use. The decision steward needed to understand if there was if the data was grounded, right? there was a approved list from which it could work. If we could trace it, meaning what version of the model, what version of the data, all of the technical capabilities that would isolate the variables in that realm and then owned. So, who is the process owner? Who is the person that can say um this is not right or this needs to be approved even though it didn't meet the criteria that we've already set? Again, this was driven toward the internal proof that this process was working. And that was the question we were answering in July. What we've added now, if we could go to the next slide, Steph is really this external consistency. So, the same list has to do more than just trace. Now we have to give real explainable and comprehendable uh results to our external clients. The check is now external instead of internal. So this is the July standard but faced outward. How we can do that is really the decision record that Stephanie's been talking about. So Steph, why don't you take us through what a decision record looks like? >> All right. So here it is the decision the infamous decision record that we've been discussing so far today or part of what it will look like. Um and this is really how we operationalize the ability for a model to be explained. And the decision record is a structured set of information captured to provide guidelines for AI at the moment that a consequential decision is made by the system. So at the moment that it is deciding and whether or not that decision is going to result in some sort of output like a judgment or a figure or a piece of collateral or it's going to then result in an action in more of an agentic system. Um there is a moment of consequential decision that is being made and that consequential decision reflects a business decision that is made. And so the decision record is really helping us to govern the process that we have from transferring human accountability over a decision to an to AI. And so when we look at a decision record, if we look at it, not compiling it after the fact of building the system, but if we walk through these fields and we consider them beforehand, it's also going to become configuration requirements um as well as that interpretation guide for the output. Um, and this is going to be a very valuable tool because one of the the challenges that we see with clients all the time now is sort of the idea or the this um constant challenge or moving goalpost of the definition of done or scope of a system. And the decision record is is an a tool that you can use that is going to help force you to scope a particular use case for an AI system. So it can be super powerful. It is a it is more work upfront to be able to fill something out, but it's going to save considerable time downstream when it comes to being able to design the build, define done, show progress, um, show meaningful progress to your stakeholders, as well as being able to provide, you know, one consolidated record that's going to serve multiple stakeholders to help them be able to interpret the results however they need. So they'll pull on different fields in this record to compile their story, but they're going within this record itself and then transfer it into the AI system is also going to be approved language and vocabulary. So coming from a governed lexicon again to make sure that we're really all singing from the same himsh sheet, if you will, with the approved language. Um, so this is where glosseries are super important. taxonomies, ontologies. Um, these are where the the the payoff for that work really begins to shine. >> Steph, I have a quick question. >> Yes. >> For decision records, is there will each agent only have one? >> So, that depends. Um what I each use case will have one. So each use case is going to be aimed at a specific decision that the organization wants made that is being transferred to AI. Um and that decision can be the result of a business process. It can be an individual task depending on whatever level of granularity. But we tie the decision record to the use case and not to the underlying model because in one model could serve multiple use cases particularly if the organization is going to pursue a componentization of strategy for AI um where they're going to sort of build individual components and then be able to deploy those across the organization um rather than building bespoke solutions from start to finish. Um, so I would say that the decision record is should be tied to the use case. >> All right, that's a thank you for that. And one question that came up in the chat that I wanted to throw out here as well. um in your context of of data a and the the data and IT world I is exp is explainability similar to to the data IT world of traceability or audit trails in order to get to how I have a thought on that but if you please you go first. So I would say yes. Again remembering that we the the model's explanation of its mechanics is going to be a plausible reconstruction rather than that deterministic log of what happened. So what we can do is we can provide sort of that root cause analysis of we we had this decision and this is the steps that we took to get to that decision or maybe not root cause necessarily but that chain of thought to get there >> right and and I think that makes a lot of sense because the chain of thought is really the explanability the the traceability is is what happened right it is the record of what exactly happened um when we think of audit logs and IT and we think of all of those things. So that's the kind of subtle difference. Um and again kind of on the continuum the the audit trail would come first the explanability and then you know and then so on. So thank you for the question. Hopefully that answers your question. >> All right. So the decision record therefore answers or serves three purposes and becomes a key human-based control for AI even if it isn't something that is necessarily interjected into the actual process of AI running. It is still a human-based control that is overseeing the entire solution. So it serves design and configuration purposes, execution purposes and then accountability. Uh so what is happening is that we are providing upfront what we want the solution to look like, what we want it to answer and how we want it to answer. at the execution time when it comes to run the what are the um key guard rails that are in place at runtime. So what data is it allowed to to access? Um what metadata does it have access to? What reason codes is it going to apply to the output? And then accountability is what are the human-based decisions that were made on the performance and the output of the model. So this is sort of that human in the loop or human on the spectrum whether it's over the loop on the loop in the loop. um it's how we are interjecting humans into this decisioning process that we've now taken from human decisioning and we've replaced it with AI decisioning and there this concept of decision stewardship that we've been talking about this entire series this is the decision record is the key artifact of that um it is helping to define those AI system boundaries it's helping to the organization to carefully consider what are the decisions that it is making about AI and the decisions that it's allowing AI to make itself. And so we can't again can't govern everything. We can't govern the model. So we govern what's around it. And in doing so, we're providing the operating conditions that help with explanability and then the information that's needed for stakeholders to interpret that consistently. >> Steph, there's a question in the chat. the decision record is tied to the use case rather than the model, which makes sense when one model can support multiple use cases. But when that shared model changes, how do you determine which decision records need to be revalidated and who owns that trigger? >> So, we did have the model inversion that needed to be documented at the time that this was built. And so when the model does get updated, there would need to be some sort of stewardship process to go back and identify which use cases were built off of that model that has then been changed. I don't I don't know if that answers the question if I understood the question enough to answer it. >> Yeah, I think that is that makes a lot of sense. And if the model inversion is stored on the decision record, it can be read um electronically and it can be part of the semantic layer so that it could be an onoff switch for the actual decision. I would think to say that if it's not running on this model in this version, route it to the steward for further review. So I think that could also happen. >> Yes. Yes. All of the fields on the decision record are also going they're not just going to be used for humans to read and interpret, but they're actually going to end up becoming system guardrails in and of themselves. So, some of these can get translated into almost a deterministic code in a way that can provide hard guard rails like a a kill switch if you will. And some of these are going to be provided to um AI through metadata and ontologies. um wherein then the model is going to sort of reason and infer over them rather than be constrained to a specific rule. >> Thanks. >> All right. So where can we start today or tomorrow with something that we have already perhaps already have built? So what we are going to do is look at this. Excuse me. Okay. Sorry. I think we had a little bit of a technical glitch on my end. I apologize. Um, it looks like a slide has dropped. So, I will go do my best to fill in the blank from here. Um, what we're So, the decision record, we go back here. It can be a little bit difficult to just sit down on a blank sheet of paper and try to come up with all of these answers at once. So we need somewhere we really need somewhere to start and you know to provide shape and direction for how we are going to fill out this decision record and that is where this concept of the competency question comes into play and the competency question is going to do two things. It is going to first define what are the questions that a stakeholder might have of the system. So how it is performing, how it is coming to its output really what are the mechanisms against how it is operating and then it's going to be competency questions to the system. So how are you actually expected to interact with it? How what are you what types of questions are you going to ask it as part of this use case? So we have those those two groups of of questions and when you sit down and we're you know we're actually actively doing this with a client now and we're working through an AI use case and we're enumerating all of the specific questions that the key stakeholder or the key users of that AI solution are going to be asking as part of that use case. And then we can go back to those regulator questions that we the board questions that we started with and say what are the audit questions that we typically get. And what's going to once we ask those questions and we build out a detailed answer form, what we're able to then do is um decompose those answers into the particular um pieces of the decision record that are going to help us answer those questions. It's going to help us identify what are the data domains that this answer is pulling on. What are the authoritative sources for those data domains and the data fields that we're going to need to pull on to answer that question. Are they governed? Are they approved? What are the particular reasons that we are allowing a um that uh there can be for a particular output? And you know, are is that a constrained list of reasons or are we allowing some sort of free flowing response? Um so those competency questions are the key for where uh we can begin with that decision record rather than sitting down and trying to fill out a form blank. Um you know they provide that starting point in shape and they also again help provide that scope and the boundaries for you know how do we define done and how do we define what is an acceptable question for this AI system to answer and what is something that is out of scope and not what this solution has been um trained or sanctioned for. So where do we start if we wanted to apply all of this today or tomorrow? So what we can do is we can go back and we can look at what we have running today and we can pick one decision flow. Let's write down who asks about it and what they're actually asking. So who are those stakeholders? Who are those? Um whether it's the um the internal user, the ultimate end user if that's external to the organization, a regulator, a board member, and what are the questions that they're asking? And write out those answers. Again, this is work upfront, but it is going to save considerable time downstream. And then anywhere where you have a gap where you're not able to answer that question or you're not able to um come up with the right questions to for that particular system or what that system is intended to be used for. There's a gap and that gap presents an opportunity and that is a real opportunity right now for that decision record to be deployed and for a decision stewardship to be deployed. And taking a step back and zooming out across all of the sessions that have happened thus far, there's a consistent framework that is has begun to emerge and it says this locate, assign, record, and scale. So locate or identify you know where that decision is being made by whom and how. So what is the problem that we need to solve for? Then the the second step is to assign who are the owners for that decision and the consulted parties who is especially responsible for any sort of exception queue or issue escalation. And then we record what needs to be captured. And we want to again we what we are trying to do is we are trying to influence here what AI can reason over in order to produce a trustworthy output. And then lastly we scale. And we're scaling this by governing the decision flow via the decision record not the individual decisions or outcomes each time that AI is queried. So we're able to sort of um take this up to an aggregate level and look across that entire use case to say this use case has been defined. It has been scoped. We have we are comfortable with the guidelines and so therefore we can deploy this rather than treating each and every output as sort of a spot check to see if this solution is working and then using some sort of qualitative scale to judge. Okay. So to recap today um we like to say you know governance is a combination of technology, people, process, policy and data. So from we have real technical limitations that are existing and that's something that we cannot change. We need to work with that. Um you know we cannot get that full explanability because we cannot actually see into the model itself. And so then therefore because explanability has been impacted therefore interpretability is impacted. And the way that we help mitigate those risks um you know is that we introduce the decision record and that's going to give us our our defensible basis for why the model behaved that it did. We're going to defend the decision record instead of needing to defend the the neural network layers. And then the competency questions become the input to that decision record. They're what um are going to help us be able to find a place to start to identify the use case to scope the use case. And no approved answers for those competency competency questions means that that decision flow isn't fully explainable yet. And then lastly, it is that decision stewardship is what produces that decision record. Um there decision stewards are first ensuring that that use case is valid for AI decisioning from the get-go and then it is facilitating the completion of that decision record and they're maintaining it over time. So what the decision stewardship here is really doing is it is helping to govern that transition of human authority over decisioning organizational decisioning to AI and I'll hand it back to Lisa. >> All right. So back to the boardroom we go and here is our mic drop answer for today's session. We can answer both of these questions honestly by saying we can't tell you how the model reasoned again but we can tell you the basis for decisioning the operating environment and the human involved every time not just once. And I think that is the honest and best answer you can give to the people at the top of your organization and to your regula your regulators and your internal auditors. All right, Steph, there's a question about um where these records live and you know I think I we've had some conversations about this and um you know what we what we've thought about and how we've kind of worked through that. So I think at first we thought about a cat the catalog for say the left side of the decision record right. Um and then the right side would be in an appendon data store somewhere right outside of the model itself uh for obvious you know reasons. Um I wonder is that still the way we would direct people or do you have some new ideas on that? tough questions today. Thanks everyone. >> I think that that is that is still the case. Um because we need to be able to separate that front end that that business reasoning from sort of that those functional those those technical realities or requirements that are happening particularly when we start getting into queries that are going to have a temporal aspect to them. We are going to need a robust system that is going to be able to hold and and log essentially the the the state of reality that existed every single time a decision record was approved for use. >> That makes that makes sense. Um I'm looking So, someone asked the question that define done piece really caught me. If the decision record establishes what the AI is sanctioned to answer, does that mean that an AI system could technically per be performing perfectly while still failing governance because it's answering questions outside its authorized decision boundary? So that is where can you read that again because I think there's a couple of components to that question. >> Okay. >> Define the define done piece really caught me. If a decision record establishes what the AI is sanctioned to answer, does that mean an AI system could be per technically performing perfectly while still failing governance because it's answering questions outside its authorized decision boundary? >> Okay, I'm sorry I misheard that. Um, so this scope flag, Whoopsies, excuse me. This scope flag is what is meant to constrain and really provide the governance around what this particular model is sanctioned to do for that use case. Um, and so this is where that the this is where language is so important. Um, especially because so many of these models were interacting with them through a large language model interface essentially. Um, so what we would then do is we would look at those competency questions and those answers and we're going to pick apart the pieces of that that represent specific uh data domains, data concepts and we're going to um enumerate those in the decision record as we walk through that chain of logic and as we've documented those those questions and those answers. When that gets handed off to the technical teams, um, what they're able to then do is pull from an established ideally an established semantic layer that is going to serve as that translation layer that connects those that that approved language with the physical instantiation of the data. So that when the technical teams for the AI solution go to grab it, it now has sort of these parameters around look for these keywords. If those keywords are not present in the question or in the query or if we can't reasonably find a close enough relation to that, we're going to then, you know, sort of provide this response of this is out of scope for this particular solution. Got it. Okay. Trying to go through some of these other questions. Thank you so much everyone for filling up the qu the Q&A and the chat. Um overall, here's another question. Overall, how does explanability and interpretability differ for third-party AI products versus those developed internally within a company? Um, so that is ultimately going to come down to how and you're going to need to work with your technical teams and your engineering teams on this because it will depend on the model the model version and the the platform as well as the vendor. So if you're building something inhouse, you have really sort of full um ability to customize like we we were talking about that semantic layer. And so that semantic layer is going to enable um because all of this decision record needs to get translated into metadata and that metadata then feeds that semantic layer and that's what's allowing that um the connection of this language as what to the physical data. And so you are um can you ask can you ask a question? Can I completely just dropped my train of thought trying to think through this? >> How does explanability and interpretability differ from third party apps versus what you develop in terms? >> Yes. So, um, you're going to have more ability to then to translate these decision records into pieces of metadata that can then get tagged to the physical data itself. Whereas, um, different vendors are going to have restrictions on what metadata that you're allowed to access and what data metadata you're allowed to pass over to their systems. >> All right, I'm gonna read one and answer it so you can take a break. Um, so Ed asked, "It seems this approach is solid for providing specific answers. However, it does not seem sustainable or even feasible at large scale without some kind of implementation architecture or structural framework and automation to get all of these details, answers, traceability, interpretation ready and available. So yes, I would agree with we would agree with that Ed um that this is the way that this is the process but without doing this in a way that is sustainable and scalable um it's not going to work. Uh this is where we roll back to the conversation about where do these things get these um types of records get stored and how do they get used. um you you still need to in your organization I think you know way back in the beginning we talked about the critical decisions you're making about AI and some of those are which are what are the decisions we will never let the organization we will never let AI make on behalf of this organization which are the ones that we will always let AI make on behalf of this organization and and which ones fall in the middle and that gives you your that gives you your road map for installing these decision records and the PE the stewards around them. Um it's it's really that criticality and the blast radius if it goes wrong which is something else we've talked about that you need to to use to determine where you know where this has its place um and where it can autogenerate if you will. But I would even say if the decision record is autogenerated, first of all, it needs to be kept separate from the uh the actual AI doing the work like we've discussed, but second of all, at some point there has to be a parameter in which someone takes a look at it. Um but would not argue that scale needs to happen. um but within a contained environment that is completely aligned to the risk pro profile that your organization has. I think there's also >> to add >> I think there's also another uncomfortable reality in that AI is has s the ability to perfectly govern AI has surpassed where we are able what we are able to do today and it's forcing real innovation and real investment in organizations in order to be able to to use this responsibly at scale And we see all of our our clients struggling with this. All right, let's see what else do we have here. Mark, what did I what did I miss? >> Well, there was just so much debate in chat and everybody was uh was so uh engaged today, which is wonderful. Uh there's still a handful in Q&A. Um let's let's tackle this one. Do you think the explanability and interpretability still come down to what data really is in the era of uh large language models? Uh is data which I think is a a bit of an interesting take. Yes and no. Um, which is a terrible answer usually for people um to hear but um I I think yes and no because I think from what we can understand now is if you are running a semantic layer that is where we can deterministically tie language to physical data. Yeah, >> if you are using models that either do not allow configuration of the semantic layer or you are just using base models off the shelf, then no, it's going to rely completely on inference um in order to make that connection. >> Yeah, I I love that answer, Stephanie. I think you're you're right on. Somebody in chat just said it feels like we're in the realm of philosophy. What is this anyway? >> Yes. Yes. It it um it's really interesting when you begin to sort of um really explore the mechanics of AI, you're really getting also into um psychology because it really AI is artificial intelligence. It is meant to mimic the thought process of humans. Um so when we talk about you know chain of thought and being able to train AI and work with AI um you really need to first understand your business problem and your business process >> and you need to be able to model that and you need to be able to represent that both in language and in data in order to translate it over into AI. Well, then in LLM land, we have uh uh these generative models that speak with a confidence and fluency that isn't really anchored by any authoritative reality. So, it really gets into philosophy when we talk about like a manual can't like it doesn't think, therefore it isn't. >> The Dunan Krueger. >> Yeah. Yeah. Yeah. Yeah. Exactly. And that's where one of the the words that I brought up at the beginning um recently read a fascinating article about something called persuasion bombing. And this is where you question the the LLM and it starts to try to persuade you that you're wrong. Um and and almost goes to the point of like fiction to uh to to do that. um because it's learning to think like a human and therefore it's trying to persuade you like a human would that its point of view is correct. So uh it you know it's governing an immature capability is fascinating in that you think you know it and then the next day uh something changes your thought process entirely. I uh we we only have a couple minutes left and and one of the question this isn't a question. One of the questioners just put this statement in chat. People are upvoting it. So uh I want to get your take on this. Um and their statement is we should not transfer human accountability to AI. What are your thoughts? >> Do you want to go first, Lisa? >> No, go ahead. Um I completely agree. I think it depends on what degree of separation between human and AI. That's where we could have a debate. Um ultimately humans should absolutely be accountable for the decision on whether or not to deploy AI in what scenario under what conditions and what to do if something goes wrong. Um, but this is again where I we go back to that chain of thought and that business process modeling. If you truly understand your business process and you've truly broken that down and you understand the risk risk aspect of that decision, then you're able to then say with a certain level of confidence, I will let AI be responsible for generating a particular out outcome or output rather, but I'm still accountable for how this particular output gets applied or used. Um, so there I I personally believe that there are decisions that AI should never be making. Um, but then I think if you say, well, AI can be responsible for summarizing my meeting notes and sending them to my colleagues, I have no problem with that. And I think that again goes back to what the law the case law is showing us is that um that is absolutely true that that the the these cases are showing us that um that I the AI did it is not an acceptable answer in any case, right? Um so and and it's kind of like what you learn in school, right? Plagiarism is plagiarism and and you you need to be able to defend what you create. So, um it seems like it's cut and dried, but I I agree with you, Stephanie. There can be shades of gray in certain in certain places, and that's what we're still looking to discover and figure out. >> All right. Well, that brings us pretty much to the end. We've got Scant Nar a second left. Um, thank you again everybody in the community for being so electric and both chat and Q&A. There's a ton we didn't get to. Um, uh, so that'll make some interesting reading for all of us later as we go through uh the the transcripts. Um, any final thoughts before I hit the end webinar button team? >> Just thank you. Thanks. Thank you all for making this a really interactive session. These are the best kind. So, I appreciate your your feedback and your questions. >> Thank you. >> Bye.