Leading IAGOV: The Right to Explanation—Meeting Regulatory Demands for Interpretable AI
Watch on YouTubeVideo summary
The webinar "Leading IAGOV: The Right to Explanation—Meeting Regulatory Demands for Interpretable AI," hosted by Mark Horseman with Lisa Winrich and Stephanie Parody, explores the critical distinction between technical explainability and stakeholder interpretability within the context of growing regulatory pressures. A central argument presented is that organizations cannot fully govern or inspect probabilistic AI models because requesting logic from them yields only a reconstructed "best guess" rather than a deterministic log, making full explainability unattainable. Consequently, interpretability remains inherently subjective to the specific stakeholder, necessitating a shift in focus from the internal mechanics of the model to the environment surrounding it. To address these limitations and mitigate associated risks, the speakers introduce the concept of a "decision record," a structured artifact that serves as a design specification, execution guide, and accountability mechanism. This record captures essential business reasoning, enforces approved vocabulary through governed glossaries, and clearly defines knowledge limits to ensure consistent interpretation across all stakeholders.
To operationalize this decision record effectively, organizations are encouraged to utilize "competency questions" that define exactly what stakeholders expect the system to answer, thereby scoping use cases and establishing boundaries for AI responses. This approach aligns with the NIST Risk Management Framework, which outlines four core principles: providing evidence and reason, ensuring outputs are meaningful and understandable, defining the scope of knowledge limits, and acknowledging the inherent inability to verify internal model mechanics. The discussion also highlights ongoing legal precedents, such as litigation involving Workday and Navy Federal, where courts are currently deliberating whether proprietary internal variables must be exposed to applicants. In response to these challenges, a scalable implementation strategy involves identifying critical decisions, assigning dedicated decision stewards, and establishing system guardrails that translate human-based controls into deterministic code or metadata constraints, moving away from inefficient spot-checking of individual outputs.
The conversation further examines the nature of Large Language Models (LLMs) and whether they rely on data or inference to connect language with physical reality. The consensus indicates that while semantic layers can create deterministic ties between language and data, base models lacking such configuration rely entirely on inference, resulting in outputs not anchored in authoritative reality. This reliance raises philosophical questions about AI mimicking human thought processes, exemplified by the concept of "persuasion bombing," where LLMs simulate human reasoning to convince users they are wrong. Regarding accountability, the speakers agree that humans must remain responsible for deploying AI and managing outcomes, even when the technology generates specific outputs like meeting summaries. This stance mirrors legal precedents establishing that claiming "the AI did it" is not an acceptable defense, similar to plagiarism in academic settings, though a nuanced view allows for scenarios where AI bears responsibility for generation while humans retain accountability for application.
In conclusion, the session acknowledges the complex shades of gray surrounding these emerging issues and emphasizes the need for a balanced approach that respects both technical limitations and regulatory demands. By integrating decision records, competency questions, and robust governance frameworks, organizations can navigate the evolving landscape of interpretable AI without compromising on safety or accountability. The discussion ultimately thanks the audience for their engagement, leaving them with a clear understanding that while full transparency into model mechanics may be impossible, structured human oversight and defined boundaries provide a viable path forward for meeting regulatory expectations in an increasingly automated world.
Read the full video transcript
Hello and welcome. My name is Mark
Horseman and I am the data evangelist
for data. We would like to thank you for
joining today's data webinar, the right
to explanation, meeting regulatory
demands for interpretable AI. It is the
latest in the monthly webinar series
leading AI governance with first San
Francisco partners. Just a couple of
points to get us started. Due to the
large number of people that attend these
sessions, you will be muted during the
webinar. For questions, we will be
collecting them by the Q&A section. If
you would like to chat with us or chat
with each other, we certainly encourage
you to do so. To open the Q&A or the
chat panel, you'll see the icons for
those features in the bottom middle of
your screen to answer the most commonly
asked question. As always, we will send
a follow-up email to all registrants
within a couple of business days
containing links to the slides. And yes,
we're recording and we'll likewise send
a link to the recording as well as any
additional information requested
throughout the webinar. Joining us today
is Lisa Winrich and Stephanie Parody. Uh
Lisa is the executive adviser at First
San Francisco Partners where she advises
Sea Level's executives on data and AI
governance and leads the firm's
innovation and service department
practice focused on AI governance and
semantics. Her work centers on the
decisions and accountability structures
that determine whether data and AI
investments hold up, who owns which
decisions, what data means across the
business, what data is fit for a given
AI use, and how governance operates as a
working practice rather than a policy
document. She specializes in building
capabilities where none exist. Before
FSFP, Alisa spent over two decades at
Sherman Williams building data
management, analytics, and AI
capabilities. She made analytics and
data science directly usable by business
teams without IT intervention. Backed by
the data literacy and governance
programs that made it safe at scale, she
built and led teams behind those
capabilities with a sustained focus on
career pathing, upskilling and coaching.
She also held technology leadership
roles across data network application
development and point of sale. Her work
spanned manufacturing, finance,
marketing, sales, and procurement.
Stephanie is a principal consultant
specializing in building and maturing
data governance programs through the
effective unification of business and
technology through data. She brings
hands-on experience from previous
marketing and data stewardship roles in
addition to CDMP certification.
Stephanie is an accomplished data
governance leader highly adept at
helping clients transition between
strategic vision and tactical execution.
and her experience spans across the
various subd discciplines of data
governance and data management with a
focus on how to activate right-sized
culturally relevant data enablement
programs. With the uptick in focus on AI
governance, Stephanie is actively
responding to the required paradigm
shifts in traditional government and
management frameworks and approaches.
And with that, let me hand it over to
Lisa and Stephanie. Welcome, my friends.
>> Welcome. Thank you. Welcome to everyone
here. Good afternoon, good morning, and
good evening to some of you. Um, we're
so excited to be here. Uh, as you can
see, our fearless leader is not here
today, but she's here in spirit. She is
taking some well-deserved time off. So,
um, she has left you in what we hope are
capable hands uh, of Stephanie and I.
So, um, today we are going to dive into
a number of things. Um, but as I was
looking back, uh, stuff, if we could go,
we're seeing the, um,
the Could you put up the, uh,
>> that could be my fault.
>> Yeah,
>> because I had to reboot Zoom earlier.
>> Messing up. All right, it is mean minute
two and Mark has already messed up our
presentation. Kelly is going to be
furious.
>> It's It's all my fault though. Try again
on the share. My apologies.
Okay,
let me
>> Okay, here we go again.
>> Awesome. I see I see the deck.
>> Looks wonderful.
>> Okay, perfect.
>> Now it looks perfect. Thank you,
Stephanie, for cleaning up for Mark's
mistake there. Um, so if we go to the
next slide, uh, Steph please, and we
talk about where we've been. This is
session nine, which I can't believe. Um,
thank you to those of you who have been
with us through nine sessions, and those
of us those of you that are just here
for the first time, don't worry, we'll
catch you up. Um, as I was looking
through some of these prior titles, one
of the things that caught my eye was the
very first session when we talked about
cutting through the noise and thinking
about the um words, the the vocabulary
words, the bingo card we created in that
session and how far we've come just in
lexicon across the industry over the
last nine months. Now we have words like
slap economy and persuasion bombing and
botsitting. All very interesting new
terms. So maybe we'll do another bingo
card soon because the the field keeps
changing. Um so today we're going to
talk about the right to explanation and
we're going to talk about the difference
between interpretability and
explainability which Stephanie will
explain because I can't interpret it. Um
if we go to the next slide we can talk
let's talk about where this fits. So
if you recall in July in July we talked
about decision clarity. We talked about
owning tracing and naming as a way for
the internal audience to understand what
they're signing up for. And this is
where the decision steward comes in um
and is the kind of coaleser and owner of
this process.
When we went to August, we talked a
little bit about decision interrogation
and we also had a stellar guest speaker
in Amy Bojack from Baker Tilly uh who is
the CISO there and she took us down this
path as well as several others um just
in a really great flea free flowing
conversation. If you haven't seen that
episode, it's one that I would highly
recommend. She really broke things down
in a way that was accessible for those
of us outside of her realm.
So today we're going to talk about
regulating decisions who for the people
that are affected by it, the people on
the other side. Um what do we owe to
them? And we're going to look at some
case law that is setting precedent right
now for this very topic.
Next month we're going to talk about the
mechanics. So demonstrating how things
work. So I just want to give you that
little preview and know that we will
talk about that next time.
All right. So getting into it. Here we
are back in our boardroom where we've
been kind of at the beginning of the
last couple of sessions. But here's
where we have a couple of different
questions um being asked the same
question being asked two different ways.
Right? So, if we have a regulator ask
asking us why we declined this
applicant, can we answer it for that
specific case? But then the board wants
to know this was one applicant. How do
we know that we can answer this question
every time? Um, and that's really what
we're going to break down here and be
able to understand if we can say yes
with confidence.
>> Right. So today we're going to talk
about two sides of the same coin,
interpretability and explanability.
And there are many similar words around
these topics that are all used
interchangeably. But today we're going
to dig into the nuances and why that
matters. We're going to look at how we
can improve and defend explanability
using a method that defines its
requirements once so that we can
reference that same output for multiple
stakeholders and we're not going back
and reinventing the wheel or trying to
reexplain something multiple times. And
we're also going to look at real
technical limitations that exist so that
we're not chasing pipe dreams, but we're
also not shifting organizational
responsibility onto the vendors. we know
exactly what we are responsible and
accountable for and we're able to defend
that and then we're able to push back on
the industry where it makes sense
because that's where the real
limitations lie.
So the whole premise of the session
today rests on what is really an
uncomfortable reality in the industry
today and that is you cannot see into
the model itself to know how it is
operating. You can't open it up and
explore hundreds of thousands or
millions of layers of neural networks
and compounding that uncomfortableness.
If you ask a model how it got to its
decision or its output, you're actually
getting its best guess at reconstructing
its logic, you're not getting an actual
real time decision log, which a lot of
us assume that we are. Um, and so
therefore, we can never really achieve
full explanability of a model because we
can never fully explain the mechanics of
how it's operating.
And because of that then in consequence
to that the ability for this to then
interpret that output is impeded. We
can't say with a 100% certainty that
we've interpreted something correctly or
we've interpreted all the possible
facets and inputs or if we don't know
how that output was created and how it
led to that particular output or
decision.
So essentially what we're saying is that
we can't govern the model and because of
that there are real risks that
organizations are forced to accept when
they adopt AI particularly anything that
is probabilistic in nature.
So to counterbalance this is where
governance steps in and really shines
and really has its value moment.
governance is going to help govern
everything outside of and around the
model in the absence of being able to
actually govern the model itself.
And we do this through something called
the decision record. And the decision
record is an incredibly powerful tool
that is not only going to help us in
explanability and interpretability of
the solution, but also it's going to
serve as a design and configuration
specification, which is going to be
super useful when it comes to helping to
close that gap between business need and
any assumptions that the technical teams
might be making or may be forced to make
based on technical limitations.
So, we'll get into the decision record
more in a little bit, but first let's
hone in on how and why explainability
and interpretability aren't the same
thing, but are related.
So, before I said that they were two
sides of the same coin, but probably
more accurately is that they're really a
sequence of events that go hand in hand.
For this session, we're going to anchor
on NIST and we're going to use that as
our authoritative framework throughout
the entire session. So all of the
definitions and all of the concepts that
we reference are going to be rooted in
that framework.
And in the riskmanagement framework,
NIST presents seven characteristics of
trustworthy AI systems.
And explainable is the first in that
sequence of bucketed, explainable, and
interpretable because it forms the
foundation for that entire
characteristic.
Explainable refers to how a decision was
made in the system. It's really
referring to the mechanics of how the AI
got to that particular output or why it
made that decision that resulted in the
output is how we can think of that.
Interpretable is then the consequence of
explainable and it's going to answer why
a decision was made by the system. Um
it's really going to be on the
stakeholder of how they interpret the
output and how they interpret the
supporting information in order to be
able to make a determination on whether
or not that output or the decision that
was made is acceptable for their use.
So explainable is really a capability of
the model itself. It's something that's
inherent to it. So it's a little bit
more objective in nature. Whereas on the
other hand, interpretable since it rests
on the stakeholder, it's really more
subjective in nature and so is therefore
a user capability.
And since we can't control how a
stakeholder is going to interpret the
model's explanation,
we can't get into their mind and force
that explanation. But what we can do is
we can influence it. And that's what
we're going to use the decision record
for.
So let's talk through an example of this
to sort of um solidify these these two
related terms. So explanability is going
to answer the question, how was this
decision made? And so the answer is
going to be something along the lines of
debt to income ratio came out at 30% was
checked against the approval threshold
which is 35% and was flagged as being
acceptable.
Interpretable is then going to answer
the question of why does that matter and
what does it mean to me? And so
interpretability is going to be what
someone is able to interpret or make
sense of based off of that explanation
and the supporting information. So
you're they're going to be able to
interpret that it was the ratio and not
someone's credit history, for example,
for why the loan was approved.
So, a little bit of a
illustration there,
but those are in explanability and
interpretability. If they're
characteristics of an AI system, their
qualities or their properties, they
don't explicitly say what needs to
happen in order to achieve them, which
can make creating a roadmap to achieving
those difficult for organizations.
And this is where NIST comes in again to
help. So an accompanying document to
that riskmanagement framework is a
document called four principles of um
explanability.
And what it does is it takes that
characteristic and breaks it down into
four specific requirements and what we
can do as an organization is we can
actually execute against each of those
requirements. So we first need to
baseline requirement is we need to be
able to provide evidence or reason for
how and why AI came with its output and
what this is going to do is this is
going to mirror the human process of how
we would work through solving this
particular problem.
It then has to be meaningful so
understandable to those intended
consumers. And this is where we are can
be very careful on the language that we
choose when we are writing these
explanations. Um, this is where a
governed glossery comes in to really
make sure everyone is interpreting this
language the same so it's consistent
across the organization or across
stakeholders and we're not coming to
different conclusions or making
different assumptions on what particular
terms in an explanation could mean that
could deviate.
Um then we have knowledge limits which
are going to be really the scope and the
boundaries for the operation of the
system. So how do we know that a
particular model is only answering
questions that it has been built or
sanctioned to answer? And so providing
knowledge limits to that um to those
teams is going to help when they're
building the solution to make sure that
they are flagging for in scope keywords
and out of scope keywords that will help
filter out any sort of input that is
coming into the model that it may
inference off of.
And then fourth is explanation accuracy.
And this can be further split into two
two parts. The first is what we is that
uncomfortable truth of today that we we
can't actually see into the model. So we
can't really take a look at for at a
forensic level whether or not the
explanation that for how the model came
up with the output actually matches what
we intended it to be. Um so instead what
we need to be able to do is to provide
something that is consistent that is
governed and approved and is defensible
record for all of the guidelines that
we've put around the model in that
system in order to influence its
outcome.
Hey Steph, looking at this for the
uninitiated like myself, it looks like
we can get pretty far down the to the
principles without ever really looking
at AI. It looks like the first three
things here at least are governance
activities that could happen perhaps
well ahead of the actual build. Is that
correct?
It is which is leading into
how the decision record becomes part of
the configuration requirements and the
design of the system itself not just an
explanation for the output.
Um so as we said we can't govern the
model itself but we govern everything
that is surrounding the model and NIST
is giving us um a framework with which
to do that.
Okay. So, like I said at the beginning,
let's see what the courts are saying
about this because we all know that um
AI regulation is coming fast and furious
and some of these court cases are really
directly setting the precedent for the
kinds of conversations we're having
today. We've talked about Moy versus
Workday um a couple of times in this
series and it kind of where it stands
today, it kind of closes the loop on the
it wasn't us defense that workday was
trying to employ. Um and the courts have
said it most certainly is. Uh and that
litigation is ongoing. Um, but what we
want to look at today is the to
Stephanie's point about the other side
of the coin, what's happening with Navy
Federal because what's what is different
about this is that we are asking to see
those things that Stephanie just
outlined. Particularly the fourth thing,
what were the what is the reason that
you decided and how you decided it? And
the courts are now deliberating on
whether that um internal proprietary
variable set is um ex should be and can
be exposed to the applicants. So taking
this one step further from we can
protect ourselves to how are we
protecting our clients.
So, if you'll remember in July, we
talked about the credit decline. Um,
Stephanie rephrased it as a credit
approval, which is probably better. Uh,
but we said that there were three things
that needed to happen. And this is for
internal use. The decision steward
needed to understand
if there was if the data was grounded,
right? there was a approved list from
which it could work. If we could trace
it, meaning what version of the model,
what version of the data, all of the
technical capabilities that would
isolate the variables in that realm and
then owned. So, who is the process
owner? Who is the person that can say um
this is not right or this needs to be
approved even though it didn't meet the
criteria that we've already set? Again,
this was driven toward the internal
proof that this process was working. And
that was the question we were answering
in July. What we've added now,
if we could go to the next slide, Steph
is really this external consistency. So,
the same list has to do more than just
trace. Now
we have to give real explainable and
comprehendable uh results to our
external clients.
The check is now external instead of
internal.
So this is the July standard but faced
outward.
How we can do that is really the
decision record that Stephanie's been
talking about. So Steph, why don't you
take us through what a decision record
looks like?
>> All right.
So here it is the decision the infamous
decision record that we've been
discussing so far today or part of what
it will look like. Um and this is really
how we operationalize the ability for a
model to be explained. And the decision
record is a structured set of
information captured to provide
guidelines for AI at the moment that a
consequential decision is made by the
system. So at the moment that it is
deciding and whether or not that
decision is going to result in some sort
of output like a judgment or a figure or
a piece of collateral or it's going to
then result in an action in more of an
agentic system. Um there is a moment of
consequential decision that is being
made and that consequential decision
reflects a business decision that is
made. And so the decision record is
really helping us to govern the process
that we have from transferring human
accountability over a decision to an to
AI.
And so when we look at a decision
record, if we look at it, not compiling
it after the fact of building the
system, but if we walk through these
fields and we consider them beforehand,
it's also going to become configuration
requirements
um as well as that interpretation guide
for the output. Um, and this is going to
be a very valuable tool because one of
the the challenges that we see with
clients
all the time now is sort of the idea or
the this um
constant challenge or moving goalpost of
the definition of done or scope of a
system. And the decision record is is an
a tool that you can use that is going to
help force you to scope a particular use
case for an AI system. So it can be
super powerful. It is a it is more work
upfront to be able to fill something
out, but it's going to save considerable
time downstream when it comes to being
able to design the build, define done,
show progress,
um, show meaningful progress to your
stakeholders, as well as being able to
provide, you know, one consolidated
record that's going to serve multiple
stakeholders to help them be able to
interpret the results however they need.
So they'll pull on different fields in
this record to compile their story, but
they're going within this record itself
and then transfer it into the AI system
is also going to be approved language
and vocabulary. So coming from a
governed lexicon again to make sure that
we're really all singing from the same
himsh sheet, if you will, with the
approved language. Um, so this is where
glosseries are super important.
taxonomies, ontologies. Um, these are
where the the the payoff for that work
really begins to shine.
>> Steph, I have a quick question.
>> Yes.
>> For decision records, is there will each
agent only have one?
>> So, that depends.
Um
what I each use case will have one. So
each use case is going to be aimed at a
specific decision that the organization
wants made that is being transferred to
AI. Um and that decision can be the
result of a business process. It can be
an individual task depending on whatever
level of granularity. But we tie the
decision record to the use case and not
to the underlying model because in one
model could serve multiple use cases
particularly if the organization is
going to pursue a componentization of
strategy for AI um where they're going
to sort of build individual components
and then be able to deploy those across
the organization
um rather than building bespoke
solutions from start to finish. Um, so I
would say that the decision record is
should be tied to the use case.
>> All right, that's a thank you for that.
And one question that came up in the
chat that I wanted to throw out here as
well. um in your context of of data a
and the the data and IT world I is exp
is explainability similar to to the data
IT world of traceability or audit trails
in order to get to how
I have a thought on that but if you
please you go first. So I would say yes.
Again remembering that we the
the model's explanation of its mechanics
is going to be a plausible
reconstruction rather than that
deterministic log of what happened. So
what we can do is we can provide sort of
that root cause analysis of we we had
this decision and this is the steps that
we took to get to that decision or maybe
not root cause necessarily but that
chain of thought to get there
>> right and and I think that makes a lot
of sense because the chain of thought is
really the explanability the the
traceability is is what happened right
it is the record of what exactly
happened um when we think of audit logs
and IT and we think of all of those
things. So that's the kind of subtle
difference. Um and again kind of on the
continuum the the audit trail would come
first the explanability and then you
know and then so on. So thank you for
the question. Hopefully that answers
your question.
>> All right.
So the decision record
therefore answers or serves three
purposes and becomes a key human-based
control for AI even if it isn't
something that is necessarily
interjected into the actual process of
AI running. It is still a human-based
control that is overseeing the entire
solution. So it serves design and
configuration purposes, execution
purposes and then accountability.
Uh so what is happening is that we are
providing upfront what we want the
solution to look like, what we want it
to answer and how we want it to answer.
at the execution time when it comes to
run the what are the um key guard rails
that are in place at runtime. So what
data is it allowed to to access? Um what
metadata does it have access to?
What reason codes is it going to apply
to the output?
And then accountability is what are the
human-based decisions that were made on
the performance and the output of the
model. So this is sort of that human in
the loop or human on the spectrum
whether it's over the loop on the loop
in the loop. um it's how we are
interjecting humans into this
decisioning process that we've now taken
from human
decisioning and we've replaced it with
AI decisioning
and there this concept of decision
stewardship that we've been talking
about this entire series this is the
decision record is the key artifact of
that um it is helping to define those AI
system boundaries it's helping to the
organization to carefully consider what
are the decisions that it is making
about AI
and the decisions that it's allowing AI
to make itself.
And so we can't again can't govern
everything. We can't govern the model.
So we govern what's around it. And in
doing so, we're providing the operating
conditions that help with explanability
and then the information that's needed
for stakeholders to interpret that
consistently.
>> Steph, there's a question in the chat.
the decision record is tied to the use
case rather than the model, which makes
sense when one model can support
multiple use cases. But when that shared
model changes, how do you determine
which decision records need to be
revalidated and who owns that trigger?
>> So, we did have the model inversion
that needed to be documented at the time
that this was built. And so when the
model does get updated, there would need
to be some sort of stewardship process
to go back and identify which use cases
were built off of that model that has
then been changed. I don't I don't know
if that answers the question if I
understood the question enough to answer
it.
>> Yeah, I think that is that makes a lot
of sense. And if the model inversion is
stored on the decision record, it can be
read um electronically and it can be
part of the semantic layer so that it
could be an onoff switch for the actual
decision. I would think to say that if
it's not running on this model in this
version, route it to the steward for
further review. So I think that could
also happen.
>> Yes. Yes. All of the fields on the
decision record are also going they're
not just going to be used for humans to
read and interpret, but they're actually
going to end up becoming system
guardrails in and of themselves. So,
some of these can get translated into
almost a deterministic code in a way
that can provide hard guard rails like a
a kill switch if you will. And some of
these are going to be provided to um AI
through metadata and ontologies.
um wherein then the model is going to
sort of reason and infer over them
rather than be constrained to a specific
rule.
>> Thanks.
>> All right. So where can we start today
or tomorrow with something that we have
already perhaps already have built? So
what we are going to do is look at
this.
Excuse me.
Okay.
Sorry. I think we had a little bit of a
technical glitch on my end. I apologize.
Um, it looks like a slide has dropped.
So, I will go do my best to fill in the
blank from here. Um,
what we're So, the decision record, we
go back here. It can be a little bit
difficult to just sit down on a blank
sheet of paper and try to come up with
all of these answers at once. So we need
somewhere we really need somewhere to
start and you know to provide shape and
direction for how we are going to fill
out this decision record and that is
where this concept of the competency
question comes into play and the
competency question is going to do two
things. It is going to first define what
are the questions that a stakeholder
might have of the system. So how it is
performing, how it is coming to its
output really what are the mechanisms
against how it is operating and then
it's going to be competency questions to
the system. So how are you actually
expected to interact with it? How what
are you what types of questions are you
going to ask it as part of this use
case? So we have those those two groups
of of questions and when you sit down
and we're you know we're actually
actively doing this with a client now
and we're working through an AI use case
and we're enumerating all of the
specific questions that the key
stakeholder or the key users of that AI
solution are going to be asking as part
of that use case. And then we can go
back to those regulator questions that
we the board questions that we started
with and say what are the audit
questions that we typically get. And
what's going to once we ask those
questions and we build out a detailed
answer form, what we're able to then do
is um decompose those answers into the
particular um pieces of the decision
record that are going to help us answer
those questions. It's going to help us
identify what are the data domains that
this answer is pulling on. What are the
authoritative sources for those data
domains and the data fields that we're
going to need to pull on to answer that
question. Are they governed? Are they
approved? What are the particular
reasons that we are allowing a um that
uh there can be for a particular output?
And you know, are is that a constrained
list of reasons or are we allowing some
sort of free flowing response? Um so
those competency questions are the key
for where uh we can begin with that
decision record rather than sitting down
and trying to fill out a form blank. Um
you know they provide that starting
point in shape and they also again help
provide that scope and the boundaries
for you know how do we define done and
how do we define what is an acceptable
question for this AI system to answer
and what is something that is out of
scope and not what this solution has
been um trained or sanctioned for.
So where do we start if we wanted to
apply all of this today or tomorrow? So
what we can do is we can go back and we
can look at what we have running today
and we can pick one decision flow.
Let's write down who asks about it and
what they're actually asking. So who are
those stakeholders? Who are those? Um
whether it's the um the internal user,
the ultimate end user if that's external
to the organization, a regulator, a
board member, and what are the questions
that they're asking?
And write out those answers. Again, this
is work upfront, but it is going to save
considerable time downstream.
And then anywhere where you have a gap
where you're not able to answer that
question or you're not able to um come
up with the right questions to for that
particular system or what that system is
intended to be used for. There's a gap
and that gap presents an opportunity and
that is a real opportunity right now for
that decision record to be deployed and
for a decision stewardship to be
deployed.
And taking a step back and zooming out
across all of the sessions that have
happened thus far, there's a consistent
framework that is has begun to emerge
and it says this locate, assign, record,
and scale. So locate or identify you
know where that decision is being made
by whom and how.
So what is the problem that we need to
solve for? Then the the second step is
to assign who are the owners for that
decision and the consulted parties who
is especially responsible for any sort
of exception queue or issue escalation.
And then we record what needs to be
captured. And we want to again we what
we are trying to do is we are trying to
influence here what AI can reason over
in order to produce a trustworthy
output.
And then lastly we scale. And we're
scaling this by governing the decision
flow via the decision record not the
individual decisions or outcomes each
time that AI is queried. So we're able
to sort of um take this up to an
aggregate level and look across that
entire use case to say this use case has
been defined. It has been scoped. We
have we are comfortable with the
guidelines and so therefore we can
deploy this rather than treating each
and every output as sort of a spot check
to see if this solution is working and
then using some sort of qualitative
scale to judge.
Okay. So to recap today um we like to
say you know governance is a combination
of technology, people, process, policy
and data. So from we have real technical
limitations that are existing and that's
something that we cannot change. We need
to work with that. Um you know we cannot
get that full explanability because we
cannot actually see into the model
itself. And so then therefore because
explanability has been impacted
therefore interpretability is impacted.
And the way that we help mitigate those
risks um you know is that we introduce
the decision record and that's going to
give us our our defensible basis for why
the model behaved that it did. We're
going to defend the decision record
instead of needing to defend the the
neural network layers.
And then the competency questions become
the input to that decision record.
They're what um are going to help us be
able to find a place to start to
identify the use case to scope the use
case. And no approved answers for those
competency competency questions means
that that decision flow isn't fully
explainable yet.
And then lastly, it is that decision
stewardship is what produces that
decision record. Um there decision
stewards are first ensuring that that
use case is valid for AI decisioning
from the get-go and then it is
facilitating the completion of that
decision record and they're maintaining
it over time. So what the decision
stewardship here is really doing is it
is helping to govern that transition of
human authority over decisioning
organizational decisioning to AI
and I'll hand it back to Lisa.
>> All right. So back to the boardroom we
go and here is our mic drop answer for
today's session. We can answer both of
these questions honestly by saying we
can't tell you how the model reasoned
again but we can tell you the basis for
decisioning the operating environment
and the human involved every time not
just once.
And I think that is
the honest and best answer you can give
to the people at the top of your
organization and to your regula your
regulators and your internal auditors.
All right,
Steph, there's a question
about
um where these records live and you know
I think I we've had some conversations
about this and um you know what we what
we've thought about and how we've kind
of worked through that. So I think at
first we thought about a cat the catalog
for say the left side of the decision
record right. Um and then the right side
would be in an appendon data store
somewhere right outside of the model
itself uh for obvious you know reasons.
Um
I wonder is that still the way we would
direct people or do you have some new
ideas on that?
tough questions today. Thanks everyone.
>> I think that that is
that is still the case. Um because
we need to be able to separate
that front end that that business
reasoning
from
sort of that those functional those
those technical realities or
requirements that are happening
particularly when we start getting into
queries that are going to have a
temporal aspect to them. We are going to
need a robust system that is going to be
able to
hold and and log essentially the the the
state of reality that existed every
single time a decision record was
approved for use.
>> That makes that makes sense. Um
I'm looking So, someone asked the
question that define done piece really
caught me. If the decision record
establishes what the AI is sanctioned to
answer, does that mean that an AI system
could technically per be performing
perfectly while still failing governance
because it's answering questions outside
its authorized decision boundary?
So
that is where
can you read that again because I think
there's a couple of components to that
question.
>> Okay.
>> Define the define done piece really
caught me. If a decision record
establishes what the AI is sanctioned to
answer, does that mean an AI system
could be per technically performing
perfectly while still failing governance
because it's answering questions outside
its authorized decision boundary?
>> Okay, I'm sorry I misheard that. Um, so
this
scope flag, Whoopsies, excuse me. This
scope flag is what is meant to constrain
and really provide the governance around
what this particular model is sanctioned
to do for that use case.
Um, and so this is where that the
this is where language is so important.
Um, especially because so many of these
models were interacting with them
through a large language model interface
essentially. Um, so what we would then
do is we would look at those competency
questions and those answers and we're
going to pick apart the pieces of that
that represent specific uh data domains,
data concepts and we're going to um
enumerate those in the decision record
as we walk through that chain of logic
and as we've documented those those
questions and those answers. When that
gets handed off to the technical teams,
um, what they're able to then do is pull
from
an established
ideally an established semantic layer
that is going to serve as that
translation layer that connects
those that that approved language with
the physical instantiation of the data.
So that when the technical teams for the
AI solution go to grab it, it now has
sort of these parameters around look for
these keywords. If those keywords are
not present in the question or in the
query or if we can't reasonably find a
close enough relation to that, we're
going to then, you know, sort of provide
this response of this is out of scope
for this particular solution.
Got it.
Okay. Trying to
go through some of these other
questions. Thank you so much everyone
for filling up the qu the Q&A and the
chat. Um
overall, here's another question.
Overall, how does explanability and
interpretability
differ for third-party AI products
versus those developed internally within
a company?
Um, so that is
ultimately going to come down to how and
you're going to need to work with your
technical teams and your engineering
teams on this because it will depend on
the model the model version and the the
platform as well as the vendor. So if
you're building something inhouse, you
have really sort of full um ability to
customize like we we were talking about
that semantic layer. And so that
semantic layer is going to enable um
because all of this decision record
needs to get translated into metadata
and that metadata then feeds that
semantic layer and that's what's
allowing that um the connection of this
language as what to the physical data.
And so
you are
um
can you ask can you ask a question? Can
I completely just dropped my train of
thought trying to think through this?
>> How does explanability and
interpretability differ from third party
apps versus what you develop in terms?
>> Yes. So, um, you're going to have more
ability to then to translate these
decision records into pieces of metadata
that can then get tagged to the physical
data itself. Whereas, um, different
vendors are going to have restrictions
on what metadata that you're allowed to
access and what data metadata you're
allowed to pass over to their systems.
>> All right, I'm gonna read one and answer
it so you can take a break.
Um, so Ed asked, "It seems this approach
is solid for providing specific answers.
However, it does not seem sustainable or
even feasible at large scale without
some kind of implementation architecture
or structural framework and automation
to get all of these details, answers,
traceability, interpretation ready and
available.
So yes, I would agree with we would
agree with that Ed um that this is the
way that this is the process but without
doing this in a way that is sustainable
and scalable um it's not going to work.
Uh this is where we roll back to the
conversation about where do these things
get these um types of records get stored
and how do they get used. um you you
still need to in your organization I
think you know way back in the beginning
we talked about the critical decisions
you're making about AI and some of those
are which are what are the decisions we
will never let the organization we will
never let AI make on behalf of this
organization which are the ones that we
will always let AI make on behalf of
this organization and and which ones
fall in the middle and that gives you
your that gives you your road map for
installing these decision records and
the PE the stewards around them. Um it's
it's really that criticality and the
blast radius if it goes wrong which is
something else we've talked about that
you need to to use to determine where
you know where this has its place um and
where it can autogenerate if you will.
But I would even say if the decision
record is autogenerated, first of all,
it needs to be kept separate from the uh
the actual AI doing the work like we've
discussed, but second of all, at some
point there has to be a parameter in
which someone takes a look at it. Um but
would not argue that scale needs to
happen. um but within a contained
environment that is completely aligned
to the risk pro profile that your
organization has.
I think there's also
>> to add
>> I think there's also
another uncomfortable reality in that
AI
is
has s the ability to perfectly govern AI
has surpassed where we are able what we
are able to do today and it's forcing
real innovation and real investment in
organizations in order to be able to to
use this responsibly at scale
And we see
all of our our clients struggling with
this.
All right, let's see what else do we
have here. Mark, what did I what did I
miss?
>> Well, there was just so much debate in
chat and everybody was uh was so uh
engaged today, which is wonderful.
Uh there's still a handful in Q&A. Um
let's let's tackle this one. Do you
think the explanability and
interpretability still come down to what
data really is in the era of uh large
language models? Uh is data
which I think is a a bit of an
interesting take.
Yes and no.
Um, which is a terrible answer usually
for people um to hear but um I I think
yes and no because I think from what we
can understand now is
if you are running a semantic layer that
is where we can deterministically tie
language to physical data. Yeah,
>> if you are using models
that either do not allow configuration
of the semantic layer or you are just
using base models off the shelf, then
no, it's going to rely completely on
inference um in order to make that
connection.
>> Yeah, I I love that answer, Stephanie. I
think you're you're right on. Somebody
in chat just said it feels like we're in
the realm of philosophy. What is this
anyway?
>> Yes.
Yes. It it um it's really interesting
when you begin to sort of um really
explore the mechanics of AI, you're
really getting also into um psychology
because it really AI is artificial
intelligence. It is meant to mimic the
thought process of humans. Um so when we
talk about you know chain of thought and
being able to train AI and work with AI
um you really need to first understand
your business problem and your business
process
>> and you need to be able to model that
and you need to be able to represent
that both in language and in data in
order to translate it over into AI.
Well, then in LLM land, we have uh uh
these generative models that speak with
a confidence and fluency that isn't
really anchored by any authoritative
reality. So, it really gets into
philosophy when we talk about like a
manual can't like it doesn't think,
therefore it isn't.
>> The Dunan Krueger.
>> Yeah. Yeah. Yeah. Yeah. Exactly.
And that's where one of the the words
that I brought up at the beginning um
recently read a fascinating article
about something called persuasion
bombing. And this is where you question
the the LLM and it starts to try to
persuade you that you're wrong. Um and
and almost goes to the point of like
fiction to uh to to do that. um because
it's learning to think like a human and
therefore it's trying to persuade you
like a human would that its point of
view is correct. So uh it you know it's
governing an immature capability is
fascinating in that you think you know
it and then the next day uh something
changes your thought process entirely.
I uh we we only have a couple minutes
left and and one of the question this
isn't a question. One of the questioners
just put this statement in chat. People
are upvoting it. So uh I want to get
your take on this. Um and their
statement is we should not transfer
human accountability to AI. What are
your thoughts?
>> Do you want to go first, Lisa?
>> No, go ahead.
Um I completely agree. I think it
depends on what degree of separation
between human and AI. That's where we
could have a debate. Um ultimately
humans should absolutely be accountable
for the decision on whether or not to
deploy AI in what scenario under what
conditions and what to do if something
goes wrong. Um, but this is again where
I we go back to that chain of thought
and that business process modeling. If
you truly understand your business
process and you've truly broken that
down and you understand the risk risk
aspect of that decision,
then you're able to then say with a
certain level of confidence, I will let
AI be responsible for generating a
particular out outcome or output rather,
but I'm still accountable for how this
particular output gets applied or used.
Um, so there I I personally believe that
there are decisions that AI should never
be making. Um, but then I think if you
say, well, AI can be responsible for
summarizing my meeting notes and sending
them to my colleagues, I have no problem
with that.
And I think that again goes back to what
the law the case law is showing us is
that um that is absolutely true that
that the the these cases are showing us
that um that I the AI did it is not an
acceptable answer in any case, right? Um
so and and it's kind of like what you
learn in school, right? Plagiarism is
plagiarism and and you you need to be
able to defend what you create. So, um
it seems like it's cut and dried, but I
I agree with you, Stephanie. There can
be shades of gray in certain in certain
places, and that's what we're still
looking to discover and figure out.
>> All right. Well, that brings us pretty
much to the end. We've got Scant Nar a
second left. Um, thank you again
everybody in the community for being so
electric and both chat and Q&A. There's
a ton we didn't get to. Um, uh, so
that'll make some interesting reading
for all of us later as we go through uh
the the transcripts. Um, any final
thoughts before I hit the end webinar
button team?
>> Just thank you. Thanks. Thank you all
for making this a really interactive
session. These are the best kind. So, I
appreciate your your feedback and your
questions.
>> Thank you.
>> Bye.