Mastering Reference Data An AI Essential for Reliable Business Information
Watch on YouTubeVideo summary
The webinar features Dr. Peter Aken, an expert with over four decades of experience in data management who argues that Artificial Intelligence has made Master Data Management (MDM) both essential for reliable business information and increasingly achievable through automation. Despite the fact that 99% of companies prioritize investments in AI, a stark paradox exists where 95% of generative AI pilot projects fail to deliver financial impact due to poor underlying data quality, talent shortages, and scalability issues rather than technology limitations alone. To overcome these challenges, organizations must address deep-seated "data debt" involving legacy structural problems instead of simply purchasing new tools, recognizing that MDM is the foundational bedrock required for any successful AI initiative without which advanced systems cannot function effectively.
Central to this discussion are precise definitions and hierarchical approaches to data management, distinguishing between reference data, which consists of controlled vocabularies defining domain values like country names or order statuses, and master data, representing authoritative records about business entities that provide necessary context for transactional information. The presentation applies Maslow's hierarchy of needs to the digital realm, asserting that foundational elements such as governance, architecture, quality, and metadata must be established before attempting advanced AI initiatives, often described metaphorically as a "three-legged stool" comprising data governance, data quality, and actual MDM capabilities. Success relies on building custom reference structures first to understand organizational needs rather than relying on silver-bullet technology solutions, while also treating AI not merely as a database but as an improvisational actor that thrives on role-playing prompts within a well-structured environment.
Achieving success in this domain requires shifting focus from reactive fixes to proactive quality assurance supported by automated tools and strong executive sponsorship, acknowledging that approximately 80% of MDM problems stem from people and processes rather than technology itself. Organizations are advised against heavy customization of master data systems and should utilize clear role definitions via CRUD matrices while appointing dedicated individuals over fractional teams to maintain focus on the ongoing nature of this process as a business function rather than just an IT project. The rise of Agentic AI further underscores the necessity for robust traditional practices, where effective autonomous agent performance depends entirely on clean reference data and proper governance to prevent catastrophic failures at scale, necessitating new operating models like federated structures with strict prime directives alongside existing guardrails against major threats.
Ultimately, integrating these elements creates a virtuous cycle where using AI to improve MDM leads to more capable systems that leverage existing processes effectively while avoiding the pitfalls of hype-driven implementation failures. Reference data architectures are shown to be essential for enterprise knowledge graphs and taxonomies, enabling AI to understand context without building everything from scratch in favor of leveraging existing standards. By ensuring their systems produce machine-readable outputs aligned with a single golden source, organizations can transform MDM practices deeply intertwined with metadata and knowledge management into a strategic advantage that yields high-quality data feeds for AI systems without immediate commercial costs, proving that addressing the human and process elements is just as critical as adopting new technologies to unlock reliable business information.
Read the full video transcript
You got it.
>> Hello and welcome. My name is Mark
Horseman and I am the data evangelist
for data. We would like to thank you for
joining today's data webinar, Mastering
Reference Data, an AI essential for
reliable business information. It is the
latest installment in a monthly series
called Data Ed Online with Dr. Peter
Aken. Um, just a couple of points to get
us.
>> I knew I was going to get you.
[laughter] Yeah,
>> number of people that attend these
sessions, you will be muted during the
webinar. For questions, we will be
collecting them by the Q&A section. If
you would like to chat with us or chat
with each other, we certainly encourage
you to do so to open the Q&A or the chat
panel. You'll find the icons for those
features in the bottom middle of your
screen to answer the most commonly asked
question. As always, we will send a
follow-up email to all registrants
within a couple of business days
containing links to the slides. And yes,
we're recording and will likewise send a
link to the recording of this session as
well as any additional information
requested throughout the webinar. Now,
let me introduce to you our speaker for
today, Dr. Peter Aken. Uh Dr. Akin is an
acknowledged data management authority
and associate professor at Virginia
Commonwealth University, president of
Damon International, and associate
director of the MIT International
Society of Chief Data Officers. For more
than 40 years, Peter has learned from
working with hundreds of data management
practices in more than 30 countries.
Among his many books are the first on
making the case for data leadership, the
first focusing on data monetization and
modern strategic data thinking, and the
first to objectively specify what it
means to be data literate. International
recognition has resulted from these and
an intensive worldwide events schedule.
Peter also hosts the longest running
data management webinar series right
here. this one that you're in. Data
online on data.net.
Before Google was big and before data
was big and before data science, Peter
founded several organizations that have
helped more than 200 businesses leverage
data. Specific savings have been
measured at more than $1.5
billion. His latest venture is anything
awesome. And with that, let me turn
everything over to my good friend Dr.
Peter Aken to get today's webinar
started. Hello and welcome my friend
>> and uh welcome to you sir. I do
apologize for getting it but I was just
trying to get you to break character and
you did. So we'll we'll explain this to
everybody else when we get to the top of
the hour and uh you guys can jump back
in on this but uh yes a pleasure as
always Mark. Thank you. Um good
afternoon everybody. This one evolved as
it went through. Uh Shannon you probably
know does most of the I don't know maybe
Shannon farms it out to Mark. We should
ask him that question of break too. Um
but anyway, the uh the topic on it and
[clears throat]
the more I worked with this material,
the more I decided that really what this
webinar is about is how AI makes master
data management both more critical and
more achievable. And that doesn't happen
very often where you get two things that
are going in the right direction and and
helping you out sort of in there. So our
pathway today, if you will, we'll start
out with a quick data management
overview. We'll talk about what is
reference and master data management
because everybody talks about MDM but
they usually always mean reference and
master and you can see from the little
diagram here that'll be explained in a
minute. That's kind of important. Um why
is it important? Well, we'll talk about
that as well. Look at some building
blocks and some guiding principles
that'll get us to best practices. Uh and
as we get in but as we do this think
about this just as a vision. um AI ready
data for master data. If you're trying
to do something with AI, wouldn't it
make sense that the main the primary
people, places, and things of the
organization were all referred to by the
same uh types and and had the same uh
understanding of them. And the answer to
that of course is yes. But this becomes
a data bottleneck that people run into
and over again because they don't have
quality data. They don't have the type
of IT talent that they need to have. If
I say it, it's IT AI and and limited
scalability. And some of you may have
heard of this forward deployed
engineering things. I I just finished a
program on that that's got a piece that
works into this, but not a whole lot. So
the the vision of the future of course
is to have really automated stewardship
around master data. And if master data
is automated and well enough understood
that it can be automated, that's a
primary key, you'll end up with
augmented metadata that will allow you
to do intelligent data cleansing and get
onto this. Now, the reason this is such
a compelling vision for me is because
when I speak to people about data, they
encounter it like the blind people and
the elephant. And uh, of course, you
know, depending on which part you run
into, the elephant looks a different
way. Well, it's the same way in it.
We've we've come into it through
different ways and some people think
it's stories and some people think it's
pipes and the answer is it's yes, but
because most have approached it with
these differing knowledge and skills, it
means that we have different
perspectives and we don't have a a good
grounding. The grounding of course comes
from my high school knowledge of Maslo
uh which started out by saying if your
food, clothing, and shelter needs are
unmet then you will never be safe. If
you're never safe, you uh will never
belong to something that is part of uh
something larger than yourself, which
means you'll never be able to tell
yourself from it. And if you're never
able to tell yourself, then you'll never
be able to hold yourself in esteem.
Pretty good thing to do on a regular
basis. And and get to self-actualization
uh in order to do that. So this is
Maslo's hierarchy of needs. I transposed
it by saying it's relatively the same in
data management in the sense that most
everybody wants to start with all those
wonderful buzzwords that you see there
in that golden triangle that I have
labeled technologies. The rest of this
is the below the iceberg. It starts out
with governance. Then we build
architecture, quality, metadata, uh
integration of things. And then if you
have these foundational things in
practice, then and only then does it
make sense to do reference and master
data. and then after that those other
things. Well, so again, think about this
because when I do explain that to
people, they go, "Okay, these
capabilities are are important. I
understand that." Um, but I want you to
to to to go to go faster. And I said,
"Well, if I go faster, it will take
longer. It'll cost more. It'll deliver
less. And it'll present greater d risk
in there in order to do it." The reason
is because we just don't really
understand data management. uh we've
used to use a definition here where we'd
said everything that happens between
when data is sourced on one end of this
and when data is used on the other side
but where this forgot was the idea that
we are going to try to make leverage to
achieve leverage by reusing it and so
that piece was completely missing from
almost all of the early definition
around this. Um so perhaps better
definition is that you have a number of
different sources of data and that this
is going to give you specialized data
engineering skills grouped in two areas.
One around data engineering and the
other around data exploitation and again
you can chop them up and put different
nouns on them but you get the idea in
terms of of how people are doing this
just to give you the idea that this
portion of the elephant is relatively
complex. Remember all of this that I've
talked about just for this diagram is
still for data use. You still have to
have formal data reuse management in
there in order to be able to do what
most organizations really want to do,
which is to exploit to leverage their
data. Whether you're going to do it for
AI or just business in general, these
reference and master data structures
that you're going to build are essential
to the entire architecture that goes on.
Now, I'm going to to give you guys a
quick 90 seconds. And Mark, if you're
still listening, tell me if this doesn't
work, so I'll stop it afterwards. But I
want to try to give you what I call the
AI reality check here. So hopefully you
can hear this.
>> All right, welcome to this explainer.
Today we are diving right into one of
the most fascinating and honestly
jarring paradoxes in the tech world
right now. We are looking at a massive,
massive disconnect. On one hand, we have
this absolutely incredible executive
investment and excitement surrounding
artificial intelligence. But on the
other hand, we're seeing a crushing
reality of realworld implementation
failures. So if you want to understand
what actually separates the AI hype from
functional AI success and what happens
when these powerful systems are turned
on us, well, you are in exactly the
right place. So let's kick things off
with just a truly staggering number.
According to the data, 99.1% of
companies explicitly state that
investment in data and AI is a top
organizational priority. I mean just
think about that for a second. That is
essentially total consensus across the
board right up to the seauite. The
mandate is crystal clear here.
Businesses believe they need AI and they
need it yesterday. The optimism is
literally unprecedented and the checks
are quite literally being written as we
speak. But brace yourselves for some
serious whiplash here. A massive 95% of
generative AI pilot projects completely
failed to deliver any significant
financial impact. 95%. No way. Right? So
on one side of the coin, everyone is
prioritizing this technology as the
absolute key to their entire future. But
on the flip side, almost all of these
initial highly funded pilot projects are
just crashing and burning on the runway.
>> But where is this actually going?
Hopefully that came through. I found
that a a really nice way of articulating
it. And the reason was because first of
all, you're not hearing my voice. You're
hearing somebody else's. So I I hit you
with somebody different. Delivering a
message that's pretty hard. 99%
of everybody agrees on this. Now, when
was the last time we agreed as a society
on 99% of anything, right? So, what this
is going to do, again, of the tension, I
talked about it a little bit, it's it's
it's not working well. There are also
some challenges because the way for AI
to work at least so far according to the
data the the way you get in that 5% that
does work in in a successful GNAI
project is to understand that you have
integrated well and complemented an
existing work process well guess what
that's a basic requirement for how
you're going to do MDM in this context
here as well there's also a lot of
grumbling about the verification tax
where people are being put on this I
wonder if any of you all have any
experience being put on on
double-checking in AI and of course then
these AI winters. So that's our our
quick data management overview. Again,
AI only benefits from better data
management. So now let's talk about what
is reference and MDM. Let's dive into
it. The first thing to understand is
that everybody's under pressure because
there's an awful lot of things that are
going on where people are not
double-checking what I call promise
auditing. you you go in and you say,
"This is going to save us a million,"
and you find out later on that it only
saved you a hundred. Um, you know, the
next time you get up and say, "I think
things are going to save us a million,"
you might be in a a little bit of a
ticklish kind of a situation. But for
the most part, we don't educate people
about the hype cycle. And it's important
to understand the hype cycle in the
context of any type of technology, but
in particular with MDM. And just in case
you didn't know, the hype cycle was
invented by somebody named Lady Augusta
Ada King. I was in uh Prague,
Czechoslovakia not too long ago and I
found one of her first programs up on
the wall of one of the museums there. Uh
but she said in considering any new
subject, there's a frequency to there's
frequently a tendency to first overrate
what we find to be already interesting
or remarkable. And secondly, by a sort
of natural reaction to undervalue the
true state of the case. In other words,
it's great, it sucks. The answer is it's
somewhere in the middle. Now, I love
that. By the way, that is the entire
song of the Eagles, the new kids in
town, if you happen to know that
particular song. But here's how it works
out in practice. We find something that
works really well in technology for a
particular reason for a particular item
and we get to the peak of peak of
exploed inflated expectations. We get to
the height of our excitement and then
find out it's not as great as we thought
it was. And the question is, where does
it actually fit in on this? And there's
a lot of context around this trying to
figure out specifically what's going on.
But the hype cycle around what's
[clears throat] been going on in MDM is
just unfortunate. Um it's just too many
things that are happening. And yet one
of the things that has not changed is
that it has been a pillar of the data
management practice in there for a long
long time. This is the uh Dembach
version two if you haven't seen it. Uh
working very hard on version three
that's coming out. uh but uh anyway you
can see reference and master data is
always been an important part. So what
do we mean by reference and master data?
These are the practice areas that we're
talking about. All right, let's see. Oh,
I'm sorry. Got to go there and there.
And so the definition of master data
control over the defined excuse me
reference I start out to do that bad.
We're going to start with reference
data. Then we'll go to master data
controller for the defined domain
values, the vocabularies including again
you can see the standard terms and
things. For example, you may want to
divide customers up into current
customers and potential customers. And I
use that as an example to show managers
that taxonomies can be quite useful
sometimes. And then they go, okay, so
how does the concept of an ex customer
come in? Do we fit them as a current or
do we fit them as a potential? Right?
And then what if we decided we wanted to
subdivide our customers down into
different areas or that we wanted to
implement a customer VIP program? All of
a sudden, this reference data is really
really particular. But here's sort of
the best way to think of it. Um, when
you're trying to set up a website or a
business, if you don't decide absolutely
upfront that you're going to do a multi-
uh language business, you're going to
have some bigger problems as we go
forward. So the reference data here
comes in and says, well, okay, we could
we could look at this and call this
Czechoslovakia. Oh, wait, it was called
the Czech Republic and then it's called
Cexia. So again, you can see different
dates mean that you're going to have to
use different ways of classifying that
data, which means the context has to be
not just that it's the name of a
country, but in this case, the name of a
country at a particular point in time.
Uh, an order status could be new, in
progress, closed, and canceled. I've
seen a lot of organizations, go great,
but I want to put it on hold. I'm sorry,
we can't do that because we don't have
that as a part of our process. Uh,
two-state UPS state abbreviation, USPS
abbreviations that are there. Reference
data sets in here. When you look, for
example, through data, you'll see a lot
and a lot of uh UK's showing up as
addresses. And of course, the proper
designation for it in most organizations
is going to be Great Britain rather than
that. So, let's take a look and see
where the concept of master data came
from. And there's another one. We're
we're changing slowly because of DEI
push back and things. It's called
primary data. Think of it like the
bedrooms. You no longer have a master
bedroom. So, you now have a primary data
rather than master data, but it still is
the same concept in here. And and here's
where the concept comes from. In
general, uh I remember doing this
processing when I was working on uh
mainframes in the 70s. So, there was
data about these business entities. In
this case, it's about uh how much I have
on my pot belly sandwich shop card.
Okay. And the business rules dictate
again the parties locations and uh
provide context for these transactions.
The term master file popped out of
exactly that process because this is the
way we used to do it. Here we go. The
balance of $100 on my pot belly card.
And I obviously did this slide some time
ago because I thought you could get a
sandwich for $5 at the time. I know
that's funny. Um, nevertheless, the
balance comes off of my pot belly card
and by the end of the day, the new
balance has been updated to the new
master file of $95. And if there were
problems, hopefully they get written out
to a error log of some sort. Keep that
error log in mind. It's going to be
quite useful to us in just a few minutes
on this. So, here's reference data
versus master. Again, control over the
domains. The reference data in this
slide is the fact that the for a period
of time the FBI and the Canadian Social
Security uh kept nine gender codes for
nine possible entries that you could
have uh as a as an entry on the Social
Security system for Canadian systems and
for the FBI. Uh and these nine gender
codes of course are fascinating uh all
sorts of stuff. The master data part of
it then the reference [clears throat]
data is the allowable values. the the
maf the excuse me the reference data for
it the master data for it is going to be
the fact that mine says male right and
that you'll have a golden source on
gender for your customer Pat whatever it
is that we're trying to do but what
doing is providing context for the
transaction data so you can't just say
it's a customer it never works right
it's too simple you're always going to
qualify that in some way and you're
going to find out more about it so these
definitions again are trying to get to
this golden version on Here Gartner
really comes back and I think correctly
categorizes master data management as a
discipline or a strategy but as I said
it's also a pillar for requiring good AI
as well. The [clears throat] problem is
it's sold as these silver bullet
solutions and this is just a perennial
problem we have in technology but in the
in this category I've seen organizations
spend an awful lot of money on these
things parties places things sounds
pretty easy if we can get those pieces
we'll do a really good job. So let's
take a look. First of all, what's
happening right now is that because AI
is of such keen interest to literally
everybody, there is an incredible
increase in data. Let's capitalize on
that and try to make use of it uh in
order to do that. And let's start by
educating our users about data debt. Now
these are areas in which AI can be
extremely helpful. first of all about
looking at the investments and trying to
come up with the uh projections and
things like that but also for how to
approach the process of making sure that
you don't just put out a new set of uh
data structures but this data debt is
really something that you have to work
through in order to get that. Uh and and
finally the the ultimate mantra is
you're you're going to get your data the
better you treat your data the better
the data is going to treat the AI. Uh I
just wanted to show you this one
illustration of a history. I was doing a
a project on a separate notebook. Okay,
so I have a notebook for my work at VCU
and I have a notebook for my work
outside of VCU. And somehow they got
crossed and uh it started to hallucinate
a little bit here on this last slide. So
I just left that in because it was kind
of funny. Um I'm going to stop here
though and play a little short clip that
I want you to sort of inculcate. And the
idea is get your AI to take on a role.
>> Okay, I'm sure you've seen all sorts of
posts telling you what prompt to use to
get the most out of your language model.
I think you can pretty much forget all
of that because there's only really one
important thing to remember, which is
that the AI that we have now is really,
really, really good at roleplaying. Um,
so you shouldn't talk to it like it's a
Google search or like you're trying to
extract information from it, as though
it's Wikipedia. Instead, you should talk
to it as though it's an improvisational
actor that can be any character
imaginable. And the reason for this is
that large language models, they don't
store facts like it's a database, right?
They they generate responses dynamically
based on having read everything that
humans have ever written and the prompt
that you give it. And that means that AI
the stuff we have now it doesn't have
this stable identity right it doesn't
have a fixed worldview it doesn't have
personal beliefs and so if you prompt it
in a particular way it will respond in
kind. If you try and prompt it as though
it's a Shakespearean bard then it will
give you a flowery response. If you try
and prompt it as though it is an
extremely effective and smart scientist,
it will give you the relevant response.
But if you prompt it as though it's an
encyclopedia, it's going to try and
sound like one. Uh, but it's still just
performing a role and you've effectively
just restricted what it can do. So
instead, what you should do is you
should you should imagine that you're a
film director, right? And that you have
a character in mind and then you should
prompt your AI accordingly. So don't
say, give me three interesting facts
about science. Say you are a
world-renowned scientist with PhDs in
biology and chemistry and physics and
and and your nephew says that science is
boring. You only have a few minutes and
you've got to you've got to give him
counter examples. What what
extraordinary stories do you use?
>> I find that to be some of the most
useful advice that I've gotten about AI
guidance uh in cases. I tend to be a
Google fan, but that's because I'm doing
a lot with YouTube videos and the
interface there is phenomenal here. And
you can see in the upper right hand
corner I put a lot of uh AI generated
content in here because I want you to
see how good it is. They have some real
knowledge that this has managed to
scrape up and it doesn't look like we
have disagreements about facts around
reference and master data. And so it
becomes a very good source of trying to
find what's actually in there and how
you can get it to work. And again it's
nice steps to take you through the
process. It can be a very much of a help
uh in order to do that. So next section,
why is reference and master data
management important? Again, I keep
hinting around with this little picture,
and this is what you're going to see, of
course, right now.
Reference values control this accessible
data value. I think I've said that three
times. You can see it's a tiny little
yellow dot in the upper right hand
corner of the screen saying, for
example, what countries do we do
business in? What types of accounts are
available? um what are the controlled
vocabulary items that we're going to be
using throughout this particular domain
of the project that goes on. Those
control the master data items. The
master data items control the access to
the system capabilities. Are you a
member of our premium club? Uh you can
see if you want to add a premium club or
a VIP sometime, you better build it in
the first time because adding it later
is never an easy task and that's where
you have to fight data debt all the way
through. Are you authorized to use this?
Are you using sharing common data
structures that go back and forth? But
you can see here again the reference
data has leverage over more master data
in terms of what's doing that. But that
where the bullet efficiencies really
come through and I want to thank Chris
Bradley for allowing me to use his
example here which is such an articulate
piece. Uh this is the idea that these
master data items control the
transactions. So that's where the $5 for
my sandwich or the fact that I've been
authorized or that I can even make a
like on a particular system. Uh this is
the kind of thing that uh uh really
helps to understand with this reference
in here. And now MDM can make data
governance much easier because it gives
you a much narrow target to focus upon.
Let's think about that. What are we
going to focus upon? The answer is what
is strategically important? I can't tell
you how many organizations I I work with
and sometimes they just don't seem to
think that's part of the picture but
gosh it is. So the word strategy didn't
start to get used until the around 1950
when the management consultants
discovered it from me coming from uh the
use in the military and the current
definition according to the management
consultants is you can see a master plan
a game plan most importantly it's a
thing it becomes a PowerPoint deck I've
had some companies where I've gone to
work for them and they say don't you
you're not going to write a report what
we want is a PowerPoint deck describing
X Y and Z could have done that from home
but okay you know we'll get it done but
I go back on the word strategy And it
really came from the use of the word
military uh in there. The military
invented the term strategy, which is a
pattern in a stream of decisions. You
can see that's very different from being
a thing. You don't ever consult the
PowerPoint to see what's happened. But
if your goal is to do something
specific, and I'll just give you one
example. It's worked very well for one
organization for years and years.
Everyday low prices, you understand who
I'm talking about and why. because
they've done a great job of making sure
that pattern guided their people
throughout the development of this
behemoth that we call Walmart. Let's
this is not a Walmart story, by the way.
I'm going to tell you a story, but it
wasn't Walmart
on here. But at the first year, they
were implementing the MDM, and there was
real confusion because they'd really
only put up one leg of the three-legged
stool that you need to have. uh users
didn't know they had spent literally, it
was a joke, but they had a three plane
loads of consultants that would come in
uh 60 consultants coming from different
parts of the country to to work on this
master data management system that they
were working. They spent $60 million on
it and the business did not know how to
use the MDM. So, we had a bad transfer
of whatever was supposed to come on. I
don't know whose fault it was. That was
not part of what I was doing, but there
was general agreement that we should go
back and restart the effort. So they
went back and did a root cause analysis
and found out that poor quality data
existed in the system. That said, you
can pretty much say that's the case for
many systems out there. So it's not that
hard of a bet uh to win. That gives you
a little bit more, but you also need to
roll in that in the adequate training.
You have to get people to understand
what it is you're trying to do or it
won't work uh on this. And I'm just
going to drop in right here. Even though
I put the button on the speaker, they
went and they started doing data
quality, right? they say, "Okay, we'll
get let's get data quality going." That
provided more of a stool leg, but still
not the three-legged stool that we're
looking for. Um, the point I was going
to make here, though, is that many
organizations try to do this and then
they come along and try to get this done
without actually investing. So, if what
I say is people will come to me and say,
I'm investing a million dollars in a
data quality stool tool stool. There's a
slip a data quality tool. I'm not pooing
data quality tools. Um, but if you have
the data quality tools and a million
dollars, you should still plan to invest
$4 million to make sure that the
organization understands how to use it
properly. So if you have only a million
dollars to invest, invest 200,000 in the
tool and invest 800,000 in training your
people how to use the tool and you'll
find it actually works a whole lot
better. You can see there are a lot of
interdependencies in this case as well.
Data governance almost always three
legs. That's why we're getting the
three-legged stool. data governance
makes a case for and is responsible for
the data quality and the data quality is
a necessary but insufficient
prerequisite for this excessive master
data items which then go into master
data uh capabilities that constrain
governance effectiveness. Remember we
have over here on our diagram the
consultant says our methodology I don't
particularly like that it's a fancy word
and really if you look up the word
methodology means the study of methods.
So, that's not terribly useful to any
customer uh that's going to do it, but a
realistic way of practicing it is
probably a better way to look at it.
Select three data management practice
areas because you're probably going to
need all three of them in order to make
it work. Uh again, I've done it with a
number of different combinations, but
here is a reasonable way to do it. We're
going to put all this together with
reference data, data quality, and data
governance in combination of three. If
you add a fourth, it's more difficult.
If you keep a third uh keep it just a
two, it's kind of hard to make it work.
So three really turned out to be the
right number there. Similarly, let me
tell you a quick story about an MDM
success that was an organization that
had purchased an ERP. That stands for an
enterprise resource planning
organization. They were buying something
like people software SAP in order to get
it to work uh in there. They found out
that the problem was every time a price
of their product, which was a liquid
product, was transferred from one tank
to another tank, it counted as a retail
sale. And they said, "We can't do that."
So they decided that rather than modify
the ERP, they would be very careful.
Nobody could use the word tank anymore.
All the tanks had to be qualified. And
each qualified tank had a set of
business rules that were associated with
it. For example, transferring product
from a pickup truck, I'm sorry, a a
tinkerer truck to a tank was not
considered a retail sale and they just
made sure that they didn't count those
things as tanks. They counted them out.
This company also did transfer air as
you can see uh fuel from one plane to
another plane flying through the air.
Very very uh interesting way of
accounting for all of that. But again
partial of course there's always the
word tank as well. Just if you didn't
know, when you buy a tank, you also buy
about 32,000 mil, excuse me, 32 million
pieces of data that control all
throughout there. What's happened, of
course, is that over time, multiple
sources of master and reference data
have grown up to each of these pieces
because that's the way that finance
work. Finance would say, why would we
send the bills to anybody other than the
master bill? Uh, the place that says to
send the bills, but there have been lots
of people who work outside of finance
that get it to different places. because
of course it's what happens as data
start to proliferate through the
organization. You end up with real
challenges around all of that. So here's
a wonderful piece of architecture here
and I'm just going to walk through it.
You start out with the business data
stewards and they have a code management
system. This becomes sort of their
lingua franken and that code management
system is of course metadata that is
focused on making sure the reference
database of record follows those
particular codes that you're using. Then
those codes eventually are pushed out to
the OLTP systems usually on a system by
system basis eventually working their
way into whatever databases that you
have associated with that and finally
gets to your weight data warehouse which
can involve pushes to dimensions
depending on how you've structured it.
Finally you can add in here as well the
externally sourced data. Now from a
reference data architecture this is
pretty much how most people do it and
this is how the AI has learned that it's
done as well. So if you try to follow
thing like this and say hey this is the
plan that I'm trying to follow help me
do it and help me avoid mistakes where
other people have made mistakes it can
be quite helpful. So that's for your
reference data that's in other words the
gender codes that go throughout your
organization or whatever is equivalent
to a gender code in order to do that.
Same thing happens for your master data
you end up with a system of record here
that follows through from the master
database of records. uh in order to pull
that together. Then it goes to the OOLTP
systems uh again following similar
pathway getting proliferated throughout
the system. And you can see of course
you use the same absolute hard uh uh
infrastructure in order to do this.
You're just coding the data slightly
differently getting to the data models
and your externally sourced data in
there. In other words, it's very very
much of a parallel operation in order to
see all of that. You can combine them of
course into one and most organizations
have uh in order to do this but you need
to do additional thinking around this as
well. It's not just a technology play.
You've got your technology. You want to
take a task orientation really. And the
task orientation is that we used to make
pins by putting them together in 12
steps. But eventually somebody looked
around and said, "Okay, there's probably
an easier way to do this. Maybe I could
combine steps one, seven, and nine and
come up with a faster one. So this is of
course the whole process of business
re-engineering or going digital or in
this case AI because you're taking steps
that are nonvalue added out of the
process and reducing cost and increasing
your revenue in order to do that. You
still have to of course go through all
of the standard business rule analysis
that you would go through as you're
doing it. Here's a an interesting
example that we found one time. All the
information was on one screen on the
source system, but we were putting in
also an ERP package software and the
same information was spread across 23
screens. I want you to imagine being the
customer service representative
responsible for directly helping the
customer who literally was standing
across the desk and trying to do that by
basing it on 23 screens.
These things fail because people don't
understand the processes. When I say
these things, I mean both master data
and artificial intelligence.
So, lots of help around gathering
different types of things that can go
into here. Tell you a secret though, you
won't find this in the slides anywhere.
If you take this information from your
log files and get AI to analyze it, you
will get a pretty good reverse
engineered process. It'll give you quite
a lot of information about your business
that is immediately useful. uh in order
to do that. You want to know more, come
see me after the event and we'll talk.
Uh anyway,
here's another way to think about
process understanding and that is that
most people looked at the traditional
engine gasoline powered engine and said
that's the infrastructure that we have.
That's how we're going to make things
work. We will put either a gas engine in
a car or an electric engine in a car.
And if you remember the early days of
electricity pre Tesla and everything
else, there were a couple of uh uh
entries into there that were quite
interesting electric cars that came on.
But they were looking at as if I have an
electric engine, I can't have a gas
engine. If I have a gas engine, I can't
have an electric engine. What if you
could have both? Well, of course, both
was what Toyota came up with when they
invented the Prius. They had the engine,
they had the electric motor both in the
same car and the battery. Now, I have
the most recent version of the Prius. I
believe it's version five. And those
engine and electric motor have now been
combined into a single structure. That
makes it even more efficient in terms of
what you're seeing. But you can see here
that what was going on in the Prius
world was that the engine and the
battery electric motor were switching
back and forth sometimes multiple times
in a minute, sometimes not. But they had
the flexibility to be able to do that.
So you didn't have to choose, do I run
the engine or do I run the other? And
that understanding of the process is
what let Toyota dominate the market in
order to do that. When you go ask
[clears throat] questions about how
master data can work, think of how it
can work in your domain context. And it
turns out the AI understands this as
well. So here are just some general
scenario if you happen to be in a retail
setting that you can look and see in
order to do this. Now, how would you use
AI in order to do this? Well, you can
use AI both to find the problems and to
help you resolve the problems. Uh, in
order to do this, in order to do that, I
urge you to look at AI as an extension
of the knowledge workers capabilities
rather than as a replacement for it
because it does also seem to be
providing much better work uh, in order
to do that. All right, here's another
little AI summary again driving this
data investment which gives us lots and
lots of things that can go into this but
we also understand the data dead
actually I think we had the slide in
here twice I might have uh duplicated
that one sorry about that guys we'll
keep rolling on here all right we get to
the building blocks this is the other
part this discipline is reasonably
mature uh Mark will be able to tell you
that he was building master data
management systems for his customers
when he was still a baby in a high chair
eating
uh baby food, right? But uh you know the
goals of these are very very
straightforward. We want to provide all
of this type of information and again I
[clears throat] ask you what AI system
would not want to have authoritative
source of high quality master and
reference information lower cost
complexity and support for the
integration efforts that come into this.
There's another whole set of categories
here. These are mainly reference slides
for you. Remember, you get the slides so
you can take them away and and take a
look at them, but understanding exactly
how these activities correlate to what's
going on in your environment. And then
specifically looking at what can likely
be helpful in order to get up with this.
And this gives you these primary
deliverables which allows you to cleanse
data, tells what your requirements are.
There's a whole series of roles and
responsibilities that go through all of
these as well. uh in order to come up
with it. Again, you need to have it
takes a village, right? We get all that
sort of thing. But there are lots and
lots of people. You don't have to have
all of them at once. However, this has
been the part that most people have
failed with. And as I said, it's sold as
a technology first solution. So, you
likely have somebody in your
organization doing ETL or ELT already at
this point. You may have tried and
hopefully had a good experience with
getting reference and master data
applications uh in place. uh this allows
you to get them. You can buy them
separately, but most people buy them
together because they are somewhat
complimentary and similar. Uh again,
some people say you don't need data
modeling tools. I say good luck with
them as well. I also include process
modeling tools. However, again, AI is
proving that it can largely supplement
some of the needs of those uh uh pieces
in there. Uh again, metadata
repositories are going to be in there.
And again to just remind us that the
metadata repository is a thing that data
cannot go into but it doesn't mean it
has to go into. So one of your first
questions should be not is that metadata
but is in fact the metadata worth
maintaining formally because that's a
very big decision. It will help
tremendously in terms of your leveraging
opportunities in here. If you haven't
heard of data profiling tools, there's
another whole uh topic we do on these,
but they are a wonderful uh set of
technologies. There's a data cleansing
tools set of things, integration tools,
all sorts of things that can go on. The
challenge, of course, gets to be that
oh, anybody ever seen business rule
engines? Boy, they are fun. That the
technology again, a fool with a tool is
still a fool. And that's an unfortunate
way to say it, but as I said, I saw this
one organization and it's not just the
one. They had lots of failures that that
happened in these things because people
generally didn't understand what was
going on. So, here's a
just a summary of reasonably good
vendors in the sense that you've seen
these folks out there and they have uh
definitely had pieces that are uh useful
looking at that. Somebody may want to
ask a question about Oracle as we get to
the end of this because it's been in the
news kind of lately uh in order to look
at that. That said, I would always start
by building yours first.
And people go, "Whoa,
you want us to build something?" Let me
give you an example.
If you build your first version of it,
it almost always will tell you so much
more about what you're attempting to do
at so much lower of a cost that it it
just I have certain conferences I've
been banned of because I tell people
things like this, right? So, let's just
take a a hospital situation. They have a
number of hospitals around the regional
area
and they want to make sure that they
have doctors that have admitting
privileges. And you may think that was
an easy thing to do, but actually turns
out to be kind of a a messy piece, but
it was implemented on a SQL Server
database because everybody has somebody
they can program in SQL Server. And all
you're doing is you're making a database
of the golden data. The golden data in
this case about physician ID and and
admission. By the way, if you have a
question about sharing data, there's a
really great website called
fiduciarycoms.org
org that you'll find has some really
interesting pieces there. And this
system is connected and they get it to
work and they practiced with it without
spending any money. They had this all
inhouse
and and they started to understand the
process of admitting a physician and
then the process of actually getting
into a hospital which is different from
admitting a physician to practice at a
hospital and they could extend this to
another system and with this other
system they learned even more. or it was
a different sort of a processing and
again you can see they're gradually
growing this and at some point it
becomes quite obvious where we need to
replace these systems and give them
instead something where they are
connected to a generalized bus
technology and that technology here in
this case connects up with these systems
because as long as they got three it's
okay but we have the fourth and the
fifth and you can see how it gets more
and more and more on here and eventually
they're going to get to the point where
this is not the right technology for it.
So you fix that by then going out and
buying your commercial master data
management package. But by doing that if
you take three years and three years
sounds like a long time but if you take
three years to do this well it will help
your AI because the same governance can
provide this you don't care that you
don't have a master data management
package on the other end of this. You
care that you have good quality data
that is feeding your AI and all the rest
of your systems. Why wouldn't you? So
there's your win-win. Okay, let's take a
look at now the idea of how this is
intertwined. And if we look here, you'll
see master data management practices are
highly intertwined with this
implication. Knowledge management in
this case, the metadata management
pieces, the data quality pieces. There's
a larger picture that gets told in all
of these, but you get the sense that
that's intertwined. Here's another one.
This was a different organization, but
they had specific pieces where they were
looking at operational data and
nonoperational data and looking in both
cases finding instances of where master
data enabled them to leverage their
efforts and feed their AI with good
quality information. Couple of important
considerations again it is a technology
it is sold as a technology. you need to
think of it as as a way of designing
systems and there's lots of good
information that you can get from it. By
the way, this is supplied by the AI. So,
very very good guidance as far as that
goes. Here's another one uh in here and
again this is the idea of just saying
the old way we did this was very much
dependent on a lot of heroic efforts and
really good specialization. Now, we can
both use AI to get it this way. we can
shift our reaction to proactive data
quality and saw earlier I was even
saying automated metadata quality which
is a real possibility and finally
starting to move towards automated
policies and standards where it starts
to figure something out and ask
questions. Now you have to have good
people on the other side in order to do
this or you will definitely not have uh
success in terms of what goes on that.
Okay, couple final pieces on this one
here, which will take us a little bit to
get through, but uh got quite a quite a
chunk on this one here. First of all,
while you'll see these crazy crazy rates
of failure, and if you just Google,
you'll find enough uh things that are
problematic out there,
take it with a grain of salt, they're
still being used, they're still being
sold. There's a a gentleman who's been
running a conference very successfully
in that area for many years, but they
are challenging. On the other hand, I
had the same challenges when I was doing
business process re-engineering. Does
anybody remember that from the late 80s
and the early 90s? Uh, done well, it
could make a significant difference.
Done poorly, it made a huge mess out of
things that were just not necessary to
be messed with. So, here we have less
than fully satisfied with their data
programs and 70 cent were less than just
satisfied, right? uh in terms of there's
another one here can't track and
consolidate where these things come from
they have no idea of the spender that's
why you put one throat to choke that's
why you have a chief data officer often
times or chief AI officer uh 25% of the
clients spend is wasted duplicated
effort I have good numbers that say
between upwards of 40% of organizational
IT spend can be reduced by better data
management practices in order to do that
as well as your cloud bills too which is
Another wonderful way to train. By the
way, should these things be cloud-based
or should they be uh on prem based?
Well, again, it depends on what you're
doing and what you're trying to
accomplish, but there are equal good
versions on cloud as well as on prem uh
in order to look at these things. Uh 64%
are planning to rearchitect the
reference data. That's a big big lift.
Imagine taking the pipes in your house
and moving them from one wall to the
next or one floor to the next or
changing a bathroom or other things like
that. These are not good. And over half
of these companies were spending $4
million a year on this reference data.
Uh so this again a market in order to
look at this. What were the basic
causes? The things you have to look at
and I'm sorry about the root cause
thing. Uh the scariest movie I ever saw
just so that you get a little bit of
picture in your mind here was a PBS
documentary about somebody flossing the
first time as closeup as this diagram is
showing you there. It was terrifying. Uh
I've been an avid flosser ever since,
but that's not information you guys need
to know. Anyway, 30% of these in this
instance was failures. Uh
serviceoriented architectures were
similarly challenged and again what were
the problems? Well, bad leadership uh
plagued many many of them again
implemented as a technology or as a
project. Uh again, get it in your
people's heads that they are not going
to need their data program when they do
not need their HR program, right? And
that will actually help them understand
this is going to be with us for a long
time because data is kind of like HR in
the sense that if we don't know what
we're doing, we're probably going to
have a mess. Uh in order to look at
that, again, the MDM was either the
enterprise data warehouse or an ERP.
Sometimes it's just too easy a solution
to put in place in order to come up with
that or it's run as an IT effort. Now
again, this is not to blame it. it has
an awful lot going on it and it's
software so why shouldn't it do it well
they should of course be responsible for
installing the software however in order
to install the software they can put it
in but unless they want to pay for
something that doesn't get used they
should insist on having good
understanding of the processes it is
designed to support master data
management needs to support the
processes the process with the Prius was
that you weren't sure whether you were
going to need an electric or a gasoline
motor at any point in time and they
invented something that you could switch
back and back and forth between them in
minutes. That meant that MDM was uh
really focused in that right area. In
this case, MDM part of an IT effort
generally about one in 10 IT shops I see
do this really really well and they
wonder what everybody else's problem is.
But uh uh for the most part it's it's
definitely been a challenge and they
don't know again the software goes in,
they got to air zero return code, so it
must be fine. No, not definitely not the
way you want to think about this thing.
If we separate governance and quality,
it becomes even more problematic. MDM
provides us the ability to focus on
these things and to provide it to us in
terms that the business understands and
that's critical for us that these MDM
initiatives are implemented oftentimes
with inappropriate technology. If I kept
growing the build your own example that
I showed you, eventually it would crash
and that would not be good for the
hospital system. However, if you monitor
this and keep in progress and understand
it and realize that those three years
that you're doing the understanding
about your environment and learning
what's going on there will give you a
much faster pathway to value when you do
go buy this stuff then buying this stuff
at first and trying to figure out how it
works is definitely the better way to
go. And finally, let's eliminate the
silos as far as that goes. There's a
number of success factors here. Again,
just briefly run through them. you're
more likely to understand these
strengths and limitations and then
therefore have success because if you
want MDM to fix everything, it doesn't.
It's a secondary effect. That said, if
you want your MDM to constantly score
high on its AI governance areas, it can
be done very very nicely. Small steps,
right? Crawl, walk, run, set
expectations, communicate. You can't
can't undercommunicate. Uh in this case,
you're going to have to have this
integrated with your architecture. or if
it's just hanging out there by itself,
it's not going to work. You need to have
incentives to make sure that the master
data is desirable. I um once was
involved with a a piece where um a
company was getting subscription
information from another company. So, in
other words, they were dependent on
subscriptions and the subscriptions only
came from the other company. But by
goodness, they were sending duplicate
data and bad data to it and we didn't
have any terms of service agreement. It
was just send us the data. It's like h
we could be a little more specific about
I won't go read the rest of these. You
get the ideas in terms of of what's
going on here. Uh again, each of these
requires specific practice area focus,
but at this point in time, AI has the
capabilities to be a good solution for
you. So, by all means, uh I would
suggest uh this is a really
extraordinarily good area to look at to
practice. Uh again, just like anything,
if you don't have executive sponsors,
it's not going to be reasonable. The
business has got to own the set the the
context for this. It just doesn't work
without it. If you have just an ITled
project, build it and they will come.
Does not tend to work for these things.
Again, the stronger your project
management, organizational change
management skills, uh the better they
are. People process technology and
information all the way around. I insist
on having this detailed information as a
CRUD matrix to show the support for the
process that we're doing. Uh sometimes
it works out really well, sometimes it
doesn't. But uh the ones that have the
CRUD matrices tend to work out much
better than the ones that don't. Uh
again, if you've got documentation on
your existing processes, use it. Uh
support this continuous improvement
process. It's a a great example of of
how to use it properly. uh management
needs to understand the importance of
these dedicated individuals. I always if
they tell you I'll give you 10% of 10
individuals, I say give me the one
individual because the one individual is
going to provide much better focus and
much better results than 10% of 10
people or 30 30% of three people, right?
Let's let's go for those dedicated
pieces in there. understand how your
systems MDM works if you have one and if
it doesn't make sure that you understand
where it integrates uh and what your
external connections to it are uh in
order to do this. Resist the urge to
customize. Uh again this is master data.
It's your business things about which
you create, read and delete information.
Uh it's it's pretty straightforward that
you shouldn't need to do any
customization on this but I see almost
everybody does it. uh stay current with
the patches that are there and then of
course lots of testing going on with it.
Final set of guiding principles uh in
order to do this is make sure that this
belongs to the organization that
everybody understands they own it and
they actually own it now but they'd
rather own it when it's in good shape.
So we're going to make sure that they
get all the way to the goal there.
Again,
it's almost synonymous with quality
improvement. There's just no way that
you can do this one time and expect it's
going to work. But management wants to
understand why it's not done by Friday.
So, you've got a lot of communicating to
work on there. Again, hopefully this is
helpful in order to do that. Um, the
business data stewards are the
authorities.
They get to determine the golden values.
Um, yes, we'll just leave it at that.
The golden values represent these
correct sources. We shouldn't be getting
information from places that are not
designated as golden values because the
data is of less known or unknown quality
whereas we want it to be of known
quality. And finally, we're only going
to be replicating I should say finally
uh replicating these master values only
from the golden sources. Uh so when
somebody gets a copy of this, we go to
the correct place in order to get it and
put it all together. Uh there's the
finally it's data changes require formal
change management. Uh yes we always used
to run to the DBAs in the past but uh
that that is going to be a a more
controlled process. It doesn't mean that
they still can't react rapidly. They
still can but there's additional guard
rails around it. I estimate it's about
10% additional overhead to do this
properly and most organizations can uh
afford to do that. Now let me give you a
what I consider to be a real treat. I
have lots of artifacts that I've
obtained over the years and this is one
I obtained from a guy named Dave Evans
at British Telecom. Hi Dave, if you're
still out there uh you didn't know this
was going to last so long. So they were
putting data into master data and they
needed to communicate this process to
people uh to the employees in the
organization to the folks that were in
the organization that were part of what
was going on to to understand the
vendors the employees the tellers you
everything right bishcom had a bunch of
everything in order to do this I've put
a copy of this out there if you're
interested at the anything awesome uh
email that you can just click right
there and go to but I'm going to play it
for you and I just want to tell you what
I'm going to play you so that you get
it. [snorts] Uh it starts out again in
an email that comes from the managing
director of the bank. So that would be
like the president of the United States
emailing me personally uh in order to do
this. And I might look at this and say,
"Wow, there's a a thing to click on."
Now, of course, in the old days, it was
okay to click on them because they had
embedded viruses in them. It can still
be done safely. uh in this case you put
it on a controlled video player or
something like this but this was done as
an email and they had good emails
tracking statistics so they could tell
of the 60,000 employees at British
Telecom X number had opened it x number
had played it x number of times uh it
was really quite good information and
most importantly they achieved the
objective because when they tested
people not just immediately after they
had seen it but months after they had
seen it and then years after they had
seen it they were able to identify some
of the seven sisters that you'll see
which was their catch name for their
master data management solution that
they used. They never talked about
reference as part of this. It's not good
or bad. Uh they just didn't. But they
they I'm sure did piece of it. Maybe it
was integrated. Let me just step back
and play you this information real
quick.
[music]
Bong
Boo. [music]
[singing]
[music]
[singing]
[music]
[singing]
>> [music]
>> talk about a simple explanation of what
went on. I just I you guys may think I'm
crazy, but I really really like that it
tells what it's doing. It did it really
well. You're welcome to reuse it. Dave
put it out there in the public domain
for you all to to do it. Um, here's
another little AI piece that came out
mostly pretty good. The real question
here is when you're looking at this, you
can build it yourself or you can create
a custom version of it. There are some
instances that are custom makes sense,
but again find out what are the actual
complexity of the requirements. Do
comparison back and forth between your
relative uh uh competition in the
organization in order to do this because
it is a challenge uh to do this and you
if you don't have control over these
pieces, it's absolutely not worth it. I
in general have aired 90% of the time on
the path that I described you which is
to start out by building a very
simplified one yourself to learn about
the process of supporting master data
management and over time growing with it
and then you can have a conversation
with the vendors and put it in after
[clears throat] the end of that but
rather than before it. Uh so we've spent
some time uh again about 55 minutes so
far looking at a data management
overview in this and what is the
reference and master data. I hope you
understand the leverage reference has
over master data and the master data has
over all your transactions. You
understand that they are the best parts
of what you have in data in the sense
that I say the best parts is the bones
of a garden. If you're a gardener you'll
understand that particular reference. um
they they provide this and there are
really good building blocks all the way
around out there to use this. You do not
need to start from zero. There's so much
information and it's mostly known
correctly by the AI. Now AI is still
errors. So don't be silly with it. But
you've got some ideas of what these
guidances and best practices and it's
almost time to bring Mark back in here
and uh see what he's got to allow in
terms of his experiences with this
because he comes at it from a a little
bit more of a semantic perspective than
I have been able to put my time into. So
I learn every time I speak with him. But
anyway, finishing this up. Yes, culture
eat strategy for breakfast. 80% of these
problems are people and processbased not
technological. So let's make sure that
we use them correctly uh when we do
that. Shifting the ethics component of
it as well. Just because we can doesn't
mean we should do it. And and these
guidelines can be implemented in the AI
guard rails which will give us a little
bit of safety uh around these things.
What you're looking at is the overall
process model for master data
management. In this case, here's the
definition. Planning, implementation,
and control activities to ensure the
consistency with a golden version of
contextual data values. Again, people,
places, things. That's what we want to
think about. If we get those right, we
can start to expand to other areas. What
are the goals? Authoritative source uh
lowered cost support for the integration
and intelligence types of efforts on
this. Your organization must put your
own goals through this filter. Go back
to the strategic uh piece that we talked
about a little bit. What is the focus of
this? The MDM is going to be a major
piece of it. You can architect it so it
supports your efforts or you can
architect it so it doesn't support your
efforts. Again, look at your inputs that
are going into this. The various
suppliers that you have, participants. I
know one organization has 200 people
doing this kind of work. Uh and for the
organization, we're glad they do it and
they do a wonderful job. Um
all sorts of activities in here in terms
of what it is and you can see they're
also subdivided as to planning,
controlling, development, operational
around all of this. And then we get to
the deliverables, the requirements that
come out of it. What's the master
planning, the golden record data
lineage, what are the metrics and
reports? What are the consumer uses of
it? Who's going to be using this? Make
sure that you express those things. If
your company makes tanks, make sure that
your uh uh uh reports express address
tanks otherwise you will uh have a
translation issue that goes in there.
Look at the measures that are around
this. Uh again, just a a start of a
number here. When I went to work for the
Department of Defense in the late 80s,
early 90s um they had 1400 connections
to the internet. By the time uh I
finished with them, they had 14. You
know, it can be done. One organization I
look worked with started a process of
eliminating 400,000 access production
databases that were running in there. It
was crazy. and uh that took them 10
years but uh there was a strategic piece
so they use tools use all sorts of
things to to focus this these are the
the union of the ven diagrams you do not
need all of this in order to do it and a
lot of it can be done with lowweight uh
stuff so hopefully what I've done here
is is taken some time to shift the
narrative away from the traditionalness
of the MDM challenges to say what can we
use with AI as a forgive the expression
co-pilot but it's a good description of
this and that doing MDM AI makes it
easier to do the MDM and doing the MDM
well makes it more uh likely that the AI
is going to be useful around this. I've
left you with a couple of references uh
to take a look at uh look at the uh of
course the books pieces that goes up and
we've got pieces that coming up and
hopefully you'll be able to to join
March uh Mark in the uh uh event in
Providence coming up and we're right
back at the top of the hour.
Perfect timing as always, Peter.
>> One of these days, Mark, you're going to
say to me, you know, you've been talking
to yourself for the last hour. There's
nobody out there, actually. So, I
apologize for catching you up short.
Should we explain to people what was
going on with that?
>> Yeah. No. Uh, so Peter was trying to
take me off script with his animated
video of me and and as I'm going through
my script, I could see out of the corner
of my eye that that I was moving around.
So yeah, which is hilarious.
>> And we never did this. This is the
amazing thing. You should see what the
the actual vocals are on this. We talk
about going and getting a drink,
[laughter] which you know, maybe not far
off, but I just sent it a picture and
said, "An animate this." And it's
scarily realistic, but that's not what
you guys want to talk about. What sort
of questions have we got that two of us
can dive into?
>> Well, we've got some wonderful questions
in Q&A and then there was some wonderful
conversations happening in chat. So, I
might jump around a bit. Please, please,
please
>> good at this stuff.
>> So, top of the question list, uh, which
I love, is a Gentic data management even
real or is it just hype? Are there
actual tools, architectures, and
implementations happening today?
>> I'm certain that people have I have not
seen it. Have you?
>> Yeah. Well, yeah, I'm doing it.
>> Okay, [laughter]
there you go. So, so jump in. And
>> so as a plug for our conference in Rhode
Island in November, um um me and a good
friend of mine are doing uh AI data
quality lab and it's all about how to
build AI agents and use AI tools as part
of a data quality initiative which is a
huge part of of data management as you
as you know Peter and and there's a lot
of places where you can build agents and
and it really depends how deep you go as
well like are we orchestrating agents or
having agents uh orchestrate amongst
themselves to solve uh data management
pieces? What are the guard rails uh that
we're putting in place for that? What is
the AI governance that we have in place
to manage those uh agents and control
them so we don't end up in the news for
some catastrophic failure like happened
to that one company? I think it was back
in April uh where his uh agents had
basically destroyed his production
database, removed all the backups and
like fried everything. [laughter]
>> It was kind of hilarious to read.
>> Yeah. Oh, yeah. And and like I know
exactly how it happened and like the the
the person who set that up like deserved
all the failures that they got because
they they missed out on so many just
common data management practices. But
yeah, there there's a lot of agentic
stuff that's happening out there in the
real world. Um, would I say that it's
popular enough to be a framework uh or
something that everybody can just grab
and follow? I I don't think so. I think
we're not quite generic enough yet. Uh
so I think if you've got a fairly AI
mature organization and understand how
to manage agents I think I think there's
a lot you can do um with uh data
management that anyway that's that's my
my uh my take on that
>> and is it is it not true that in order
to manage the agentic performance you
have to do pretty good at data
management?
>> Oh yeah a thousand%. Like [laughter] all
all of all of everything that we do in
AI requires so much traditional data
management to happen to be useful
>> and and so this is where a lot of
failures have been happening especially
around master and reference data. uh
this this conversation was happening in
chat a little bit and it's like what
how's my data dictionary work like is it
reference data is it is it is it
metadata uh uh how is AI using it and
and it's so true
in that we have to train these things
and before we had started recording
Peter I was talking to you about context
layers and semantic layers and um and
people using languages like sparkle and
shackle
um to uh be able to knowledge graph out
their content so that AI can learn from
it. There's so much traditional data
management that happens in that universe
uh to to make those things function. I I
think I put a link in chat uh to
enterprise knowledges uh uh book that
they recently published um a very very
good book uh that talks about all of
that. And sorry, you're getting me on a
tangent here, Peter.
>> So, this this one organization that I've
been talking to, one of their executives
had come up to me and he was like, I you
know, I kind of want to just replace our
entire analytics function with AI
agents. How can we make that happen? And
I'm like, well, how many sparkle and
shackle experts do you have? economics
of it realistically.
>> Yeah. So, it's like probably not a good
idea if you don't have the right
expertise to actually train the AI on
what your data means.
>> And then you're going to fire them, too.
So, they're going to be real
enthusiastic about cooperating.
>> Exactly.
>> Yeah. Not enough people have uh learned
that lesson yet. [laughter]
>> Well, we've we've aligned on this issue
that the more you do these things that
are on the screen in front of you, well,
the easier your AI will be. and and most
AI projects that I see, they don't have
these these concepts in there. They're
they're like, "Wow, I got this great way
of optimizing X, you know, and it's very
algorithmic centric uh in order to do
this." And uh I've just finished writing
a section for my online uh data
preparation class where the admonition
is don't start with the algorithm, start
with the data, right? figure out what
your data is going to tell you before
you go to say try to square peg round
hole the thing into a you know some sort
of other relationship in the algorithm
for sure
>> that dead I was gonna say [laughter]
>> we have another question in here that I
uh love as well um so how should
organizations rethink reference data
governance when agents rely on it for
autonomous decisionmaking
So what do you think Peter? Do you think
do you think we need to rethink
>> Well, they just articulated exactly what
the problem is, right? If you have your
governance set up so that you can't tell
Czechoslovakia from, you know, another
country uh in there, none of what you're
trying to do is going to be trustworthy
uh to do this. So yes, absolutely. I
think that more we can set this up,
there's a a a complete side issue on
this. The two state post code uh was
done as a comedy routine by a guy named
Gary Gullman that's on YouTube. It's
hilarious. My wife called me one day and
said, "They're talking about data on t
in a comedy stand-up routine. You got to
come see this." But it's true. If we
don't get those right, right, none of
the rest of this stuff is going to be
usable in here. So, absolutely, the
governance around this thing, this is
where it's going to go read this. If it
goes and reads everything as being in
Great Britain, it won't understand what
Great Britain is and you won't be able
to have a total sales adding up to 100%
or whatever the the the piece is. You
must have seen examples of this
incorrectly done.
>> The the words I use for this, Peter, is
failure at scale.
>> Oh, wonderful. Yes, we can move fast.
>> Like in the before times, if we had a
reference data problem, somebody would
look at a report and go, "Oh, that's
wrong." and go fix it. And now if you're
automating a bunch of stuff with AI, it
fails everywhere simultaneously almost
like it just it it fails everything.
Like it just has no hope to begin with.
Um like going back to how the questioner
worded this uh do we need to rethink
reference data governance? Sort of.
>> My first gut reaction would be no
because this is still reference data
governance. However, I had the epiphany
not too long ago, Peter, and and I'd
love your thoughts on this.
We write and produce content differently
for AI to learn from than we do for
humans to read it.
>> Which means now you're diverging in two
two different pathways.
>> So how do we present reference data for
AI to learn from it?
>> If it should be done, you have to make
sure that your AI gets contextually with
the the concept of you know reference
leveraging meta. And if it's they got
sorry um master uh then we shouldn't
even say master. remember supposed to
say primary uh all the way through but
uh uh if it doesn't understand that
leverage piece I was going to ask you on
this on your agentics component are you
able to replace case tool functionality
yet with the AI
>> I've had students playing with it and
it's kind of is but I don't trust it as
much as I
>> yeah that's a technical term for that by
the way
>> I want to go back and and take one of
these case tools and augment it in a way
that they haven't thought about yet
because every time I see what they're
doing to it. It just drives me crazy.
>> It's it it's not very consistently good.
And it it's depends so much on the model
that you use, right? Your mileage may
vary.
>> Yes.
>> All right. We got some other MDM
questions out there. Did we answer the
last one?
>> Do you see Agentic AI changing the way
organizations must structure their data
management operating models?
>> Yes. and to everybody's benefit by
putting them all together and making
them all machine machine readable so
that we know what the agents are going
after as well as what we're going after
and it's all the same golden source.
Everything will be better at that point.
The challenge is in order to show that
it's better, you're going to have to
connect it to business results and the
business people are going to have to
understand that if they do X, they'll
have more of why, whatever it is, and
prove that to them so that they can get
there. You'll you'll end up with a very
Why is it working? Well, because we did
all this work, right? We harmonized. We
made sure that these things were
accessible. We took the stuff that they
shouldn't be getting accessed. You
remember the ones that broke out, Mark?
You know why they broke out? The AIS?
>> Well, like the the the chatbot AI, the
LLM.
>> Yeah. They left they left the
instructions in the sandbox.
>> Yeah. Yeah. Yeah. Oh my gosh. Yeah.
>> Only once, right?
>> Yeah. Only once.
Um what what I like about how this
question is worded um because like one
of my favorite operating models that
I've had the most success with in a wide
variety of organizations is a flavor of
a federated model.
>> And when people ask me what is a
federated model really? I I just love
talking about Star Trek. So we've got
the United Federation of Planets and
they have their prime directives, but
then you've got the Klingon Empires got
to fit in there, right? And so the
Klingan empire has its own culture and
it has to adhere to the prime
directives. So how is it doing that? Uh
but when we think of agentic AI, if
we've got a bunch of semi-autonomous
agents with human on the loop or human
out of the loop even, then how are they
adhering to the prime directives? It'd
be like if you tried to incorporate the
Sylons from Battlestar Galactica into
>> into the United Federation of Planets,
right? Um,
>> you lose me at Battle Star. So,
unfortunately, that's how old I am.
>> Or android, Commander Data. And
>> I was with you on the on the Star Trek
stuff. Absolutely.
>> Yeah.
To the Prime Directives and and how do
we build those guard rails? And for
Aentic stuff when it can just spin up
and run around and and do all sorts of
crazy whatever, how are we controlling
that? I I read a book uh by Dale Martin
not too long ago ago called Bears and
Mosquitoes. And so he talks about bears
as being analogous to traditional data
management and traditional data
governance. Um you go to a campsite, you
see a warning for bears, you do bear
safe things with your food, you put it
in your car, there's fences to keep the
bears out. uh there's protocol to be
safe around bears so you can understand
the risks of bears and protect yourself
uh from bears. It makes [clears throat]
a lot of sense. But then he talks about
agents and like the first
implementations of agents,
>> they still kind of fell into those guard
rails. But as we kind of do more and
more with Aentic AI and software that we
buy has agents running in the background
that does something. Maybe you're using
everybody's favorite uh customer
relationship management system and it's
taking in email and moving a customer
through the funnel in an automated way.
It's spinning up an agent to draft an
email to the customer to to introduce
them to some products. the customer
replies and an agent reads that reply
and works with it and then uh uh the
salesperson is is finally the human's
finally in the loop at that point. The
customer wouldn't have even known that
they're talking to an AI. So what is the
operating model for that? Like you've
got all of these things that just ran um
and and uh and that's where he has the
mosquitoes uh analogy. So the mosquitoes
have gotten into your tent. They're
already there.
>> Uh so you can have all the mosquito
netting and bug spray that you want, but
is it protecting you from the
mosquitoes? So the nature of how we
govern, manage, create, and let agents
run
um uh requires us to think about how do
we federate them into our operating
model. And I think that's
>> scanning my brain for a picture to put
this behind you. It's a I'm with you on
this. I hope everybody else is.
>> It's a great book, by the way. Like I I
absolutely adored it. It it it's a
perfect analogy for the the struggles
that we face in managing AI versus
managing data and where the intersection
are between those. So, um it's it's a
book I recommend for sure.
>> Interested. That's one that's in the
link in the chat.
>> Yeah.
Um, I lost track of chat because there's
so many messages in there. Um, does
anybody have any other questions
>> uh to add in?
>> We'veated thoroughly
>> chat or Q&A.
>> Going going.
Well, Peter, do you have any Oh, for
companies who are interested in
developing taxonomy and ontology
foundations. Oh, I love you already for
asking this question. Uh, for knowledge
graphing at the enterprise level, do the
same reference, uh, architecture
practices apply to terms similar to how
reference data is used?
Um, so when I was in higher ed, um, one
of the questions that I'd run into all
the time, you'll love this too, Peter,
how many students do you do we have? I'm
like, well, how many students do you
want? know what a student is defined as,
right?
>> Could be anywhere between 500 and
30,000.
>> I say I'm not even going to talk to you
about it unless you got a half hour to
let me sit down and go through all the
variations of it,
>> all the different types of students. Uh
so like
and when we're we're dealing with
knowledge graphing and using a taxonomy
to categorize our data, especially for
master data, we need to understand those
things, right? And it's all powered by
reference data, which is why the DMBach
kind of combines reference and master
data into the same chapter because
they're they're so interrelated.
>> Like the taxonomy that defines people,
places, and things, especially products,
too, Peter.
>> I mean, it's powered off of reference
data. And the taxonomy is is reference
data. But our ability to master and
survive data
>> is uh
>> I go there and say there's all these
reference models though that we can get
a hold of pretty easily
>> and and use those as starting places for
goodness sakes. I mean it's not going to
be perfect but it's better than building
from scratch.
>> Yeah. And as we get into knowledge
graphing that's basically what helps the
AI learn in a way that's not ridiculous
>> because your AI can learn off of
unstructured whatever or your current
structure. But if you're on the cloud,
uh say goodbye to your your your compute
budget. Uh but knowledge graphing really
puts things next to each other and helps
the AI understand the context as it's
training. And that's why everybody's
talking about knowledge graphing right
now.
>> Yep. Finally.
Some days it takes feels like it just
takes forever to get people's eyes open
sometimes.
>> All right. Well, guys, thanks very much.
Uh hope this has been helpful.
Mark, pleasure as always.
>> Peter, do you have any uh closing
thoughts before I hit the end webinar
button?
>> Again, it enables what's going on. So,
we look at this as you can use AI to
build your MDM and reference data
better. And that if you do that, your AI
is going to [clears throat] be better on
top of that as well. Win-win. Grab it
while you can. Thanks, Mark.
>> Wonderfully said. Have a wonderful day,
everybody, and we'll see you next time.
>> Cheers.