Submind YouTube summaries
Thumbnail for Mastering Reference Data  An AI Essential for Reliable Business Information

Mastering Reference Data An AI Essential for Reliable Business Information

Watch on YouTube

Video summary

The webinar features Dr. Peter Aken, an expert with over four decades of experience in data management who argues that Artificial Intelligence has made Master Data Management (MDM) both essential for reliable business information and increasingly achievable through automation. Despite the fact that 99% of companies prioritize investments in AI, a stark paradox exists where 95% of generative AI pilot projects fail to deliver financial impact due to poor underlying data quality, talent shortages, and scalability issues rather than technology limitations alone. To overcome these challenges, organizations must address deep-seated "data debt" involving legacy structural problems instead of simply purchasing new tools, recognizing that MDM is the foundational bedrock required for any successful AI initiative without which advanced systems cannot function effectively. Central to this discussion are precise definitions and hierarchical approaches to data management, distinguishing between reference data, which consists of controlled vocabularies defining domain values like country names or order statuses, and master data, representing authoritative records about business entities that provide necessary context for transactional information. The presentation applies Maslow's hierarchy of needs to the digital realm, asserting that foundational elements such as governance, architecture, quality, and metadata must be established before attempting advanced AI initiatives, often described metaphorically as a "three-legged stool" comprising data governance, data quality, and actual MDM capabilities. Success relies on building custom reference structures first to understand organizational needs rather than relying on silver-bullet technology solutions, while also treating AI not merely as a database but as an improvisational actor that thrives on role-playing prompts within a well-structured environment. Achieving success in this domain requires shifting focus from reactive fixes to proactive quality assurance supported by automated tools and strong executive sponsorship, acknowledging that approximately 80% of MDM problems stem from people and processes rather than technology itself. Organizations are advised against heavy customization of master data systems and should utilize clear role definitions via CRUD matrices while appointing dedicated individuals over fractional teams to maintain focus on the ongoing nature of this process as a business function rather than just an IT project. The rise of Agentic AI further underscores the necessity for robust traditional practices, where effective autonomous agent performance depends entirely on clean reference data and proper governance to prevent catastrophic failures at scale, necessitating new operating models like federated structures with strict prime directives alongside existing guardrails against major threats. Ultimately, integrating these elements creates a virtuous cycle where using AI to improve MDM leads to more capable systems that leverage existing processes effectively while avoiding the pitfalls of hype-driven implementation failures. Reference data architectures are shown to be essential for enterprise knowledge graphs and taxonomies, enabling AI to understand context without building everything from scratch in favor of leveraging existing standards. By ensuring their systems produce machine-readable outputs aligned with a single golden source, organizations can transform MDM practices deeply intertwined with metadata and knowledge management into a strategic advantage that yields high-quality data feeds for AI systems without immediate commercial costs, proving that addressing the human and process elements is just as critical as adopting new technologies to unlock reliable business information.
Read the full video transcript
You got it. >> Hello and welcome. My name is Mark Horseman and I am the data evangelist for data. We would like to thank you for joining today's data webinar, Mastering Reference Data, an AI essential for reliable business information. It is the latest installment in a monthly series called Data Ed Online with Dr. Peter Aken. Um, just a couple of points to get us. >> I knew I was going to get you. [laughter] Yeah, >> number of people that attend these sessions, you will be muted during the webinar. For questions, we will be collecting them by the Q&A section. If you would like to chat with us or chat with each other, we certainly encourage you to do so to open the Q&A or the chat panel. You'll find the icons for those features in the bottom middle of your screen to answer the most commonly asked question. As always, we will send a follow-up email to all registrants within a couple of business days containing links to the slides. And yes, we're recording and will likewise send a link to the recording of this session as well as any additional information requested throughout the webinar. Now, let me introduce to you our speaker for today, Dr. Peter Aken. Uh Dr. Akin is an acknowledged data management authority and associate professor at Virginia Commonwealth University, president of Damon International, and associate director of the MIT International Society of Chief Data Officers. For more than 40 years, Peter has learned from working with hundreds of data management practices in more than 30 countries. Among his many books are the first on making the case for data leadership, the first focusing on data monetization and modern strategic data thinking, and the first to objectively specify what it means to be data literate. International recognition has resulted from these and an intensive worldwide events schedule. Peter also hosts the longest running data management webinar series right here. this one that you're in. Data online on data.net. Before Google was big and before data was big and before data science, Peter founded several organizations that have helped more than 200 businesses leverage data. Specific savings have been measured at more than $1.5 billion. His latest venture is anything awesome. And with that, let me turn everything over to my good friend Dr. Peter Aken to get today's webinar started. Hello and welcome my friend >> and uh welcome to you sir. I do apologize for getting it but I was just trying to get you to break character and you did. So we'll we'll explain this to everybody else when we get to the top of the hour and uh you guys can jump back in on this but uh yes a pleasure as always Mark. Thank you. Um good afternoon everybody. This one evolved as it went through. Uh Shannon you probably know does most of the I don't know maybe Shannon farms it out to Mark. We should ask him that question of break too. Um but anyway, the uh the topic on it and [clears throat] the more I worked with this material, the more I decided that really what this webinar is about is how AI makes master data management both more critical and more achievable. And that doesn't happen very often where you get two things that are going in the right direction and and helping you out sort of in there. So our pathway today, if you will, we'll start out with a quick data management overview. We'll talk about what is reference and master data management because everybody talks about MDM but they usually always mean reference and master and you can see from the little diagram here that'll be explained in a minute. That's kind of important. Um why is it important? Well, we'll talk about that as well. Look at some building blocks and some guiding principles that'll get us to best practices. Uh and as we get in but as we do this think about this just as a vision. um AI ready data for master data. If you're trying to do something with AI, wouldn't it make sense that the main the primary people, places, and things of the organization were all referred to by the same uh types and and had the same uh understanding of them. And the answer to that of course is yes. But this becomes a data bottleneck that people run into and over again because they don't have quality data. They don't have the type of IT talent that they need to have. If I say it, it's IT AI and and limited scalability. And some of you may have heard of this forward deployed engineering things. I I just finished a program on that that's got a piece that works into this, but not a whole lot. So the the vision of the future of course is to have really automated stewardship around master data. And if master data is automated and well enough understood that it can be automated, that's a primary key, you'll end up with augmented metadata that will allow you to do intelligent data cleansing and get onto this. Now, the reason this is such a compelling vision for me is because when I speak to people about data, they encounter it like the blind people and the elephant. And uh, of course, you know, depending on which part you run into, the elephant looks a different way. Well, it's the same way in it. We've we've come into it through different ways and some people think it's stories and some people think it's pipes and the answer is it's yes, but because most have approached it with these differing knowledge and skills, it means that we have different perspectives and we don't have a a good grounding. The grounding of course comes from my high school knowledge of Maslo uh which started out by saying if your food, clothing, and shelter needs are unmet then you will never be safe. If you're never safe, you uh will never belong to something that is part of uh something larger than yourself, which means you'll never be able to tell yourself from it. And if you're never able to tell yourself, then you'll never be able to hold yourself in esteem. Pretty good thing to do on a regular basis. And and get to self-actualization uh in order to do that. So this is Maslo's hierarchy of needs. I transposed it by saying it's relatively the same in data management in the sense that most everybody wants to start with all those wonderful buzzwords that you see there in that golden triangle that I have labeled technologies. The rest of this is the below the iceberg. It starts out with governance. Then we build architecture, quality, metadata, uh integration of things. And then if you have these foundational things in practice, then and only then does it make sense to do reference and master data. and then after that those other things. Well, so again, think about this because when I do explain that to people, they go, "Okay, these capabilities are are important. I understand that." Um, but I want you to to to to go to go faster. And I said, "Well, if I go faster, it will take longer. It'll cost more. It'll deliver less. And it'll present greater d risk in there in order to do it." The reason is because we just don't really understand data management. uh we've used to use a definition here where we'd said everything that happens between when data is sourced on one end of this and when data is used on the other side but where this forgot was the idea that we are going to try to make leverage to achieve leverage by reusing it and so that piece was completely missing from almost all of the early definition around this. Um so perhaps better definition is that you have a number of different sources of data and that this is going to give you specialized data engineering skills grouped in two areas. One around data engineering and the other around data exploitation and again you can chop them up and put different nouns on them but you get the idea in terms of of how people are doing this just to give you the idea that this portion of the elephant is relatively complex. Remember all of this that I've talked about just for this diagram is still for data use. You still have to have formal data reuse management in there in order to be able to do what most organizations really want to do, which is to exploit to leverage their data. Whether you're going to do it for AI or just business in general, these reference and master data structures that you're going to build are essential to the entire architecture that goes on. Now, I'm going to to give you guys a quick 90 seconds. And Mark, if you're still listening, tell me if this doesn't work, so I'll stop it afterwards. But I want to try to give you what I call the AI reality check here. So hopefully you can hear this. >> All right, welcome to this explainer. Today we are diving right into one of the most fascinating and honestly jarring paradoxes in the tech world right now. We are looking at a massive, massive disconnect. On one hand, we have this absolutely incredible executive investment and excitement surrounding artificial intelligence. But on the other hand, we're seeing a crushing reality of realworld implementation failures. So if you want to understand what actually separates the AI hype from functional AI success and what happens when these powerful systems are turned on us, well, you are in exactly the right place. So let's kick things off with just a truly staggering number. According to the data, 99.1% of companies explicitly state that investment in data and AI is a top organizational priority. I mean just think about that for a second. That is essentially total consensus across the board right up to the seauite. The mandate is crystal clear here. Businesses believe they need AI and they need it yesterday. The optimism is literally unprecedented and the checks are quite literally being written as we speak. But brace yourselves for some serious whiplash here. A massive 95% of generative AI pilot projects completely failed to deliver any significant financial impact. 95%. No way. Right? So on one side of the coin, everyone is prioritizing this technology as the absolute key to their entire future. But on the flip side, almost all of these initial highly funded pilot projects are just crashing and burning on the runway. >> But where is this actually going? Hopefully that came through. I found that a a really nice way of articulating it. And the reason was because first of all, you're not hearing my voice. You're hearing somebody else's. So I I hit you with somebody different. Delivering a message that's pretty hard. 99% of everybody agrees on this. Now, when was the last time we agreed as a society on 99% of anything, right? So, what this is going to do, again, of the tension, I talked about it a little bit, it's it's it's not working well. There are also some challenges because the way for AI to work at least so far according to the data the the way you get in that 5% that does work in in a successful GNAI project is to understand that you have integrated well and complemented an existing work process well guess what that's a basic requirement for how you're going to do MDM in this context here as well there's also a lot of grumbling about the verification tax where people are being put on this I wonder if any of you all have any experience being put on on double-checking in AI and of course then these AI winters. So that's our our quick data management overview. Again, AI only benefits from better data management. So now let's talk about what is reference and MDM. Let's dive into it. The first thing to understand is that everybody's under pressure because there's an awful lot of things that are going on where people are not double-checking what I call promise auditing. you you go in and you say, "This is going to save us a million," and you find out later on that it only saved you a hundred. Um, you know, the next time you get up and say, "I think things are going to save us a million," you might be in a a little bit of a ticklish kind of a situation. But for the most part, we don't educate people about the hype cycle. And it's important to understand the hype cycle in the context of any type of technology, but in particular with MDM. And just in case you didn't know, the hype cycle was invented by somebody named Lady Augusta Ada King. I was in uh Prague, Czechoslovakia not too long ago and I found one of her first programs up on the wall of one of the museums there. Uh but she said in considering any new subject, there's a frequency to there's frequently a tendency to first overrate what we find to be already interesting or remarkable. And secondly, by a sort of natural reaction to undervalue the true state of the case. In other words, it's great, it sucks. The answer is it's somewhere in the middle. Now, I love that. By the way, that is the entire song of the Eagles, the new kids in town, if you happen to know that particular song. But here's how it works out in practice. We find something that works really well in technology for a particular reason for a particular item and we get to the peak of peak of exploed inflated expectations. We get to the height of our excitement and then find out it's not as great as we thought it was. And the question is, where does it actually fit in on this? And there's a lot of context around this trying to figure out specifically what's going on. But the hype cycle around what's [clears throat] been going on in MDM is just unfortunate. Um it's just too many things that are happening. And yet one of the things that has not changed is that it has been a pillar of the data management practice in there for a long long time. This is the uh Dembach version two if you haven't seen it. Uh working very hard on version three that's coming out. uh but uh anyway you can see reference and master data is always been an important part. So what do we mean by reference and master data? These are the practice areas that we're talking about. All right, let's see. Oh, I'm sorry. Got to go there and there. And so the definition of master data control over the defined excuse me reference I start out to do that bad. We're going to start with reference data. Then we'll go to master data controller for the defined domain values, the vocabularies including again you can see the standard terms and things. For example, you may want to divide customers up into current customers and potential customers. And I use that as an example to show managers that taxonomies can be quite useful sometimes. And then they go, okay, so how does the concept of an ex customer come in? Do we fit them as a current or do we fit them as a potential? Right? And then what if we decided we wanted to subdivide our customers down into different areas or that we wanted to implement a customer VIP program? All of a sudden, this reference data is really really particular. But here's sort of the best way to think of it. Um, when you're trying to set up a website or a business, if you don't decide absolutely upfront that you're going to do a multi- uh language business, you're going to have some bigger problems as we go forward. So the reference data here comes in and says, well, okay, we could we could look at this and call this Czechoslovakia. Oh, wait, it was called the Czech Republic and then it's called Cexia. So again, you can see different dates mean that you're going to have to use different ways of classifying that data, which means the context has to be not just that it's the name of a country, but in this case, the name of a country at a particular point in time. Uh, an order status could be new, in progress, closed, and canceled. I've seen a lot of organizations, go great, but I want to put it on hold. I'm sorry, we can't do that because we don't have that as a part of our process. Uh, two-state UPS state abbreviation, USPS abbreviations that are there. Reference data sets in here. When you look, for example, through data, you'll see a lot and a lot of uh UK's showing up as addresses. And of course, the proper designation for it in most organizations is going to be Great Britain rather than that. So, let's take a look and see where the concept of master data came from. And there's another one. We're we're changing slowly because of DEI push back and things. It's called primary data. Think of it like the bedrooms. You no longer have a master bedroom. So, you now have a primary data rather than master data, but it still is the same concept in here. And and here's where the concept comes from. In general, uh I remember doing this processing when I was working on uh mainframes in the 70s. So, there was data about these business entities. In this case, it's about uh how much I have on my pot belly sandwich shop card. Okay. And the business rules dictate again the parties locations and uh provide context for these transactions. The term master file popped out of exactly that process because this is the way we used to do it. Here we go. The balance of $100 on my pot belly card. And I obviously did this slide some time ago because I thought you could get a sandwich for $5 at the time. I know that's funny. Um, nevertheless, the balance comes off of my pot belly card and by the end of the day, the new balance has been updated to the new master file of $95. And if there were problems, hopefully they get written out to a error log of some sort. Keep that error log in mind. It's going to be quite useful to us in just a few minutes on this. So, here's reference data versus master. Again, control over the domains. The reference data in this slide is the fact that the for a period of time the FBI and the Canadian Social Security uh kept nine gender codes for nine possible entries that you could have uh as a as an entry on the Social Security system for Canadian systems and for the FBI. Uh and these nine gender codes of course are fascinating uh all sorts of stuff. The master data part of it then the reference [clears throat] data is the allowable values. the the maf the excuse me the reference data for it the master data for it is going to be the fact that mine says male right and that you'll have a golden source on gender for your customer Pat whatever it is that we're trying to do but what doing is providing context for the transaction data so you can't just say it's a customer it never works right it's too simple you're always going to qualify that in some way and you're going to find out more about it so these definitions again are trying to get to this golden version on Here Gartner really comes back and I think correctly categorizes master data management as a discipline or a strategy but as I said it's also a pillar for requiring good AI as well. The [clears throat] problem is it's sold as these silver bullet solutions and this is just a perennial problem we have in technology but in the in this category I've seen organizations spend an awful lot of money on these things parties places things sounds pretty easy if we can get those pieces we'll do a really good job. So let's take a look. First of all, what's happening right now is that because AI is of such keen interest to literally everybody, there is an incredible increase in data. Let's capitalize on that and try to make use of it uh in order to do that. And let's start by educating our users about data debt. Now these are areas in which AI can be extremely helpful. first of all about looking at the investments and trying to come up with the uh projections and things like that but also for how to approach the process of making sure that you don't just put out a new set of uh data structures but this data debt is really something that you have to work through in order to get that. Uh and and finally the the ultimate mantra is you're you're going to get your data the better you treat your data the better the data is going to treat the AI. Uh I just wanted to show you this one illustration of a history. I was doing a a project on a separate notebook. Okay, so I have a notebook for my work at VCU and I have a notebook for my work outside of VCU. And somehow they got crossed and uh it started to hallucinate a little bit here on this last slide. So I just left that in because it was kind of funny. Um I'm going to stop here though and play a little short clip that I want you to sort of inculcate. And the idea is get your AI to take on a role. >> Okay, I'm sure you've seen all sorts of posts telling you what prompt to use to get the most out of your language model. I think you can pretty much forget all of that because there's only really one important thing to remember, which is that the AI that we have now is really, really, really good at roleplaying. Um, so you shouldn't talk to it like it's a Google search or like you're trying to extract information from it, as though it's Wikipedia. Instead, you should talk to it as though it's an improvisational actor that can be any character imaginable. And the reason for this is that large language models, they don't store facts like it's a database, right? They they generate responses dynamically based on having read everything that humans have ever written and the prompt that you give it. And that means that AI the stuff we have now it doesn't have this stable identity right it doesn't have a fixed worldview it doesn't have personal beliefs and so if you prompt it in a particular way it will respond in kind. If you try and prompt it as though it's a Shakespearean bard then it will give you a flowery response. If you try and prompt it as though it is an extremely effective and smart scientist, it will give you the relevant response. But if you prompt it as though it's an encyclopedia, it's going to try and sound like one. Uh, but it's still just performing a role and you've effectively just restricted what it can do. So instead, what you should do is you should you should imagine that you're a film director, right? And that you have a character in mind and then you should prompt your AI accordingly. So don't say, give me three interesting facts about science. Say you are a world-renowned scientist with PhDs in biology and chemistry and physics and and and your nephew says that science is boring. You only have a few minutes and you've got to you've got to give him counter examples. What what extraordinary stories do you use? >> I find that to be some of the most useful advice that I've gotten about AI guidance uh in cases. I tend to be a Google fan, but that's because I'm doing a lot with YouTube videos and the interface there is phenomenal here. And you can see in the upper right hand corner I put a lot of uh AI generated content in here because I want you to see how good it is. They have some real knowledge that this has managed to scrape up and it doesn't look like we have disagreements about facts around reference and master data. And so it becomes a very good source of trying to find what's actually in there and how you can get it to work. And again it's nice steps to take you through the process. It can be a very much of a help uh in order to do that. So next section, why is reference and master data management important? Again, I keep hinting around with this little picture, and this is what you're going to see, of course, right now. Reference values control this accessible data value. I think I've said that three times. You can see it's a tiny little yellow dot in the upper right hand corner of the screen saying, for example, what countries do we do business in? What types of accounts are available? um what are the controlled vocabulary items that we're going to be using throughout this particular domain of the project that goes on. Those control the master data items. The master data items control the access to the system capabilities. Are you a member of our premium club? Uh you can see if you want to add a premium club or a VIP sometime, you better build it in the first time because adding it later is never an easy task and that's where you have to fight data debt all the way through. Are you authorized to use this? Are you using sharing common data structures that go back and forth? But you can see here again the reference data has leverage over more master data in terms of what's doing that. But that where the bullet efficiencies really come through and I want to thank Chris Bradley for allowing me to use his example here which is such an articulate piece. Uh this is the idea that these master data items control the transactions. So that's where the $5 for my sandwich or the fact that I've been authorized or that I can even make a like on a particular system. Uh this is the kind of thing that uh uh really helps to understand with this reference in here. And now MDM can make data governance much easier because it gives you a much narrow target to focus upon. Let's think about that. What are we going to focus upon? The answer is what is strategically important? I can't tell you how many organizations I I work with and sometimes they just don't seem to think that's part of the picture but gosh it is. So the word strategy didn't start to get used until the around 1950 when the management consultants discovered it from me coming from uh the use in the military and the current definition according to the management consultants is you can see a master plan a game plan most importantly it's a thing it becomes a PowerPoint deck I've had some companies where I've gone to work for them and they say don't you you're not going to write a report what we want is a PowerPoint deck describing X Y and Z could have done that from home but okay you know we'll get it done but I go back on the word strategy And it really came from the use of the word military uh in there. The military invented the term strategy, which is a pattern in a stream of decisions. You can see that's very different from being a thing. You don't ever consult the PowerPoint to see what's happened. But if your goal is to do something specific, and I'll just give you one example. It's worked very well for one organization for years and years. Everyday low prices, you understand who I'm talking about and why. because they've done a great job of making sure that pattern guided their people throughout the development of this behemoth that we call Walmart. Let's this is not a Walmart story, by the way. I'm going to tell you a story, but it wasn't Walmart on here. But at the first year, they were implementing the MDM, and there was real confusion because they'd really only put up one leg of the three-legged stool that you need to have. uh users didn't know they had spent literally, it was a joke, but they had a three plane loads of consultants that would come in uh 60 consultants coming from different parts of the country to to work on this master data management system that they were working. They spent $60 million on it and the business did not know how to use the MDM. So, we had a bad transfer of whatever was supposed to come on. I don't know whose fault it was. That was not part of what I was doing, but there was general agreement that we should go back and restart the effort. So they went back and did a root cause analysis and found out that poor quality data existed in the system. That said, you can pretty much say that's the case for many systems out there. So it's not that hard of a bet uh to win. That gives you a little bit more, but you also need to roll in that in the adequate training. You have to get people to understand what it is you're trying to do or it won't work uh on this. And I'm just going to drop in right here. Even though I put the button on the speaker, they went and they started doing data quality, right? they say, "Okay, we'll get let's get data quality going." That provided more of a stool leg, but still not the three-legged stool that we're looking for. Um, the point I was going to make here, though, is that many organizations try to do this and then they come along and try to get this done without actually investing. So, if what I say is people will come to me and say, I'm investing a million dollars in a data quality stool tool stool. There's a slip a data quality tool. I'm not pooing data quality tools. Um, but if you have the data quality tools and a million dollars, you should still plan to invest $4 million to make sure that the organization understands how to use it properly. So if you have only a million dollars to invest, invest 200,000 in the tool and invest 800,000 in training your people how to use the tool and you'll find it actually works a whole lot better. You can see there are a lot of interdependencies in this case as well. Data governance almost always three legs. That's why we're getting the three-legged stool. data governance makes a case for and is responsible for the data quality and the data quality is a necessary but insufficient prerequisite for this excessive master data items which then go into master data uh capabilities that constrain governance effectiveness. Remember we have over here on our diagram the consultant says our methodology I don't particularly like that it's a fancy word and really if you look up the word methodology means the study of methods. So, that's not terribly useful to any customer uh that's going to do it, but a realistic way of practicing it is probably a better way to look at it. Select three data management practice areas because you're probably going to need all three of them in order to make it work. Uh again, I've done it with a number of different combinations, but here is a reasonable way to do it. We're going to put all this together with reference data, data quality, and data governance in combination of three. If you add a fourth, it's more difficult. If you keep a third uh keep it just a two, it's kind of hard to make it work. So three really turned out to be the right number there. Similarly, let me tell you a quick story about an MDM success that was an organization that had purchased an ERP. That stands for an enterprise resource planning organization. They were buying something like people software SAP in order to get it to work uh in there. They found out that the problem was every time a price of their product, which was a liquid product, was transferred from one tank to another tank, it counted as a retail sale. And they said, "We can't do that." So they decided that rather than modify the ERP, they would be very careful. Nobody could use the word tank anymore. All the tanks had to be qualified. And each qualified tank had a set of business rules that were associated with it. For example, transferring product from a pickup truck, I'm sorry, a a tinkerer truck to a tank was not considered a retail sale and they just made sure that they didn't count those things as tanks. They counted them out. This company also did transfer air as you can see uh fuel from one plane to another plane flying through the air. Very very uh interesting way of accounting for all of that. But again partial of course there's always the word tank as well. Just if you didn't know, when you buy a tank, you also buy about 32,000 mil, excuse me, 32 million pieces of data that control all throughout there. What's happened, of course, is that over time, multiple sources of master and reference data have grown up to each of these pieces because that's the way that finance work. Finance would say, why would we send the bills to anybody other than the master bill? Uh, the place that says to send the bills, but there have been lots of people who work outside of finance that get it to different places. because of course it's what happens as data start to proliferate through the organization. You end up with real challenges around all of that. So here's a wonderful piece of architecture here and I'm just going to walk through it. You start out with the business data stewards and they have a code management system. This becomes sort of their lingua franken and that code management system is of course metadata that is focused on making sure the reference database of record follows those particular codes that you're using. Then those codes eventually are pushed out to the OLTP systems usually on a system by system basis eventually working their way into whatever databases that you have associated with that and finally gets to your weight data warehouse which can involve pushes to dimensions depending on how you've structured it. Finally you can add in here as well the externally sourced data. Now from a reference data architecture this is pretty much how most people do it and this is how the AI has learned that it's done as well. So if you try to follow thing like this and say hey this is the plan that I'm trying to follow help me do it and help me avoid mistakes where other people have made mistakes it can be quite helpful. So that's for your reference data that's in other words the gender codes that go throughout your organization or whatever is equivalent to a gender code in order to do that. Same thing happens for your master data you end up with a system of record here that follows through from the master database of records. uh in order to pull that together. Then it goes to the OOLTP systems uh again following similar pathway getting proliferated throughout the system. And you can see of course you use the same absolute hard uh uh infrastructure in order to do this. You're just coding the data slightly differently getting to the data models and your externally sourced data in there. In other words, it's very very much of a parallel operation in order to see all of that. You can combine them of course into one and most organizations have uh in order to do this but you need to do additional thinking around this as well. It's not just a technology play. You've got your technology. You want to take a task orientation really. And the task orientation is that we used to make pins by putting them together in 12 steps. But eventually somebody looked around and said, "Okay, there's probably an easier way to do this. Maybe I could combine steps one, seven, and nine and come up with a faster one. So this is of course the whole process of business re-engineering or going digital or in this case AI because you're taking steps that are nonvalue added out of the process and reducing cost and increasing your revenue in order to do that. You still have to of course go through all of the standard business rule analysis that you would go through as you're doing it. Here's a an interesting example that we found one time. All the information was on one screen on the source system, but we were putting in also an ERP package software and the same information was spread across 23 screens. I want you to imagine being the customer service representative responsible for directly helping the customer who literally was standing across the desk and trying to do that by basing it on 23 screens. These things fail because people don't understand the processes. When I say these things, I mean both master data and artificial intelligence. So, lots of help around gathering different types of things that can go into here. Tell you a secret though, you won't find this in the slides anywhere. If you take this information from your log files and get AI to analyze it, you will get a pretty good reverse engineered process. It'll give you quite a lot of information about your business that is immediately useful. uh in order to do that. You want to know more, come see me after the event and we'll talk. Uh anyway, here's another way to think about process understanding and that is that most people looked at the traditional engine gasoline powered engine and said that's the infrastructure that we have. That's how we're going to make things work. We will put either a gas engine in a car or an electric engine in a car. And if you remember the early days of electricity pre Tesla and everything else, there were a couple of uh uh entries into there that were quite interesting electric cars that came on. But they were looking at as if I have an electric engine, I can't have a gas engine. If I have a gas engine, I can't have an electric engine. What if you could have both? Well, of course, both was what Toyota came up with when they invented the Prius. They had the engine, they had the electric motor both in the same car and the battery. Now, I have the most recent version of the Prius. I believe it's version five. And those engine and electric motor have now been combined into a single structure. That makes it even more efficient in terms of what you're seeing. But you can see here that what was going on in the Prius world was that the engine and the battery electric motor were switching back and forth sometimes multiple times in a minute, sometimes not. But they had the flexibility to be able to do that. So you didn't have to choose, do I run the engine or do I run the other? And that understanding of the process is what let Toyota dominate the market in order to do that. When you go ask [clears throat] questions about how master data can work, think of how it can work in your domain context. And it turns out the AI understands this as well. So here are just some general scenario if you happen to be in a retail setting that you can look and see in order to do this. Now, how would you use AI in order to do this? Well, you can use AI both to find the problems and to help you resolve the problems. Uh, in order to do this, in order to do that, I urge you to look at AI as an extension of the knowledge workers capabilities rather than as a replacement for it because it does also seem to be providing much better work uh, in order to do that. All right, here's another little AI summary again driving this data investment which gives us lots and lots of things that can go into this but we also understand the data dead actually I think we had the slide in here twice I might have uh duplicated that one sorry about that guys we'll keep rolling on here all right we get to the building blocks this is the other part this discipline is reasonably mature uh Mark will be able to tell you that he was building master data management systems for his customers when he was still a baby in a high chair eating uh baby food, right? But uh you know the goals of these are very very straightforward. We want to provide all of this type of information and again I [clears throat] ask you what AI system would not want to have authoritative source of high quality master and reference information lower cost complexity and support for the integration efforts that come into this. There's another whole set of categories here. These are mainly reference slides for you. Remember, you get the slides so you can take them away and and take a look at them, but understanding exactly how these activities correlate to what's going on in your environment. And then specifically looking at what can likely be helpful in order to get up with this. And this gives you these primary deliverables which allows you to cleanse data, tells what your requirements are. There's a whole series of roles and responsibilities that go through all of these as well. uh in order to come up with it. Again, you need to have it takes a village, right? We get all that sort of thing. But there are lots and lots of people. You don't have to have all of them at once. However, this has been the part that most people have failed with. And as I said, it's sold as a technology first solution. So, you likely have somebody in your organization doing ETL or ELT already at this point. You may have tried and hopefully had a good experience with getting reference and master data applications uh in place. uh this allows you to get them. You can buy them separately, but most people buy them together because they are somewhat complimentary and similar. Uh again, some people say you don't need data modeling tools. I say good luck with them as well. I also include process modeling tools. However, again, AI is proving that it can largely supplement some of the needs of those uh uh pieces in there. Uh again, metadata repositories are going to be in there. And again to just remind us that the metadata repository is a thing that data cannot go into but it doesn't mean it has to go into. So one of your first questions should be not is that metadata but is in fact the metadata worth maintaining formally because that's a very big decision. It will help tremendously in terms of your leveraging opportunities in here. If you haven't heard of data profiling tools, there's another whole uh topic we do on these, but they are a wonderful uh set of technologies. There's a data cleansing tools set of things, integration tools, all sorts of things that can go on. The challenge, of course, gets to be that oh, anybody ever seen business rule engines? Boy, they are fun. That the technology again, a fool with a tool is still a fool. And that's an unfortunate way to say it, but as I said, I saw this one organization and it's not just the one. They had lots of failures that that happened in these things because people generally didn't understand what was going on. So, here's a just a summary of reasonably good vendors in the sense that you've seen these folks out there and they have uh definitely had pieces that are uh useful looking at that. Somebody may want to ask a question about Oracle as we get to the end of this because it's been in the news kind of lately uh in order to look at that. That said, I would always start by building yours first. And people go, "Whoa, you want us to build something?" Let me give you an example. If you build your first version of it, it almost always will tell you so much more about what you're attempting to do at so much lower of a cost that it it just I have certain conferences I've been banned of because I tell people things like this, right? So, let's just take a a hospital situation. They have a number of hospitals around the regional area and they want to make sure that they have doctors that have admitting privileges. And you may think that was an easy thing to do, but actually turns out to be kind of a a messy piece, but it was implemented on a SQL Server database because everybody has somebody they can program in SQL Server. And all you're doing is you're making a database of the golden data. The golden data in this case about physician ID and and admission. By the way, if you have a question about sharing data, there's a really great website called fiduciarycoms.org org that you'll find has some really interesting pieces there. And this system is connected and they get it to work and they practiced with it without spending any money. They had this all inhouse and and they started to understand the process of admitting a physician and then the process of actually getting into a hospital which is different from admitting a physician to practice at a hospital and they could extend this to another system and with this other system they learned even more. or it was a different sort of a processing and again you can see they're gradually growing this and at some point it becomes quite obvious where we need to replace these systems and give them instead something where they are connected to a generalized bus technology and that technology here in this case connects up with these systems because as long as they got three it's okay but we have the fourth and the fifth and you can see how it gets more and more and more on here and eventually they're going to get to the point where this is not the right technology for it. So you fix that by then going out and buying your commercial master data management package. But by doing that if you take three years and three years sounds like a long time but if you take three years to do this well it will help your AI because the same governance can provide this you don't care that you don't have a master data management package on the other end of this. You care that you have good quality data that is feeding your AI and all the rest of your systems. Why wouldn't you? So there's your win-win. Okay, let's take a look at now the idea of how this is intertwined. And if we look here, you'll see master data management practices are highly intertwined with this implication. Knowledge management in this case, the metadata management pieces, the data quality pieces. There's a larger picture that gets told in all of these, but you get the sense that that's intertwined. Here's another one. This was a different organization, but they had specific pieces where they were looking at operational data and nonoperational data and looking in both cases finding instances of where master data enabled them to leverage their efforts and feed their AI with good quality information. Couple of important considerations again it is a technology it is sold as a technology. you need to think of it as as a way of designing systems and there's lots of good information that you can get from it. By the way, this is supplied by the AI. So, very very good guidance as far as that goes. Here's another one uh in here and again this is the idea of just saying the old way we did this was very much dependent on a lot of heroic efforts and really good specialization. Now, we can both use AI to get it this way. we can shift our reaction to proactive data quality and saw earlier I was even saying automated metadata quality which is a real possibility and finally starting to move towards automated policies and standards where it starts to figure something out and ask questions. Now you have to have good people on the other side in order to do this or you will definitely not have uh success in terms of what goes on that. Okay, couple final pieces on this one here, which will take us a little bit to get through, but uh got quite a quite a chunk on this one here. First of all, while you'll see these crazy crazy rates of failure, and if you just Google, you'll find enough uh things that are problematic out there, take it with a grain of salt, they're still being used, they're still being sold. There's a a gentleman who's been running a conference very successfully in that area for many years, but they are challenging. On the other hand, I had the same challenges when I was doing business process re-engineering. Does anybody remember that from the late 80s and the early 90s? Uh, done well, it could make a significant difference. Done poorly, it made a huge mess out of things that were just not necessary to be messed with. So, here we have less than fully satisfied with their data programs and 70 cent were less than just satisfied, right? uh in terms of there's another one here can't track and consolidate where these things come from they have no idea of the spender that's why you put one throat to choke that's why you have a chief data officer often times or chief AI officer uh 25% of the clients spend is wasted duplicated effort I have good numbers that say between upwards of 40% of organizational IT spend can be reduced by better data management practices in order to do that as well as your cloud bills too which is Another wonderful way to train. By the way, should these things be cloud-based or should they be uh on prem based? Well, again, it depends on what you're doing and what you're trying to accomplish, but there are equal good versions on cloud as well as on prem uh in order to look at these things. Uh 64% are planning to rearchitect the reference data. That's a big big lift. Imagine taking the pipes in your house and moving them from one wall to the next or one floor to the next or changing a bathroom or other things like that. These are not good. And over half of these companies were spending $4 million a year on this reference data. Uh so this again a market in order to look at this. What were the basic causes? The things you have to look at and I'm sorry about the root cause thing. Uh the scariest movie I ever saw just so that you get a little bit of picture in your mind here was a PBS documentary about somebody flossing the first time as closeup as this diagram is showing you there. It was terrifying. Uh I've been an avid flosser ever since, but that's not information you guys need to know. Anyway, 30% of these in this instance was failures. Uh serviceoriented architectures were similarly challenged and again what were the problems? Well, bad leadership uh plagued many many of them again implemented as a technology or as a project. Uh again, get it in your people's heads that they are not going to need their data program when they do not need their HR program, right? And that will actually help them understand this is going to be with us for a long time because data is kind of like HR in the sense that if we don't know what we're doing, we're probably going to have a mess. Uh in order to look at that, again, the MDM was either the enterprise data warehouse or an ERP. Sometimes it's just too easy a solution to put in place in order to come up with that or it's run as an IT effort. Now again, this is not to blame it. it has an awful lot going on it and it's software so why shouldn't it do it well they should of course be responsible for installing the software however in order to install the software they can put it in but unless they want to pay for something that doesn't get used they should insist on having good understanding of the processes it is designed to support master data management needs to support the processes the process with the Prius was that you weren't sure whether you were going to need an electric or a gasoline motor at any point in time and they invented something that you could switch back and back and forth between them in minutes. That meant that MDM was uh really focused in that right area. In this case, MDM part of an IT effort generally about one in 10 IT shops I see do this really really well and they wonder what everybody else's problem is. But uh uh for the most part it's it's definitely been a challenge and they don't know again the software goes in, they got to air zero return code, so it must be fine. No, not definitely not the way you want to think about this thing. If we separate governance and quality, it becomes even more problematic. MDM provides us the ability to focus on these things and to provide it to us in terms that the business understands and that's critical for us that these MDM initiatives are implemented oftentimes with inappropriate technology. If I kept growing the build your own example that I showed you, eventually it would crash and that would not be good for the hospital system. However, if you monitor this and keep in progress and understand it and realize that those three years that you're doing the understanding about your environment and learning what's going on there will give you a much faster pathway to value when you do go buy this stuff then buying this stuff at first and trying to figure out how it works is definitely the better way to go. And finally, let's eliminate the silos as far as that goes. There's a number of success factors here. Again, just briefly run through them. you're more likely to understand these strengths and limitations and then therefore have success because if you want MDM to fix everything, it doesn't. It's a secondary effect. That said, if you want your MDM to constantly score high on its AI governance areas, it can be done very very nicely. Small steps, right? Crawl, walk, run, set expectations, communicate. You can't can't undercommunicate. Uh in this case, you're going to have to have this integrated with your architecture. or if it's just hanging out there by itself, it's not going to work. You need to have incentives to make sure that the master data is desirable. I um once was involved with a a piece where um a company was getting subscription information from another company. So, in other words, they were dependent on subscriptions and the subscriptions only came from the other company. But by goodness, they were sending duplicate data and bad data to it and we didn't have any terms of service agreement. It was just send us the data. It's like h we could be a little more specific about I won't go read the rest of these. You get the ideas in terms of of what's going on here. Uh again, each of these requires specific practice area focus, but at this point in time, AI has the capabilities to be a good solution for you. So, by all means, uh I would suggest uh this is a really extraordinarily good area to look at to practice. Uh again, just like anything, if you don't have executive sponsors, it's not going to be reasonable. The business has got to own the set the the context for this. It just doesn't work without it. If you have just an ITled project, build it and they will come. Does not tend to work for these things. Again, the stronger your project management, organizational change management skills, uh the better they are. People process technology and information all the way around. I insist on having this detailed information as a CRUD matrix to show the support for the process that we're doing. Uh sometimes it works out really well, sometimes it doesn't. But uh the ones that have the CRUD matrices tend to work out much better than the ones that don't. Uh again, if you've got documentation on your existing processes, use it. Uh support this continuous improvement process. It's a a great example of of how to use it properly. uh management needs to understand the importance of these dedicated individuals. I always if they tell you I'll give you 10% of 10 individuals, I say give me the one individual because the one individual is going to provide much better focus and much better results than 10% of 10 people or 30 30% of three people, right? Let's let's go for those dedicated pieces in there. understand how your systems MDM works if you have one and if it doesn't make sure that you understand where it integrates uh and what your external connections to it are uh in order to do this. Resist the urge to customize. Uh again this is master data. It's your business things about which you create, read and delete information. Uh it's it's pretty straightforward that you shouldn't need to do any customization on this but I see almost everybody does it. uh stay current with the patches that are there and then of course lots of testing going on with it. Final set of guiding principles uh in order to do this is make sure that this belongs to the organization that everybody understands they own it and they actually own it now but they'd rather own it when it's in good shape. So we're going to make sure that they get all the way to the goal there. Again, it's almost synonymous with quality improvement. There's just no way that you can do this one time and expect it's going to work. But management wants to understand why it's not done by Friday. So, you've got a lot of communicating to work on there. Again, hopefully this is helpful in order to do that. Um, the business data stewards are the authorities. They get to determine the golden values. Um, yes, we'll just leave it at that. The golden values represent these correct sources. We shouldn't be getting information from places that are not designated as golden values because the data is of less known or unknown quality whereas we want it to be of known quality. And finally, we're only going to be replicating I should say finally uh replicating these master values only from the golden sources. Uh so when somebody gets a copy of this, we go to the correct place in order to get it and put it all together. Uh there's the finally it's data changes require formal change management. Uh yes we always used to run to the DBAs in the past but uh that that is going to be a a more controlled process. It doesn't mean that they still can't react rapidly. They still can but there's additional guard rails around it. I estimate it's about 10% additional overhead to do this properly and most organizations can uh afford to do that. Now let me give you a what I consider to be a real treat. I have lots of artifacts that I've obtained over the years and this is one I obtained from a guy named Dave Evans at British Telecom. Hi Dave, if you're still out there uh you didn't know this was going to last so long. So they were putting data into master data and they needed to communicate this process to people uh to the employees in the organization to the folks that were in the organization that were part of what was going on to to understand the vendors the employees the tellers you everything right bishcom had a bunch of everything in order to do this I've put a copy of this out there if you're interested at the anything awesome uh email that you can just click right there and go to but I'm going to play it for you and I just want to tell you what I'm going to play you so that you get it. [snorts] Uh it starts out again in an email that comes from the managing director of the bank. So that would be like the president of the United States emailing me personally uh in order to do this. And I might look at this and say, "Wow, there's a a thing to click on." Now, of course, in the old days, it was okay to click on them because they had embedded viruses in them. It can still be done safely. uh in this case you put it on a controlled video player or something like this but this was done as an email and they had good emails tracking statistics so they could tell of the 60,000 employees at British Telecom X number had opened it x number had played it x number of times uh it was really quite good information and most importantly they achieved the objective because when they tested people not just immediately after they had seen it but months after they had seen it and then years after they had seen it they were able to identify some of the seven sisters that you'll see which was their catch name for their master data management solution that they used. They never talked about reference as part of this. It's not good or bad. Uh they just didn't. But they they I'm sure did piece of it. Maybe it was integrated. Let me just step back and play you this information real quick. [music] Bong Boo. [music] [singing] [music] [singing] [music] [singing] >> [music] >> talk about a simple explanation of what went on. I just I you guys may think I'm crazy, but I really really like that it tells what it's doing. It did it really well. You're welcome to reuse it. Dave put it out there in the public domain for you all to to do it. Um, here's another little AI piece that came out mostly pretty good. The real question here is when you're looking at this, you can build it yourself or you can create a custom version of it. There are some instances that are custom makes sense, but again find out what are the actual complexity of the requirements. Do comparison back and forth between your relative uh uh competition in the organization in order to do this because it is a challenge uh to do this and you if you don't have control over these pieces, it's absolutely not worth it. I in general have aired 90% of the time on the path that I described you which is to start out by building a very simplified one yourself to learn about the process of supporting master data management and over time growing with it and then you can have a conversation with the vendors and put it in after [clears throat] the end of that but rather than before it. Uh so we've spent some time uh again about 55 minutes so far looking at a data management overview in this and what is the reference and master data. I hope you understand the leverage reference has over master data and the master data has over all your transactions. You understand that they are the best parts of what you have in data in the sense that I say the best parts is the bones of a garden. If you're a gardener you'll understand that particular reference. um they they provide this and there are really good building blocks all the way around out there to use this. You do not need to start from zero. There's so much information and it's mostly known correctly by the AI. Now AI is still errors. So don't be silly with it. But you've got some ideas of what these guidances and best practices and it's almost time to bring Mark back in here and uh see what he's got to allow in terms of his experiences with this because he comes at it from a a little bit more of a semantic perspective than I have been able to put my time into. So I learn every time I speak with him. But anyway, finishing this up. Yes, culture eat strategy for breakfast. 80% of these problems are people and processbased not technological. So let's make sure that we use them correctly uh when we do that. Shifting the ethics component of it as well. Just because we can doesn't mean we should do it. And and these guidelines can be implemented in the AI guard rails which will give us a little bit of safety uh around these things. What you're looking at is the overall process model for master data management. In this case, here's the definition. Planning, implementation, and control activities to ensure the consistency with a golden version of contextual data values. Again, people, places, things. That's what we want to think about. If we get those right, we can start to expand to other areas. What are the goals? Authoritative source uh lowered cost support for the integration and intelligence types of efforts on this. Your organization must put your own goals through this filter. Go back to the strategic uh piece that we talked about a little bit. What is the focus of this? The MDM is going to be a major piece of it. You can architect it so it supports your efforts or you can architect it so it doesn't support your efforts. Again, look at your inputs that are going into this. The various suppliers that you have, participants. I know one organization has 200 people doing this kind of work. Uh and for the organization, we're glad they do it and they do a wonderful job. Um all sorts of activities in here in terms of what it is and you can see they're also subdivided as to planning, controlling, development, operational around all of this. And then we get to the deliverables, the requirements that come out of it. What's the master planning, the golden record data lineage, what are the metrics and reports? What are the consumer uses of it? Who's going to be using this? Make sure that you express those things. If your company makes tanks, make sure that your uh uh uh reports express address tanks otherwise you will uh have a translation issue that goes in there. Look at the measures that are around this. Uh again, just a a start of a number here. When I went to work for the Department of Defense in the late 80s, early 90s um they had 1400 connections to the internet. By the time uh I finished with them, they had 14. You know, it can be done. One organization I look worked with started a process of eliminating 400,000 access production databases that were running in there. It was crazy. and uh that took them 10 years but uh there was a strategic piece so they use tools use all sorts of things to to focus this these are the the union of the ven diagrams you do not need all of this in order to do it and a lot of it can be done with lowweight uh stuff so hopefully what I've done here is is taken some time to shift the narrative away from the traditionalness of the MDM challenges to say what can we use with AI as a forgive the expression co-pilot but it's a good description of this and that doing MDM AI makes it easier to do the MDM and doing the MDM well makes it more uh likely that the AI is going to be useful around this. I've left you with a couple of references uh to take a look at uh look at the uh of course the books pieces that goes up and we've got pieces that coming up and hopefully you'll be able to to join March uh Mark in the uh uh event in Providence coming up and we're right back at the top of the hour. Perfect timing as always, Peter. >> One of these days, Mark, you're going to say to me, you know, you've been talking to yourself for the last hour. There's nobody out there, actually. So, I apologize for catching you up short. Should we explain to people what was going on with that? >> Yeah. No. Uh, so Peter was trying to take me off script with his animated video of me and and as I'm going through my script, I could see out of the corner of my eye that that I was moving around. So yeah, which is hilarious. >> And we never did this. This is the amazing thing. You should see what the the actual vocals are on this. We talk about going and getting a drink, [laughter] which you know, maybe not far off, but I just sent it a picture and said, "An animate this." And it's scarily realistic, but that's not what you guys want to talk about. What sort of questions have we got that two of us can dive into? >> Well, we've got some wonderful questions in Q&A and then there was some wonderful conversations happening in chat. So, I might jump around a bit. Please, please, please >> good at this stuff. >> So, top of the question list, uh, which I love, is a Gentic data management even real or is it just hype? Are there actual tools, architectures, and implementations happening today? >> I'm certain that people have I have not seen it. Have you? >> Yeah. Well, yeah, I'm doing it. >> Okay, [laughter] there you go. So, so jump in. And >> so as a plug for our conference in Rhode Island in November, um um me and a good friend of mine are doing uh AI data quality lab and it's all about how to build AI agents and use AI tools as part of a data quality initiative which is a huge part of of data management as you as you know Peter and and there's a lot of places where you can build agents and and it really depends how deep you go as well like are we orchestrating agents or having agents uh orchestrate amongst themselves to solve uh data management pieces? What are the guard rails uh that we're putting in place for that? What is the AI governance that we have in place to manage those uh agents and control them so we don't end up in the news for some catastrophic failure like happened to that one company? I think it was back in April uh where his uh agents had basically destroyed his production database, removed all the backups and like fried everything. [laughter] >> It was kind of hilarious to read. >> Yeah. Oh, yeah. And and like I know exactly how it happened and like the the the person who set that up like deserved all the failures that they got because they they missed out on so many just common data management practices. But yeah, there there's a lot of agentic stuff that's happening out there in the real world. Um, would I say that it's popular enough to be a framework uh or something that everybody can just grab and follow? I I don't think so. I think we're not quite generic enough yet. Uh so I think if you've got a fairly AI mature organization and understand how to manage agents I think I think there's a lot you can do um with uh data management that anyway that's that's my my uh my take on that >> and is it is it not true that in order to manage the agentic performance you have to do pretty good at data management? >> Oh yeah a thousand%. Like [laughter] all all of all of everything that we do in AI requires so much traditional data management to happen to be useful >> and and so this is where a lot of failures have been happening especially around master and reference data. uh this this conversation was happening in chat a little bit and it's like what how's my data dictionary work like is it reference data is it is it is it metadata uh uh how is AI using it and and it's so true in that we have to train these things and before we had started recording Peter I was talking to you about context layers and semantic layers and um and people using languages like sparkle and shackle um to uh be able to knowledge graph out their content so that AI can learn from it. There's so much traditional data management that happens in that universe uh to to make those things function. I I think I put a link in chat uh to enterprise knowledges uh uh book that they recently published um a very very good book uh that talks about all of that. And sorry, you're getting me on a tangent here, Peter. >> So, this this one organization that I've been talking to, one of their executives had come up to me and he was like, I you know, I kind of want to just replace our entire analytics function with AI agents. How can we make that happen? And I'm like, well, how many sparkle and shackle experts do you have? economics of it realistically. >> Yeah. So, it's like probably not a good idea if you don't have the right expertise to actually train the AI on what your data means. >> And then you're going to fire them, too. So, they're going to be real enthusiastic about cooperating. >> Exactly. >> Yeah. Not enough people have uh learned that lesson yet. [laughter] >> Well, we've we've aligned on this issue that the more you do these things that are on the screen in front of you, well, the easier your AI will be. and and most AI projects that I see, they don't have these these concepts in there. They're they're like, "Wow, I got this great way of optimizing X, you know, and it's very algorithmic centric uh in order to do this." And uh I've just finished writing a section for my online uh data preparation class where the admonition is don't start with the algorithm, start with the data, right? figure out what your data is going to tell you before you go to say try to square peg round hole the thing into a you know some sort of other relationship in the algorithm for sure >> that dead I was gonna say [laughter] >> we have another question in here that I uh love as well um so how should organizations rethink reference data governance when agents rely on it for autonomous decisionmaking So what do you think Peter? Do you think do you think we need to rethink >> Well, they just articulated exactly what the problem is, right? If you have your governance set up so that you can't tell Czechoslovakia from, you know, another country uh in there, none of what you're trying to do is going to be trustworthy uh to do this. So yes, absolutely. I think that more we can set this up, there's a a a complete side issue on this. The two state post code uh was done as a comedy routine by a guy named Gary Gullman that's on YouTube. It's hilarious. My wife called me one day and said, "They're talking about data on t in a comedy stand-up routine. You got to come see this." But it's true. If we don't get those right, right, none of the rest of this stuff is going to be usable in here. So, absolutely, the governance around this thing, this is where it's going to go read this. If it goes and reads everything as being in Great Britain, it won't understand what Great Britain is and you won't be able to have a total sales adding up to 100% or whatever the the the piece is. You must have seen examples of this incorrectly done. >> The the words I use for this, Peter, is failure at scale. >> Oh, wonderful. Yes, we can move fast. >> Like in the before times, if we had a reference data problem, somebody would look at a report and go, "Oh, that's wrong." and go fix it. And now if you're automating a bunch of stuff with AI, it fails everywhere simultaneously almost like it just it it fails everything. Like it just has no hope to begin with. Um like going back to how the questioner worded this uh do we need to rethink reference data governance? Sort of. >> My first gut reaction would be no because this is still reference data governance. However, I had the epiphany not too long ago, Peter, and and I'd love your thoughts on this. We write and produce content differently for AI to learn from than we do for humans to read it. >> Which means now you're diverging in two two different pathways. >> So how do we present reference data for AI to learn from it? >> If it should be done, you have to make sure that your AI gets contextually with the the concept of you know reference leveraging meta. And if it's they got sorry um master uh then we shouldn't even say master. remember supposed to say primary uh all the way through but uh uh if it doesn't understand that leverage piece I was going to ask you on this on your agentics component are you able to replace case tool functionality yet with the AI >> I've had students playing with it and it's kind of is but I don't trust it as much as I >> yeah that's a technical term for that by the way >> I want to go back and and take one of these case tools and augment it in a way that they haven't thought about yet because every time I see what they're doing to it. It just drives me crazy. >> It's it it's not very consistently good. And it it's depends so much on the model that you use, right? Your mileage may vary. >> Yes. >> All right. We got some other MDM questions out there. Did we answer the last one? >> Do you see Agentic AI changing the way organizations must structure their data management operating models? >> Yes. and to everybody's benefit by putting them all together and making them all machine machine readable so that we know what the agents are going after as well as what we're going after and it's all the same golden source. Everything will be better at that point. The challenge is in order to show that it's better, you're going to have to connect it to business results and the business people are going to have to understand that if they do X, they'll have more of why, whatever it is, and prove that to them so that they can get there. You'll you'll end up with a very Why is it working? Well, because we did all this work, right? We harmonized. We made sure that these things were accessible. We took the stuff that they shouldn't be getting accessed. You remember the ones that broke out, Mark? You know why they broke out? The AIS? >> Well, like the the the chatbot AI, the LLM. >> Yeah. They left they left the instructions in the sandbox. >> Yeah. Yeah. Yeah. Oh my gosh. Yeah. >> Only once, right? >> Yeah. Only once. Um what what I like about how this question is worded um because like one of my favorite operating models that I've had the most success with in a wide variety of organizations is a flavor of a federated model. >> And when people ask me what is a federated model really? I I just love talking about Star Trek. So we've got the United Federation of Planets and they have their prime directives, but then you've got the Klingon Empires got to fit in there, right? And so the Klingan empire has its own culture and it has to adhere to the prime directives. So how is it doing that? Uh but when we think of agentic AI, if we've got a bunch of semi-autonomous agents with human on the loop or human out of the loop even, then how are they adhering to the prime directives? It'd be like if you tried to incorporate the Sylons from Battlestar Galactica into >> into the United Federation of Planets, right? Um, >> you lose me at Battle Star. So, unfortunately, that's how old I am. >> Or android, Commander Data. And >> I was with you on the on the Star Trek stuff. Absolutely. >> Yeah. To the Prime Directives and and how do we build those guard rails? And for Aentic stuff when it can just spin up and run around and and do all sorts of crazy whatever, how are we controlling that? I I read a book uh by Dale Martin not too long ago ago called Bears and Mosquitoes. And so he talks about bears as being analogous to traditional data management and traditional data governance. Um you go to a campsite, you see a warning for bears, you do bear safe things with your food, you put it in your car, there's fences to keep the bears out. uh there's protocol to be safe around bears so you can understand the risks of bears and protect yourself uh from bears. It makes [clears throat] a lot of sense. But then he talks about agents and like the first implementations of agents, >> they still kind of fell into those guard rails. But as we kind of do more and more with Aentic AI and software that we buy has agents running in the background that does something. Maybe you're using everybody's favorite uh customer relationship management system and it's taking in email and moving a customer through the funnel in an automated way. It's spinning up an agent to draft an email to the customer to to introduce them to some products. the customer replies and an agent reads that reply and works with it and then uh uh the salesperson is is finally the human's finally in the loop at that point. The customer wouldn't have even known that they're talking to an AI. So what is the operating model for that? Like you've got all of these things that just ran um and and uh and that's where he has the mosquitoes uh analogy. So the mosquitoes have gotten into your tent. They're already there. >> Uh so you can have all the mosquito netting and bug spray that you want, but is it protecting you from the mosquitoes? So the nature of how we govern, manage, create, and let agents run um uh requires us to think about how do we federate them into our operating model. And I think that's >> scanning my brain for a picture to put this behind you. It's a I'm with you on this. I hope everybody else is. >> It's a great book, by the way. Like I I absolutely adored it. It it it's a perfect analogy for the the struggles that we face in managing AI versus managing data and where the intersection are between those. So, um it's it's a book I recommend for sure. >> Interested. That's one that's in the link in the chat. >> Yeah. Um, I lost track of chat because there's so many messages in there. Um, does anybody have any other questions >> uh to add in? >> We'veated thoroughly >> chat or Q&A. >> Going going. Well, Peter, do you have any Oh, for companies who are interested in developing taxonomy and ontology foundations. Oh, I love you already for asking this question. Uh, for knowledge graphing at the enterprise level, do the same reference, uh, architecture practices apply to terms similar to how reference data is used? Um, so when I was in higher ed, um, one of the questions that I'd run into all the time, you'll love this too, Peter, how many students do you do we have? I'm like, well, how many students do you want? know what a student is defined as, right? >> Could be anywhere between 500 and 30,000. >> I say I'm not even going to talk to you about it unless you got a half hour to let me sit down and go through all the variations of it, >> all the different types of students. Uh so like and when we're we're dealing with knowledge graphing and using a taxonomy to categorize our data, especially for master data, we need to understand those things, right? And it's all powered by reference data, which is why the DMBach kind of combines reference and master data into the same chapter because they're they're so interrelated. >> Like the taxonomy that defines people, places, and things, especially products, too, Peter. >> I mean, it's powered off of reference data. And the taxonomy is is reference data. But our ability to master and survive data >> is uh >> I go there and say there's all these reference models though that we can get a hold of pretty easily >> and and use those as starting places for goodness sakes. I mean it's not going to be perfect but it's better than building from scratch. >> Yeah. And as we get into knowledge graphing that's basically what helps the AI learn in a way that's not ridiculous >> because your AI can learn off of unstructured whatever or your current structure. But if you're on the cloud, uh say goodbye to your your your compute budget. Uh but knowledge graphing really puts things next to each other and helps the AI understand the context as it's training. And that's why everybody's talking about knowledge graphing right now. >> Yep. Finally. Some days it takes feels like it just takes forever to get people's eyes open sometimes. >> All right. Well, guys, thanks very much. Uh hope this has been helpful. Mark, pleasure as always. >> Peter, do you have any uh closing thoughts before I hit the end webinar button? >> Again, it enables what's going on. So, we look at this as you can use AI to build your MDM and reference data better. And that if you do that, your AI is going to [clears throat] be better on top of that as well. Win-win. Grab it while you can. Thanks, Mark. >> Wonderfully said. Have a wonderful day, everybody, and we'll see you next time. >> Cheers.