Submind YouTube summaries
Thumbnail for Week 7 - Lecture 32 : Remaining Useful Life (RUL) Estimation Methodologies

Week 7 - Lecture 32 : Remaining Useful Life (RUL) Estimation Methodologies

Watch on YouTube

Video summary

The lecture introduces the fundamental distinction between passive and active systems within the context of fault diagnostics, explaining how each requires specific approaches to understanding failure mechanisms. Passive systems, such as electronic components or civil structures like dams, typically degrade due to material-level issues like humidity exposure or corrosion over time, whereas active systems involving relative motion, such as pumps or compressors, consume their operational life through wear and erosion. The core argument presented is that diagnostics cannot function in isolation; they must be integrated with corrective mechanisms and surveillance programs that monitor both normal signatures and abnormal deviations. This holistic view ensures that when a system fails, there is an established procedure to address the issue, whether it involves online sensor data analysis or offline testing protocols designed to catch early signs of degradation before catastrophic failure occurs. To effectively manage these risks, the lecture emphasizes the necessity of a structured framework that prioritizes diagnostic requirements based on safety significance and reliability impact, often utilizing the Pareto principle where addressing the top 20% of critical components mitigates 80% of potential risks. Various analytical methodologies are discussed as essential tools in this ecosystem, including Failure Mode Effect and Diagnostic Analysis (FMECA), Fault Tree Analysis, Event Tree Analysis, and Fishbone diagrams for root cause identification. The presentation highlights that while traditional methods like control room monitoring and process parameter surveillance have historically been the primary means of detection, modern approaches increasingly integrate machine learning and deep learning models to interpret complex patterns in sensor data. These advanced techniques allow engineers to distinguish between random failures and systematic issues, ensuring that quality assurance programs remain robust against common cause failures that could compromise redundant safety systems. A significant portion of the discussion is dedicated to the practical application of FMEA charts specifically adapted for diagnostics, which serve as a procedural guide for documenting system failures, their causes, and consequences. The lecture illustrates how an FMEA chart captures details such as the mode of failure (e.g., a gate failing to open), the direct cause (such as a control failure or relay malfunction), and the severity level based on safety implications. It is argued that identifying a direct cause like a failed relay is only the first step; a thorough investigation must delve into the physics of failure to determine if the root cause was aging, solder joint degradation, electromagnetic interference, or human error. This deep dive is crucial for developing effective corrective actions, such as replacing components or improving environmental controls, thereby moving the plant toward a fault-free ecosystem where failures are minimized through continuous learning and updated training modules for engineers. In conclusion, the lecture underscores that a successful diagnostic strategy relies on a combination of qualitative knowledge, quantitative data, and domain expertise to create a comprehensive health management system. The evolution from manual inspections to real-time signal monitoring and finally to sophisticated machine learning models represents a shift toward "smart diagnostics" that can predict remaining useful life and prevent accidents before they happen. The session ends by outlining a chronological flowchart for implementing these techniques, starting with incident definition and data collection, moving through the selection of appropriate diagnostic tools from an ensemble toolbox, and culminating in pattern recognition for prognostics. Future lectures are promised to cover deeper aspects of root cause analysis and remaining methodologies, reinforcing the idea that diagnostics is not just a technical task but a cultural shift involving safety policies, human factor modeling, and continuous improvement in plant operations.
Read the full video transcript
[music] >> So, uh um after having a a nice background of diagnostics and its relation to the prognostics, uh now we will discuss in this lecture as well as next lecture, um the fault diagnostics uh tools, methods, their features. So, um part A uh is uh the scope of this lecture. Um let us see uh first we'll see that uh you know uh it is pretty important to understand uh requirement of system approach uh to fault diagnosis. Require What a system requires in the context of fault diagnosis. So, one thing you'll uh understand from here uh that uh systems uh can have two fundamental characteristic. One, a passive system and other one is a active system. Passive system could be like electronics by and large. And um then um for mechanical system, it could be uh for piping uh and then it could be uh joint gaskets, and all. So, they are But, then they require diagnostics. How they fail uh and how when they fail, uh how to correct it, uh those kind of So, diagnostics doesn't up lane uh work in isolation. We have to have a corrective mechanism or written procedure to implement the diagnostic. Then we take the second category, active system. Diesel generator. Then we have uh centrifugal pump. Okay, we have a compressor. So, these are active systems. That means there is a relative motion between the two parts. So, one thing which is pretty clear about these two mode, passive systems, they fail less compared to active system. Why? Because there is a relative motion between two mechanical parts or there is a mechanical motion. Pump is operating, so bearing is rotating. So, there is a relative motion and then they consume the life in terms of wear in terms of erosion. So, so life consumption thing approach of physics of failure mode that apply over here also. If we have a electronic system, they degrade at material level, okay? And their mating part thing like in electronics, let us see metallization. Metallization degraded by either humidity or temperature. Of course, there is a passive layer which is protecting this thing from humidity and the humidity, but temperature is something which is built in there. So, high temperature, so that's why they are tested it and diagnostics is done at high temperature. How the how the this passive layer will fail and all. And then So, once we know that there are areas like for a civil structure even for civil structure dam for dam also prognosis and diagnosis is required because how how the dam is going to fail or the crack used to develop, what kind of crack will develop if there were in between metal components embedded into the concrete, how their corrosion will occur and it will fail. Before we commission these systems, we should have a complete idea and accordingly we should have provisions also, you know, online program online or offline testing. So, these thing surveillance test is that thing which is done for even passive components also a sensor is put there and then whether it is in normal signature you have those signature with you it is compared with the abnormal the moment any abnormal signature is found then investigation starts actually now for diagnosis techniques there are many for diagnosis techniques which are there but in this lecture we will we will be discussing about you know FMEA failure mode effect analysis there's a family in that family FMECA FMM ECA FME only FMEA also and then we have failure mode effect and diagnostic analysis so this is a special thing which is this is which is very relevant to this lecture and this will be covered root cause analysis is a very elaborate technique and it will be dealt with in a separate lecture and then we have a nails base approach it is basically a Bayesian methodology where we have prior posterior and you know evidence and based on that we build our diagnostic module so what are the requirements of many of them have been discussed and so if I'm talking at the system level I should be able to identify and prioritize my diagnostic requirements because and that will be I'll be able to do when there is a quantification of this which system failed more which system failed less which system is set safety significant which system has a relevance to reliability and that will tell us which is more important and so also a prioritized list will be formed okay Um, then machine learning approach that we have to discuss because if you have to draw inference, if something has gone wrong, what is the inference? That means what action First, what is what is gone wrong and what action to follow? And then the scope of the analysis, how detailed you want to do it. Again, it will come from identification prioritization. Only important things we have to do. And remaining things we can take it to the corrective mechanism as and when they precipitate actually. And for doing all these things, we should have a quality assurance program. It's a key to having a good diagnostics framework at the system level actually, you know? And where if you look at it, the diagnostics are basically in the plant. Even even though machine learning approaches are not there, they are there as part of surveillance, area monitoring, okay? And then control room monitoring of the process parameters. Sometimes what is wrong with the equipment, it is reflected in the process parameter. Okay? So, like suppose if a the pump is not well, it will reflect into flow and pressure. So, that sign signature alone, even if you don't have a diagnostic provisions there, will be informing control room. And then safety should be overriding metrics. So, for diagnostics also, what is safety or what are the risk components of the world, they will precipitate on top. And we know that. If we cover some higher priority 20% component, we'll be addressing 80% risk actually. And we we'll have one method on this actually in this lecture itself. And then human error root cause analysis, you know? Sometimes even humans become the subject of the discussion or diagnostics when any failure occurs. For example, if there was a spurious oil tank available in the plant lubrication tank. And quality assurance was not up to the so quality also is coming to the picture. Or if the spurious oil mixed with the water or somewhere because it has entered and if you make it make up that oil in the bearing using the same oil collected from the tank. So, all there is a potential for common cause failure. And it is our duty that before adding the oil or is a human to check its a color, it's a viscosity, even in the bull's-eye you can see if there is some contamination, bubble formation. You know, or some sort of a foam formation that will tell us if you ignore that then only component fail. Cause and effect. We should be able to relate. On a should be documented. Actually, let's imagine there is a document which gives about the passive component and active component for the probably this forms part of training. How a component will fail, how failure will precipitate and finally the they land into some sort of consequences. So, then if we do not have a quantified data, even a quality qualitative data should be enough. But nowadays we have this fault tree event tree and they are this kind of risk assessment or reliability assessments are performed at the plant level. So, quantified data is also available. If not, if your plant is at design stage, then you go into generic source data. At least you make beginning. And maybe maybe another couple of years or 10 years or so, your data will become dominating in your database, and you'll it will be reflecting your own plant's characteristic. Common cause failure, I would say, is the thing one has to watch because it kills the returns. Redundancy is provided to account to save the plant from single component failure. So, redundancy first let's say two out of three voting logic has been implemented. That means first component failed, nothing will happen, plant will keep operating. But if the second component fails, then plant has to be shut down. So, but then what happens if common cause occurs, all the three will fail. Or even if two fails, we have we have got into the failure criteria. So, one has to save. And online sensor data equation and analysis. The online data system should should be in line with the requirement of the plant equipment. If there is some time constant variation, data lag is there, you know, it is not coherently available for diagnostics and prognostics, then corrective action program should start. Or somewhere between the data is coming with noise. And we are not able to do diagnostics or prognostics, whatever. So, that has to be seen. And [snorts] let's understand that there is a reasonable overlap between diagnostics and prognostics because prognostics also I cannot implement though I'm tracking a parameter, but that tracking of parameter has got been indicated by diagnostics that which mechanism out of four mechanism is dominating, which one should be monitored, which one which one is a follow-up that we did not monitor. So, all those plannings are done and then only it will be called as smart diagnostics and prognostics. So, this is just some procedures I'm or technique I'll be presenting here. So, yeah, if if you would have heard about it, the diagnostics and prognostics their birth especially in engineering domain, conventional domain, it has started with the entry of the component on onto the shop floor and their operation. So, in those days advanced tools were not available. So, surveillance and monitoring, they were the two tools. It could be There were even no control room also. One has to go to the machine and see what what is going on. But then the slowly slowly a culture of online signal monitoring started and we built control room. So, information came to control room. Not all the signal, at least process signal, they came to the plant control room and later on some windows were reserved for even diagnostic prognostics also. So, this was the traditional technique, real-time surveillance and monitoring and almost like tens of years we operated our plants or maintained our plant using this technique. Then the formal methods of formal methods of diagnostics they have started. And one of them is a failure mode effect and diagnostic analysis. So, this will will be covering. It is like similar to FMEA. There is a FMEA form and then we go on making the entry for the for for individual component in the plant and ensure that that we have good coverage of all the plants where this basic plant component, whether it is a passive or active electrical electronics, their interconnections, okay? And their interface and all. So, and then fault tree and event tree. For a complex system, this method because fault tree event tree, let's take it take the case of a nuclear plant. For most of the plants, the probabilistic risk assessment level one study exists. So, that knowledge is available in the record. Documented knowledge and approved knowledge, you know? So, we can derive rules production rules from here. What happens What How top event like say diesel generator How it can fail. So, all those knowledge, they are built into the system. They are brought to the component level and how it reflects into failure of a system. And if we know those metrics, probably we can save the failures also. Fishbone diagram. Fishbone diagram is a in reliability paradigm. This is a well-known methodology, okay? Which is basically cause and effect analysis and root cause analysis. I think last but one lecture will be independently discussing root cause analysis. A rule-based approach and this one, they are either rules are developed with a qualitative knowledge or with a formal knowledge. So, in even the fault tree also, they have a cause and effect paradigm. And here rule-based approach also, we have to develop a paradigm and develop the production rule. Now, every time quantitative approach is not possible like fault tree event tree. So, sometimes fuzzy base approach fuzzy rules are used for developing and fuzzy rule-based I'll be discussing somewhere when we go into implementation or case studies and all. So, Uh, strategy is like this. Uh, discuss the techniques uh, to the extent possible and remaining techniques when we present the case study, they'll be discussed, you know. Then Pareto analysis, it is something like 80/20 rule, it applies here. And precursor analysis, precursor analysis and fault and event tree analysis. Because precursor analysis is very important tool if any failure or root cause analysis is to perform, so we have to map it on a uh, PR study. Uh, that is fault tree and event tree analysis. Then we have to see uh, which are the precursors required to Because, you know, failure at plant level doesn't happen because of one component failure. There are uh, there are sequence of component except common cause failure. Common cause failure in one shot it can fail all the component, but uh, normally it is individual failures of the component. So, like it could be human action also. It could be a component failure also. So, two events. Uh, so and two components and one human action or no human action, three component failure or two component failure. Various combination that only lead us to an accident sequence, you know. So, then uh, so having an overview of these techniques, uh, let us see uh, uh, like, you know, if I use uh, a fault diagnostics for a domain, let us say industry, then it's not that one technique will be suffice. We have to use host of techniques or the techniques which which I have shown on the previous slide. Uh, practically all of them they will come into picture because diagnostic depend on the nature of problem that you have. So, for some problem, it will be uh, Pareto method will suit. For some problem, fault tree and event tree methodology will be uh, available. We can make a cause-effect diagram, fishbone methodology. So, and particularly when we do a root cause analysis, uh, integrated technique has to be there actually, you know. So, uh, so and then if we say that, then accordingly, so look at the So, uh, the flowchart. And I think this will give us a chronological order of implementation, actually. So, incident failure definition. Before we think about a diagnostics, we should know what we are trying to do, what failure we are really concerned, what machine we are really concerned, and the failure definition is available with us or along with the background information. This is very important because this will only drive the other following steps. And then, data and information collection for whatever fault we are doing. And then, ensemble of diagnostics toolbox. You know, this is very important. Practically, whatever we are discussing, all of these things will come into the toolbox, and we will apply each one of them looking at the machine type and the power of that particular tool that we are want to employ. So, that will give us a machine learning model for this thing. And for what we'll feed here, whether it is diagnostics or prognostics, we'll feed pattern. You know, failure pattern, I would say. Or for prognostic, it will be a prediction pattern that we are to Pattern means Which are three, fourth, or nth component, nth signal that we should feed, and their pattern their vectors are there with us, which goes up, which goes down, which goes to water it. And like that, we form n number of patterns for each of the failure. So, that now we Now we we have done that. So, deep learning model is there. So, machine learning is done now. Deep learning model, that means artificial neural network comes into the picture. And then, final step is identification of failure pattern. And if we can do that, this input can go to prognosis prognostic again, or it will go definitely. If you are trying to perform, it will become a valuable source of knowledge for prognostics also. So, I I think pattern feature and then noise noise reduction, these are the tools which will be called as pre-processing of the information. And finally, we have input for prognostics and for own learning a diagnostic methods also. So, now uh What are the modes of uh diagnostics? It could be online, it could be offline, it could be in between. That means we we put our chart or recorder for a couple of hours or couple of minutes on the on the machine that we are trying to monitor and do the diagnostics. So, online, offline, and online it it is basically condition monitoring we call actually. We are monitoring the health of the machine and on perpetual basis. Okay. So, the scope of diagnostics will be unlimited prognostics. So, diagnostics tell me what is what is happening and tell me what is going to happen. That means how long I can depend on this particular machine or subject subject component. Then, capability of which we deal with is provision to support failure management program. So, till we are in the diagnostic domain, we can say failures are happening. But when we go to the prognostics, we can say failures are failures may not happen. Because we are having a prognostics and before the failure actually occurs, it is not translating into catastrophic failure. Degradation is happening and I'm monitoring and I'm giving you remaining useful life. And what are the things required tools and application software in probabilistic modeling or one software is there like if I want to perform the diagnostic using fault tree event tree, there are formal tools that are available in the open domain. And they are called PSA software which provides a ecosystem for doing complete modeling for risk assessment. Finite element analysis. Why finite element analysis? These are the tools which I am I'm referring to when I go for either diagnostics or root cause analysis because if I do I'll discuss this aspects into but these are the toolbox that are available and the same toolbox please remember should be available when we when we do the root cause analysis. And then human factor modeling without human factor consideration, we cannot say any diagnostic is completed or not. So plant simulator reliability labs PHM prognostics and health management labs you can find in industries or R&D institution and condition monitoring lab. Expertise Without expertise or domain knowledge, we cannot do a good diagnostics. Okay and a good diagnostic means less failures, less interruption, more safety. So knowledge is required in this area. Okay. And then resource What actually policy plant policy should be there for diagnostics a diagnostic document. So it will become into training also. So for any component or equipment when we are talking about, we'll talk about the diagnostic module also. Okay. And there should be a simulator which will be simulating all this training people in this area. For example a pump for pump what are the diagnostics that are required? If it fails, what happens? If its failure has to be avoided, what are the models and methods that are available in the plant? Direct or indirect signals that are available in the plant? And then, templates uh, should be there for guidance and follow-up. And training checklist should cover all these all these things. So, it becomes a plant policy, actually. And we become robust before a person goes to the field or a trainee goes to the field, uh, trained uh, engineer uh, goes to the field, like healthcare, uh, he knows the equipment, how diagnostics has to be done, and how he knows where to limit the scope, and where we have to extend the scope for elaborate testing and uh, analysis. Um, now, let us see this FMEA, failure mode effect and diagnostic analysis. This is a framework which is dedicated to diagnostics only. Okay? So, you know, FMEA is a family and I mentioned FMEA, FMEC, FMMECA, FMEA. FMEC, uh, criticality analysis and all that. So, uh, so, they're there, one family is FMEA. Uh, diagnostic analysis. Uh, therefore, procedural elements uh, include the flowchart, procedural aspect, data collection system. This we have discussed in the flowchart that we had saw in the previous lecture. The focus in this section will be on a specific aspect related to diagnostics and prognostics. For example, requirement of tools and methods will be reviewed in the context of mechanical systems. Okay? Uh, availability of diagnostic template for uh, specific category of component should be available as part of training and qualification. So, let us see this uh, FMEA chart. Uh, in fact, in literature, there are many formats that are available, but I thought uh, a format which I can consider for this lecture. So, this has been adopted from the available literature. Mostly, it's a features are same like FMEA, except that FMEDA will be have a diagnostic module over here, you know? So, let's take the case of like, you know, if I have to discuss this chart, it is better to understand with a case study, you know? So, this chart is first we have to write the system's name, analyst's name, date, and analyst who prepared it, and then approved by. So, that means it it is got that administrative flavor in into the plant. Then, motor operated gate wall emergency ejection sign a light. We discuss this emergency injection fire emergency injection system. Okay. So, it's failure. Now, failed to open. That means when it wall mode of failure was wall motor operated wall. A wall which is operated by motor. A wall could be operated by air. Air operated wall they are called. Or manual wall which are having handle they can operate. So, we are discussing here motor operated wall. And what we found, the incident was failed to open. Or failed to close. These two modes are there. Failed to open and failed to close. These are the modes which are relevant for our analysis. There could be other failure also. Leakages across the wall or any other thing. But, we are discussing on doing what we have chosen is failed to open. For injection, we require failed to open. And what happened? Potential failure effect. Injection failure. A wall which was supposed to open and provide the water inventory for fire fighting, that failed because injection failed. Okay? Severity level very high. Of course, we are dealing in a safety incident. So, and then direct cause was control failure. Wall did not open because there is a control failure happened and wall remain the motor did not start actually because it was a motor operated wall. So, control failure. Occurrence, it doesn't happen normally. So, it was a low-level occurrence. Okay? But, risk factor was three. This now 1 2 3 4 as we go up the risk factor increases. Since it was related to safety system fire water, it was high, you know? So, it was safety significant. And then diagnostics interlock malfunction. Now, this itself will take a book or physics of failure study if why the interlock failed. One way outcome is interlock malfunctions or interlock failed. But, why those interlock failed? What is interlock? If you I talk on electronics, it is a relay interlocks, relay based interlock. If it is solid state electronics, we can say transistor malfunctions. Okay? So, that means we have to go to the physics physics of failure study. And this title alone one can invest right from 2 days to 1 year why it failed and a solution has to be found so that this doesn't happen. So, it doesn't remain limited to this chart. There are for everything there is a one one supporting cases. I'll see I'll show you the next slide. And the corrective action failed relay replaced. So, simply relay was failed. But, why that relay has failed? Was it aging aging or was it some solder joint had failed there or was it that some induction effect happened or some electromagnetic waves they interfered with this operation? Something So, we have to go into the detail why it had failed. And we have to perform a complete diagnostics. So, just to understand the the work that goes on behind the scene for each column, the there is a work which is there actually. So, let's say So, if it is column number one that I have to fill up, okay? Requires backup information like type design details. The moment a failure happens or failures that we are happening, we have to discover the procurement thing, then the literature available on those failure modes and all that. Okay? Specification and my plant specific details. What is the accumulated experience? Is it happening frequently? You know? So, these are all part of the diagnostics. If it is a rare event, we may call it a random event, but there is a very thin line between random failure and a systematic failure. Okay? So, initially it is better we assume that it is a systematic failure. We don't If we don't find that we are giving disproportionate time to investigate, then we consider that it is a random failure, okay? >> [snorts] >> But we have to keep a watch. If if it repeats next time, then it is not a random failure. It is a systematic failure. A fault which was quality related issue, it was a fabrication related issue, it was human error, it was a design, it was to carry certain current, but it was carrying more current at the site of failure. So, all those things have to be investigated. Then only we can say we have we are approaching towards a fault free ecosystem. Now, item number two, mode of failure. Information on the mode of failures. It is leakage, stuck open, stuck close. Okay? So, relay if you are talking about. So, why it got stuck up? That means it remained in stuck closed condition or it remained in stuck open condition. So, is it that the contact resistance is in has increased due to corrosion or oxide layer formation? And these are deeper points actually. Unless until we do a root cause analysis, we'll not be able to tell. But we have to find out why it had happened. Then third one was the consequences should be provided like effect on component. So, any failure, whether it is remaining local, the component failure itself, or it is going to the system which is a part of the plant or the eight plant level. Plant level means the plant will be shut down due to this irritation or plant may not be shut down when it was required to be shut down. It becomes a safety case then. Okay. Based on the plant safety and reliability characters, there can be a severity level like insignificant, low, medium, low, high, very high. So, that that categorization, even if we we are not able to do it quantitatively level, we have to do it qualitatively qualitatively level and so that the highest significant failure, they are on the radar for monitoring and diagnosis. Now, column number five is direct cause of failure. Here we are talking about relay failed. Okay, we changed the relay, we started it. But that is a direct cause. What has caused it? Is it that temperature humidity in the plant which has gone high? It is a Is it a degraded component that So, all those things they are called root cause analysis has to be performed for the those things. And so that if we that failure should not recur. Occurrence level based on qualitative, quantitative, medium, high, low. So, the if you look at the literature available, Uh, either they give it a 1 3 4 that categorization, numerical integer numerical or they will say high low. So, it is like imprecise attitude. But, then if we have the data, uh, how it it failed, how much time was so it's reliability, then we have a superior information available on this. Okay. And then deeper level analysis machine learning approaches are required actually. And the uh, management action form part of it actually. So, in this week, uh, we saw the background of diagnostics. How it has evolved, you know. Um, and uh, how FMEA we were able to model it, okay. And then diagnostic techniques we have indicated in this um, but remaining techniques we'll discuss in the next lecture. And, uh, root cause analysis will be which is the most important tool we'll discuss in the next And then format that we have discussed for FMEA like other analysis also they will have their own documentation. >> [bell] [music]