Submind YouTube summaries
Thumbnail for Week 7 - Lecture 33 : Machine Learning and Neural Networks in Prognostics

Week 7 - Lecture 33 : Machine Learning and Neural Networks in Prognostics

Watch on YouTube

Video summary

The lecture focuses on advanced fault diagnostic techniques, specifically real-time surveillance and monitoring, which serve as the foundation for machine learning algorithms in prognostics. The primary objective of these methods is to enable online data analysis that can predict component failures and guide operators on necessary corrective actions to ensure safe system operation. Real-time surveillance involves continuous monitoring of process parameters and equipment health from both central control rooms and local areas, where dedicated staff listen for anomalies like unusual noises or vibrations. When a parameter deviates from its expected steady state, triggering an audio-visual alarm, the immediate goal is to bring the plant back to a normal operating or shutdown state through minor corrective actions, such as restarting equipment or replacing failed electronic cards. To manage complex systems where failures are inevitable but must be mitigated, the lecture introduces fault tree analysis and precursor analysis, which map events onto probabilistic risk assessment models to identify unsafe states. A critical concept discussed is common cause failure, where independent redundant components fail simultaneously due to shared vulnerabilities like flooding, fire, or electrical bus failures. The instructor illustrates this with a firefighting system example, showing how two redundant pump trains can still fail together if they share a single water tank, power supply, or are affected by human error during testing. By analyzing these dependencies using fault trees, engineers can develop production rules for expert systems that distinguish between independent failures and common cause failures, allowing for more robust diagnostic logic that accounts for environmental factors and human contributions. The lecture further explores the Pareto analysis, often referred to as the 80/20 rule, which posits that 80% of domain-specific issues stem from 20% of the causes. This principle is central to fault tree inventory techniques, where ranking failures by probability allows analysts to prioritize resources on the most critical root causes. Using a fishbone diagram, various categories such as measurement errors, machine degradation (like bearing wear or corrosion), human error, and environmental conditions are mapped to their potential effects. By plotting failure data and applying Pareto analysis, organizations can identify that addressing just two major causes—such as power supply issues and bearing failures in the example provided—can resolve 80% of total injection failures, thereby optimizing maintenance strategies and reducing catastrophic risks without needing exhaustive investigation into every minor deviation. Ultimately, these diagnostic tools transform qualitative observations and quantitative data into actionable rules for machine learning algorithms, bridging the gap between traditional safety engineering and modern prognostics. While current monitoring systems have limitations, they provide high coverage in industries ranging from chemical plants to nuclear facilities, ensuring that latent failures are detected before they escalate. The integration of fault tree knowledge with deep learning algorithms allows for a deeper understanding of subsystems, moving beyond surface-level component failures to investigate supporting systems like fuel and lubrication networks. This comprehensive approach ensures that when rare but severe events occur, the system can quickly infer the cause, assess the severity, and execute the correct procedure to maintain safety and operational continuity.
Read the full video transcript
[music] So friends, we are into the third lecture of the seventh week and uh we'll continue discussing the fall diagnostic techniques. So it is a fall diagnostics part B and uh in the next lecture that is fourth lecture we'll discuss um root cause analysis uh which is uh uh very very very useful technique when we have a system which is very complicated uh you know some sorry not complicated complex systems uh and they have a high stake in terms of uh their continuous operation and they have high very high stakes that they should operate very safely. So let us discuss some more uh techniques uh of fall diagnosis because eventually what is our objective it these techniques should be uh operating as a procedure in a machine learning algorithm that means online data and then it will tell you that component if component fails how that is something like operator advisory system uh what how how we can correct it how we can avoid the failure uh or uh what action to be needed um and what you should check to perform a uh a correct uh diagnostics and draw the inference of course. So uh uh the scope of this lecture remains into uh realtime surveillance and monitoring. uh I have firsthand experience of realtime surveillance and monitoring though it may be having some limitation because these are the techniques they have born uh with um the advanced systems like uh uh process system uh and then you know um chemical systems, nuclear systems and then uh transport systems uh so and I think the society got benefited uh through uh through realtime surveillance and monitoring um um because uh they enable very high coverage in a train. You'll find an operator uh operator uh cabin where he keeps monitoring the system uh operating parameter and safety parameter. Similarly, in a plant control room of a process or chemical or nuclear industry, a person monitors the process parameters and some of the direct equipment parameters, their health and all. So, um but then there is always uh always to go or learn and improve the system. So we have this fult tree inventory uh technique uh that we uh have and then uh in fultry eventries um from a huge document these inferences are drawn and they are converted into the rules um and suppose these are quant this could be quantified qualitative both and then fishbond diagram and we have been seeing in our reliability or management literature uh these are the diagnostic tool They are also called cause and effect tools. Uh precursor analysis. It has got its origin somewhere in the probabistic risk assessment. That means whatever event has happened, we should be able to map it on a P model uh of the plant and we should see whether it is having any uh trace of contribution to the uh plant unsafe states you know uh and that will tell us which category of diagnostics or prognostic should be performed on for this. If it has a serious consequences then we have to have a serious consequence and I'll show you uh what are the level of detail that need to be investigated. Parental analysis is there in one word if somebody has to define it is a 8020 rule which is imple implemented for cause and effect type of uh uh subjects you know. So um real time surveillance um let me spend some time on this slide uh what I am trying to put it because in literature you won't find this in a formal way um being talked about in a or delivered as a lecture actually. So we can say the present mode of monitorings they are either from control room of the plant or in the local area because the equipments are not located in the control room they are located in some areas uh where quite possibly a dedicated staff has been put to monitor them. Though in control room it is surveillance that means all the information is coming and the various trips and parameters they are brought to the control room uh and continuous monitoring is going on. Now what what to do for continuous monitoring? If any parameter is changing its state from safe state to or alarm state uh there is a audio visual signal will be there and it will draw attention of the operator. Uh what is got deviated? 90% of the time it is some minor corrective action that is required to start something, stop something, close a wall, open a wall. If it is electronics um something is there then repair the card replace the card is the easiest thing which can happen. Uh if it is a electrical system uh then some alarm is there. Then you see why the alarm has come. It is some under voltage relay has failed or uh something some problem is there because our uh the actual power supply has fail fail failed and now our plant dedicated uh power supply system uh captive power diesel generator and all they are starting. So it is a situation of transient. So this is called a surveillance and if for every case the objective is same bring the uh plant to normal operating state or normal shutdown state. So this is how uh in input for uh surveillance and then as we are telling test and surveillance test and surveillance means some systems they are remaining passive mode. So unless until uh there is some situation or activation then only they will come into. So there the periodic testing so surveillance and test and the control room these are the follow. Now timely detection and assessment of any deviation. So this is the objective function of surveillance and um test and surveillance and monitoring. Monitoring is simply uh keeping a watch on the parameter and uh steady state readings that we are and we have a mental model what should be the uh reading on instrument when the plant is operating when plant is shut down or plant is in any other state. So those templates should be matching with our mental mapping you know um so time uh time uh this is a time- tested plant surveillance. If we say that in uh today's word um my plants have operated safely the credit goes to the uh the realtime surveillance and monitoring actually uh it has got some limitation also but that that's how the new techniques are coming but it has done well actually if you say that this industry has done well that means this kind of monitoring uh realtime surveillance monitoring uh should get the due credit for this and credit should be given to the designer and operator also This involves plant control room and uh and uh respective local areas staff you know. So total staff for a mega system it could be of the order of 20 25 30 35 not more than that and they are monitoring the system from control room also and from the local area. Um even if I put some uh instrumentation in the site um it is easier for a local staff to go and hear the noise and get a sense of it uh whether why the temperature is going up why the vibrations are seen today which are not there the case. So that first end information is communicated to the control room and control room uh they will talk to the uh respective agency or people whether it is a working hours or odd hours. Uh that whole procedure starts you know for corrective action program. Um and here maj major activities are monitoring and uh plant control loom and corrective action. Um the plant states may range from normal operation to deviation. Deviation could be transient also. So that means in a in a very fraction of time plant uh changes one state from another state or it could be slow deviation uh you know where we can we can track the deviation but transients it is like zero or one plant it was operating it has come to the shutdown state and alternate plant states um uh shutdown involving normal to partulated events any complex system when they are built u they are built around the assumption that some event will occur. It cannot be avoided. They are called anticipated occurrence and there is not much bigger threat. Uh but if system goes from one state to another state, there are some inherent issues. So they are called like power supply failure. It will happen once in a year or even once in 10 years. It is it will happen but it will have its own frequency. But accident like situation like a breach in the system and all they they may may they may come during the life of the plant ranging from 40 to uh 60 years or they may happen once in a lifetime. So they are called rare events but monitoring is same for both the things uh through instrumentations you know uh and then there is something additional component walkdown in the plant there are certain 5% features which if you don't visit the plant area or if the the site staff doesn't tell you you will pick up that thing and some discussion will started and okay so that what we see there what is the deviation from normal operation noise level uh you can say or uh you know some uh change patterns uh in the plant or some new thing has been some somewhere some uh spot was found weight that means somewhere some some leakages leakages were there and that can be seen only in walkdown actually otherwise it will uh you know propagate into a catastrophic failure. So walk down to the plant uh you know um instrumentation room and area there is very important testing and uh routine safety system for routine system is very important um because latent failures will show up sometimes in a transient way. So it is very important to look for latent causes which might be whispering something or which may not whisper something but it will appear. uh the staff is trained to take periodic reading. So it it is led in failure. It might manifest in terms of some process parameter reading change in those readings and we have to understand that. So u fault diagnostic techniques real time and surveillance monitoring this is very important technique. Now faulty technique I had shown you some faulty probably intuitively you will be able to see how the fault tree knowledge can be uh converted into diagnostics because what we see there uh is a top event that means a recupant or the plant uh if I say fire injection system uh then fire injection how it can happen so that is a cause effect analysis okay and we can create we can develop the rules using this fault tree analysis and uh we can develop an expert system machine learning algorithm uh for file diagnosis uh not only for fault diagnosis but also for what action to be taken. So it is called precursor or it is called anti-erent and reaction you know. So um those kind of things we have and faultries there are some feature is not only hardware component fault trees can can take an input on human factor also common cause failure also and it's not that there is no provision for handling common cause failure. The plants are built uh considering various methods of isolation of common cause failure, treatment of common cause failure and knowing a priority that this equipments can be a potential uh can see a potential common cause failure because of environmental or ecosystems. Okay, sometimes distance factor, sometimes a separation factor, sometimes an elevation factor. If I have three redundant system and if I put them in one locality. So if the flooding uh situation comes all the three will fail. So uh engineering and safety science has evolved to the extent that that one will be located at other elevation. Okay. Number one uh that is at higher elevation. So that level rise will not affect all the three. Second thing is temperature humidity affects the electronics. Then uh there may be some s symptom or signals which will show the degradation online uh or during uh if bit testing built-in test some symptoms will be available. So uh common cause failure protections are there but it can happen in spite of all those safeguard and provisions. So fault tree analysis is a very uh uh it it it embeds all the knowledge and but again it depends on the level of detail we are going see it is up to up to an analyst to assume in a fault tree what is the basic component basic component that means where fault tree ends last leave I a diesel generator can be last node for us but diesel generators also they have many supporting system like new mobiles I'll be coming into those types uh starting air system, lubile system, fuel system, control systems and you know breaker systems. So uh we'll come to those details also. But fault tree at this point you should understand that the detailing of the fault trees may require additional efforts to reach the uh subsystem supporting system and in subsystem reporting system also individual component we should be able to go. So faulty will give you an idea. Yeah, this is the thing but you have to go further in faulty we take wherever data is available and then we end there because data is not available but when when I talk about the fall diagnosis or prognosis I have to go down further that means we have to further go deeper and that is how it can be called as a truly deeper learning and deep learning algorithms are also there with us. So, so now let us see. I was talking about the firefighting system. Uh because I chose this system because everyone knows fire its consequences and what is a firefighting system. Okay. So u this firefighting system comprised of a set that is pump motor and bearing and b that is again two trends. Why we have pro provided two trends? because we want redundant equipment. If this system fails then this system should come. But again uh we have one one uh electrical power supply panel. So in most of the practical cases even power supplies are separate but suppose if they are they are having from the same bus electrical bus two different breaker and that bus itself fails. So just for the simplicity I have taken this electrical power supplies both will fail. So then what will happen? This particular thing has become part of the common cause failure that is electrical failure and leading to both the redundant devices. Otherwise they are independent actually you know pump motor bearing maybe one separation in between and then we have um one wall. So wall A train. So wall A wall. So this is pump discharge wall. Here also pump discharge wall. And what is happening is they are joining uh both of them to one header and then it is a firewater injection starts. So but then we have given only one wall. So common injection wall. So um this has to go through the study of common cause failure category. Okay. And then this water is sucked from this tank. So tank is also common though a rare event but if the there is no water in the tank for some reason whatever so then this water fire water tank also but in sometimes in the analysis you have to take some decision that structures specific system the failure probabilities are very low. So when I do a common cost fail though to in principle it is fitting into the common cost failure thing but I I can ignore that you know because no leakage and this tanks tanks are required only when the injection occurs otherwise we see uh level and these levels are communicated to the control room any drop in this level 10% and we'll automatically some I have not shown here the tank will be brought to the normal level this is the operation So I can ignore this but let us see how the fault tree will evolve. So this is our fault injection system and they are located in what I have shown here deliberately they are located in a room and it it has got a small opening because finally they are like any other industrial systems. Okay. So firefighting room is there deliberately I have shown because the room condition may affect a common cause phenomena in electrical power supply. It could be flooding also. It could be fire al even though it is firefighting system the this also has opportunity to see some fire event. So the then the it will affect the system adversely. So if we have understood this diagram uh we go to the next slide and we see how fault tree uh looks like. So now we know that there are two trend uh I have given. So firefighting system failure top event. Okay. And independent event I'm sorry it should have been EI you know. So I independent failure and C stands for common cause failure. So we have so in fault tree how it happens this event will happen either this alone or this alone happens. So they have a path. If this acting this event is reality then the path goes up. If cause event happened because of this down the line phenomena it will here independent failure is um um like two trains are there independent train A and train B you saw in the previous diagram both the train failure A and miss both the 10 train failure can only go to the top in orgate any one happening input happening here it it goes to the top it activates this node it is called uh intermediate node. Okay. So and then finally from intermediate node it can it can but here I have given the redundancy uh in the system the way in real time we had the redundancy it has been shown over here. Now end get mean miss train A and train B should fail to activate this node and if this node is activated then we have over here. So now common cause failures will be what? Water, no water in the tank. Okay. So I have considered here no water in the tank, injection wall failure, electricity failure and human factor. Human factor means by mistake a common wall is left closed. It was not opened after testing. So it rem and injection failure occurs. uh somebody uh a team worked on the uh electrical breaker and they while testing you have to isolate the breaker. They did not remove the isolation. So human factor electricity even if it is available okay and if it fails then also it can lead to injection walls failure. uh that is in final wall failure can also lead to u you know due to some component failure or due to human error or uh you know so because it is single wall that should work if I want to remove this probability then I'll put two walls in parallel injection wall so if both the walls open if one wall open our job will be done but it depends on this design and finally we have to do optimization keeping in view the whole perspective you here the gates are orgate uh this is intermediate event uh or gate uh then end gate okay so I should write orgate here uh transfer gate that means you can transfer this input and you develop it on some other page so independent ta and t b has been developed on the next page actually so what happens independent ta pump a failure motor failure wall failure and bearing failure and pump failure means what It could be pump a failure. It is le leakage. It is a mechanical seal failure. Mechanical seal is something which doesn't allow water leakages from the pump. Pump shaft is rotating and um the socket is having a stationary. Now leakage will occur but mechanical seal type of components are there which avoid leakage of water from there. Like you might have seen generally when pumps are provided with uh gaskets and all that some leakage occur but when for certain things mechanical seal is so rotating uh shaft uh will be rubbing across the stationary shaft and provide the ceiling. This sign indicates undeveloped events because uh motor can have its own failure but I have not developed it. And then uh this same faulty is applicable for B also sort of and we got the independent DB. So if this pump fails since it is argate here also argate and top also faulty it is orate it will pump a failure itself will activate the thing but again we have end gate there. So even if this TA fails completely uh it will not resto both the things have to fail both the uh TA TA and TB has to fail uh TA and TB has to fail uh to uh give the final output. So TA alone u although I have written TC it is a TI. So uh so here TI will be activated only when TA and TB both the fails. Okay. And then if that happens then output goes to the so that means fire system failure. If one fails then output will not go on top that condition is end condition. But here one of the event happens it goes to the top. So this is how it can be converted into knowledge. So uh looking at this how I'll develop the rule. So I'll develop the first rule fire failure fire system failure uh independent C failure independent CCF failure. Okay. And how independency failure this part C failure occur? This will occur as uh rule uh independent A and independent B gets activated. So the second rule has come into picture uh CCFC failure is uh water or injection or electricity or human failure. So like that we write production rules and develop an expert system. This is simple logic that I'm inventing when you do at the plant level thousands of the rules are there and there is a inference mechanism there is a data because sometimes this may may not be qualitative input a signal a process parameter has to touch certain value and that is some uh real number okay um so if that is then failure so for all the component what we have here we have put here there is a failure definition and that definition becomes very important when we develop the rules because those rules are activated when uh anything goes wrong or any deviation occurs. So we have seen the fault tree and how we can develop the production rules for the diagnostics. Now inventory we have brought whatever injection failures are there finally they'll become part of this you know and then we'll know safe unsafe. So our criteria or our severity will be dictated. If it is safe uh be happy nothing to be done it can remain down but whatever plant requires uh you have time but if it is unsafe category you have to take immediate action. So that risk or severity things are brought in uh to uh diagnostics uh through uh inventory methodology. Okay. And here you get the probability also there is a quantification also. So these rules can take high priority when we lend them into this kind of situ injection failure you know and uh the here also it is injection failure here also it is and they are all unsafe. This is fishbone diagram u you have a cause and effect you know. So cost could be some fundamental metrics have been given. It could be contributed by measurement. It could be by men. Actually we should say human because men and women both are contributing to the uh this present ecosystems equally. Environment environment makes the whole difference uh if they are not specific and it will be different. So what can be the things? We'll see through one case study. And then machine uh has its own contribution methods. And then we have to connect the dot over here. How it becomes a reality? One machine something went wrong. Uh it was man was there and it was a environmental and material related concept that led to the undesired. But we will in a in a diagram we'll have all the three probabilities over there actually. So let's see now how how practically a fishbone diagram uh looks like you know so cause causes sorry u the causes we have uh here um on this line measurement calibration induces some problem if is not correct parameters or measurement uh creates some problem I'm just giving some example you know how the dots can be connected And then when we talk about the machine lubrication, bearing electrical electric breaker uh these can contribute to the um these are the causes that are there which can happen men training human error institutional failure. In fact this can come here or this can come here also. So um then environment is now plant policy quality and humidity. Many times quality quality related issues uh they can uh become part of the injection failure. So this applic diagram we have drawn for injection failure uh erosion aging corrosion these are routine things. So some it can this and then uh methods testing maintenance surveillance that can lead to so you can say that it can travel from this end what are the problems and finally how it it can define the injection failure here. Okay. So it has got all the attribute fishbone diagram could be complicated also. There could be one or two more um you can say header elements can be there depending on the domain that we are operating. If it is in defense switchbound diagram will be different. If it is in nuclear defense fishbone diagram will be how radiation can affect my electronic system performance that also can be seen actually. Um chemical industry they can have their own causes and failures and all. So it depends. So domain specific aspects have to be this is just generic generic injection firefighting system is a very generic example and it is better to understand from generic example. Now paralysis the the 80 8020 rule is uh is the thing that means uh 80% of domain specific issues can be traced to 20% causes. If we can find out or correct 20% causes 80% of the follow-up problems, it is a very powerful statement. Okay. Uh and parad analysis has this at the center actually. Uh it it it also provides ranking. So if I say 80% or 20%. So I should be able to rank them actually you know. So 80% failure and 20. So I should know what is important is I should know 20% causes to address 20%. So define the problem uh core parrot of philosophy 80/20 rule uh analysis simulation and bin of the causes of any identical component okay and diagnostics for failure. So this completes here and that's how per analysis perform. Okay. So output of the par uh so let's say we talk we are talking about injection failure I everywhere I'm taking injection failure as the so um power supply was the cause you can see that power supply was common probably this analysis if you do we will have separate power supply for both because there it is ranking very high seven events are there of power supply failure and bearing failure this is second five failures are there so that means uh and then they they contri contribute to this is 100% failure and they once they touch so easily they touch uh here so this this has to be reduced and this is very important information we get and quality and surveillance their events are less and um we can improvise we can be we can be but then here in this case we have to do a uh level three root cause analysis but here we can do a simply level one root cost analysis to and what is level one, level two, level three that we'll discuss at present you understand it requires very high attention and uh uh the root cause analysis is a resource consuming option. So uh but we have to apply here because the number of failures and they contribute. So let us say how we are meeting the 80/20 rule. Uh first two causes total 12 uh 12 out of total 15 causes. 12 causes this these two causes uh 7 + 5 it makes 20 12 out of 15 causes and this that means they address 80% of failure and need immediate action. The two quality and surveillance can be fixed. So this will happen only when this parto chart will be happen when when I have data with me and I plot them I rank them I see how they are there comparing and contributing to 80% of the failures actually you know so this is parto analysis and uh we um what we have discussed parto and other root cause analysis techniques and now Fourth lecture is on root cause analysis. [music] >> [music]