Submind YouTube summaries
Thumbnail for Week 10 - Lecture 49 : Identification and Prioritization for Implementing PHM

Week 10 - Lecture 49 : Identification and Prioritization for Implementing PHM

Watch on YouTube

Video summary

The lecture focuses on the critical processes of identification and prioritization within Plant Health Management (PHM) for operational plants, distinguishing this real-time scenario from the comprehensive Probabilistic Risk Assessment (PRA) conducted during the design phase. The core argument is that while safety cannot be quantified as a simple mathematical parameter, risk can be modeled using established formalisms to determine which components contribute most significantly to core damage frequency. This approach allows operators to create a priority list addressing the top 10 to 30 components based on available resources, thereby maximizing improvements in overall plant reliability and risk reduction. The identification process is not solely mathematical but integrates qualitative insights from expert judgment, operational crew feedback, and supervision reports, ensuring that even non-failed components are maintained proactively if they show signs of latent degradation or pose a threat to system integrity. In the context of operation and maintenance, prioritization is governed by performance evaluation indicators where safety systems take precedence over operational systems, which in turn are managed alongside quality assurance protocols. The lecture highlights that traditional approaches often rely on subjective qualitative analysis, such as plant area reporting and human factor assessments, lacking a systematic mathematical basis. To address this, the presentation introduces an integrated risk model that combines technological hazards, industrial safety issues like fire or flooding, security risks, delivery risks, and liability concerns into a single framework. This holistic view is further enhanced by machine learning algorithms that weigh different risk factors, allowing for a more nuanced understanding of how various threats interact to influence overall plant safety and operational continuity. The lecture concludes by detailing specific mathematical metrics used to quantify component importance, moving beyond qualitative guesses to precise calculations. Key measures discussed include Fussell-Vesely importance, which assesses the fraction of system unavailability attributable to a specific component, and Risk Reduction Worth (RRW), which calculates the benefit of making a component perfectly reliable. Additionally, Risk Achievement Worth is explained as a metric that determines how much risk increases if a component fails completely, helping operators identify critical targets for maintenance. These mathematical models directly inform decisions regarding surveillance test intervals and allowable outage times, ensuring that maintenance schedules are optimized to prevent latent failures while balancing the need for continuous plant operation against acceptable levels of unavailability.
Read the full video transcript
[music] >> So, we are on to our fourth lecture of 10th week. Identification and prioritization for implementing PHM. Um here we'll be discussing some mathematical well-established uh formalism, uh how to identify and prioritize the component. And since uh I told you that our matrix is uh risk matrix. Because safety is not a not a mathematical parameter. So, uh we cannot do uh prioritization based on safety matrix. So, we have to take a risk matrix. Um uh and then we have to identify and prioritize it. In previous lecture we you saw that uh how we brought uh um this uh system level unavailabilities. And then we were we explored uh which are the accident sequences. Accident sequences are made up of what? Components. Okay? So, component one, component two, component three fails, then the top event, core damage, occurs. Okay? So, uh now we'll be talking about uh the ex- accident sequence level, that is cut set level. Which component has got its sen- sensitivity to core damage? Core damage means uh um the damage state of and that is that adds to risk uh element of the uh so, risk element of the uh various uh accident sequences. So, uh before we we do that, uh let us see role of identification and uh prioritization operation maintenance. Because we are talking about the plant operation. Please remember that. In fact, this thing is done in design stage also, but that time we we do complete PRA and all that. But the plant stage is postulated or plant operation is that is going to happen in that way. But here it is a real-time scenario plant is operating. So, whatever insights that are available and mathematical model modeling that are available based on that the prioritization identification is done and identification and prioritization is done to understand the risk importance of each and every component so that as part of PHM they can be we can make a priority list. Okay? And first 10 20 30 components, whatever resources available, if we address and we'll see what kind of gain we have in terms of improvement in risk and reliability. >> [snorts] >> Okay, then so so you can see here in operation maintenance the the identification identification and prioritization, these are the two inherent component of even when the plant is operating, if you were to take a maintenance schedule, we were to identify components and and prioritize them which component comes so this happens because of plant configuration itself. You know, there are safety system and there are operation system. So, for these systems the risk significant component comes first and then plant operating component and then some insights from expert insight from operation maintenance crew and then annual program is drawn and the systematically all the components and one more important thing is how frequently or what should be that maintenance interval? Maintenance interval. That has to be decided. Earlier I would tell you which are the conventional procedures and how now it is done mathematically and rather the mathematics is open for their use for for for choosing the maintenance intervals and all that that will be there but so these are as I told you they are inherent component of operation maintenance. And it is governed by performance evaluation. So safety and perform reliability are the two indicators. Reliability for process system and safety for standby system. And then uh quality assurance plays role lot of role in that. The first part is the inspection interval has to be decided. So QA decides what should be that interval. Okay. Uh by and large safety system takes priority and then of course operational system. Operational system because because failure of operational system results into initiating event and the initiating events when it gets valid and then system response and you know the event tree and all how they form part of the core damage frequency. So quality assurance is very important thing. Quality assurance part uh you know part of all the activities which are happening in the plant but here I'm talking to you in terms of identification prioritization. And then super supervision and reporting. Many component may not may not appear based on your mathematical mathematics but in supervision when you visit the areas of the plant and you find oh suddenly this item now it should go into maintenance next time. Why because there can be some issues with that and it is not falling in the safety category it will be not falling in but it is important to do that. So, uh you can say a qualitative thinking goes into many of the items which have to be taken into The component is not failed, but expert decision is made. Let us do some maintenance because the the component we had the spare part we had procured in other system, they have been failing more. So, as a precautionary measure, let us replace here also. So, mathematics is one thing, but there are so many domain related aspects that come into the picture. And so, so supervision and reporting has got some some water is dripping from somewhere. You find your work done and you feel it should not get aggravated and that takes the priority then. Okay, structural priority. Then then we have the documentation and development updating. All these activities on the same updating of the document when you do based on your maintenance inside or operation inside, they form up operation and maintenance management. And finally, if the issue is caught and is corrected, then corrective action to be implemented. Okay? So, issue identified, okay, and prioritized, but corrective action program has to be it could corrective action program could be the action to be taken when the when the as part of the if I have to talk in terms of PHM, you know, the whether this component should go in the in the maintenance or repair now or replacement, repair, replacement or overall. Time being, let us buy some time. Okay? So, those kind of thing corrective action program. So, it will be ad hoc program and then it will go for a permanent solution. And feedback. Feedback is very important whether it is a PHM because the whole PHM can look or different quite different compared to when you started it actually. And that happens because of feedback which you are getting from the performance of the PHM system itself. Okay? So, as I mentioned, PHM system should have its own risk and reliability criteria, you know? And its own performance actually. So, one is plant performance, then its own performance also should go into the picture. So, this is what when when I look at the plant in a in a from the PHM implementation perspective. Now, before we go for PHM, what are what are the routine or traditional approaches? Because there is a lot to learn from routine approaches. One, why? Because some of the things will continue the way it was there. Number two, what improvement we can we can do you know, in this itself so that it not matches with the PHM. Because you know, PHM is a resource intensive activity. So, some balance has to be made because money has to be spent. And then certain thing cannot be compromised like safety, and certain thing cannot cannot be can be managed by to doing some small action. So, a qualitative analysis is required. Performance [snorts] signature analysis in control room has to These are the things which are traditionally done. Plant area reporting. Here I would I would tell something. If we had nuclear industry had demonstrated safety record for for itself, then these procedures or these are activities are none important actually. So, discarding these activities or improving upon these activities, we have to really go in a sensible way. So, so qualitative analysis was one part which was happening as part of risk and reliability program. Then performance signature analysis in control room. Plant areas reporting. Audio-visual alarm trips and all those in control room scenario, especially in the plant transients. Plant walkdown I showed you on previous slide what kind of feedback it can generate. Condition monitoring trends, schedule-based maintenance management, abnormal symptoms of precursor analysis Precursor analysis means any failure which is coming in or any failure, what is its role in the complete accident sequence that you have. Okay, so they are called accident precursors, you know. So, any symptom which can take part into the accident, it's potential. Uh daily log reporting. These are the basic approaches for identification. Uh traditionally it has been happening. Uh be it a plant area, be it human factor improvement, be it quality improvement, all these and are maybe any modification that you need to be reflected into the plant. So, uh so now if I say uh how safety used to be prioritized, it is a performance indicators, safety and reliability, uh and performance indicators they are analyzed annually, uh level of redundancy, uh system diversity, common cause potential. This is one of the major indicator and we we have to have a special provision. In fact, uh there are plants who perform common cause failure analysis other than the safety analysis. They perform common cause failure analysis because it requires a very special attention because it kills the redundancy, diversity, and rest of the provision that are existing in the plant. So, it is required and human factor also can be one of the common cause failure. So, human factor becomes very important. Uh again I'll repeat because we are on prioritization and identification. Uh human factor has been found to be one of the major contributor to the accidents. World over this observation is there in the uh international standards books and all that. So, it requires special latent degradation. Uh you know, for safety system, uh, if there is electronic channel, if there is no actuation, and if there is no provision for monitoring it, then some failures they might remain uh, latent. And to know about those failures, uh, the the best thing is you perform surveillance testing. So, uh, let us say diesel generator system, emergency power supply system. Uh, if we test it, we'll it will reveal the latent failure. That is the only way to know that. So, these testing testings are performed. Then aging symptoms and quality levels. These are for safety system. Similarly, some of the things may be exchanging. It is not that hard and fast, but by and large, these are the thing uh, it metrics that goes for safety system uh, prioritization. Um, and then process system. Uh, reliability, failure-free operation, downtime, routine overalls and repairs, aging symptoms, and common cause failure. Uh, again, here also it comes here. So, at the end of the day, here we see we are talking about safety means safety hazard, whatever was available. There are some industrial safety issues, and there are security issues. So, if I talk about the safety, I have to use a holistic statement for safety, which takes into account these things also. So, uh, but the biggest weakness is the conventional approach to is largely qualitative in nature, and tends to be subjective because these things are monitored based on the discussion, brainstorming, and some expert opinion available there. But there is no there is no systematic rational base mathematics behind them actually. So, there is a need for improvement that of course we can say. So, traditional approach, if it has to be seen in the modern development, wherein new factors are entering. Industrial safety was part of it. So, it may not produce the the domain system hazard issue. But finally, >> [snorts] >> um, like you know, industrial safety include fire, flooding, falling from a height, impact targets and all that. These are the security. It requires attention because there is a quite a bit of overlap between safety and security. So, that is could be advantage if they are segregated and handled by the means they are required actually. But security as a issue alone, they need to be addressed in a separate analysis. Okay? So, this is and then if you go for modeling because now we are talking not in terms of safety, in terms of risk we are talking about. So, we can create a integrated statement of risk. So, it can be interpreted as integrated statement of safety. And which are the component of it? Technological shift like if I talk about the nuclear plant, it is a nuclear hazard. Okay? Industrial safety, which is common to almost all the engineering system. Then delivery safety. This I have given a special name. It is output. What What is the risk that next month I'll not be able to deliver my product, whatever may be the product. So, this is called delivery risk and then security risk. So, here we have talked about the safety risk, industrial risk, security risk. It is output and then liability. Liability is a very special class of this thing for any damage. Who will be liable to pay? Who will be liable to compensate? So, this should be integral part of the complete operation of the plant, liability risk. And if I use the machine learning learning algorithm, then I have all the risk and its weight is factor. Because weight is factor for technological or hazard risk will be much more because it can have consequences, potential for consequences beyond plant plant. But industrial safety, yes, it is important, but it will have a limited consequences in there. Similarly, delivery risk, I will not be able to deliver my committed output, but it is a financial loss. But, financial loss also perpetually is not good because if I am having good finance because of my plant operation, that goes same resources goes into improving the technological risk, industrial risk, and all that. So, it's feedback thing. So, it is very important from that point of view it has to be seen. Then there's security risk. It is a new dimension I will not say relatively new dimension and it has got its own waiting factor. Probably, it should match at waiting factor of this because security issues provide a threat and then finally, it is the liability issues. Okay. And then, if I have all this input together and then I have all this hidden layer, it could be more than one hidden layer, and I can give a statement, but my problem is having parameter for this vector. So, that means what we require is a consortium of different industry or among nuclear industry for different plant, we perform this analysis and train this neural network with these vectors where R1, R2, R3, R4, and I5 they are the components, and we produce an output here. And this output could output could be because this will have highest wattage factor. So, so, and then remaining things will accordingly. So, we can have a output spectrum here how my plant what is the safety level actually. If I have quantification of all these things, what should be my safety? This is a science that has to be developed further, but it has been proposed in my book that has to go on. So, and then okay, we have talked about it. Risk is a nothing but summation of individual risk R1, R2, R3 where RI is the wattage uh, and then likelihood into consequences. Machine learning model is available here. This is from my book on risk-based risk-conscious operation management. Okay. So, a lot of R&D has to go into it uh, to provide it integrate In In In fact, it is not exhaustive. There could be any other factor which will add on. It is basically uh, what kind of ecosystem that we are operating? Earlier we used to have only technical risk. Uh, we never talked about security risk, but do you have to talk about the security risk now? Because it is directly relevant. Okay. The Now Now Now we are talking about the quantified approach for prioritization. See, prioritization is one thing. I have prioritized the component. Okay. Top five five component. But what should be the frequency? What should be the frequency? Uh, our test interval, I would say. Surveillance test interval for the the those component. Why? Because we know uh, so here we are 80/20 criteria. 20% component will give you benefit of 80%. Very good. So, I have made the list of the component using importance measure which uh, we talked talked in on a class three power supply system and loss of offsite power scenario. Uh, now let's see if we are talking about the component, but then who will tell me what should be the frequency? The frequency of test interval or maintenance is governed by performance of the system. And that performance of the system uh, can be defined how the systems is degrading or because reliability If it is a new component, fresh component, reliability will be one. But during its operation, it will reduce. It is inherent. Okay? And this we are talking about the random uh, phenomena. Okay? So, it will go reduce. So, what we do in plant? We stop the component, we shut down the reactor, and uh, like like suppose there is a loop, cooling loop. We shut down the cooling loop. We do some maintenance maintenance or surveillance job test or we do testing. Testing reveals the latent failure, we know that. So, if any latent failure was sitting over there, you know, that should be known and for that I do shutdown, I do maintenance or test interval and maybe it may be overhaul, it will be replacement, repair, whatever and then again I started back. So, it has become new. What the point here is Initially it was a reliability was availability plant availability, you know that. Uptime upon down down time total down time. Uptime upon down time per uptime. That will give me steady state availability and reliability you know that. So, if the reliability drops, so this is a criteria. We have to decide what is the criteria. So, and then it will tell me how the testing or shutdown maintenance job should be done actually. So, one thing was I identified the component, second thing was how frequently I should do it. So, this is a a theoretic theoretical model we have done, okay? And this can be done by taking into account failure characteristic of the component. But I should know this criteria and this criteria is decided by what is acceptable level of degradation. Maybe initially we start with the conservative level and later on we go down and try to finalize it and finally we see how it is reflecting in our life. Finally we are talking about plant unavailability. So, what is the unavailability and what that with that availability what core damage frequency we are getting and that will be telling me telling us, "Okay, if the standard core damage frequency international criteria is 1 into 10 to the power of minus 6 and so this availability is meeting my criteria, so my unavailability should be I'm talking about availability. Converse of this will be unavailability. So, this unavailability we are we are talking about it actually. Okay? So, then so we have discussed here one is the prior rotation and then we have talked about what we'll do with that duration. If you have to do the testing and all that. Okay? Because finally, I should know the weakness of the component. It is a hardware system. Okay? One on one thing you can say that if I reach somewhere um RUL or here okay? Then I should because now in PHM we are monitoring the degradation. So, we know that liability already going down and if I know this trend, even this time will not will not be required. We'll I'll do plant management mode I'll go. You know, and PHM will tell tell us okay, now we have to we can start we can keep it running or whatever. So, somewhere this concept they are overlapping actually, you know? But without PHM this is this is the way you shut down and then you do management. If I'm having PHM then this particular thing will not be required. So, you have come out of this slot with availability of PHM. Okay? Now, in principle risk importance. What is a general mathematical formula is importance of a component risk importance of a component is a function of uh change in availability of the individual component and change in availability at at system level. Uh you understood this we are talking about I represent the importance of the component II or a capital I is the importance of the component I and small I is the name of the component and it is a function of D is nothing but a function of unavailability of the Ith component and U stands for unavailability of at the system level. Okay. Uh I explained this terms over here. So, you you must know that because by in in a in a academic domain, we talk a lot lot about the reliability and all that. So, that means reliability means successful domain. Probability that component will operate by a given period of time under given condition for a certain time failure-free operation of the So, that that we say. So, it is in success domain. So, um sometimes you you have a mathematical expression you will find in literature based on the reliability being more important of the that we talk about. But in in risk analysis, we always deal in failure domain. So, here we have some risk important measure and then objective and scope based on the PRA. We have three risk important measure and we'll see their mathematical model and all. So, risk important measure Fussell-Vesely important, then risk reduction worth important, risk achievement worth important, and inspection important measure. This will not be covering these three. If you cover, probably by default, you can you can study yourself on this because of positive the time I have to skip this, but keep it there for the your self-study. So, Fussell-Vesely important. What is the important? Um you we understood the term IIT. importance of the ith component. Now, Fussell-Vesely importance is a superscript that we have over here. And then is nothing but So, normally these things they are they talk in terms of ratio, in terms of this thing also. So, availability of the ith component divide by system availability. So simple. You have fraction of these two. So, you will get the this important measure. Okay? But when we when we have the minimal cut set from the PRA. There we have a importance of the component in terms of the cut set. You know, so what is what is minimal cut set plant level along with the quantification frequency per year core damage frequency. The first of basically importance talks in terms of core damage frequency. So it is at the system level this thing and this is at the highest component level. IIT and IST are the frequency of cut set. We have put the instead of unavailability we have put frequency of the cut set plant level for which importance measure is being actuated and system level summation of the this thing. So we have now there are two more importance level risk reduction worth and risk achievement worth. Here we talk about the risk reduction worth and here it is manifest in two forms RRW risk reduction worth ratio mode and in difference mode. Analysts they choose different models for different So if I talk in terms of risk reduction worth ratio level then it is core damage frequency for the plant and CDF U that means unavailability for the component is made zero. What what does it mean? If I make unavailability is equal to zero that means the component is availability is uh is available. Okay? So so and that effect given that this one. So the core damage frequency given that the subject component it's unavailability is zero. So that means it is available completely available. Okay? So we have this core damage frequency at plant level and then this one at to go. So, see see F P CDF is the frequency of the plant CDF. And CF BDF due to the CDF frequently obtained either setting the element I with zero, that is component is highly available or reliable. I should use the word highly available. And in difference term, we have the same term, but we have this uh difference criteria. Similarly, for risk achievement worth, we set the component unavailability to one. That means component is unavailable. Completely unavailable. The procedure involves setting the element to availability one. So, UIT There it was zero, that is here it is one, okay? And this importance measure is is ratio of plant and then individual components. Sorry. Individual components in the plant level UI is equal to one and then plant UT normal UT. Difference model is again same here. We can see that the effect of risk increase by by setting the system level availability to one will result in a risk increase. Okay? So, so at one time we reduce risk then another time by setting it to one, it is unavailable and then we see that it is increasing risk. So, risk achievement worth, you know? And it is useful when we when when we have a target. If we have a target of 1 into 10 to minus 3 per year in core damage frequency, then we can achieve by manipulating our different unavailability and setting a target for those particular SSC where the target comes from here and whatever by maintenance practice by PHM or whatever, we try to reduce. So, that means Uh, directly we got the list of the component uh, where uh, from risk achievement part we have to work more to reduce their the uh, unavailability. Okay? So, overview of week four is we have seen uh, different importance measures uh, for component and then uh, we discussed role of uh, uh, identification prioritization uh, what what was the traditional approaches? They were all qualitative in nature and significance of scheduling SSC. This was a very important point that we made out. Even if you identify importance, what how frequently you do it? Uh, there is a procedure uh, which is called surveillance test interval. Okay? For safety system. And for operating system, allowable outage time. How far you can keep it? And they sometimes form the part of risk monitors, you know? Because risk monitors are there to advise you on maintenance management. So, um, probably we don't discuss in or in the my book one of the book you can find this uh, factors. Thank you. >> [music] [music] [bell] [music]