Submind YouTube summaries
Thumbnail for Week 10 - Lecture 48 : Risk-Based Approach for System Modeling

Week 10 - Lecture 48 : Risk-Based Approach for System Modeling

Watch on YouTube

Video summary

The lecture introduces a risk-based approach for system modeling, which integrates traditional deterministic engineering knowledge with modern probabilistic methodologies to address uncertainties in complex systems. This integrated framework is designed to ensure both plant reliability and safety within a competitive global ecosystem by combining established safety cultures with advanced technologies like Integrated Maintenance Logistics (IML). The core of this approach involves understanding the hierarchical structure of a system, ranging from the overall plant down to subsystems and individual components, while accounting for interactions between human actions, machines, tools, and methods. A key paradigm shift highlighted is the role of power electronics in replacing traditional passive systems like UPS units, necessitating new modeling strategies that handle the conversion processes from AC to DC and back again to maintain uninterrupted power supply. To effectively model these complex systems, the speaker utilizes fault trees and event trees as central tools for visualizing interrelationships between components rather than analyzing them in isolation. This graphical method allows for the inclusion of critical factors often missed in standard table-based analyses, such as common cause failures, human intervention issues, and machine interface problems. For instance, in a nuclear plant scenario involving a loss of offsite power event, the model demonstrates how multiple redundant layers—such as diesel generators, batteries, and passive containment systems—interact to maintain safety functions like decay heat removal. By mapping these logical connections, engineers can diagnose potential failure points and understand how a single component failure might propagate through the system, ultimately leading to either a safe state or core damage depending on the success of subsequent safety mechanisms. The practical application of this risk-based approach involves quantifying probabilities to prioritize System Structures Components (SSCs) for Prognostics and Health Management (PHM) initiatives. Through cut-set analysis and Boolean logic, the lecture illustrates how independent failures and common cause failures are combined to calculate the overall probability of emergency power supply failure. In the provided example, the calculated risk was found to be close to regulatory limits, indicating that while the system is robust, there is still room for improvement. By summing the contributions of various accident sequences, such as loss of offsite power and loss of coolant accidents, the analysis reveals specific percentages of core damage frequency attributed to different systems. This data-driven insight allows operators to identify which systems contribute most significantly to risk—such as a system contributing 13% or a loss of coolant accident contributing 32%—and prioritize maintenance and monitoring efforts accordingly. Ultimately, the lecture concludes that a holistic, risk-conscious approach is essential for moving beyond traditional methods to achieve higher levels of safety and operational efficiency. By focusing on the most critical systems identified through probabilistic risk assessment, organizations can allocate resources effectively to prevent core damage and ensure plant availability. The process operates in a closed loop where goals are validated against both deterministic requirements like structural integrity and probabilistic targets set by regulatory bodies; if these goals are not met, the system design or operational procedures must be redefined. This systematic prioritization ensures that PHM efforts are directed toward components and systems that offer the greatest impact on overall plant safety, providing a clear roadmap for managing mega-systems where multiple agencies and complex interactions define the operational landscape.
Read the full video transcript
[music] So we discussed uh so far uh the system modeling uh or system issues uh that was very important because the because the niche of this lecture is uh that we are talking about systems approach. Um but then there is a very important component uh in systems approach is system modeling because unless until we do system modeling or first unless until we understand the system and we if we have a domain knowledge also uh they then we cannot even attempt system modeling and now we have understood the the the aspects associated with the system uh and then uh it's a different form form manifestations. Uh now let us say the system modeling system modeling may first thing will be you have to be an expert in uh that system uh design information it's operational information uh uh human involvement in that are main machine interface uh issues uh and many more uh aspect assess associated uh with the services um and then procurement event all those things without that we cannot do system modeling. Uh so let us discuss uh a riskbased approach uh for system modeling. Uh riskbased approach is a riskbased engineering as part of that I have developed a riskbased approach uh for system modeling. Essentially it is uh it is using lot of probabistic risk assessment methodology along with the uh some new paradigm that are available uh in terms of uncertainty characterization uh in terms of human factor uh handling and all that. But here the keeping in view the availability of time we'll discuss only the central character of riskbased approach and as and when it comes to uh implementing PHM uh we will explore other areas also. So complex uh system we have discussed enough now probably you know all and the key metrics of uh systems approach uh that we'll be discussing role of riskbased approach uh here uh and then identification of SSC's prioritization uh and then uh role of riskbased risk conscious approach uh actually I have written two books this book is uh riskbased engineering um and here how we can design operate the system uh with the uh risk or safety as the as the overriding factor you know and um how we can we move away from the traditional approaches or how we integrate rather I would say uh the traditional approach the best part of the traditional approach and the um advances in the u last 20 to 30 years um and then have a combination of which gives us better results uh and better insights and ensures both plant reliability because you know the times have changed now uh in this uh time of globalization competitive uh uh ecosystem the plants have to operate also okay and at the same time ensure the safety also how it can be done it can be done by what best part of what we had as a as a safety culture and best part of what the technology has enabled us including a IML and all and provide the uh integrated approach uh for system design and operations. So um the it is basically sort of a perspective that I'm talking about uh when I talk about the system I actually I'm talking about a plant that I have told you uh and then a system uh you know literally has can have a subsystem is a subp part of the plant so plant comes on top systems subsystems another subsystems and level of subsystem then the components that's how it flows. So that is a hierarchy we if we see the plant hierarchy over there then operation management ecosystem unless until we uh understand understand the uh complete ecosystem how the maintenance uh human actions um machines tools methods they interact uh you know how communication occur occurs among different agencies and then we realize and now we have uh the micro electronics and power electronics. This was not there earlier but now power electronics has been playing and it is replacing the traditional uh electrical tools and method with passive uh passive like uh UPS you know uninterrupted power supply system. Uh they are they are the uh brought in paradigm shift into how we look processing electrical power from AC to DC to AC and finally AC to DC and again AC. So it becomes uninterrupted power supply. So all these things they they have to be addressed. Okay. When we talk about a systems approach um we we are calling uh nuclear plant we are taking it as a reference because as you saw on my book also I written reference and nuclear plant. So it is easier to explain and understand the best of the systems uh and from there to draw the references. And then we have uh the objective here is uh to maintaining highest level of safety and then uh reliable oper ensuring reliable operation. This this is a broad perspective how we look at the systems. Now um let's say in in this reference plant we have considered a lot because every plant has got their own uh characterizing the name of the system the name of the event. So here we have for reference plant this nuclear plant loss of offset power that means this is one of the event though it is a anticipated occurrence uh but it it was found that it contributes to risk from 10% to uh you know 10 to 8% or 13% you know uh because uh somebody might ask how loss of offset power when loss of offset power occurs the plant is shut down but the decay heat is is production at very fractional level. Let's say let's say 1% 2% 5% it goes on and core has to be cooled. Now if the on-site uh dedicated power supplies diesel generator or any other they have to produce the power that means they have to come uh they have to start and uh meet the whatever 10% plant loads uh essential load requirement. If they do not start there is a problem. So um then then you have a this is a class 3 power supply then class 2 and class one batteries and all that. So there smooth operation uh I would sum up is required and here we have this integrated approach uh integrated which is uh something uh graphically nature that means we can see visuals and we can understand how the components are interconnected what happens if one component because you know they have used logic gate and all that. uh so we can by looking at the things we can provide a very good diagnostics how the plant operates and when it will fail or when the system will fail or when the component will fail. So uh so it's a and there uh the advantage is like there are traditional approaches like failure mode effect effect analysis failure mode effect and criticality analysis there we choose one component and we see the consequences till and last but then on the same row but here it's not that it shows the inter relationship between the component this fault treeries and all that they are very very um elegant mechanisms that enable plant representation and They show plant characteristic also they and there are some aspect which cannot be covered in table methods and they are like common cause failure, human factor in integration or main machine interface how human uh can intervene all those things are uh can be can be shown in a integrated way into the fault tree and then the event tree will show um I think we have discussed fault tree eventry so I'm directly talking about then when I provide the case study probably you'll get a better idea Now ens symbol of functions you know the when we talk of phm component wise either it is electron electronic components mechanical component but here ensemble of methods are required right from uh correct initiation to correct propagation okay then we we require damination electronics okay of course the the mode of failure is mechanical only but then those failures have to be se seen and modeled okay when we talk about The uh power electronics it is the capacitor which has got a risk component which is which has got a reliability component both IGBT integrated circuits field programmable gate arrays each one will have their own characteristic and modeling requirement that's why systems approach is required it is way challenging compared to the component specific approaches multi- agencies condition and uh and here uh if you have a systems approach or especially mega system approach lot of agencies are involved and they are taking part into operation maintenance logical things and all that and on the top of that there is a regulatory body which oversees FL operation or gives provides intervention level if required if any safety uh case is being made out and then they will uh they will have stipulation on those things. So uh executive body alone is not uh working in isolation. There are some independent and that's how the safeties uh safety uh uh is taken care of. So well uh safety risk and security as I mentioned now we have to bother about safety and security risk also. So our plant model should have all these features. So as I mentioned in riskbased approach uh why we are using riskbased approach it is allowing the proven knowledge deterministic compon know knowledge of the component [snorts] and then probabistic relatively new knowledge how to represent plant system interconnectivity and then integrate both of them. Okay. So we have best of the things in terms of safety and reliability validation with the deterministic and probabistic uh goals and criteria uh probabistic goals are defined at regulatory level or at international level whether we are able to meet them in a holist holistic way. uh uh and probably uh deterministic goals in terms of uh minimal flow requirement, minimal pressure requirement, uh minimal structural integrity requirement. So those things are validated and enhanced monitoring and serless that gives safety and availability goals and um and you ask yourself a question whether the goals are met. If it is yes uh uh job is done. If it is no no then uh then you have to redefine the goals and call it is it operates in a uh closed loop and that's how we have this integrated displacement approach and uh here um fault tree event trees they are at the core of it along with the deterministic model and methods and makes lot lot of effort goes into the common cause failure modeling because common cause is something even if you have a lot of redundancy built into the system common cause can knock off which is could be external parameter it could be internal parameter uh but it can knock off uh the complete redundancy and diversity. So simple example one seismic event and all our redundant system uh they can get adversely affected. Of course things are designed keeping in you the seismic requirements and all that. Uh but then uh but then um their fragility uh against this uh uh this seismic event um has to be tested and it should be uh shown to demonstrate that it will meet in this uh requirement. So common cause components have to be mapped on to those events. Quantified approach to and this is a quantified approach. quantification we understand better and all that and this is as I mentioned in previous slide it is based on my book one uh integrated riskbased engineering. So just for the sake of the the plant uh we you have a plant and then we have safety systems. Okay. So uh as far as the uh safety system, primary reactor shutdown system, secondary reactor shutdown system, any disturbance in the plant, plant should be brought to the safe state. And the cooling function reactor primary cooling I have shown here uh in a simple way very elegant this thing. Uh then uh secondary shutdown system uh decay heat removal system. You see what kind of layers we have in emergency core cooling system. primary secondary decay uh decay removal and then emergency core cooling system. So uh this is a pump or primary regular primary system. Then we have a decay heat removal system and I have shown two two only emergency cooling system I have not shown here. This is the fuel which produces heat hot water is goes and it goes to the heat exchanger and it operates in a closed loop. If the this pump because of loss of offside power we'll be talking in this lecture. So loss of upside power means this bigger pump almost like uh 700 kilowatt or you know uh that size this will not be operating because there's no power supply uh coming from offside power. So small pumps will operate and cater to the decay heat requirement. Plant is shut down heat production has come down but even 5% heat has to be removed and uh that again has to operate in a closed loop. I have shown a very very simplified view of just to give you an idea and when the steam is produced here it operates a turbine and the thing condens condensed water it goes to uh condenser and then finally it is uh you know fed to the uh this room this is primary coolant pump this is decay heat removal pump reactor cold water so then we have different type of power supplies class three power supply as I told you class 4 power supply means normal power supply grid power supply this pump operates but diesel generator power supply is available in shutdown. So these pumps will operate and then then we have a UPS uninterrupted power supply class one DC batteries are there and the then passive contaminate isolation system primary containment secondary containment tertiary containment probably you can visual visualize this as a part of the containment actually okay so there are two walls two lay two boxes I have shown only one primary containment only I have shown there can be two containment and then for their ventilation uh emergency exhaust system human recovery action and all that. So we'll see the loss of offsite power event modeling uh just to understand how to model a system for prognostics per prioritization you know at system level. Okay. So uh and then our we will take the example of class 3 power diesel generator. So this is a U class 4 power failures. That means class 4 means offside power feeding these buses. In the moment of loss of offside power, okay, these feeders will stop feeding. Then these breakers will open and then we have diesel generator 1, diesel generator 2. They will automatically sensing the under voltage on these buses. They will start automatically breaker closes. And C and D are emergency buses. So they will supply power to the emergency buses. Then class three loads. So these loads itself are called you can say emergency loads but it is not there is no emergency because the plant is it it's normal operation state only. Uh so they will feed our class 3 loads only uh loads you know from the both the buses. There are many loads actually and like one of the thing is I showed you the deep heat revols they should operate. Yeah. In fact the other more than class they are on class two. Why class two? Because the class two will take AC convert into DC and again will feties. So that means uh as long as batteries are available in the plant that AC power will be available and on which those pumps will be operating. So these are the defense in depth which goes into uh design of these systems that we have our uh there. So now if I have to create the fault tree uh for this uh particular system you you please remember DG1 DG2 CV breaker uh they are named 7 8 uh 7 8 and then uh they are feeding the class 3 loads. So interconnecting type breaker is CB6. So how I developed the faulty here? So class 3 power failure will happen only when under voltage on bus C and bus D. Okay, you understood this this and this and they fail then only it will happen. If one bus fails the under voltage is not there a power supply will be available on bus D and there is it is not a class 3 power. So that means both the DG should fail and their breakout should fail. How will we translate them? Under voltage relay the power supply from DG1. Power supply from DG1 will fail when when either DG has failed or circuit breaker 7. You can look back into the previous figure or common cause failure has occurred. Okay, common cause failure DG. This is a common event. This event will lead to both the DG's failure like uh then then we have bus D the second bus fail DG2 power supply and it will happen DG2 fails CV8 fails or common cause failure happens okay now now under voltage bus D so we we have talked about DG1 power supply bus from D because why we I'm talking about this supply this supply is coming from DG2 uh you can you can see there yeah um Suppose if my DG fails you know but DG1 is operate DG1 fails DG2 is available so I'll close this breaker and this power supply will be available from here. Similarly if this DG uh uh if uh if this DG fails uh then my power supply will be available but if both the DG fail then only there will not be any power supply. So that kind of interconnection we are able to and this only can be provided on the fault tree actually if I want to go for failure mode effect analysis critical analysis I cannot do that because I cannot reflect the plant logic how it operates actually. So under voltage on bus D. Now the uh DG2 that is uh bus D is fed by DG2 uh DG2 failure CB at failure or D circuit breaker. Okay. And then power supply from bus C. Okay. So that will be only when DG1 is avail. Now now the cuts set analysis minimal cuts set. This is a very elaborate fry. It explains me plant logic but I have to develop a reduced fault tree for reducing the computational time and to get a better idea of the system. So I do a cuts set analysis. You know that how uh using uh using this uh this cutset analysis tools uh we can reduce the fault tree the class 3 power supply failure independent failure I name it and common common cost failure of DG because common cost failure of DG was there everywhere it can be reflected on. So this if common cause occurs then DG fails. So how class 3 occurs either through independent failure or or through uh common cause failure. Okay. Now DG1 train fails. DG1 CB7 DG2 trains C DG2 C and these two trains should fail to res to uh enable class 3 power supply. Okay. So either this or this. So that's why orgate this uh DG1 and DG2 end. So end gate these two should fail to realize. So this logic we have implemented and we have cuts set please note this we have cuts set single order cuts set DG uh common ground cross failure DG and then DG1 DG2 second uh second order cuts set uh again CB7 and DG2 and DG1 and CB8 and CB7 and CB8. These are the cuts set. You can see minimal cuts set generation we have done and we found uh this uh uh and all these things are done by boolean logic. We have discussed this earlier. So apply the uh like if this and this DG into D uh CCF DG into CCF DG it will be one DG a into a is a A + A is equal to A. All those things you would have seen that. So you can apply uh boolean logic and we got the the reduced fault tree at the same time this these are the cut set you know. So uh now suppose if I have failure probability of this uh DG from qualitative I can go to quantitative fry analysis. So I got suppose I had a data of DG1 failure, DG2 failure uh from plant specific data or from generic source whatever but this I assume that the data is available even I have a class common cause failure data also with me if you don't have data for independent failure you take 10% of that 0.1 and you'll get common cost failure data this is the uh as in normal uh this thing normal domain we talk that common cost failure conservatively you should keep 10% point one of the independent failure. So we have taken DG independent failure and there this is test minus three. Okay. Now I got through this. So I assigned this probabilities from here to all this DG and I got the emergency power supply failure probability is 3.2 to 10 - 3 per demand. Now I got the two 3.2 into 10^us 3. How I I would know that it is good, bad or what it tells the numbers cannot make much. But then if I have a regulatory stipulation that the class 3 or emergency power supply failure probability should be at least 1 into 10^us 3. So we are we have some scope for improvement. That means the message to us is improve this. But since it is very close to 10^ minus 3 um I can say yes little efforts are required but I'm in the same domain actually. Okay. So this is how we should read the complex systems actually you know and then I have done the faulty analysis. I have analyzed power supply failure probability of emergency uh power supply 3.2 - 3. Let us say these are other safety systems. Okay. and their probability have been analyzed through again faulty only. So now I am analyzing the inventory for loss of offset power. I know that the frequency of loss of offset power is 1.2 1.2 per year at my site. Okay, you can have your value at your side. Primary shutdown system I told you secondary shutdown system first reactor should trip shut down. Okay, so if it is doesn't trip then I should see whether the secondary shutdown has uh came into action. So I will go on asking question and I'll draw the event tree on top it is success on bottom it is failure that means primary shutdown system success up uh the system has not responded not come or not actuated down okay again the secondary shutdown system has started so I'll go up and then I'll ask emergency power supply is available then we we don't have about human action and all that we have either success or failure decay heat removal should operate you So this is how the uh and then u inventory output is one is number of sequences they are 1 2 3 4 they are labeled and the uh you can see loss of offside power is safe if all these things. So if it is event one you go here that means power failure has occurred but all my safety system they operated safely. If it is a loss of offset power and loop and DHR even if loss of offset power is there for some reason if my DK heat removal system is not there too then it is a then it is a problem it is a damage state you can say uh then same way we have created this sequence over here if I add all these sequence and their quantification whatever we have I'll get the code image frequency See uh here so uh I would add 2 4 6 8 10 that is unsafe state only I'll add uh 2 4 6 and and you add and you arrive at the core image frequency statement now I got here uh so CDF loss of upside power is equal to 1 to n uh CD um uh code damage state so CD are this number stages and sedd also will come there actually Okay. So now I got this is the scenario for me and I got this uh code damage statement. If I sum up summation of risk statement for level one loop 6.2 there some correction is there here you can just note it you add yourself and go to the next stage. So these are the things they are called accident sequence uh and whether they are safe uh safety threat or not. uh and then uh finally uh you add these things together. Um and then now if I somebody asked me a question uh because I finally I have to focus on whether how important the uh power supply system uh for for for implementation of PHM. So I should know how much it is contributing to the code damage. Okay, then that will be one indicator. Remaining indicator risk importance and all we'll see in the next lecture. So loop contribution is 7.2 into 10^ - 6. Okay. Um assume that the plant CDF is 10^ - 5. This because this we are not shown. Let us say a complete P was done and this particular figure was available with us and we know that 5.25 is the plant uh uh CDF total. So what is the contribution of loop? 7.2 - 6 5.25 total I'll get 13%. Uh and then loop contribution is this. If for similarly I have estimated for other events also so I find 13% is like other than the ATWS anticipated transient all abbreviation you will not know it is a symptomatic for you yes just understand loss of offset power and loss of coolant accident they are contributing uh the minor loca is contributes 32%. So this comes somewhere in between there and loss of regulation con 14%. So we have here 13% and it cannot be neglected. So that means there are some power supply system especially diesel generator said which is contributing to to the uh emergency power uh that for that PSM should be employed and uh that that is what our reading uh so I given you a system level view many of the terms uh you would not have understood what you have to see that I have analyzed one loop other uh events are part of a complete book which has been developed as a probabistic risk assessment. ment of the plant and that we have we have not shown we have just given even these numbers I have assumed to show you that what is the contribution of power supply and how things uh SSC's system structures component they can be prioritized okay so this lecture I think he's g given you practical hands-on uh for uh for you to model your own system and try to understand the risk importance of the system and according accordingly delay system level there are 10 system which system is more important uh from PHM point of view which system is second to third fourth fifth like I told you one system was contributing almost 35%. So that is the most important system. It was a um uh small loca or medium loca whatever loss of coolant accident. So that will give us the rating from that rating you can go to the component because component only build the system and for you then then you'll be using because you know the when you perform proistic assessment that you have the minimal uh uh core damage frequency in terms of the component uh not in terms of systems. So you got the which component fails and how it results into core damage or which two component fail it result into core damage. So this is we have seen at the system level. Now we'll in next lecture we'll see how from CC uh CCF we'll go to the component level. [music] [bell] [music] >> [music] [music]