Video summary
The lecture introduces a risk-based approach for system modeling, which integrates traditional deterministic engineering knowledge with modern probabilistic methodologies to address uncertainties in complex systems. This integrated framework is designed to ensure both plant reliability and safety within a competitive global ecosystem by combining established safety cultures with advanced technologies like Integrated Maintenance Logistics (IML). The core of this approach involves understanding the hierarchical structure of a system, ranging from the overall plant down to subsystems and individual components, while accounting for interactions between human actions, machines, tools, and methods. A key paradigm shift highlighted is the role of power electronics in replacing traditional passive systems like UPS units, necessitating new modeling strategies that handle the conversion processes from AC to DC and back again to maintain uninterrupted power supply.
To effectively model these complex systems, the speaker utilizes fault trees and event trees as central tools for visualizing interrelationships between components rather than analyzing them in isolation. This graphical method allows for the inclusion of critical factors often missed in standard table-based analyses, such as common cause failures, human intervention issues, and machine interface problems. For instance, in a nuclear plant scenario involving a loss of offsite power event, the model demonstrates how multiple redundant layers—such as diesel generators, batteries, and passive containment systems—interact to maintain safety functions like decay heat removal. By mapping these logical connections, engineers can diagnose potential failure points and understand how a single component failure might propagate through the system, ultimately leading to either a safe state or core damage depending on the success of subsequent safety mechanisms.
The practical application of this risk-based approach involves quantifying probabilities to prioritize System Structures Components (SSCs) for Prognostics and Health Management (PHM) initiatives. Through cut-set analysis and Boolean logic, the lecture illustrates how independent failures and common cause failures are combined to calculate the overall probability of emergency power supply failure. In the provided example, the calculated risk was found to be close to regulatory limits, indicating that while the system is robust, there is still room for improvement. By summing the contributions of various accident sequences, such as loss of offsite power and loss of coolant accidents, the analysis reveals specific percentages of core damage frequency attributed to different systems. This data-driven insight allows operators to identify which systems contribute most significantly to risk—such as a system contributing 13% or a loss of coolant accident contributing 32%—and prioritize maintenance and monitoring efforts accordingly.
Ultimately, the lecture concludes that a holistic, risk-conscious approach is essential for moving beyond traditional methods to achieve higher levels of safety and operational efficiency. By focusing on the most critical systems identified through probabilistic risk assessment, organizations can allocate resources effectively to prevent core damage and ensure plant availability. The process operates in a closed loop where goals are validated against both deterministic requirements like structural integrity and probabilistic targets set by regulatory bodies; if these goals are not met, the system design or operational procedures must be redefined. This systematic prioritization ensures that PHM efforts are directed toward components and systems that offer the greatest impact on overall plant safety, providing a clear roadmap for managing mega-systems where multiple agencies and complex interactions define the operational landscape.
Read the full video transcript
[music]
So we discussed uh so far uh the system
modeling uh or system issues uh that was
very important because the because the
niche of this lecture is uh that we are
talking about systems approach. Um but
then there is a very important component
uh in systems approach is system
modeling because unless until we do
system modeling
or first unless until we understand the
system and we if we have a domain
knowledge also uh they then we cannot
even attempt system modeling and now we
have understood the the the aspects
associated with the system uh and then
uh it's a different form form
manifestations. Uh now let us say the
system modeling system modeling may
first thing will be you have to be an
expert in uh that system uh design
information it's operational information
uh uh human involvement in that are main
machine interface uh issues uh and many
more uh aspect assess associated uh with
the services um and then procurement
event all those things without that we
cannot do system modeling. Uh so let us
discuss uh a riskbased approach uh for
system modeling. Uh riskbased approach
is a riskbased engineering
as part of that I have developed a
riskbased approach uh for system
modeling. Essentially it is uh it is
using lot of probabistic risk assessment
methodology along with the uh some new
paradigm that are available uh in terms
of uncertainty characterization uh in
terms of human factor uh handling and
all that. But here the keeping in view
the availability of time we'll discuss
only the central character of riskbased
approach and as and when it comes to uh
implementing PHM uh we will explore
other areas also. So complex uh system
we have discussed enough now probably
you know all and the key metrics of uh
systems approach uh that we'll be
discussing role of riskbased approach uh
here uh and then identification of SSC's
prioritization uh and then uh role of
riskbased risk conscious approach uh
actually I have written two books this
book is uh riskbased engineering um and
here how we can design operate the
system uh with the uh risk or safety as
the as the overriding factor
you know and um how we can we move away
from the traditional approaches or how
we integrate rather I would say uh the
traditional approach the best part of
the traditional approach and the um
advances in the u last 20 to 30 years um
and then have a combination of
which gives us better results uh and
better insights and ensures both plant
reliability because you know the times
have changed now uh in this uh time of
globalization competitive uh uh
ecosystem the plants have to operate
also okay and at the same time ensure
the safety also how it can be done it
can be done by what best part of what we
had as a as a safety culture and best
part of what the technology has enabled
us including a IML and all and provide
the uh integrated approach uh for system
design and operations.
So um the it is basically sort of a
perspective that I'm talking about uh
when I talk about the system I actually
I'm talking about a plant that I have
told you uh and then a system
uh you know literally has can have a
subsystem
is a subp part of the plant so plant
comes on top systems subsystems another
subsystems and level of subsystem then
the components that's how it flows. So
that is a hierarchy we if we see the
plant hierarchy over there then
operation management ecosystem
unless until we uh understand understand
the
uh complete ecosystem how the
maintenance uh human actions um machines
tools methods they interact uh you know
how communication occur occurs among
different agencies and then we realize
and now we have uh the micro electronics
and power electronics. This was not
there earlier but now power electronics
has been playing and it is replacing the
traditional uh electrical tools and
method with passive uh passive like uh
UPS you know uninterrupted power supply
system. Uh they are they are the uh
brought in paradigm shift into how we
look processing electrical power from AC
to DC to AC and finally AC to DC and
again AC. So it becomes uninterrupted
power supply. So all these things they
they have to be addressed. Okay. When we
talk about a systems approach
um we we are calling uh nuclear plant we
are taking it as a reference because as
you saw on my book also I written
reference and nuclear plant. So it is
easier to explain and understand the
best of the systems uh and from there to
draw the references. And then we have uh
the objective here is uh to maintaining
highest level of safety and then uh
reliable oper ensuring reliable
operation. This this is a broad
perspective how we look at the systems.
Now um let's say in in this reference
plant we have considered a lot because
every plant has got their own uh
characterizing the name of the system
the name of the event. So here we have
for reference plant this nuclear plant
loss of offset power that means this is
one of the event though it is a
anticipated occurrence uh but it it was
found that it contributes to risk from
10% to uh you know 10 to 8% or 13% you
know uh because uh somebody might ask
how loss of offset power when loss of
offset power occurs the plant is shut
down but the decay heat is
is production at very fractional level.
Let's say let's say 1% 2% 5% it goes on
and core has to be cooled. Now if the
on-site uh dedicated power supplies
diesel generator or any other they have
to produce the power that means they
have to come uh they have to start and
uh meet the whatever 10% plant loads uh
essential load requirement. If they do
not start there is a problem. So um then
then you have a this is a class 3 power
supply then class 2 and class one
batteries and all that. So there smooth
operation uh I would sum up is required
and here we have this integrated
approach uh integrated which is uh
something uh graphically nature that
means we can see visuals and we can
understand how the components are
interconnected what happens if one
component because you know they have
used logic gate and all that. uh so we
can by looking at the things we can
provide a very good diagnostics how the
plant operates and when it will fail or
when the system will fail or when the
component will fail. So uh so it's a and
there uh the advantage is like there are
traditional approaches like failure mode
effect effect analysis failure mode
effect and criticality analysis there we
choose one component and we see the
consequences till and last but then on
the same row but here it's not that it
shows the inter relationship between the
component this fault treeries and all
that they are very very um elegant
mechanisms that enable plant
representation and They show plant
characteristic also they and there are
some aspect which cannot be covered in
table methods and they are like common
cause failure, human factor in
integration or main machine interface
how human uh can intervene all those
things are uh can be can be shown in a
integrated way into the fault tree and
then the event tree will show um I think
we have discussed fault tree eventry so
I'm directly talking about then when I
provide the case study probably you'll
get a better idea
Now ens symbol of functions you know the
when we talk of phm component wise
either it is electron electronic
components mechanical component but here
ensemble of methods are required right
from uh correct initiation to correct
propagation okay then we we require
damination electronics okay of course
the the mode of failure is mechanical
only but then those failures have to be
se seen and modeled okay when we talk
about The uh power electronics it is the
capacitor which has got a risk component
which is which has got a reliability
component both IGBT integrated circuits
field programmable gate arrays each one
will have their own characteristic and
modeling requirement that's why systems
approach is required it is way
challenging compared to the component
specific approaches multi- agencies
condition and uh and here uh if you have
a systems approach or especially mega
system approach lot of agencies are
involved and they are taking part into
operation maintenance logical things and
all that and on the top of that there is
a regulatory body which oversees FL
operation or gives provides intervention
level if required if any safety uh case
is being made out and then they will uh
they will have stipulation on those
things. So uh executive body alone is
not uh
working in isolation. There are some
independent and that's how the safeties
uh safety uh uh is taken care of. So
well uh safety risk and security as I
mentioned now we have to bother about
safety and security risk also. So our
plant model should have all these
features.
So as I mentioned in riskbased approach
uh why we are using riskbased approach
it is allowing the proven knowledge
deterministic compon know knowledge of
the component [snorts] and then
probabistic relatively new knowledge how
to represent plant system
interconnectivity and then integrate
both of them. Okay. So we have best of
the things in terms of safety and
reliability validation with the
deterministic and probabistic uh goals
and criteria
uh probabistic goals
are defined at regulatory level or at
international level whether we are able
to meet them in a holist holistic way.
uh uh and probably uh deterministic
goals in terms of uh minimal flow
requirement, minimal pressure
requirement, uh minimal structural
integrity requirement.
So those things are validated and
enhanced monitoring and serless that
gives safety and availability goals and
um and you ask yourself a question
whether the goals are met. If it is yes
uh uh job is done. If it is no no then
uh then you have to redefine the goals
and call it is it operates in a uh
closed loop and that's how we have this
integrated displacement approach and uh
here um fault tree event trees they are
at the core of it along with the
deterministic model and methods and
makes lot lot of effort goes into the
common cause failure modeling because
common cause is something even if you
have a lot of redundancy built into the
system common cause can knock off which
is could be external parameter it could
be internal parameter uh but it can
knock off uh the complete redundancy and
diversity. So simple example one seismic
event and all our redundant system uh
they can get adversely affected. Of
course things are designed keeping in
you the seismic requirements and all
that. Uh but then uh but then um their
fragility uh against this uh uh this
seismic event um has to be tested and it
should be uh shown to demonstrate that
it will meet in this uh requirement. So
common cause components have to be
mapped on to those events. Quantified
approach to and this is a quantified
approach. quantification we understand
better and all that and this is as I
mentioned in previous slide it is based
on my book one uh integrated riskbased
engineering.
So just for the sake of the the plant uh
we you have a plant and then we have
safety systems. Okay. So uh as far as
the uh safety system, primary reactor
shutdown system, secondary reactor
shutdown system, any disturbance in the
plant, plant should be brought to the
safe state. And the cooling function
reactor primary cooling I have shown
here uh in a simple way very elegant
this thing. Uh then uh secondary
shutdown system uh decay heat removal
system. You see what kind of layers we
have in emergency core cooling system.
primary secondary decay uh decay removal
and then emergency core cooling system.
So uh this is a pump or primary regular
primary system. Then we have a decay
heat removal system and I have shown two
two only emergency cooling system I have
not shown here. This is the fuel which
produces heat hot water is goes and it
goes to the heat exchanger and it
operates in a closed loop. If the this
pump because of loss of offside power
we'll be talking in this lecture. So
loss of upside power means this bigger
pump almost like uh 700 kilowatt or you
know uh that size this will not be
operating because there's no power
supply uh coming from offside power. So
small pumps will operate and cater to
the decay heat requirement. Plant is
shut down heat production has come down
but even 5% heat has to be removed and
uh that again has to operate in a closed
loop. I have shown a very very
simplified view of just to give you an
idea and when the steam is produced here
it operates a turbine and the thing
condens condensed water it goes to uh
condenser and then finally it is uh you
know fed to the uh this room this is
primary coolant pump this is decay heat
removal pump reactor cold water so then
we have different type of power supplies
class three power supply as I told you
class 4 power supply means normal power
supply grid power supply this pump
operates but diesel generator power
supply is available in shutdown. So
these pumps will operate and then then
we have a UPS uninterrupted power supply
class one DC batteries are there and the
then passive contaminate isolation
system primary containment secondary
containment tertiary containment
probably you can visual visualize this
as a part of the containment actually
okay so there are two walls two lay two
boxes
I have shown only one primary
containment only I have shown there can
be two containment
and then for their ventilation uh
emergency exhaust system human recovery
action and all that. So we'll see the
loss of offsite power event modeling uh
just to understand how to model a system
for prognostics per prioritization you
know at system level. Okay. So uh and
then our we will take the example of
class 3 power diesel generator.
So
this is a U class 4 power failures. That
means class 4 means offside power
feeding these buses. In the moment of
loss of offside power, okay, these
feeders will stop feeding. Then these
breakers will open and then we have
diesel generator 1, diesel generator 2.
They will automatically sensing the
under voltage on these buses. They will
start automatically breaker closes. And
C and D are emergency buses. So they
will supply power to the emergency
buses. Then class three loads. So these
loads itself are called you can say
emergency loads but it is not there is
no emergency because the plant is it
it's normal operation state only. Uh so
they will feed our class 3 loads only uh
loads you know from the both the buses.
There are many loads actually and like
one of the thing is I showed you the
deep heat revols they should operate.
Yeah. In fact the other more than class
they are on class two. Why class two?
Because the class two will take AC
convert into DC and again will feties.
So that means uh as long as batteries
are available in the plant that AC power
will be available and on which those
pumps will be operating. So these are
the defense in depth which goes into uh
design of these systems that we have our
uh there. So now if I have to create the
fault tree uh for this uh particular
system you you please remember DG1 DG2
CV breaker uh they are named 7 8 uh 7 8
and then uh they are feeding the class 3
loads. So interconnecting type breaker
is CB6. So how I developed the faulty
here? So class 3 power failure will
happen only when under voltage on bus C
and bus D. Okay, you understood this
this and this and they fail then only it
will happen. If one bus fails the under
voltage is not there a power supply will
be available on bus D and there is it is
not a class 3 power. So that means both
the DG should fail and their breakout
should fail. How will we translate them?
Under voltage relay the power supply
from DG1. Power supply from DG1 will
fail when when either DG has failed or
circuit breaker 7. You can look back
into the previous figure or common cause
failure has occurred. Okay, common cause
failure DG. This is a common event. This
event will lead to both the DG's failure
like uh then then we have bus D the
second bus fail DG2 power supply and it
will happen DG2 fails CV8 fails or
common cause failure happens okay now
now under voltage bus D so we we have
talked about DG1 power supply bus from D
because why we I'm talking about this
supply this supply is coming from DG2 uh
you can you can see there yeah um
Suppose if my DG fails you know but DG1
is operate DG1 fails DG2 is available so
I'll close this breaker and this power
supply will be available from here.
Similarly if this DG uh uh if uh if this
DG fails uh then my power supply will be
available but if both the DG fail then
only there will not be any power supply.
So that kind of interconnection we are
able to and this only can be provided on
the fault tree actually if I want to go
for failure mode effect analysis
critical analysis I cannot do that
because I cannot reflect the plant logic
how it operates actually. So under
voltage on bus D. Now the uh DG2 that is
uh bus D is fed by DG2 uh DG2 failure CB
at failure or D circuit breaker. Okay.
And then power supply from bus C. Okay.
So that will be only when DG1 is avail.
Now now the cuts set analysis minimal
cuts set. This is a very elaborate fry.
It explains me plant logic but I have to
develop a reduced fault tree for
reducing the computational time and to
get a better idea of the system. So I do
a cuts set analysis. You know that how
uh using uh using this uh this cutset
analysis tools uh we can reduce the
fault tree the class 3 power supply
failure independent failure I name it
and common common cost failure of DG
because common cost failure of DG was
there everywhere it can be reflected on.
So this if common cause occurs then DG
fails. So how class 3 occurs either
through independent failure or or
through uh common cause failure. Okay.
Now DG1 train fails. DG1 CB7 DG2 trains
C DG2 C and these two trains should fail
to res to uh enable class 3 power
supply. Okay. So either this or this. So
that's why orgate this uh DG1 and DG2
end. So end gate these two should fail
to realize. So this logic we have
implemented and we have cuts set please
note this we have cuts set single order
cuts set DG uh common ground cross
failure DG and then DG1 DG2 second uh
second order cuts set uh again CB7 and
DG2 and DG1 and CB8 and CB7 and CB8.
These are the cuts set. You can see
minimal cuts set generation we have done
and we found uh this uh uh and all these
things are done by boolean logic. We
have discussed this earlier. So apply
the uh like if this and this DG into D
uh CCF DG into CCF DG it will be one DG
a into a is a A + A is equal to A. All
those things you would have seen that.
So you can apply uh boolean logic and we
got the the reduced fault tree at the
same time this these are the cut set you
know. So uh now suppose if I have
failure probability of this uh DG from
qualitative I can go to quantitative fry
analysis. So I got suppose I had a data
of DG1 failure, DG2 failure uh from
plant specific data or from generic
source whatever but this I assume that
the data is available even I have a
class common cause failure data also
with me if you don't have data for
independent failure you take 10% of that
0.1 and you'll get common cost failure
data this is the uh as in normal uh this
thing normal domain we talk that common
cost failure conservatively you should
keep 10% point one of the independent
failure. So we have taken DG independent
failure and there this is test minus
three. Okay. Now I got through this. So
I assigned this probabilities from here
to all this DG and I got the emergency
power supply failure probability is 3.2
to 10 - 3 per demand.
Now I got the two 3.2 into 10^us 3. How
I I would know that it is good, bad or
what it tells the numbers cannot make
much. But then if I have a regulatory
stipulation that the class 3 or
emergency power supply failure
probability should be at least 1 into
10^us 3. So we are we have some scope
for improvement. That means the message
to us is improve this. But since it is
very close to 10^ minus 3 um I can say
yes little efforts are required but I'm
in the same domain actually. Okay. So
this is how we should read the complex
systems actually you know and then I
have done the faulty analysis. I have
analyzed power supply failure
probability of emergency uh power supply
3.2 - 3. Let us say these are other
safety systems. Okay. and their
probability have been analyzed through
again faulty only. So now I am analyzing
the inventory for loss of offset power.
I know that the frequency of loss of
offset power is 1.2 1.2 per year at my
site. Okay, you can have your value at
your side. Primary shutdown system I
told you secondary shutdown system first
reactor should trip shut down. Okay, so
if it is doesn't trip then I should see
whether the secondary shutdown has uh
came into action. So I will go on asking
question and I'll draw the event tree on
top it is success on bottom it is
failure that means primary shutdown
system success up
uh the system has not responded not come
or not actuated down okay again the
secondary shutdown system has started so
I'll go up and then I'll ask emergency
power supply is available then we we
don't have about human action and all
that we have either success or failure
decay heat removal should operate you So
this is how the uh and then u inventory
output is one is number of sequences
they are 1 2 3 4 they are labeled and
the uh you can see loss of offside power
is safe if all these things. So if it is
event one you go here that means power
failure has occurred but all my safety
system they operated safely. If it is a
loss of offset power and loop and DHR
even if loss of offset power is there
for some reason if my DK heat removal
system is not there too then it is a
then it is a problem it is a damage
state you can say uh then same way we
have created this sequence over here if
I add all these sequence and their
quantification whatever we have I'll get
the code image frequency See uh here so
uh I would add 2 4 6 8 10 that is unsafe
state only I'll add uh 2 4 6 and and you
add and you arrive at the core image
frequency statement now I got here
uh so CDF loss of upside power is equal
to 1 to n uh CD um uh code damage state
so CD are this number stages and sedd
also will come there actually Okay. So
now I got this is the scenario for me
and I got this uh code damage statement.
If I sum up summation of risk statement
for level one loop 6.2
there some correction is there here you
can just note it you add yourself and go
to the next stage. So these are the
things they are called accident sequence
uh and whether they are safe uh safety
threat or not. uh and then uh finally uh
you add these things together. Um and
then now if I somebody asked me a
question uh because I finally I have to
focus on whether how important the uh
power supply system uh for for for
implementation of PHM. So I should know
how much it is contributing to the code
damage. Okay, then that will be one
indicator. Remaining indicator risk
importance and all we'll see in the next
lecture. So loop contribution is 7.2
into 10^ - 6. Okay. Um assume that the
plant CDF is 10^ - 5. This because this
we are not shown. Let us say a complete
P was done and this particular figure
was available with us and we know that
5.25 is the plant uh uh CDF total. So
what is the contribution of loop? 7.2 -
6 5.25 total I'll get 13%. Uh and then
loop contribution is this. If for
similarly I have estimated for other
events also so I find 13% is like other
than the ATWS anticipated transient all
abbreviation you will not know it is a
symptomatic for you yes just understand
loss of offset power and loss of coolant
accident they are contributing uh the
minor loca is contributes 32%. So this
comes somewhere in between there and
loss of regulation con 14%. So we have
here 13% and it cannot be neglected. So
that means there are some power supply
system especially diesel generator said
which is contributing to to the uh
emergency power uh that for that PSM
should be employed and uh that that is
what our reading uh so I given you a
system level view many of the terms uh
you would not have understood what you
have to see that I have analyzed one
loop other uh events are part of a
complete book which has been developed
as a probabistic risk assessment. ment
of the plant and that we have we have
not shown we have just given even these
numbers I have assumed to show you that
what is the contribution
of power supply and how things uh SSC's
system structures component they can be
prioritized
okay so this lecture I think he's g
given you practical hands-on uh for uh
for you to model your own system and try
to understand the risk importance of the
system and according accordingly delay
system level there are 10 system which
system is more important uh from PHM
point of view which system is second to
third fourth fifth like I told you one
system was contributing almost 35%. So
that is the most important system. It
was a um uh small loca or medium loca
whatever loss of coolant accident. So
that will give us the rating from that
rating you can go to the component
because component only build the system
and for you then then you'll be using
because you know the when you perform
proistic assessment that you have the
minimal uh uh core damage frequency in
terms of the component uh not in terms
of systems. So you got the which
component fails and how it results into
core damage or which two component fail
it result into core damage. So this is
we have seen at the system level. Now
we'll in next lecture we'll see how from
CC uh CCF we'll go to the component
level.
[music]
[bell]
[music]
>> [music]
[music]