Week 10 - Lecture 49 : Identification and Prioritization for Implementing PHM
Watch on YouTubeVideo summary
The lecture focuses on the critical processes of identification and prioritization within Plant Health Management (PHM) for operational plants, distinguishing this real-time scenario from the comprehensive Probabilistic Risk Assessment (PRA) conducted during the design phase. The core argument is that while safety cannot be quantified as a simple mathematical parameter, risk can be modeled using established formalisms to determine which components contribute most significantly to core damage frequency. This approach allows operators to create a priority list addressing the top 10 to 30 components based on available resources, thereby maximizing improvements in overall plant reliability and risk reduction. The identification process is not solely mathematical but integrates qualitative insights from expert judgment, operational crew feedback, and supervision reports, ensuring that even non-failed components are maintained proactively if they show signs of latent degradation or pose a threat to system integrity.
In the context of operation and maintenance, prioritization is governed by performance evaluation indicators where safety systems take precedence over operational systems, which in turn are managed alongside quality assurance protocols. The lecture highlights that traditional approaches often rely on subjective qualitative analysis, such as plant area reporting and human factor assessments, lacking a systematic mathematical basis. To address this, the presentation introduces an integrated risk model that combines technological hazards, industrial safety issues like fire or flooding, security risks, delivery risks, and liability concerns into a single framework. This holistic view is further enhanced by machine learning algorithms that weigh different risk factors, allowing for a more nuanced understanding of how various threats interact to influence overall plant safety and operational continuity.
The lecture concludes by detailing specific mathematical metrics used to quantify component importance, moving beyond qualitative guesses to precise calculations. Key measures discussed include Fussell-Vesely importance, which assesses the fraction of system unavailability attributable to a specific component, and Risk Reduction Worth (RRW), which calculates the benefit of making a component perfectly reliable. Additionally, Risk Achievement Worth is explained as a metric that determines how much risk increases if a component fails completely, helping operators identify critical targets for maintenance. These mathematical models directly inform decisions regarding surveillance test intervals and allowable outage times, ensuring that maintenance schedules are optimized to prevent latent failures while balancing the need for continuous plant operation against acceptable levels of unavailability.
Read the full video transcript
[music]
>> So, we are on to our fourth lecture of
10th week. Identification and
prioritization for implementing PHM.
Um here we'll be discussing some
mathematical well-established uh
formalism, uh how to identify and
prioritize the component. And since uh
I told you that our matrix is
uh risk matrix. Because safety is not a
not a mathematical parameter. So,
uh we cannot do uh prioritization based
on safety matrix. So, we have to take a
risk matrix. Um
uh and then we have to identify and
prioritize it. In previous lecture we
you saw that uh how we brought uh um
this uh system level unavailabilities.
And then we were we explored uh which
are the accident sequences. Accident
sequences are made up of what?
Components.
Okay? So, component one, component two,
component three fails, then
the top event, core damage, occurs.
Okay? So, uh now we'll be talking about
uh the ex- accident sequence level, that
is cut set level. Which component has
got its sen- sensitivity to core damage?
Core damage means uh
um
the damage state of and that is that
adds to risk uh
element of the
uh
so, risk element of the
uh various uh accident sequences.
So, uh before we we do that,
uh let us see role of identification and
uh prioritization operation maintenance.
Because we are talking about the plant
operation. Please remember that. In
fact, this thing is done in design stage
also, but
that time
we we do complete PRA and all that. But
the plant stage is postulated or plant
operation is that is going to happen in
that way. But here it is a real-time
scenario plant is operating. So,
whatever insights that are available and
mathematical model modeling that are
available
based on that
the prioritization identification is
done and
identification and prioritization is
done
to understand the risk importance of
each and every component so that as part
of PHM
they can be we can make a priority list.
Okay? And
first 10 20 30 components, whatever
resources available, if we address
and we'll see what kind of gain we have
in terms of
improvement in risk and reliability.
>> [snorts]
>> Okay, then so so you can see here in
operation maintenance the the
identification identification
and prioritization, these are the two
inherent component of
even when the plant is operating, if you
were to take a maintenance schedule, we
were to identify components and
and prioritize them
which component comes so this happens
because of plant configuration itself.
You know, there are safety system and
there are operation system. So, for
these systems the risk significant
component comes first and then plant
operating component and then some
insights from expert insight from
operation maintenance crew and then
annual program is drawn and the
systematically all the components and
one more important thing is
how
frequently or what should be that
maintenance interval? Maintenance
interval. That has to be decided.
Earlier I would tell you which are the
conventional procedures and how now it
is done mathematically and rather the
mathematics is open for their use for
for for choosing the maintenance
intervals and all that that will be
there but so
these are as I told you they are
inherent component of operation
maintenance.
And it is governed by performance
evaluation. So safety and perform
reliability are the two
indicators.
Reliability for process system and
safety for standby system. And then
uh
quality assurance plays role lot of role
in that. The first part is the
inspection interval has to be decided.
So QA decides what should be that
interval. Okay. Uh
by and large safety system takes
priority and then of course
operational system.
Operational system because
because failure of operational system
results into initiating event and the
initiating events when it gets valid and
then system response and you know the
event tree and all how they form part of
the core damage frequency. So
quality assurance is very important
thing. Quality assurance part uh
you know part of all the activities
which are happening in the plant but
here I'm talking to you in terms of
identification prioritization. And then
super supervision and reporting. Many
component may not may not appear
based on your mathematical mathematics
but in supervision when you visit the
areas of the plant and
you find oh suddenly this item now it
should go into maintenance next time.
Why because
there can be some issues with that and
it is not falling in the safety category
it will be not falling in but it is
important to do that. So, uh
you can say a qualitative thinking goes
into many of the items which have to be
taken into
The component is not failed, but expert
decision is made. Let us do some
maintenance because the
the component we had the spare part we
had procured in other system, they have
been failing more. So, as a
precautionary measure, let us replace
here also. So,
mathematics is one thing, but there are
so many domain related aspects that come
into the picture. And so, so supervision
and reporting has got some some water is
dripping from somewhere. You find your
work done and you feel it should not get
aggravated and that takes the priority
then. Okay, structural priority. Then
then we have the documentation and
development updating. All these
activities on the same
updating of the document when you do
based on your maintenance inside or
operation inside, they form up operation
and maintenance management. And finally,
if the issue is
caught and is corrected, then corrective
action to be implemented. Okay? So,
issue identified, okay, and prioritized,
but corrective action program has to be
it could corrective action program could
be the action to be taken when the when
the
as part of the
if I have to talk in terms of PHM,
you know, the whether this component
should go in the in the maintenance or
repair now or replacement, repair,
replacement or overall. Time being, let
us buy some time. Okay? So, those kind
of thing corrective action program.
So, it will be ad hoc program and then
it will go for a
permanent solution.
And feedback. Feedback is very important
whether it is a PHM because the whole
PHM can look or different quite
different compared to when you started
it actually. And that happens because of
feedback which you are getting
from the performance of the PHM system
itself. Okay? So, as I mentioned, PHM
system should have its own
risk and reliability criteria, you know?
And its own performance actually. So,
one is plant performance, then its own
performance also should go into the
picture. So, this is what when when I
look at the plant in a in a from the PHM
implementation perspective.
Now,
before we go for PHM,
what are what are the routine or
traditional approaches? Because there is
a lot to learn from
routine approaches. One, why? Because
some of the things will continue the way
it was there. Number two, what
improvement we can we can do
you know, in this itself so that it not
matches with the PHM. Because you know,
PHM is a resource intensive activity.
So, some balance has to be made because
money has to be spent. And then certain
thing cannot be compromised like safety,
and certain thing cannot cannot be can
be managed by to doing some small
action. So, a qualitative analysis is
required. Performance [snorts] signature
analysis in control room has to These
are the things which are traditionally
done. Plant area reporting.
Here I would I would tell something.
If we had nuclear industry had
demonstrated
safety record for for itself, then these
procedures or these are activities are
none important actually. So, discarding
these activities or improving upon these
activities, we have to really go in a
sensible way. So, so qualitative
analysis was one part which was
happening as part of risk and
reliability program.
Then performance signature analysis in
control room. Plant areas reporting.
Audio-visual alarm trips and all those
in control room scenario, especially in
the plant transients. Plant walkdown I
showed you on previous slide what kind
of feedback it can generate. Condition
monitoring trends,
schedule-based maintenance management,
abnormal symptoms of precursor analysis
Precursor analysis means any failure
which is coming in or
any failure, what is its role in the
complete accident sequence that you
have. Okay, so they are called accident
precursors, you know. So, any symptom
which can take part into the accident,
it's potential. Uh daily log reporting.
These are the basic approaches for
identification.
Uh
traditionally it has been happening.
Uh be it a plant area, be it human
factor improvement, be it quality
improvement, all these and are maybe any
modification that you need to be
reflected into the plant.
So,
uh so now if I say
uh how safety
used to be prioritized, it is a
performance indicators, safety and
reliability,
uh and performance indicators they are
analyzed annually, uh level of
redundancy,
uh system diversity, common cause
potential. This is one of the major
indicator and we we have to have a
special provision. In fact,
uh there are plants who perform common
cause failure analysis other than the
safety analysis. They perform common
cause failure analysis because it
requires a very special attention
because it kills the redundancy,
diversity, and rest of the provision
that are existing in the plant. So, it
is required and human factor also can be
one of the common cause failure. So,
human factor becomes very important. Uh
again I'll repeat because we are on
prioritization and identification.
Uh human factor has been found to be one
of the major contributor to the
accidents. World over this observation
is there in the uh international
standards books and all that. So, it
requires special latent degradation.
Uh you know, for safety system, uh, if
there is electronic channel, if there is
no actuation, and if there is no
provision for monitoring it, then some
failures they might remain
uh, latent. And to know about those
failures, uh, the the best thing is you
perform surveillance testing. So,
uh,
let us say
diesel generator
system, emergency power supply system.
Uh, if we test it, we'll it will reveal
the latent failure. That is the only way
to know that. So, these testing testings
are performed. Then aging symptoms and
quality levels. These are for safety
system. Similarly, some of the things
may be exchanging. It is not that hard
and fast, but by and large, these are
the thing
uh, it metrics that goes for safety
system
uh, prioritization.
Um, and then process system.
Uh, reliability, failure-free operation,
downtime, routine overalls and repairs,
aging symptoms, and common cause
failure.
Uh, again, here also it comes here. So,
at the end of the day, here we see we
are talking about safety means
safety hazard, whatever was available.
There are some industrial safety issues,
and there are security issues. So, if I
talk about the safety, I have to use a
holistic statement for safety, which
takes into account these things also.
So,
uh, but the biggest weakness is the
conventional approach to is largely
qualitative in nature, and tends to be
subjective because these things are
monitored based on the discussion,
brainstorming, and some expert opinion
available there. But there is no
there is no systematic rational base
mathematics behind them actually. So,
there is a need for improvement that of
course we can say. So, traditional
approach, if it has to be seen in the
modern development, wherein new
factors are
entering. Industrial safety was part of
it. So, it may not produce the the
domain system hazard issue. But finally,
>> [snorts]
>> um, like you know, industrial safety
include fire, flooding,
falling from a height, impact targets
and all that. These are the security. It
requires attention because there is a
quite a bit of overlap between safety
and security. So, that is could be
advantage if they are segregated and
handled
by the means they are required actually.
But security as a issue alone,
they need to be addressed in a separate
analysis. Okay? So,
this is
and then if you go for modeling because
now we are talking not in terms of
safety, in terms of risk we are talking
about. So, we can create a integrated
statement of risk. So, it can be
interpreted as integrated statement of
safety. And which are the component of
it? Technological shift like if I talk
about the nuclear plant, it is a nuclear
hazard. Okay?
Industrial safety, which is common to
almost all the engineering system.
Then delivery safety. This I have given
a special name.
It is output.
What What is the risk that next month
I'll not be able to deliver my product,
whatever may be the product. So, this is
called delivery risk and then security
risk. So, here we have talked about the
safety risk, industrial risk, security
risk. It is output and then liability.
Liability is a very special
class of this thing for any damage. Who
will be liable to pay? Who will be
liable to compensate? So, this should be
integral part of the complete operation
of the plant, liability risk. And if I
use the machine learning learning
algorithm, then I have all the risk and
its weight is factor. Because weight is
factor for technological or hazard risk
will be much more
because it can have consequences,
potential for consequences
beyond plant plant. But industrial
safety, yes, it is important, but it
will have a limited consequences in
there. Similarly, delivery risk, I will
not be able to deliver my committed
output, but
it is a financial loss. But, financial
loss also perpetually is not good
because if I am having good finance
because of my plant operation, that goes
same resources goes into improving the
technological risk, industrial risk, and
all that. So, it's feedback thing. So,
it is very important from that point of
view it has to be seen. Then there's
security risk. It is a new dimension
I will not say relatively new dimension
and it has got its own waiting factor.
Probably, it should match at waiting
factor of this because security issues
provide a threat and then finally, it is
the liability issues. Okay.
And then, if I have all this input
together and then I have all this hidden
layer, it could be more than one hidden
layer, and I can give a statement, but
my problem is having parameter for this
vector.
So, that means
what we require is a consortium of
different industry or among nuclear
industry for different plant, we perform
this analysis and
train this neural network with these
vectors
where R1, R2, R3, R4, and I5 they are
the components, and we produce an output
here. And this output could output could
be
because this will have highest wattage
factor. So, so, and then remaining
things will accordingly. So, we can have
a output spectrum here how my plant what
is the safety level actually. If I have
quantification of all these things, what
should be my safety? This is a science
that has to be developed further, but it
has been proposed in my book
that has to go on. So,
and then
okay, we have talked about it.
Risk is a nothing but summation of
individual risk R1, R2, R3 where RI is
the
wattage uh,
and then likelihood into consequences.
Machine learning model is available
here. This is from my book on risk-based
risk-conscious operation management.
Okay.
So,
a lot of R&D has to go into it uh, to
provide it integrate In In In fact, it
is not exhaustive. There could be any
other factor which will add on. It is
basically uh, what kind of ecosystem
that we are operating? Earlier we used
to have only technical risk. Uh, we
never talked about security risk, but do
you have to talk about the security risk
now? Because it is directly relevant.
Okay. The Now Now Now we are talking
about the quantified approach for
prioritization. See, prioritization is
one thing. I have prioritized the
component. Okay. Top five five
component. But what should be the
frequency?
What should be the frequency? Uh, our
test interval, I would say. Surveillance
test interval for the the those
component. Why? Because we know uh, so
here we are 80/20 criteria. 20%
component will give you benefit of 80%.
Very good. So, I have made the list of
the component using importance measure
which uh, we talked talked in on a class
three power supply system and loss of
offsite power scenario. Uh, now let's
see if we are talking about the
component, but then who will tell me
what should be the frequency? The
frequency of test interval or
maintenance is governed by performance
of the system. And that performance of
the system uh, can be defined how the
systems is degrading or because
reliability If it is a new component,
fresh component, reliability will be
one. But
during its operation, it will reduce. It
is inherent. Okay? And this we are
talking about the random uh, phenomena.
Okay? So, it will go reduce. So, what we
do in plant? We stop the component, we
shut down the reactor, and uh, like like
suppose there is a loop, cooling loop.
We shut down the cooling loop. We do
some maintenance maintenance or
surveillance job test or we do testing.
Testing reveals the latent failure, we
know that. So, if any
latent failure was sitting over there,
you know,
that should be known and for that I do
shutdown, I do maintenance or test
interval and
maybe it may be overhaul, it will be
replacement, repair, whatever and then
again I started back. So, it has become
new. What the point here is Initially it
was a reliability was availability plant
availability, you know that. Uptime upon
down down time total down time.
Uptime upon down time per uptime.
That will give me steady state
availability and reliability you know
that. So,
if the reliability
drops, so this is a criteria. We have to
decide what is the criteria.
So, and then it will tell me how
the
testing or
shutdown maintenance job should be done
actually. So, one thing was I identified
the component, second thing was how
frequently I should do it. So, this is a
a theoretic theoretical model we have
done, okay? And this can be done by
taking into account failure
characteristic of the component. But I
should know this criteria and this
criteria
is decided by what is acceptable level
of degradation. Maybe initially we start
with the conservative level and later on
we go down and try to finalize it and
finally we see
how it is reflecting in our life.
Finally we are talking about plant
unavailability. So, what is the
unavailability and what that with that
availability what core damage frequency
we are getting and that will be telling
me telling us, "Okay, if the standard
core damage frequency international
criteria is 1 into 10 to the power of
minus 6 and so this availability is
meeting my criteria, so my
unavailability should be I'm talking
about availability. Converse of this
will be unavailability. So, this
unavailability we are we are talking
about it actually. Okay? So,
then so we have discussed here one is
the prior rotation and then we have
talked about what we'll do with that
duration. If you have to do the testing
and all that. Okay? Because finally, I
should know the weakness of the
component. It is a hardware system.
Okay? One on one thing you can say that
if I reach somewhere um RUL or here
okay? Then I should because now in PHM
we are monitoring the degradation. So,
we know that liability already going
down and if I know this trend, even this
time will not will not be required.
We'll I'll do plant management mode I'll
go.
You know,
and PHM will tell tell us okay, now we
have to we can start we can keep it
running or whatever. So, somewhere this
concept they are overlapping actually,
you know? But without PHM this is this
is the way you shut down and then you do
management. If I'm having PHM then this
particular thing will not be required.
So, you have come out of this slot with
availability of PHM.
Okay? Now,
in principle
risk importance. What is a general
mathematical formula is
importance of a component risk
importance of a component is a function
of
uh
change in availability of the individual
component
and change in availability at at system
level.
Uh you understood this we are talking
about I represent the importance of the
component II or a capital I is the
importance of the component I and
small I is the name of the component and
it is a function of D is nothing but a
function of
unavailability of the Ith component and
U stands for
unavailability of at the system level.
Okay.
Uh I explained this terms over here. So,
you you must know that
because by in in a in a academic domain,
we talk a lot lot about the reliability
and all that. So, that means reliability
means successful domain. Probability
that component will operate by a given
period of time under given condition for
a certain time
failure-free operation of the So, that
that we say. So, it is in success
domain. So, um sometimes you you
have a mathematical
expression
you will find in literature
based on the reliability being more
important of the
that we talk about. But in
in risk analysis, we always deal in
failure domain. So, here we have some
risk important measure and then
objective and scope based on the PRA.
We have three risk important measure and
we'll see their mathematical model and
all. So, risk important measure
Fussell-Vesely important, then risk
reduction worth important, risk
achievement worth important, and
inspection important measure. This will
not be covering these three. If you
cover, probably by default, you can you
can study yourself on this because of
positive the time I have to skip this,
but keep it there for the
your self-study.
So, Fussell-Vesely important. What is
the important?
Um you we understood the term IIT.
importance of the ith component. Now,
Fussell-Vesely importance is a
superscript that we have over here. And
then is nothing but So, normally these
things they are they talk in terms of
ratio, in terms of this thing also. So,
availability of the ith component divide
by system availability. So simple. You
have fraction of these two. So, you will
get the this important measure. Okay?
But when we when we have the
minimal cut set
from the PRA.
There we have a
importance of the component in terms of
the cut set.
You know,
so
what is what is
minimal cut set plant level along with
the quantification frequency per year
core damage frequency. The first of
basically importance talks in terms of
core damage frequency. So
it is at the
system level
this thing and this is at the highest
component level. IIT
and IST are the frequency of cut set. We
have put the instead of unavailability
we have put frequency of the cut set
plant level for which importance measure
is being actuated and system level
summation of the this thing.
So
we have now there are two more
importance level risk reduction worth
and risk achievement worth. Here we talk
about the risk reduction worth and here
it is manifest in two forms
RRW risk reduction worth ratio mode and
in difference mode.
Analysts they choose different models
for different
So if I talk in terms of
risk reduction worth ratio level then it
is core damage frequency for the plant
and CDF
U that means unavailability for the
component is made zero.
What what does it mean? If I make
unavailability
is equal to zero that means the
component is availability is uh
is available.
Okay?
So so
and that effect given that this one. So
the core damage frequency given that the
subject component it's unavailability is
zero. So that means it is available
completely available. Okay? So we have
this core damage frequency at plant
level and then this one at to go. So,
see see F P CDF
is the frequency of the plant CDF. And
CF BDF due to the CDF frequently
obtained either setting the element I
with zero, that is component is highly
available or reliable. I should use the
word highly available.
And in difference term, we have the same
term, but we have this
uh difference criteria. Similarly, for
risk achievement worth,
we set the component unavailability to
one.
That means component is unavailable.
Completely unavailable. The procedure
involves setting the element to
availability one. So, UIT There it was
zero, that is here it is one, okay? And
this importance measure is is ratio of
plant and then individual components.
Sorry.
Individual components in the plant level
UI is equal to one and then plant
UT normal UT.
Difference model is again same here.
We can see that the effect of risk
increase by by setting the system level
availability to one will result in a
risk increase.
Okay?
So, so
at one time we reduce risk then
another time by setting it to one, it is
unavailable and then we see that it is
increasing risk. So, risk achievement
worth, you know?
And it is useful when we when when we
have a target. If we have a target of 1
into 10 to minus 3 per year in core
damage frequency, then we can achieve by
manipulating our
different unavailability and setting a
target for those particular SSC where
the target comes from here and whatever
by maintenance practice by PHM or
whatever, we try to reduce. So, that
means
Uh, directly we got the list of the
component uh, where uh, from risk
achievement part we have to work more to
reduce their the
uh, unavailability. Okay?
So, overview of week four is we have
seen
uh, different importance measures
uh, for component and then
uh, we discussed role of uh,
uh,
identification prioritization uh, what
what was the traditional approaches?
They were all qualitative in nature and
significance of scheduling SSC. This was
a very important point that we made out.
Even if you identify importance, what
how frequently you do it? Uh, there is a
procedure uh, which is called
surveillance test interval.
Okay? For safety system. And for
operating system, allowable outage time.
How far you can keep it? And they
sometimes form the part of risk
monitors, you know? Because risk
monitors are there to advise you on
maintenance management. So, um,
probably we don't discuss in or in the
my book one of the book you can find
this uh, factors. Thank you.
>> [music]
[music]
[bell]
[music]