Submind YouTube summaries
Thumbnail for Week 4 - Lecture 19 : Sensor Systems for Electronics PHM and Common Cause Failure

Week 4 - Lecture 19 : Sensor Systems for Electronics PHM and Common Cause Failure

Watch on YouTube

Video summary

The lecture focuses on Prognostics and Health Management (PHM) specifically applied to electronic systems, highlighting how this industry is actively adopting these technologies despite them currently being largely at a laboratory level. Unlike massive mechanical systems like diesel generators that require vast resources for monitoring, electronics offer a relatively manageable environment where PHM can be effectively implemented. The primary drivers of degradation in electronic components are environmental factors such as temperature and humidity, which induce thermal fatigue and other severe mechanisms. To counteract these effects and ensure system reliability, the lecture introduces two main classes of monitoring: life consumption monitoring based on physical models and data-driven health monitoring, often integrated with fusion techniques to reduce uncertainty in predicting remaining useful life. A critical aspect discussed is the management of common cause failures within redundant systems, where environmental stressors like humidity or temperature can simultaneously affect multiple channels. To mitigate this risk, electronic safety-critical systems employ redundancy combined with diversity; for instance, using three sensors monitoring the same parameter ensures that a single failure does not compromise system integrity, while diverse systems operate on fundamentally different principles to prevent simultaneous failures from identical causes. Design strategies include physical separation of redundant modules in different locations or elevations to guard against flooding and fire, staggering maintenance schedules to avoid human error affecting all units simultaneously, and utilizing independent power sources for each channel. Additionally, rigorous quality assurance through military-grade specifications (Mil-Std) and continuous periodic testing are essential to detect faults even when safety systems are in passive modes. The lecture concludes by detailing the specific precursor parameters monitored for various electronic components to track degradation mechanisms before failure occurs. For electrolytic capacitors, key indicators include equivalent series resistance increases, leakage current changes, and internal electrolyte levels measurable via X-rays or other methods. Power electronics rely heavily on derating strategies where units are operated below their maximum capacity to enhance safety, while monitoring parameters for Insulated Gate Bipolar Transistors (IGBTs) involve collector-emitter voltage thresholds and thermal resistance. Other components like ceramic chip capacitors and CMOS devices have unique failure signatures such as dissipation factor shifts, RF noise variations, and logic level deviations. Finally, the presentation outlines a Markov model approach to mathematically analyze system reliability by defining states for healthy operation, compromised margins with one failed channel, and total system failure due to either multiple individual failures or common cause events, ultimately aiming to meet strict safety targets like a failure probability of less than 10^-5 per year.
Read the full video transcript
So uh we are in the fourth lecture of uh this week uh that is fourth week and now we are talking about a special application area that is sensor. system uh for uh phm phm of electronics. In fact um um the electronic industry has caught up with this development though not at a uh advanced rate but it is still initiated actually use of uh as I mentioned uh PHM still is at the laboratory level and uh electronics industry is one of the active u uh industry which is trying to use it. uh you know um one more one more reason for electronics is it is u uh comparatively easier to go for phm of electronics u because uh compared to the mega system like uh diesel generators uh uh you know or any uh generating station generator u there are there are huge systems so there resources also are required but In electronics can be managed uh relatively easily though they have at material level they also have a lot of complexity understanding of upgradation mechanism u you know so but then uh our fourth lecture um uh will be on phm of electronics in fact I am uh I am planning to have a phm of electronics and phm of u mechanical engineering so mechanical engineering when we have a separate week for that we'll discuss uh and phm of electronics an overview I'll be give you this in lecture many case studies will be covered in our later uh lectures you know so um as I mentioned um environment in fact environment is even for mega structure it is there actually why because suppose if we have a uh you know humidity and uh Uh and then there there is a structure at the coastal line you know and uh then we have uh along with the humidity the CO2 concentration you know then uh carbide you know the carbon formation at the layer and that is degrading the uh the structure uh strength to the mm level or millimeter level you know uh deep into the concrete um that affect the strength of the but we uh as far as electronics is concerned the temperature, humidity uh apart from quality they are the major contributor to uh electronics and of course thermal fatigue these are the some parameter thermal fatigue is a operation mechanism. Okay. But then environment is a one of the these two parameters that is temperature and humidity they are one of the uh severe degradation mechanism are they electronics are susceptible to these two mechanisms uh maybe more but these two are so uh that we talk about the uh electronics uh often there are two classes uh of monitoring um it is one is life consumption monitoring and then second one is datadriven health monitoring of Of course, there is third one also that is integrated uh fusion technique where we we we integrate uh fusion uh technology with a datadriven uh method. So um so as I said the PF model and datadriven model they they in principle uh tend to reduce the uncertainty in prediction of uh remaining useful life. So that's how it is and for common cause failure uh dedicated sensors are required because even uh humidity can induce um a common cause failure um into into the uh electronics redundant channels redundant and diverse channels you know diverse channels have less probability compared to redundant channel. Uh similarly increase in temperature the electronics uh uh cubicles uh where it is like if we cannot maintain a good um ground benign condition that is uh 22° centigrade temperature and 50% humidity uh the performance of the electronics will uh degrade uh and that will affect our target functions adversely actually and u uh that's why um That's why there is a due consideration uh in the reliability and safety for safety in the design of electronic system. If I have to like we have one sensor and one channel uh to increase the reliability of the complete monitoring signal there will be three sensors. Okay. And though right from sensor to data equation and actu actuation function these three channels will be monitoring the same parameter. If it is temperature, it is only monitoring the same parameter. So u so why? Because if one sensor fails, it should not affect the uh you know uh system adversely. Of course, capacity will come down. Earlier three sensors is monitoring now two sensors are monitoring uh but then uh reliability or safety is not affected but if two sensor fail uh then yes the system should go to the safe mode. Okay. So that's why they they are called having and then if there is a common cause failure then all the three or at least two will go to fail failure mode. So um uh that's why for common cause failure in the panel itself there will be temperature monitoring uh and humidity moni monitoring systems are there. In fact air flow monitoring systems are also there. There should not be uh reduction in air flow. So uh and that is again a process parameter. So that is a kind of uh built-in uh uh safety reliability is uh there in the system especially for a complex system because the uh because the consequences uh um of safety and reliability uh especially safety being the overriding factor uh it is a no no for uh any uh electronic system because our subject is electronic system as of now. So uh what are the potential precursors uh in um in electronic system uh that we have uh see and the component so let's say for our uh electrolytic capacitor I would say um particularly for aluminum electrolytic capacitor let us say let us choose this component and uh then if this is a component what are the precursor that we need to monitor to track the degradation So one is that capacitors alone you know uh reduction in capacitors is an indication that capacitor uh uh so capability is coming down equivalent series resistance you know of the capacitor if that that increases uh that is another parameter. So and it has been demonstrated demonstrated uh you know that if you monitor these parameters you can track the degradation mechanism and that in turn the remaining useful life of the cap leakage current and electrolyte level. In fact these two are also extensively being used. uh in fact I had a very good experience on measuring the level of electrolyte uh through X-rays or whatever or any uh internal buildup uh of you know uh inside the capacitor that also and that reflect in terms of uh uh leakage current or electro uh equivalent series resistance. Okay. Now we have insulator get bipolar transistor IGBT. This is very important component. uh in fact I have listed out only those component which are which are uh relatively having more failure rate compared especially if you talk about the power electronics you know um power electronics is little tricky game micro electronics we have still lot of built-in quality reliability and all but same thing is not true with the power electronics power electronics if you want to have a safe mode derating is the only mode if If you require a unit of 20 kilowatt, you go for 40 kowatt or 50 kowatt so that the components are less stressed. So that is called the derating to increase the reliability or safety of the system. So IGBT collector emitter on voltage get threshold voltage uh transductance collector leakage current uh chip thermal uh the resistance. Uh so these are the things for IGBT and they are very important device um for power electronics as as well as micro electronics ceramic chip capacitor leakage current resistance dissipation factor and RF noise um and and the increase in these parameters we'll know that our ceramic chip capacitor is a giving problem thing though both are capacitor but both have other than the uh uh you can Say leakage current there is no other parameter maybe we can say equivalent series resistance is common here it is resistance here we call equivalent series resist across the capacitor then uh complimentary metal oxide cos uh devices um supply leakage current um supply current variation operating signature and then current noise and then logical level variation uh because all simos devices they have logical operations on the through the transistors on board and these kind of things gets reflected here at logic level variation and then RF noise radio frequency noise and then RF power supply um there are something but somehow I have not given anything because I was not getting pretty confident about those kind of precursor failure but maybe in future classes when I do more research probably for RF power supply we'll I'll give you something Uh then solder joints delamination thermal fatigue these are the two major modes and uh they sometime appear as a crack or you know and the they are monitored by either voltage or current v variation in the circuit at appropriate location cable connector cable and connector change in impedance leakage current insulation embritment loss of fitment. So these are the things in fact um uh cable and connectors they uh when when we talk of couple of years of operation because they are not because of degradation because they are plugged in plugged out uh uh comparatively more frequently they are the cause of uh electronics module uh failure. So sometimes it is replacing certain isolation component or sometimes you have to go for large maintenance replacing the card itself you know and then switch mode power supply uh which is so common then all the computers and all output current and voltage and ripple and efficiency these three similarly field effect transistor also uh I I'm still trying to work on what are the degradation mechanism probably in the uh maybe next lecture uh next lecture mean next week I'll cover when we talk about the realtime application of electronics PHM then criteria for PHM sensor like here we saw for mechanical also but we here we have measurement temperature humidity dust and is for uh mostly they are for environmental monitoring then output is voltage current and uh capacitors these are the things so input output uh we have seen uh and the measurement quality matrix is accuracy again on of sensitivity, precision u because they are miniature components. So precision is very important for them and uh ranges and resolution. These are the things sensor electrical parameters power uh power consumption uh this particular thing that is power consumption uh to have a solution for this. Now the modern sensors they are having a onboard uh power management that there will be external supply but still onboard itself there will be in autonomous mode they will have the power management system uh you know uh to to continue the ability of the uh sensors uh to perform its intended function. Now operating environment uh requirement ground benign industrial so depending on which environment we are operating our robustness should be built accordingly. uh if uh I'm having a control room my instrumentation should room should have uh the ground ban condition if it is a factory the class component and then the component quality itself matters actually if I have military grade component then I can expect a huge good performance you know u but if an industrial setup we have to have those kind of specification otherwise the reliability will not be the same actually in industrial environment. So online sensor monitoring as as I mentioned for common cause failure uh is uh should be there actually typical configuration of electronic channels in a complex safety critical system. Now we have come to a very specific requirement unless until you are able to imagine what is a complex engineering system. So simply I have given the definition of complex engineering system earlier. Um it is safety critical system. Earlier the safety is overriding factor. Then there are uh number of component they are interacting with each other. Uh they are rel much higher compared to a normal system. Human interaction is a is a uh there is a lot of inter human machine interaction uh that that uh helps some decision making and all. So it is there and then there there are redund because it is a complex system single component or single channel failure should not affect my operation the redundancy diversity uh you know and fail safe criteria they should be built. So that's how visualizing a uh not only number of component it's interconnection okay and its operation uh and they are not linear in all uh uh domain uh like a simple component they will have a linear relationship but they they don't have a linear relationship so we have to have that kind of diagnostic mechanism electronic system so redundancy in monitoring and production by so this is one of the characteristic of K out of N that means at least k number of channels or components are required out of n uh for for having a noble operation. Okay. So, so success criteria if failure criteria if two component fails out of n component that there then the system fails that is so it depends how you define the configuration in failure domain or in success domain then diversity uh diversity uh uh diverse system diversity form part of the uh thing diversity means redundancy is one thing similar line at three but diversity means you are having components or channels or system operating on fundamentally two different principle. Okay. So that is called diversity. That means the cause for failure here and here it will not be uh the same phenomena will not be it is thing you know at plant level if I have built two system and another one is diverse system that means it's a operation of mechanism totally totally different from the first principle. Let's say a redundant system is operating by fall of safety devices. A diverse system will be operating not by fall. It will be injection of certain uh certain uh fluid into the system to bring the state of the system to the safe state. So this is called diverse system operating entirely on different principle. Then provision of defense against perceivable and postulated common cause failure. So on one hand we say that common cause failure is uh is is a serious thing and it should be looked into it but that doesn't mean there are provisions are not there in the plant there are provisions made like uh separate separation between two redundant module so that uh if any fire or flooding is there both the modules will not get affected so there is a physical separation I'll I'll I'll show you in some way or the other how we implement this independent source source of power. Suppose if the all the redundant component they are being fed from the same source then common cost failure is the moment that particular source fail. So uh in redundant system the similar items they operate from different source of power actually or different location of power you know. So then shoulding maintenance to reduce the human error. In fact maintenance is staggered maintenance or the philosophy maintenance philosophy they they are such that they say simple in one day all the three redundant system will not be maintained or maintenance activity will be carried out because mistake happening at one can happen at two and three also. So after if you do it in a staggered manner in the sense that uh on one equipment you do it today next next month or next 15 days third one for so same common cause will not affect or same crew or same tools or same thinking process will not affect all the three redundant system location module to guard flooding fire I mean if you if I want to safeguard against common pass I'll be locating equipment in different location not only different location but different elevations also because let's say flooding is there in the lowest location if I have one redundant unit at the other loc top location that that will not fail at least that will work actually and fire there should be separation you know so these are the things and this is how safety philosophies are implemented in complex engineering system imagine we we we had one parameter that is uh CP1 uh primary system. Okay, let us say um first is we are monitoring the environment but because it is electronics and then uh the these are the uh sensors here this is a sensor. So it is a channel A, channel B, channel C. So same parameter is being monitored by three channels. Okay. So then we have a like usual sensor after sensor we have data equation system channel A channel B channel C. So here we are having three channels monitoring the same parameter. So we can say there is a redundancy the this channel is also doing the same job. This channel is also doing the same job and this channel is also doing the same job. advantage is one channel failure doesn't affect my operation because my logic is two out of three control protection logic that means as long as two channels are operating it doesn't get second advantage is I'm able to do maintenance also on one channel without affecting my plant operation so and then safety of course u one one channel failure will not affect our system because two channels are in majority voting two are working properly out of three. So my system output will go and it will not get affected because the moment this channel is B the output will be generated from these two channels. Okay. But we said all the three channels are uh uh same that means they might get affected by a single cause. Let us say um so what what I should do I should have separation between these channels. So this is a physical separation to reduce the chances of common cause failure and this is done in realtime plant also these channels either they are located at different location uh different rooms so that there is a isolation okay so now imagine uh if suppose there is increase in humidity in this uh room let us say this before this partition up uh then only this channel will be affected but if this separation is not there then it will be affect all the channels so that will lead to a common cause. So we are trying to uh to separate out common cause failure. Okay. So this is between among these three channel but the same function uh it is a primary we have a diverse function because suppose if this because of some common cause reason this channel fails. So I I as I mentioned the second safety system is built upon the diverse mode different totally different uh uh you know fundamental principle. Okay. So this is there. So with this common cause this will not be get affected. If there is a common cause here but this is primary. So if there there is a failure here automatic signal will go to take care or it is online always. So these two diverse system monitoring the same parameter are these two system uh the if this fails then this will actuate my control output will go. So there is a red redundancy and there is a diversity for this system and this is how uh systems are built uh for complex engineering system. you you won't find this in uh normal systems or I would say the systems which are not um uh safety oriented or you know u where the safety is the overriding factor you know so this is the kind of caution being design provisions are made uh dur uh in fact uh in many plants you'll find not only three four channels also okay because if one channel fails we totally depending on two two channels. So um so um it is something like you are trying to have more resilience against uh failure. So uh three out of four. Now it depends how you get the your reliability get saturated whether in two out of three or two out of one uh but reliability or safety is the prime importance you know. Now let us do some uh common cost failure mapping. Uh you know uh before that let us understand electronic system quality ensured through specification like American military standard. Okay. Uh if you don't have national standard mil grade component mostly quality is assured through buying a component which is milgrade. So I'm in some sense I'm I'm reducing the chances of common cause failure because procurement has to be done to build a system but if I buy quality system the chances of uh failure or affecting the u reliability of the system is less then testing as part of procurement is also measured. Now uh like you saw pre previous system uh you know uh where yeah u there is one more uh thing which is there systems they are tested by system alone automatic signals are sent and uh they are when there so uh that means there is a continuous testing which goes on into the safety system especially electronic system you will generate a signal at this this end and you send it here and response will come back here. So this complete channel will be monitored that this channel is functioning. This is especially true if this channel channel channel is doing a function of safety in safety uh system they they are in passive mode till the demand comes. So for safety system monitoring of the uh let's say if I give one small impulse and test it the in between component or modules I obtain the desired output and then I'll say it is healthy. If the desired output doesn't come uh this is called periodic testing module here also periodic testing module. So you send a small impulse and test the complete channel. You you have found a beautiful way of checking the availability of the safety channels also when they are not active. Okay. So that's how the safety is ensured uh even for passive mode uh especially safety system because process process system they are on always if any deviation is there you will come to know but in safety system you will not come to know because they are not having active channel of course arrangements are made in the design that some uh some active uh phenomena is happening around here but testing periodically and this periodicity is it is not something like uh minutes or hours or days. It is frau some second one second one pulse will go. So perpetually uh we are monitoring the health of the system in continuous mode actually you know if any any fault is there it will get displayed on the control room or the panel itself. So, so here we are trying to say is that testing the system uh in in institute in the location. Okay. Failure mode effect analysis. This is one more thing there should not be any failure mode left out which did not form part of our coverage. Reliability analysis at the plant level and meeting reliability at risk level. So for all these systems it is very much required in complex engine system that reliability target and safety target should be met. So at system level at the plant level also individual production uh system unavailability uh should be so it's just a quantitive figure for some standard I would have got it so 1 into 10^ - 5 if it is a single channel the failure probability you can imagine should be less than or equal to 1 into 10^ minus 5 per demand it can be around also but it should not deviate too much similar similarly for a complex uh Let's say uh uh nuclear plant the s the failure of this thing should be 10 - 5 per year one. uh now you can com if you compare with other systems probably you'll you'll find that the uh kind of safe uh inbuilt safety uh into the system is so high that we get this probability you know and they all make up from from component to system like electronic system from system to plant and that's how it is reflected here per year 10 to minus 5 even though the plant design ensures resilience please note all of against the common cause failure through various defense mechanisms uh like uh redundancy, diversity, independence, power supply separation, human factors uh consideration, software reliability. Uh in many safety critical systems uh the digital uh production uh is used along with some redundancy or diversity. Why? Because a software failure uh can disable because imagine if the same software on a digital three digital channels are operating any combination of input fault and some some uh failure occurs it will be a common cause failure for all the three. So always one has to guard and you provide additional arrangement either through analog or maybe even other another digital system built on different principles altogether. So that is the thing one has to ensure that when electronics or this thing. So now let us say we are talking about plant we carried out a uh risk assessment of the plant and so many common cause u uh list was there with us. Common cost failure list was there which has come up on top you know importance very high or high rather uh though it is rare but relatively high um so human affect which group this group this group and this group they are called common cause component group so I have just written here I'm not giving the specific details here and group one to group n these are the so many things and for some temperature is they are giving so see here here I'm saying it is not getting affected by temperature some common cause group let us say hardware component they will they will don't get affected by temperature but if it is electronic system yes temperature also will will play a role in some group and then humidity again temperature and humidity high sensitivity for electronics actually flooding so I have not given the complete the this thing I can map it here and what will happen if I map I have a physics of failure approach. I can work on see common cause failure uh is a phenomena defense against that is again a phenomena we have reduced the possibility but at the end of the day I to understand all the common cause failures and through physics of failure approach so that this thing become irrelevant irrelevant I will through robustness through some production at that level or through some separation at that level I will reduce the whatever PF model or root cause analysis model says. So what an elegant mechanism we have here only uh probabistic mode we got this information. Now this input we give it to physics of failure uh consideration or root cause analysis and we find this is the cause uh humidity affects these three things. So what we should do? Do we should we provide any barrier or we should locate in the case or I mean there the sky is open if we understood the cause for failure. If it is a quality related failure then we will check the quality of the component at the end before it goes into the uh plant. If it is a maintenance related issue like maintenance means either human or some parts which have not gone wrong and which is potential for common cause failure and there's some experience is there if it is there then we take action institutional failure electromagnetic electromagnetic is a phenomena which can affect the system in adverse manner and it doesn't have to intro introduce the system and uh so one has to um so in common cause failure the in the configuration management itself these things are ruled out you know the the signals if it is there it will get attunated and the effect will be negligible or known uh failure so fault tolerant approach for electron design and all so I'm just giving you one example how common cause failures are handled so marcom model so is a popular approach and u and then I am making this two out of three u failure I'm do going that means handling the all the channels in for you saw the three channels they are operating and all that so two out of three I consider when the two channel out of three channel fail I call it a failure so suppose if I have three channels and then the assumption is protection and control system comprise of three redundant channel okay the failure uh criteria is two out of three two channel out of three channel or only one repair station. Okay. And I have assumed here one more thing that the failure rates are say average failure rate I am applying for you for simplifying my model. So in marco there are you'll understand when you see the next slide actually the there are one P1 P2 P3 poor condition P1 condition means all three out of three they are operating and system is healthy. In P2 one out of three failed. So that means we have two operating system it is failed but compromised margins are reduced and P3 is two out of three channel failure system failure and three out of three failure system. So that means this three out of three failure means there is something common cause. So common cause and two out of three I declare the system to be failed. Okay. So if I have to model this system how what I do? So as I mentioned P1 is the healthy state, P2 is the um P2 is the compromised state. Uh okay uh and I have so let us see P1 P2 are operating states okay because even if one component failure we have operating state and uh there are three channels so there can be three any of the channel can fail. So three mode and then P2 one channel is repaired we can only repair one channel. So that's why one mu is one only. Uh P3 is a failed state. So I have convert concentrated them into three state module P1, P2 and P3 because common cause also takes system to failed state and two out of three also to takes to the to the failed state. So the P3 is called as a failed state. Exponential distribution V used uh and all the channels have same failure rate that is average failure rate. MU is repair rate and lambda C is common cause failure rate. So this information and we saw how the three channels they operate and what is the failure rate two out of three. So I'll write a state specific equation uh probability of being in state one and if rate of change of migrating from that state to the other state. So dp 1 by dt that is probability of change in probability of this state is minus 3 lambda. So minus that mean it is departing from this state. So minus 3 lambda plus mu returning to this state. So this is the uh system we have. So first equation of the system that is they are called maron equation or state space equations dp2 by dt. So dp2 by dt for this state this is a outgoing state state that is and this is again outgoing and this is incoming. So 3 lambda will be plus incoming state and minus 2 lambda minus mu minus uh 2 lambda minus mu. So that will define uh rate of change of state uh dp2 by uh dt and this one and third one also dp3 by dt. So the three state means here these two states are joining it is called absorbing state you know dp3 by d2 2 lambda plus 3 lambda c 2 lambda plus common cause failure three lambda c actually this should be uh lambda c only there should not be three lambda c because we are saying common cause failure all the three channels have failed okay so here this three will go out please correct in your this thing this equation can be solved for arriving at probability of system failure rate. I have written the differential equation. You can use any method um algebraic method or you can say uh you can these are very simple equation you can say and find out the probability of uh system being in P1, P2 and P3. Our interest will be failed state. This P3 state so that we can find out. Okay. You can use matrix method also you can use algebraic methods also uh any method you can use and solve it. Okay. So week four electronics phm. We have seen a interesting uh interesting uh features of uh redundant system diverse system. We have seen precursor parameter how they are um for which the mode what is the precursor parameter for for which uh sensor what are the um uh degradation attribute we should monitor role of common cost failure in complex engineering system sensor failure and common cost failure. Why envir environmental modeling and finally the common uh a marco model uh to account for failure repair and common cause failure. So probably you would have understood now by this time that it is not only a electronic component but it is basically a channel that means starting from sensor to final actuations you can monitor and their combination uh and then common cause failure what effect and human factor what effect it has got probably it was pretty clear in this now I'll touch upon this electronics PM when I have dedicated week for uh electronics prognostics and health management because there are many subtle aspect that we need to discuss uh over there as part of the system. Thank you.