Submind YouTube summaries
Thumbnail for MLFPM Symposium 2022: Bowen Fan

MLFPM Symposium 2022: Bowen Fan

Watch on YouTube

Video summary

Bowen Fan from ETH Zurich presented his research on predicting recovery from multiple organ dysfunction syndrome (MODS) in pediatric sepsis patients, highlighting the critical need for early intervention. Sepsis is a life-threatening condition caused by the body's response to infection, which can rapidly progress to severe organ damage and death; notably, it disproportionately affects children under five years old. MODS occurs when two or more organs fail concurrently, significantly increasing mortality rates compared to standard sepsis cases. The primary goal of this study was to develop a machine learning framework capable of predicting whether a patient would recover to a state with zero or single organ dysfunction within one week, thereby enabling clinicians to provide personalized care and timely interventions for those at risk of poor outcomes. The project utilized data from the Swiss Pediatric Sepsis Study (SPSS), a retrospective national multi-center database involving ten major hospitals, alongside an external validation dataset from Robert Children's Hospital of Chicago to ensure generalizability. The study formulated recovery prediction as a binary classification task using daily electronic health records collected up to six days after sepsis onset. To avoid data leakage, only variables available on the first day of admission were used, including vital signs, laboratory results, clinical scores, and demographic information. Various machine learning models were tested, with Random Forest emerging as the top performer due to its superior accuracy and ability to generalize across different clinical settings compared to baseline clinical rules and other algorithms like logistic regression or support vector machines. A key finding of the research was that cardiovascular and respiratory systems are the most critical factors in predicting recovery, as evidenced by the top ten features identified by the model, which were heavily associated with these organ systems. The study demonstrated that the Random Forest model could accurately predict recovery status with high precision and recall, even when applied to unseen data from a different hospital. Furthermore, the model's interpretability was enhanced using SHAP values, which allowed clinicians to understand specific risk factors, such as low oxygen saturation or high lactate levels, that drive the prediction. The presenter emphasized that while the model is not perfect, it is preferable to generate false negatives rather than false positives, ensuring that patients predicted to recover are still monitored closely to prevent unexpected deterioration. In conclusion, this work represents a significant step toward personalized prognostics in pediatric sepsis by creating a robust, interpretable machine learning framework that performs well across different institutions. The presenter noted that while the current dataset size limits the ability to stratify results by bacterial strain or extend the prediction horizon beyond one week, future efforts will focus on expanding data collection and potentially utilizing federated learning to collaborate without sharing sensitive patient data. By integrating genomic and proteomic profiles in subsequent studies, the team aims to discover novel clinical phenotypes and develop even more effective personalized treatment strategies for children suffering from sepsis.
Read the full video transcript
so we continue with more Talks by our doctoral students the esrs the next one is Bowen fan from eth Zurich I'm his advisor and Bowen will talk about his work on predicting recovery from multiple organ dysfunction in Pediatrics access patients go on the philosophers thanks thanks a lot for the introduction and as cousin said my name is R2 and from the Department of biosystem Science and Engineering engineering from adher Strike I started in the first March of 2020 and and this work is one of our recent work that has been published in ISB 2020 this year okay so first of all I would like to say thank you for all these collaborators as this project is a joint project between adhdrick and also some clinical Institute that's interesting oh no sorry so University Hospital offer and NR Robert Children's Hospital of Chicago and also University Children's Hospital of Zurich and it's not possible to get this work done without their food spot so I would like to say thank you here so there are two main focuses of our project the first one is pediatric sepsis and I would like to give your proper definition of water sepsis is and sepsis is a life-threatening organized function caused by this regulated to the host response to infection and this sepsis can progress rapidly that leads to severe organ damage in a test and also on the right hand side you can see a factory of surface which is from the word sepsis Day event and function you can see uh there's a 47 to 50 million cases per year and at least 11 million deaths per year and in one in five thousands very well is associated with sepsis and also sepsis is number one cause of deaths in the hospital and number one cause of hospital readmission and also number one cause of the health care and up to 50 of the sepsis survivors suffered from long-term fiscal or psychological effects and also what's even worse is that sepsis is disproportionately affects children as uh 40 of the cases are children under five years so so that's why pediatric sepsis definitely works our additional attention and the second Focus I would like to say is that the multiple organ dysfunction syndrome shortest mods and mods is defined as two or more concurrent organ decisions dysfunctions and mods is highly likely to be developed from sepsis in pediatric patients and comparing to regular sepsis cases a moss has even increasing mobility and mortality and here in this study we use the organized function definition from the international pediatric steps consensus conference and this definition includes six different uh major organ systems in respiratory cardiovascular Central Nervous renal haptic and hemphological systems and for those patients who accepted mods at this episode that is much less likely for them to recover to a rather myostay like zero or single organized function so even those through treatment and even for those who survive the Mars they may still suffer from lifelong consequences to patients to themselves and also their families so here in a study for those patients who were already in Mars where crying sepsis is important to investigate their chance to recover to buy the better State like zero single organization and here this is also this uh the main motivation behind our study that we do early prediction of the recovery from Mars and its pediatric patients with sepsis and hopefully with this work we could provide the clinicians the opportunity to to take extra care and also enable Tamil interventions and so here we imply some machine learning models for this prediction and although there are already many studies that have focused on the the early recognition of sepsis but this uh prediction or reproduction of the most recovery and services patient has never been attempted with machine learning and also the current standard of care in clinic practice is still rule-based so with this effort with this work we hopefully develop something framework which product could allow for a personalized prognostics for those patients with pediatric patients with sepsis so this works mainly developed on the Swiss pediatric study showed us SPSS and this SPSS is a retrospective National and multi-centered core study including children up to 17 years old with blood blood culture proven bacterial infection and involves 10 major pediatric hospitals in Switzerland and this SPSS database includes electronic health records at a daily level which keep track of the most abnormal measurements within every 24 hours and these records present up to six days of a broad culture sampling or we get the day of sepsis at the day of the reception set and a recent systematic review shows that for the early prediction of sepsis more than 85 percent of study only validate their their models on one single database using a cross-validation scheme and even though their models performed quite well in their own settings the performance will never really verified in the other database and also for this clinical research external validation is extremely important for the uh for assessing this model's generalizability so that's why we managed to assess to another Pediatric Services patient database from NN proper education Children's Hospital of Chicago showed as lchc here that's for the purpose of external validation so let's have a proper proper definition of this mods recovery so we formulate this as a binary classification task that on the first day of predict the patient using the data on uh from the first day of stepson said with a Time Horizon of seven days here I show a few examples of the most recovery and day zero here is indicates the day of sepsis onset and we have the data until day six so there's a one week and we do the prediction on day six and here the blue arrow indicates the mod State and red arrow indicator mod free state so the first one is the recovery example that the patient was in the mod was the mods and after day three he or she got better and eventually on day six uh he was your smart free and the second one as a now recovery example that's the patient starting with mouse and got better after day two by eventually turning into Mods again and the third one was excluded because the mod the page email the patient started with mods free so that's that's not what we're very considered in the study and also one thing to notice is that uh if the patient deceased before they said we still disease before day six we still consider the patient is a non-recovery so after removing those invalid samples for the SPSS data set we have a non-recovery samples of 138 and Recovery samples of 118 that translates to a positive class prevalence of 46.1 percent and for the other external validation cell we have now recovery example of 210 and Recovery example of 181 this has a very similar positive Club class prevalence and also the color size is more or less similar in the second augural order of magnitude and with the help of the clinicians we also collected some relevant variables for this prediction task that includes the vital science laboratory test results clinical scores chronic disorder information and also demographics there's 44 of them in total and the data are only collected on the first day of service as long said that without the risk of future information leakage and here is the result here we for the internal validation that we implemented and compared several machine learning models including our logistic regressions for the vector machine random Forest multi-layer perception and librarian posting machine and also we implement the clinical Baseline model that using the decision tree which with the second pediatric logistic organ dysfunction score which this score has been verified to be a vertical indicator of modality and also organized function in large pediatric patient database so we also adopted a nested cross validation scheme that would split the data in your training validation and test it and we repeat this for independent rounds and with the model development and Hyper parameter tuning was done on the training application set and we only report the performance on tests only and on the right hand side you can see the internal validation results here we use two metrics area on the rockef and area on the Precision curve so all five models the five main machine learning models performs comparably well and all of them up from the clinical Baseline model by a large margin and what the winner of these five model is actually the render forest model where with a a rock of 79.1 percent Au Rock of 73.6 percent that's why we chose the random force model as our main model and for the later external validation and also the render forest model is uh highly non-linear model and 10 degree out of it and not doesn't really generalize very well so it's particularly important to do an external validation for this random force model here then we do this external validation actually for the completeness of the study we did this validation uh in both centers and also tested on the other so here you see two plots and the red curves with the uncertain events are the internal validation results of for the 10 10 folder cross validation and the blue curves are the external radiation results that means the model was trained on one Center and tested on the other so basically we can see that the generalization of this random forest model is quite good successful in both directions with there is a every round the rock curve uh over 75 for all the cases and also the area around the Precision curve is over 70 percent with the Positive class prevalence of around 46 percent and also if we look at the high record region whereas the of more clinical interest that most of the uh the event can be captured can be captured uh for example if you look at Rico at 80 that means uh four out of five recovery events can be correctly predicted and we still have a Precision score around 60 60 that means for every five uh predictions the model made okay three of them are attract so we see this render forest model here have both High uh prediction performance and also High generalizability in terms of this uh oh in terms of this most recovery prediction task next and later we also analysis whether the render forest model can predict the most recovery earlier than one week like one one to five days in advance so here we show the result of the relative Au PRC area on the Precision recoil curve uh for different time points the relatively uprc is the absolute aoprc normalized by the positive class prevalence on each date using this Matrix we could have a rather fair comparison of different time points so the we developed a random forest model on the SPSS data set and evaluate them on both so he's the director for the SPSS data set the performance kept going down as we increasing the the time Horizon and for this lchc data sets external validation we see the performance goes up a bit in the first two days and then drops from the day three Awards and on day six uh both external validation and internal validation are achieved very similar performance and also this has a it's something we will be expected as for a shorter Horizon we should have a better performance and also for uh for the temporal limitation of data sets we could only do the prediction for the first week but for like longer Horizon like two or three weeks we could probably just expect a bit lower performance here and but not here multi and uh pediatric sepsis patients they are mostly the highly Dynamics disease that usually with the progression or the recovery scene within very few uh for the first few days of the admission so and also the Persistence of Mouse uh for the for a whole week is already a better sign that is highly associated with high mortality and Mobility and also we want to have a model that is have a both High prediction performance and also high generalizability so that's why we chose day six as our main focus and also another thing we decided is the high interpretability of the model so that's why we chose to use this shape values to to explain the risk factors discovered by this random forest model here uh I already lost count how many times you have seen this so hopefully you know get too tired of that so anyway so on the left hand side you see a this one plot showing the top 10 features here and on the right hand side you see the full names of these two uh 10 top 10 features and Roxas is the top 10 the full names of them and how do you interpret this this one plot that basically each dot is represent uh represents uh each uh patient and also the Red Dot indicates a higher value of the feature and blow down indicated lower value of this feature and the closer of the dot to the right hand side of the x-axis that means the feature is driving the model to give you a positive prediction and the closer to the left hand side that means the feature is driving the model to give you a negative prediction so for example if we look at the top one the lowest oxygen saturation level that means that here there's a lot of red dots here that means that the patients were with high oxygen saturated saturation level then the model tends to assign them high chance to for the recovery well for the second one the lactate and we see this the the red also are located on the left-hand side on the on the axis x-axis and usually that means heightened likely values you're already indicating this organ damage going on so that's why the model thinks these patients are not going to recover from us and for uh for this problem what we found interesting is that we found two organ systems the cardiovascular and respiratory systems are critical for this most recovery prediction because we find this top 10 features most of these top 10 features are more or less associated with these two organ systems so in clinical practice we could probably pay more attention to most patients with these two type of types of other dysfunction rather than treating an equally compared to patients which there are other types of organizations so finally I would like to summarize my presentation here in this study we developed a machine learning framework for predicting most recovery in Pediatric Services patients for one week in advance and we conduct a comprehensive experiment to show that the proposed model can not only just predict most recovery with hierarchy but also can be transferable across different clinical sites to those unseen patient data even in inter Continental setup and also the prediction of the model is interpretable from a clinical point of view so this means the model we developed here could provide a more insights to the clinicians or maybe make they are more understandable for them and also we believe that our model have certain clinical utilities that it could probably assist clinicians for better patient assessment and triage on Day Zero the day of sepsis Stone set and now we are doing something else using the same data set that we including the genomics and the proteomics profiles of this patient and we try to discover novel clinical phenotypes of pediatric sepsis and by doing a characterization of the Asian cyber groups for hopefully with this effort we could develop something for better personalized treatment and that's all for my presentation if you want to have a further information of this paper you can scan the QR code here thanks a lot I'm ready [Applause] thank you Bowen for the questions we're going to be asked foreign for the great talk uh one question regarding the clinical interpretability that you mentioned in the end so uh how would this look like is that these these graphs that you show or is there more behind it so how would a clinician interpreter results yeah I think this will be probably helped to help to support their decisions like patients with like for example low oxygen saturation or high lactic value they should pay more attention to that or maybe all right but it's probably if you have a certain uh patient right it's um most probably the case that there will be certain points which will be on the blue and that's on the red right so otherwise it would be a clear decision so if I I get it if I see a picture and it's it's always in the red then it's clear and I can see also as a clinician um yeah how the system comes to it yeah we also actually provide some examples for like how model how the model like make the prediction for each individual samples but I'm not here here it's in the paper already so it's like a survival score for each individual patient now there will be much easier interpretable as in not in the cover level but individual level did you have many more features then you said these are the top 10 features that is top then we have 44 features in order okay and um did you also like quantify like how much each feature kind of contributes to the solution the the do you say like with these 10 features I capture pretty much I don't know 99 percent or oh yeah we actually I think we didn't really qualify like how much exactly this top 10 contribute to the model District uh the prediction now it's a physical point if we could probably include this in future work but but you said it's the first two right that have the I only talk about this first two here to just to Showcase like how this model like take uh these features into account when they're making prediction overall if you look at the resulting performance of the system right you could already show that basically it doesn't matter which machine learning message you use but it's always better than the clinical status right but do you think there's a way to push the performance much further because there was not at least not from the first view that was not so significant differences between the individual methods right yeah that's right that's right and performance is well it's below 0.8 right so it's it's it's not super reliable this is yeah do you think there's a way to push it further or I think yeah of course for machine learning models the primary way will always be to have more data since for this uh Pediatric Services agents we only have like a cover size of less than 300 which is a pretty little for this data driven approaches yeah probably if possible equally includes more data but it could be very difficult or we do a more exhausted research or could be also this could lead to some overfitting issues okay thanks a lot Leslie thank you really talk um I have a question on Slide 11. could you explain again the so you trained on the SPSS and could you explain again why the the Chicago data for the relative EU PRC first increases and then decreases yeah or what's the interpretation of that I think um it's a little bit difficult to interpret this or it's kind of a mystery of why that phenomenon happens yeah we didn't really try to interpret this at the beginning so just because there's external validation then we could already access to the data and just that pass the model and let it run so if anyone could give a proper guess of what is happening it's a interesting artifact then okay thank you hmm bonus one another question um so we always took for granted in this project that the the sepsis had been confirmed as a bacterial infection is this actually known at the point of time in which we would make our prediction so in other words there's an inclusion Criterion for this data set which where it's for me not totally clear that this is like given or determined is the right word that this is determined at the time at which we we make this prediction so the is this the bacterial infection we know at which time points this is typically confirmed this is maybe three days into the stay or something yeah that's the the things that for this data set is actually that we use the blood contraction proven uh proven bacterial infection that we do the black culturally sampling and we once we found the bacteria we make this as a day exactly and that's the inclusion Criterion of this this Pediatric subsystem um but I think an open point is when exactly is this determined is this maybe determined after we make our prediction because then there would be another like latent set of patients on which we could also make our predictions but which do not fulfill this inclusion criteria I think that would be an extremely interesting external validation question to test the model on patients who are at risk for developing sepsis for example and see whether your model will predict them to recover basically because the recovery likelihood will be higher I guess because it's not uncertained yet whether they have sepsis or not so other data sets like that available this is the question Maybe so they have our National Data stream for Pediatric research in in Switzerland in fact and headed by one of our collaborators on on the paper so they are going to collect data sets like that like in intensive care data sets for for children for Pediatric patients and there you could do this I mean what what this data here represents is a very clean data set in the sense that that the sepsis has really been confirmed to be connected to a bacterial infection here on these data sets it's not not necessarily the case for all things or for all patients that are labeled aseptic in the in a database yeah and then then exactly the situation how how much the system generalized if you train on a high quality expenses to obtain labels yeah um I was just wondering if this Swiss pediatric sepsis data set is public ah so far it's not really for some privacy issues okay and also the externalization is also not accessible to public we wished even not not to us maybe just an additional command what you said is extremely important because if we think about the translational perspective of your modeling efforts these models tend to be applied as the as we go to more poorly defined samples and therefore I think it's crucially important to establish this right away because to determine what the operational window of these models is another idea that came to my mind is have you looked into decision curve analysis to establish the value of your model so basically doing that benefit analysis how much do a clinicians gain or a patient gain from a certain prediction so by taking in the the costs and the opportunities basically of a prediction yeah I think for this one it's actually a retrospective study so it's super difficulty whether to really compare against like a clinical condition decisions so hopefully we can do some prospective study like really put this into clinical practice and we see really compared to like the true decision made by the extensions as this clinical Baseline still sounds in a proxy to those hey uh thanks a lot Pauline it was a great presentation I was just wondering um because you were talking about how you had to basically hand in the model and get the predictions back and you couldn't have access to the data when doing this external validation could it be could this be a nice use case for something like Federated learning or swarm learning that we had presenting in the network like early on like uh maybe using something like that you can collect data from multiple sites and yeah I think it is a very good point and even in this case it was extremely difficult for us to like to communicate and everything for the model development also evaluation if we could like do this anonymously using a federal learning I think they'll be very useful and for also for the other collaborations in medical field thanks thanks thank you a lot but if I may add this project was interesting in the sense that for us this was the really the first example of what you could call Federated learning so we sent the models to yeah yes and they ran the model there so so in all our other projects we negotiate data access at some point we get several data sets and then we harmonize them and then we run it this this was really done decentralized so we only sent the code we never got the inverse data from this project who was first I'll start in the back yes um do you know if there is any stratification in yours in your sample especially regarding the kind of bacterial strain that causes Celsius I'm sorry I didn't really get it um do you know if there is maybe an effect like like a strain specific effect in terms of like the features that you can that you get from your patient that could have an impact on the performance of your model in the end maybe I can help you to interpret it so whether we could stratify foreign limited cover size we didn't do this eventually for the subgroup analysis or any stratification of patient or based on whatever the side infection or the pathogens so no no oh no we didn't use those as feature because this may not be available for this external validation set I'm coordinating the adult sepsis study in Switzerland and we looked into this point very much but the case numbers are too small to to stratify for the type of bacterial in fact for the pathogen that causes sepsis so in in a like realistic time Horizon so like a bigger like collection of areas that is needed in just a Swiss ball Juliano with Mike so I mean this phone said the the size of the data set did not allow to stratify but what we did too is like look for enrichments of certain like statistical enrichment of certain strains in for example the like false positives or false negatives just to see whether for some strange the the model was like struggling more than for others but we didn't find anything like exceptional thank you and final question by Giovanni thank you boys for the talk I have a small curiosity also in addition to my question apologies if I missed it but why is the rate of Subs is so much higher in children like is it because of the immune system that's not really developed more reassistant and also they're more vulnerable in general thank you um so my question is since somebody has mentioned the model is not like 100 reliable um where is the situation in which you will you would get some predictions wrong and in the clinical setting not all errors are made equal in this case would it be worse to get like a false positive or a false negative like what would be the consequences of both and I think yeah in this case I think that's a very good question in this case between the recovery protection not like with your immortality prediction so I would say it's it's better to give a false negative so we consider it was still like pay attention these negative examples that I would still pay attention to this patients because the models think that speech will now recover but hopefully the patient eventually recovers they will be good but if not we can still have more additional attention on it thank you thanks a lot good thank you again Bowen and we move on to the next speaker [Applause]