Video summary
Bowen Fan from ETH Zurich presented his research on predicting recovery from multiple organ dysfunction syndrome (MODS) in pediatric sepsis patients, highlighting the critical need for early intervention. Sepsis is a life-threatening condition caused by the body's response to infection, which can rapidly progress to severe organ damage and death; notably, it disproportionately affects children under five years old. MODS occurs when two or more organs fail concurrently, significantly increasing mortality rates compared to standard sepsis cases. The primary goal of this study was to develop a machine learning framework capable of predicting whether a patient would recover to a state with zero or single organ dysfunction within one week, thereby enabling clinicians to provide personalized care and timely interventions for those at risk of poor outcomes.
The project utilized data from the Swiss Pediatric Sepsis Study (SPSS), a retrospective national multi-center database involving ten major hospitals, alongside an external validation dataset from Robert Children's Hospital of Chicago to ensure generalizability. The study formulated recovery prediction as a binary classification task using daily electronic health records collected up to six days after sepsis onset. To avoid data leakage, only variables available on the first day of admission were used, including vital signs, laboratory results, clinical scores, and demographic information. Various machine learning models were tested, with Random Forest emerging as the top performer due to its superior accuracy and ability to generalize across different clinical settings compared to baseline clinical rules and other algorithms like logistic regression or support vector machines.
A key finding of the research was that cardiovascular and respiratory systems are the most critical factors in predicting recovery, as evidenced by the top ten features identified by the model, which were heavily associated with these organ systems. The study demonstrated that the Random Forest model could accurately predict recovery status with high precision and recall, even when applied to unseen data from a different hospital. Furthermore, the model's interpretability was enhanced using SHAP values, which allowed clinicians to understand specific risk factors, such as low oxygen saturation or high lactate levels, that drive the prediction. The presenter emphasized that while the model is not perfect, it is preferable to generate false negatives rather than false positives, ensuring that patients predicted to recover are still monitored closely to prevent unexpected deterioration.
In conclusion, this work represents a significant step toward personalized prognostics in pediatric sepsis by creating a robust, interpretable machine learning framework that performs well across different institutions. The presenter noted that while the current dataset size limits the ability to stratify results by bacterial strain or extend the prediction horizon beyond one week, future efforts will focus on expanding data collection and potentially utilizing federated learning to collaborate without sharing sensitive patient data. By integrating genomic and proteomic profiles in subsequent studies, the team aims to discover novel clinical phenotypes and develop even more effective personalized treatment strategies for children suffering from sepsis.
Read the full video transcript
so we continue with more Talks by our
doctoral students the esrs the next one
is Bowen fan from eth Zurich I'm his
advisor and Bowen will talk about his
work on predicting recovery from
multiple organ dysfunction in Pediatrics
access patients go on the philosophers
thanks thanks a lot for the introduction
and as cousin said my name is
R2 and from the Department of biosystem
Science and Engineering engineering from
adher Strike I started in the first
March of 2020 and
and this work is one of our recent work
that has been published in ISB 2020 this
year
okay so first of all I would like to say
thank you for all these collaborators as
this project is a joint project between
adhdrick and also some clinical
Institute that's interesting oh no sorry
so University Hospital offer and NR
Robert Children's Hospital of Chicago
and also University Children's Hospital
of Zurich and it's not possible to get
this work done without their food spot
so I would like to say thank you here
so there are two main focuses of our
project
the first one is pediatric sepsis and I
would like to give your proper
definition of water sepsis is and sepsis
is a life-threatening organized function
caused by this regulated to the host
response to infection and this sepsis
can progress rapidly that leads to
severe organ damage in a test
and also on the right hand side you can
see a factory of surface which is from
the word sepsis Day event and function
you can see uh there's a 47 to 50
million cases per year and at least 11
million deaths per year and in one in
five thousands very well is associated
with sepsis and also sepsis is number
one cause of deaths in the hospital and
number one cause of hospital readmission
and also number one cause of the health
care
and up to 50 of the sepsis survivors
suffered from long-term fiscal or
psychological effects
and also what's even worse is that
sepsis is disproportionately affects
children as uh 40 of the cases are
children under five years so so that's
why pediatric sepsis definitely works
our additional attention
and the second Focus I would like to say
is that the multiple organ dysfunction
syndrome shortest mods and mods is
defined as two or more concurrent organ
decisions dysfunctions and mods is
highly likely to be developed from
sepsis in pediatric patients and
comparing to regular sepsis cases a moss
has even increasing mobility and
mortality
and here in this study we use the
organized function definition from the
international pediatric steps consensus
conference and this definition includes
six different uh major organ systems in
respiratory cardiovascular Central
Nervous renal haptic and hemphological
systems
and for those patients who accepted mods
at this episode that is much less likely
for them to recover to a rather myostay
like zero or single organized function
so even those through treatment and even
for those who survive the Mars they may
still suffer from lifelong consequences
to patients to themselves and also their
families
so here in a study for those patients
who were already in Mars where crying
sepsis is important to investigate their
chance to recover to buy the better
State like zero single organization and
here this is also this uh the main
motivation behind our study that we do
early prediction of the recovery from
Mars and its pediatric patients with
sepsis and hopefully with this work we
could provide the clinicians the
opportunity to to take extra care and
also enable Tamil interventions
and so here we imply some machine
learning models for this prediction
and although there are already many
studies that have focused on the the
early recognition of sepsis but this uh
prediction or reproduction of the most
recovery and services patient has never
been attempted with machine learning and
also the current standard of care in
clinic practice is still rule-based so
with this effort with this work we
hopefully develop something framework
which product could allow for a
personalized prognostics for those
patients with pediatric patients with
sepsis
so this works mainly developed on the
Swiss pediatric study showed us SPSS
and this SPSS is a retrospective
National and multi-centered core study
including children up to 17 years old
with blood blood culture proven
bacterial infection and involves 10
major pediatric hospitals in Switzerland
and this SPSS database includes
electronic health records at a daily
level
which keep track of the most abnormal
measurements within every 24 hours
and these records present up to six days
of a broad culture sampling or we get
the day of sepsis at the day of the
reception set
and a recent systematic review shows
that
for the early prediction of sepsis more
than 85 percent of study only validate
their their models on one single
database using a cross-validation scheme
and even though their models performed
quite well in their own settings the
performance will never really verified
in the other database
and also for this clinical research
external validation is extremely
important for the uh for assessing this
model's generalizability so that's why
we managed to assess to another
Pediatric Services patient database from
NN proper education Children's Hospital
of Chicago showed as lchc here
that's for the purpose of external
validation
so let's have a proper proper definition
of this mods recovery so we formulate
this as a binary classification task
that on the first day of predict the
patient using the data on uh from the
first day of stepson said
with a Time Horizon of seven days
here I show a few examples of the most
recovery and day zero here is indicates
the day of sepsis onset and we have the
data until day six
so there's a one week and we do the
prediction on day six
and here the blue arrow indicates the
mod State and red arrow indicator mod
free state
so the first one is the recovery example
that the patient was in the mod was the
mods and after day three he or she got
better and eventually on day six uh he
was your smart free and the second one
as a now recovery example that's the
patient starting with mouse and got
better after day two by eventually
turning into Mods again
and the third one was excluded because
the mod the page email the patient
started with mods free so that's that's
not what we're very considered in the
study and also one thing to notice is
that uh if the patient deceased before
they said we still disease before day
six we still consider the patient is a
non-recovery
so after removing those invalid samples
for the
SPSS data set we have a non-recovery
samples of
138 and Recovery samples of 118 that
translates to a positive class
prevalence of 46.1 percent
and for the other external validation
cell we have now recovery example of 210
and Recovery example of 181 this has a
very similar positive Club class
prevalence and also the color size is
more or less similar in the second
augural order of magnitude
and with the help of the clinicians we
also collected some relevant variables
for this prediction task that includes
the vital science laboratory test
results clinical scores chronic disorder
information and also demographics
there's 44 of them in total and the data
are only collected on the first day of
service as long said that without the
risk of future information leakage
and here is the result here we for the
internal validation that we implemented
and compared several machine learning
models
including our logistic regressions for
the vector machine random Forest
multi-layer perception and librarian
posting machine and also we implement
the clinical Baseline model that using
the decision tree which with the second
pediatric logistic organ dysfunction
score which this score has been verified
to be
a vertical indicator of modality and
also organized function in large
pediatric patient database
so
we also adopted a nested cross
validation scheme that would split the
data in your training validation and
test it and we repeat this for
independent rounds
and with the model development and Hyper
parameter tuning was done on the
training application set and we only
report the performance on tests only
and on the right hand side you can see
the internal validation results here we
use two metrics area on the rockef and
area on the Precision curve
so all five models the five main machine
learning models performs comparably well
and all of them up from the clinical
Baseline model by a large margin and
what the winner of these five model is
actually the render forest model where
with a
a rock of 79.1 percent Au Rock of
73.6 percent
that's why we chose the random force
model as our main model and for the
later external validation
and also the render forest model is uh
highly non-linear model and 10 degree
out of it and not doesn't really
generalize very well so it's
particularly important to do an external
validation for this random force model
here
then we do this external validation
actually for the completeness of the
study we did this validation uh
in both centers and also tested on the
other
so here you see two plots
and the red curves with the uncertain
events are the internal validation
results of for the 10 10 folder cross
validation
and the blue curves are the external
radiation results that means the model
was trained on one Center and tested on
the other
so basically we can see that the
generalization of this random forest
model is quite good successful in both
directions with there is a every round
the rock curve uh over 75 for all the
cases and also the area around the
Precision curve is over 70 percent
with the Positive class prevalence of
around 46 percent
and also if we look at the high record
region whereas the of more clinical
interest
that most of the uh the event can be
captured
can be captured uh
for example if you look at Rico at 80
that means uh four out of five recovery
events can be correctly predicted and we
still have a Precision score around 60
60 that means for every five uh
predictions the model made
okay three of them are attract
so we see this render forest model here
have both High uh prediction performance
and also High generalizability
in terms of this uh
oh in terms of this most recovery
prediction task
next
and later we also analysis whether the
render forest model can predict the most
recovery earlier than one week like one
one to five days in advance
so here we show the result of the
relative Au PRC area on the Precision
recoil curve uh for different time
points the relatively uprc is the
absolute aoprc normalized by the
positive class prevalence on each date
using this Matrix we could have a rather
fair comparison of different time points
so the we developed a random forest
model on the SPSS data set and evaluate
them on both
so he's the director for the SPSS data
set the performance kept going down as
we increasing the the time Horizon and
for this lchc data sets external
validation we see the performance goes
up a bit in the first two days and then
drops
from the day three Awards
and on day six uh both external
validation and internal validation are
achieved very similar performance
and also this has a it's something we
will be expected as for a shorter
Horizon we should have a better
performance
and also for uh for the temporal
limitation of data sets we could only do
the prediction for the first week but
for like longer Horizon like two or
three weeks we could probably just
expect a bit lower performance here
and but not here multi and uh pediatric
sepsis patients they are mostly the
highly Dynamics disease that usually
with the progression or the recovery
scene within very few uh for the first
few days of the admission
so and also the Persistence of Mouse uh
for the for a whole week is already a
better sign that is highly associated
with high mortality and Mobility
and also we want to have a model that is
have a both High prediction performance
and also high generalizability so that's
why we chose day six as our main focus
and also another thing we decided is the
high interpretability of the model so
that's why we chose to use this shape
values to to explain the risk factors
discovered by this random forest model
here
uh I already lost count how many times
you have seen this so hopefully you know
get too tired of that
so anyway so on the left hand side
you see a this one plot showing the top
10 features here
and on the right hand side you see the
full names of these two uh 10 top 10
features and Roxas is the top 10 the
full names of them
and how do you interpret this this one
plot that basically each dot is
represent uh represents uh each uh
patient and also the Red Dot indicates a
higher value of the feature and blow
down indicated lower value of this
feature
and the closer of the dot to the right
hand side of the x-axis that means the
feature is driving the model to give you
a positive prediction and the closer to
the left hand side that means the
feature is driving the model to give you
a negative prediction so for example if
we look at the top one the lowest oxygen
saturation level
that means that here
there's a lot of red dots here that
means that the patients were
with high oxygen saturated saturation
level then the model tends to assign
them high chance to for the recovery
well for the second one the lactate
and we see this
the the red also are located on the
left-hand side on the on the axis x-axis
and usually that means heightened likely
values you're already indicating this
organ damage going on so that's why the
model thinks these patients are not
going to recover from us
and for uh for this problem what we
found interesting is that we found two
organ systems the cardiovascular and
respiratory systems are critical for
this most recovery prediction because we
find this top 10 features most of these
top 10 features are more or less
associated with these two organ systems
so
in clinical practice we could probably
pay more attention to most patients with
these two type of types of other
dysfunction rather than treating an
equally compared to patients which there
are other types of organizations
so finally I would like to summarize my
presentation here in this study we
developed a machine learning framework
for predicting most recovery in
Pediatric Services patients for one week
in advance
and we conduct a comprehensive
experiment to show that
the
proposed model can not only just predict
most recovery with hierarchy but also
can be transferable across different
clinical sites to those unseen patient
data even in inter Continental setup
and also the prediction of the model is
interpretable from a clinical point of
view
so this means the model we developed
here could provide a
more insights to the clinicians or maybe
make they are more understandable for
them
and also we believe that our model have
certain clinical utilities that it could
probably assist clinicians
for better patient assessment and triage
on Day Zero the day of sepsis Stone set
and now we are doing something else
using the same data set that we
including the genomics and the
proteomics profiles of this patient and
we try to discover novel clinical
phenotypes of pediatric sepsis and by
doing a characterization of the Asian
cyber groups for hopefully with this
effort we could develop something for
better personalized treatment
and that's all for my presentation if
you want to have a
further information of this paper you
can scan the QR code here thanks a lot
I'm ready
[Applause]
thank you Bowen for the questions we're
going to be asked
foreign
for the great talk uh one question
regarding the clinical interpretability
that you mentioned in the end so uh how
would this look like is that these these
graphs that you show or is there more
behind it so how would a clinician
interpreter results yeah I think this
will be probably helped to help to
support their decisions like patients
with like for example low oxygen
saturation or high lactic value they
should pay more attention to that
or maybe all right but it's probably if
you have a certain uh patient right it's
um most probably the case that there
will be certain points which will be on
the blue and that's on the red right so
otherwise it would be a clear decision
so if I I get it if I see a picture and
it's it's always in the red then it's
clear and I can see also as a clinician
um yeah how the system comes to it yeah
we also actually provide some examples
for like how model how the model like
make the prediction for each individual
samples but I'm not here here it's in
the paper already so it's like a
survival score for each individual
patient now there will be much easier
interpretable as in not in the cover
level but individual level did you have
many more features then you said these
are the top 10 features that is top then
we have 44 features in order
okay and um did you also like quantify
like how much each feature kind of
contributes to the solution the the do
you say like with these 10 features I
capture pretty much I don't know 99
percent
or
oh yeah we actually I think we didn't
really qualify like how much exactly
this top 10 contribute to the model
District uh the prediction
now it's a physical point if we could
probably include this in future work but
but you said it's the first two right
that have the I only talk about this
first two here to just to Showcase like
how this model like take uh these
features into account when they're
making prediction
overall if you look at the resulting
performance of the system right you
could already show that basically it
doesn't matter which machine learning
message you use but it's always better
than the clinical status right but do
you think there's a way to push the
performance much further because there
was not at least not from the first view
that was not so significant differences
between the individual methods right
yeah that's right that's right and
performance is well it's below 0.8 right
so it's it's it's not
super reliable this is yeah do you think
there's a way to push it further or
I think yeah of course for machine
learning models the primary way will
always be to have more data since for
this uh Pediatric Services agents we
only have like
a cover size of less than 300 which is a
pretty little for this data driven
approaches yeah probably if possible
equally includes more data but it could
be very difficult
or we do a more exhausted research or
could be also this could lead to some
overfitting issues
okay thanks a lot
Leslie
thank you really talk um I have a
question on Slide 11.
could you explain again the so you
trained on the SPSS and could you
explain again why the the Chicago data
for the relative EU PRC first increases
and then decreases yeah or what's the
interpretation of that
I think
um it's a little bit difficult to
interpret this
or it's kind of a mystery of why that
phenomenon happens
yeah we didn't really try to interpret
this at the beginning so just because
there's external validation then we
could already access to the data and
just that
pass the model and let it run
so
if anyone could give a proper guess of
what is happening
it's a interesting artifact then okay
thank you
hmm
bonus one another question
um
so we always took for granted in this
project that the the sepsis had been
confirmed as a bacterial infection
is this actually known at the point of
time in which we would make our
prediction
so in other words there's an inclusion
Criterion for this data set
which where it's for me not totally
clear that this is like
given or determined is the right word
that this is determined at the time at
which we we make this prediction
so the is this the bacterial infection
we know at which time points this is
typically confirmed this is maybe three
days into the stay or something yeah
that's the the things that for this data
set is actually
that we use the blood contraction proven
uh proven bacterial infection that we do
the black culturally sampling and we
once we found the bacteria we make this
as a day exactly and that's the
inclusion Criterion of this this
Pediatric subsystem
um
but I think an open point is when
exactly is this determined is this maybe
determined after we make our prediction
because then there would be another like
latent set of patients on which we could
also make our predictions but which do
not fulfill this inclusion criteria
I think that would be an extremely
interesting external validation question
to test the model on patients who are at
risk for developing sepsis for example
and see whether your model will predict
them
to recover basically because the
recovery likelihood will be higher I
guess because it's not uncertained yet
whether they have sepsis or not
so other data sets like that available
this is the question
Maybe
so they have our National Data stream
for Pediatric research in in Switzerland
in fact and headed by one of our
collaborators on on the paper so they
are going to collect data sets like that
like in intensive care data sets for for
children for Pediatric patients and
there you could do this I mean what
what this data here represents is a very
clean data set in the sense that that
the sepsis has really been confirmed to
be connected to a bacterial infection
here on these data sets it's not not
necessarily the case for all things or
for all patients that are labeled
aseptic in the in a database yeah
and then then exactly the situation how
how much the system generalized if you
train on a high quality expenses to
obtain labels yeah
um I was just wondering if this Swiss
pediatric sepsis data set is public ah
so far it's not really for some privacy
issues okay
and also the externalization is also not
accessible to public
we wished even not not to us
maybe just an additional command what
you said is extremely important because
if we think about the translational
perspective of your modeling efforts
these models tend to be applied as the
as we go to more poorly defined samples
and therefore I think it's crucially
important to establish this right away
because to determine what the
operational window of these models is
another idea that came to my mind is
have you looked into decision curve
analysis to establish the value of your
model so basically doing that benefit
analysis how much do a clinicians gain
or a patient gain from a certain
prediction so by taking in the the costs
and the opportunities basically of a
prediction
yeah I think for this one it's actually
a retrospective study so
it's super difficulty whether to really
compare against like a clinical
condition decisions
so hopefully we can do some prospective
study like really put this into clinical
practice and we see really compared to
like the true decision made by the
extensions as this clinical Baseline
still sounds in a proxy to those
hey uh thanks a lot Pauline it was a
great presentation I was just wondering
um because you were talking about how
you had to basically hand in the model
and get the predictions back and you
couldn't have access to the data when
doing this external validation could it
be could this be a nice use case for
something like Federated learning or
swarm learning that we had presenting in
the network like early on like uh maybe
using something like that you can
collect data from multiple sites and
yeah I think it is a very good point and
even in this case it was extremely
difficult for us to like to communicate
and everything for the model development
also evaluation if we could like do this
anonymously using a federal learning I
think they'll be very useful and for
also for the other collaborations in
medical field
thanks thanks thank you a lot but if I
may add this project was interesting in
the sense that for us this was the
really the first example of what you
could call Federated learning so we sent
the models to yeah yes and they ran the
model there so so in all our other
projects we negotiate data access at
some point we get several data sets and
then we harmonize them and then we run
it this this was really done
decentralized so we only sent the code
we never got the inverse data from this
project
who was first
I'll start in the back yes
um do you know if there is any
stratification in yours in your sample
especially regarding the kind of
bacterial strain that causes Celsius
I'm sorry I didn't really get it um do
you know if there is maybe an effect
like like a strain specific effect
in terms of like the features that you
can that you get from your patient that
could have an impact on the performance
of your model in the end
maybe I can help you to interpret it
so whether we could stratify foreign
limited cover size we didn't do this
eventually for the subgroup analysis or
any stratification of patient or based
on whatever the side infection or the
pathogens
so
no no oh no we didn't use those as
feature because this may not be
available for this external validation
set
I'm coordinating the adult sepsis study
in Switzerland and we looked into this
point very much but the case numbers are
too small to to stratify for the type of
bacterial in fact for the pathogen that
causes sepsis so in in a like realistic
time Horizon so like a bigger like
collection of areas that is needed in
just a Swiss ball Juliano
with Mike
so I mean this phone said the the size
of the data set did not allow to
stratify but what we did too is like
look for enrichments of certain like
statistical enrichment of certain
strains in for example the like false
positives or false negatives just to see
whether for some strange the the model
was like struggling more than for others
but we didn't find anything like
exceptional
thank you and final question by Giovanni
thank you boys for the talk I have a
small curiosity also in addition to my
question apologies if I missed it but
why is the rate of Subs is so much
higher in children like is it because of
the immune system that's not really
developed more reassistant and also
they're more vulnerable in general
thank you
um so my question is since somebody has
mentioned the model is not like 100
reliable
um where is the situation in which you
will you would get some predictions
wrong and in the clinical setting not
all errors are made equal in this case
would it be worse to get like a false
positive or a false negative like what
would be the consequences of both
and I think yeah in this case I think
that's a very good question in this case
between the recovery protection not like
with your immortality prediction
so I would say it's
it's better to give a false negative
so we consider it was still like pay
attention these negative examples that I
would still pay attention to this
patients because the models think that
speech will now recover but hopefully
the patient eventually recovers they
will be good but if not we can still
have more additional attention on it
thank you thanks a lot good thank you
again Bowen and we move on to the next
speaker
[Applause]