SIB Bioinformatics Award ceremony 2025 - Early Career Award - Michael Skinnider
Watch on YouTubeVideo summary
Michael Skinnider, the recipient of the 2025 Early Career Award from SIB Bioinformatics, has established himself as a distinguished computational biologist with significant achievements across diverse fields including genomics, transcriptomics, proteomics, and metabolomics. His career trajectory began at Hamilton University in Canada, where he developed tools to discover bacterial secondary metabolites, before expanding his research into protein interaction networks and single-cell analysis during his MD-PhD studies. Notably, his work on identifying specific subpopulations of neurons contributed to breakthroughs in restoring walking ability for paralyzed patients through electrical stimulation of the spinal cord. His versatility was further highlighted by his successful defense of the MD-PhD program in Canada and his immediate appointment as a faculty member at Princeton University following his doctoral studies.
The core of Skinnider's award-winning research addresses a fundamental gap in biological science: while scientists can routinely measure entire genomes, transcriptomes, and proteomes, comprehensively identifying small molecules in metabolomics remains a formidable challenge. He describes this unidentified portion of the metabolome as "chemical dark matter," which his lab aims to illuminate using advanced computational methods. To tackle this problem, Skinnider and his team developed DeepMet, a chemical language model trained on the structures of all known human metabolites. This innovative approach treats chemical structures as text strings, allowing neural networks to learn biosynthetic logic similar to how they process natural language, thereby predicting the existence of previously unknown metabolites with remarkable accuracy.
The effectiveness of the DeepMet model was validated through a rigorous testing process that mirrored their earlier success in forensic chemistry with the DarkNPS model for detecting designer drugs. In one notable experiment, the team waited six months after training their model to see if it could predict emerging illicit drugs; remarkably, the model successfully anticipated over 90% of new substances discovered during that period. Similarly, DeepMet generated structures for novel metabolites that were subsequently synthesized and confirmed in human urine samples and mouse tissue data. This capability has already led to the discovery of more than 40 previously unrecognized mammalian metabolites, demonstrating how computational predictions can guide experimental efforts to map the complete landscape of mammalian metabolism.
Looking toward the future, Skinnider emphasizes that while chemical language models have successfully narrowed the vast search space of possible molecules, further improvements are needed to directly link generated structures with mass spectrometry data. His lab is currently exploring strategies such as self-supervised learning and utilizing unlabeled data to enhance these connections. He concludes by reiterating that the inability to fully measure small molecules in biological systems represents a unique opportunity to address fundamental scientific challenges through computation, a mission that continues to drive his research group forward.
Read the full video transcript
Uh it's my pleasure to introduce uh the
next winner of the early career award.
uh this one is typically given to the
young researchers in the bionformatics
or computational biology who completed
their PhD uh in the last six years
and this year winner uh is Michael
Skinned and he was selected because of
the great his great achievements in the
very different fields of binformatics
and also by his uh support or maybe
contributions to the binformatics
community.
Uh a few words about his career. So he
actually started with the single cell uh
RNA seek data. He did nice benchmarking
of the methods for differential
expression analysis. There he also came
up with the machine learning approach
agur uh to identify um specific cells
which are critical for the healing
process after uh um after walking
issues.
uh he also uh then moved to uh masspec
data uh to protoomics and uh later to
the metabolomics and there he developed
a a new let's call it metabolic language
model deep math to dive into the deep to
the dark matter of the metabolics data
and identify uh the new or until now not
known metabolites in the in the spectra.
for the community. He also developed
several art packages and um from the
other achievements uh he also helped to
defend and maybe rescue the MD PhD
program in Canada. And as a last uh kind
of piece of information
uh um he was offered a faculty position
at the Princeton University directly
after a PhD. So, please welcome to the
stage uh Michael Skinn Skinner.
>> Thank you very much for that uh very
kind introduction. Um I want to start
just by saying, you know, how grateful I
am to be to be speaking to you today. um
this is really an enormous honor and
it's made more meaningful for me by the
fact that many of the previous
recipients um of this award are people
who I've really looked up to and whose
work has been very uh meaningful and
inspiring to me. So um
I I thought I would tell you um a little
bit about um our work on developing
computational methods to help discover
previously unknown uh mamalian
metabolites. Um before I jump right into
that though uh I thought I'd just talk a
little bit about um you know my own
background and how we came to be uh
interested in working on this particular
problem. Um I am of course a
computational biologist by training. Um
that training really encompassed many
different kinds of biological and
biochemical data from genomics to
transcrytoics to proteomics in addition
to metabolomics.
So uh for instance I got my start in
research as an undergraduate student um
at Hamilton or a university in Hamilton
Canada uh where I worked with uh Nathan
McGarvey and my work there focused on
developing computational tools to
discover bacterial secondary metabolites
or natural products. Uh these methods
used as input genomic and metabolomic
data sets.
And so this is where I really first
developed a deep interest in metabolism
and not just metabolism but in
technologies for measuring metabolism.
Um notwithstanding that interest uh when
I moved to the west coast of Canada to
do my MDPhD I also moved into a
different field. Uh so my work with
Leonard Foster focused on mapping
protein interaction networks using a
technique called co-fractionation mass
spectrometry.
uh and this work culminated in the first
uh tissue specific protein interaction
networks in mammals.
Um during that time I also spent quite a
bit of time uh at EPFL where I worked in
Gregar Cortine's lab uh developing
computational tools to identify
subpopuls of neurons that uh produce a
particular behavior. this time using
single cell and spatial transcrytoics
data as input. And ultimately these
methods pointed us towards a subt type
of spinal cord neurons uh which can be
uh activated through electrical
stimulation of the spinal cord in order
to restore walking in human patients who
were previously paralyzed.
Um at Princeton as I said my lab has
really come full circle to to my
undergraduate interest in metabolism. Um
but now instead of uh focusing on
bacterial secondary metabolism uh we're
really much more interested in the
mamalian metabolom and particularly in
discovering metabolites associated with
the development progression or treatment
of cancer.
So I tell you all this in part to
explain what it is about uh metabolism
that I just could not stay away from.
um you know I I told you about very
different um themes and ideas in
genomics, transcrytoics and proteomics
and these are of course three very
different fields but um I would argue
that they all have one important thing
in common which is that in these fields
it's become routine to measure
essentially the complete complement of
DNA, RNA or proteins in any given
biological sample. And the situation in
metabolomics is really very different.
uh it is still today essentially
impossible to comprehensively identify
the small molecules in that same sample.
So this seems like a very profound gap
in our ability to study and understand
living systems.
Um it's a gap that I think is
particularly exciting uh from a
computational perspective because I
would argue that this gap is not for
lack of an appropriate uh experimental
technique. Uh in fact with mass
spectrometry we already have an
analytical technique um that can detect
or acquire data from thousands of
metabolites in any given biological
sample. And not only that but massspec
uh collects very rich and multifaceted
data for every one of these uh small
molecule metabolites.
So this seems much more like a
computational problem. How can we take
this rich mass spectrometry data and
deduce from it the structure of the
metabolite that was observed by the mass
spectrometer?
And that unfortunately turns out to be
uh sufficiently difficult that today uh
in a uh routine metabolomic study really
only a very small fraction of the
signals that are required by the mass
spectrometer will actually be identified
that is to say linked to the chemical
structure of the corresponding
metabolite. And so the rest of that data
is is commonly referred to as the
chemical dark matter of the metabolome.
So this uh dark matter is really what
we're uh focused on illuminating. Um
we're broadly interested in using uh
computational tools, many of them
leveraging advances in AI um to um
discover this hidden chemistry that's
sort of hiding in our routinely
collected metabolomics data and really
uh do that as a means to finally uh put
together a complete map of the mamalian
metabolism.
So that is quite a grand uh challenge
and you know you might ask how can we
possibly make progress uh on this. Um
I'll tell you a little bit about um the
logical starting point for an approach
that we've had a lot of success with
over the past uh 2 years or so. And uh
that that thought is sort of as follows.
Um over the last century uh scientists
have amassed an incredibly rich
understanding of the known mamalian
metabolom. So can we now leverage that
understanding in a systematic way to
predict the composition of the as of yet
unknown metabolom
uh and the computational approach that
we've been using to explore this idea is
based on uh chemical language models. So
the core uh concept here is that we
represent the chemical structures of
known metabolites as short strings of
text. We particularly like a format
called smiles that I'm showing you here.
And of course, we do this because
representing metabolites as text allows
us uh to repurpose the same kinds of
neural networks that have been so
successful at learning the syntax and
semantics of natural language. Uh and
instead of applying these language
models to the words in a sentence, we
can instead effectively apply them to
model the atoms and bonds in a chemical
structure. And so our hope was that this
approach might allow us to as I said
learn from the structures of known
metabolites uh to generate structures uh
for as of yet undiscovered metabolites
which we could then target for discovery
experimentally.
Uh I should mention that this was an
idea that we already had a certain
amount of proof of concept for uh in a
related field specifically that of
forensic chemistry uh where we had
previously used language models to uh
helped discover emerging uh elicit drugs
of abuse also known as designer drugs.
So a few years prior to this uh work on
the mamalian metabolom we had trained a
chemical language model on the
structures of all known designer drugs a
model that we called dark nps and after
we trained this model um we did
something a little bit unusual in the
field of machine learning which is that
we waited um we actually waited for 6
months and over that period of course um
new drugs of abuse were emerging on the
market around the world and they were
being discovered by forensic
laboratories.
And so at the end of that six-month
period, we asked how many of these
emerging drugs of abuse did our language
model successfully anticipate.
And remarkably, we found that uh dark
NPS successfully generated more than 90%
of these emerging elicit drugs in this
perspective test uh including the uh
structures that I'm showing you here.
We then also showed that we could
integrate the language models
predictions with mass spectrometry data
and that this allowed us to discover
previously unknown elicit drugs uh in
clinical samples as well as in law
enforcement seizures um with really a
remarkable degree of accuracy. And in
fact, we worked with the Danish National
Forensic Lab uh to discover or help
discover a new uh dissociative uh drug,
this derivative of uh the well-known
street drug PCP that I'm showing you
here.
So um all of this you know past work
really led to a lot of optimism that we
could use essentially the same approach
and rather than learning the medicinal
chemistry logic of illicit drug
synthesis we could instead learn the
biosynthetic logic of metabolism.
So uh we implemented that idea in a
chemical language model that we named
deep. uh deepmet is a model that's been
trained on the structures of all known
human metabolites and which as I'll show
you can guide the discovery of novel
mamalian metabolites uh and much of the
work that I'm going to show here uh was
led by Tony a PhD student in my lab.
So I mentioned that um we had trained
deep so that we could generate new
metabolite like chemical structures
which we could then target for discovery
and um we found that deepmet did an
excellent job of generating metabolite
like structures. Uh so good in fact that
a second machine learning model could
not tell apart the generated structures
from those of known metabolites.
And just as we had seen with dark NPS,
um we found that deepmet successfully
generated the vast majority of newly
discovered metabolites that were added
to metabolic databases after we had
trained our model.
Um these predictions in fact were good
enough that we could buy or synthesize
chemical standards uh for compounds that
Demet predicted ought to exist. Uh and
then we could use the data that we had
acquired from these standards to
actually discover these metabolites uh
in human samples. So here I'm showing
you two examples of metabolites that
DeepMet predicted likely existed as
human metabolites which we were able to
experimentally confirm in human urine.
Um we also again developed approaches to
integrate these predictions more
directly with metabolomics data so that
we could uh assign structures to many of
the unidentified uh peaks or signals uh
within the mouse metabolism. So for
instance here I'm showing you a uh mouse
metabolite that's very specific to the
kidney and pancreas. Uh we synthesized
the predictive structure shown at the
far left. Uh and uh we found that the
experimental data from this chemical
standard uh matched almost perfectly to
the corresponding signal in the mouse
kidney.
Um this is an approach that we've now um
been able to uh scale up to some extent
uh to the point that we've now used deep
to discover more than 40 previously
unrecognized mamalian metabolites. And
this is really exciting to me because uh
it suggests that um you know we could
really uh use this technology to
accelerate the mapping of the mamalian
metabolism.
So you know why does this approach work
so well? Um, I think a good explanation
is that chemical language models
effectively allow us to narrow the
search space. Um, when we're hunting for
a particular category of small molecule,
chemical space as a whole is sort of
incomprehensibly vast, but metabolite
like chemical space or designer drug
like chemical space are small subsets.
Uh, and so by narrowing our search to
consider only metabolite like
structures, we've essentially made this
uh problem a lot easier.
That being said, I I think that there
are still really important opportunities
to do a better job of linking these
generated structures to the mass
spectrometry data itself. And so this is
really uh a major focus of the lab going
forward. We have a few different ideas
about how to do this that are now
underway. Uh for instance, collecting
more data ourselves, learning from
unlabelled data, for instance, with
self-supervised learning. uh learning
from new types of data that are not
really typically considered in structure
annotation in metabolomics and then of
course uh simply developing better
machine learning models.
So I'll stop there um maybe just by
reiterating this idea that um our
inability to comprehensively measure
small molecules in biological systems is
I think um an opportunity to address a
fundamental challenge using computation
and this is really the the challenge
that uh motivates uh my lab. So um I'll
I'll stop just by thanking um the people
in my lab who worked on this of course
our funding sources and once again I
really want to express my gratitude to
the Swiss Institute of Bioinformatics.
Uh so thank you.
>> Thank you Michael. Congratulations. Uh
any question the audience? Yes please.
>> Yeah thank you very much for the
interesting presentation. Um I like the
test that you did you know just waiting
for design of drugs to be created and
see whether or not you could have
detected them and 90% is quite
impressive but then my question
automatically is actually how many of
them did you generated and linked to
that um since those are L&Ms I guess you
are able to rank which are the most
likely also in term of metabolites
>> yes that those are all that's a a great
series of questions. Um I sort of you
know uh simplified for the purpose of uh
just time but in fact we do um
you know assign uh a score to each of
the generated structures that sort of
correlates well with its uh likelihood
of uh occurring in the future. So the
90% number uh you know it reflects uh
the proportion of generative structures
that were detected but the denominator
here is huge. It's uh on the order of
millions. Um that being said in in both
that work and then in the deep work we
found that we could actually use that
score to sort of prioritize without any
analytical data at all which of the
generated structures are most likely uh
to emerge or be discovered in the
future.
And so I I imagine that's how you're
trying to detect the new metabolites
just looking at the most likely ones.
>> Exactly. And then using that same uh
idea to sort of uh reank our
interpretation of the mass spectrometry
data particularly the MSMS data itself.
Yeah.
>> Um sorry our schedule is quite packed.
We don't have time for another question.
I think that you could discuss during
the lunch. Thank you. Um thank you again
Michael.